Skip to main content

writing

what the stopped projects know

gartner thinks 40% of these projects get canceled. the failures teach more than the wins, because they show you what the tools were amplifying.

the most useful ai project to study is one that got killed halfway through. the industry is about to hand you plenty of them.

gartner predicted in june 2025 that more than 40% of agentic-ai projects would be canceled by the end of 2027, pointing to escalating costs, unclear business value, and inadequate risk controls. set that beside mit nanda's finding that roughly 95% of enterprise generative-ai pilots showed no measurable p&l impact, and the next two years get hard to misread. most of this work stops, or fails to matter. cancellations don't get their own press release.

it's worth being honest about why they die, because the ready-made explanation is wrong. the story people reach for is that the tools weren't ready. what we've watched is closer to the reverse. the tools worked. they did the one thing tools do: take what they're handed and make more of it, faster. the trouble was in what they were handed.

the old, slow way of working did two jobs at once, and only one of them was ever visible. it produced the deliverable, and underneath that it quietly caught things. a vague brief got firmed up over a couple of review cycles. a shaky grasp of process got corrected by the people downstream who had to build on it. work passed through enough hands that most of the weak parts were intercepted long before anything shipped. the friction everyone complained about was the organization's immune system.

ai took the friction out, and everyone was glad to see it go. the interception went with it. there is now nothing standing between a person's real fluency and the thing that ships. so when the output is bad, calling it ai slop misses what happened. the ai faithfully amplified an input that was already weak, and no one was left in the middle to catch it.

that is the mechanism under most of the cancellations. a pilot gets scoped by someone who can't yet tell a working demo from a working system. it gets promised on a timeline by someone whose sense of what's feasible was never earned by shipping anything. and it gets evaluated by no one. evaluating it well takes the exact craft the org decided it could automate away: process, standards, a real feel for what good looks like. two quarters in, the gap reaches someone with budget authority. pulling the plug turns out to be the first well-informed decision anyone has made on the project.

elizabeth stone, who runs product and engineering at netflix, sees this from the other direction. on lenny's podcast recently, she said some of the code these agents write is genuinely hard for her to follow. she can see that it performs better. she can't always see why. that sounds like the same problem the failed pilots have, but it's the opposite. stone knows what good work looks like, so she knows exactly where her own understanding stops. that is the safe way to not understand something. the dangerous way is trusting the output with no way to check it.

this is why a reversal teaches more than a win. a win worked inside one company's data, its constraints, its tolerance for being wrong. you can admire it. you can't lift it into your own situation, though nearly everyone tries. a reversal carries over. it shows you what the team believed at kickoff, what they found once the thing was real, and the belief they had to let go of. that belief is usually the one you're holding right now.

some reversals are public enough to study from your desk. klarna spent early 2024 telling everyone its support assistant was doing the work of 700 agents. by may 2025 its ceo was saying the cost focus had produced lower quality, and the company started hiring humans back for support. mcdonald's ran ai voice ordering with ibm at more than a hundred drive-thrus and ended the test in 2024. zillow shut down its home-buying business in 2021 after the pricing model kept overpaying for houses. none of these companies were careless. they believed something reasonable, tested it at scale, and stopped when the evidence came in.

most teams save this material for the postmortem, which is the one moment it can't help anyone. it belongs at kickoff, as a premortem: sit the team down before the work starts and list the risks that could show up during the project, while there's still time to do something about them. before the budget is committed and everyone is sold on their own optimism, put two or three halted projects from your own sector on the table. write down, in advance and in plain language, what would make you stop this one. name the person allowed to make that call. it reads as pessimism in the room. really it's the cheapest chance you'll get to agree on a stop condition, before anyone's invested in not stopping.

there's a tell in how projects end, and it's worth knowing before you start one. not every cancellation is a competence problem. gartner's own list includes cost escalation and unclear business value, and those kill capable teams too. but the two failures die differently. a team that knew what it was doing tends to stop early, on a condition it set for itself in advance. a team that was in over its head tends to drift, quarter after quarter, until finance stops it from the outside. how a project dies tells you which kind of team was running it. if it's your project, it tells you which kind you are.

there's a regulatory reason to formalize the stop condition, too. the eu ai act's post-market monitoring duties for high-risk systems assume you have a defined way to notice that a deployment has gone wrong and act on it. a written stop condition, agreed at kickoff, used to be good practice. it's starting to look like a document a regulator will expect you to already have.

**Sources**

- Gartner, "Over 40% of Agentic AI Projects Will Be Canceled by End of 2027" (press release, June 2025) — a prediction, cited as such. - MIT NANDA initiative, "The GenAI Divide: State of AI in Business 2025" — the 95% figure (attributed). - Klarna — the Feb 2024 "work of 700 agents" claim is the company's own; the May 2025 reversal and "lower quality" admission are CEO Sebastian Siemiatkowski's (Bloomberg interview and subsequent coverage). - McDonald's–IBM automated order taking — test at 100+ drive-thrus ended June 2024 (CNBC, Restaurant Business). - Zillow Offers shutdown — company announcement, November 2021. - EU AI Act, Regulation (EU) 2024/1689 — post-market monitoring obligations for high-risk systems.

The campaign diagnosed the *symptom* (accountability unassigned, infrastructure skipped, stop conditions unwritten). Chris named the *cause*: **ai is a magnifier with nothing to magnify.** The old, slow way of working laundered competence — friction was the org's immune system, and enough hands intercepted weak work before it shipped. ai removed the friction, and with it the interception, so for the first time nothing sits between a person's real fluency and the output. what reads as "ai produced slop" is usually "ai amplified an input that was always weak, and no one was left to catch it." Stone is the proof from the top: she can't always follow the agents' code but she knows what good looks like, so she knows what she can't see — that's the dividing line.

- **Where it's woven:** essay 2 (the spine + "how a project dies is the tell"), essay 1 (the lighter "owning output requires being able to judge it"), and Beat 4 LinkedIn (the tell, in miniature — flagged as diverged, above). - **Standalone option — DRAFTED 2026-08-03** at `magnifier-standalone-draft.md` (nothing scheduled, nothing published).

Chris put the open question plainly in a /today walk: *"is it exposing less competent people or enabling fluent SMEs and ICs?"* His own lean, same session: the interesting story is **enablement** — skill-carrying agents multiplying people who already have judgment — not single-agent replacement of SMEs, and not an indictment of anyone.

That re-points the thesis without discarding it. The mechanism is unchanged: no friction left between a person's real fluency and what ships. What changes is which population leads the telling. An SME with real judgment was never short of ideas, only of hands and calendars; several agents, each carrying a skill she would otherwise hire or wait for, is a real change in what she can do in a week. The exposure case is the same mechanism pointed at someone with no one left to check them. Stated once, as diagnosis, not as a verdict on anybody.

The practical consequence, and the reason it matters commercially: uniform rollouts distribute licenses evenly and judgment not at all. Sequencing (who gets the agents first, and who writes down what good looks like for the review step) is the lever.

**Where the re-spined version lives:** `magnifier-standalone-draft.md`. The scheduled beats still carry the original exposure-led spine — re-pointing them means editing already-scheduled posts, which stays parked under Open decision #3.

1. **AI-content flagging risk (Chris's flag).** Substack may start notifying readers about AI-generated content. The magnifier weave + a real lived detail from you is the best defense (specific beats generic). Still your call per essay: post as-is / hold / personalize before scheduling. 2. **Your swing + the two anecdote slots.** Essay 1 and essay 2 now carry the thesis; each still has one natural slot for a real thing you've lived (essay 2 has an HTML-comment marker at the reversal). Drop it in or cut it — both stand. 3. **Scheduler sync.** All 5 LinkedIn beats + X Beats 1/3/5 diverge from this doc — see the "Scheduler sync list" at the top. On your go: sync those 8 posts + schedule the 2 essays in one pass (Sat Aug 22 8am · Sat Sep 24 8am). 4. **Held beat 3.5 / "the magnifier" standalone.** See above — greenlight to build.