This isn't another "AI will replace you" post. It's an attempt to describe where things actually stand right now, what the evidence says, and what it might mean for those of us who ship products for a living.

Let me say the quiet part first: I bought into some of the hype. I tried building a few things with coding agents because everyone around me was, and the experience was a strange mix of genuinely impressive and quietly expensive. That tension, between what the tools can do and what they cost to run, turns out to be the whole story. So here is the longer version, with the receipts.

The pitch, and the boomerang

AI was sold as an almost omnipotent stand-in: something that could do anyone's job faster and cheaper than a human doing it. For a while, the pitch looked plausible, and a wave of companies acted on it. Then reality arrived.

Klarna is the headline example. The Swedish fintech replaced roughly 700 customer service roles with an OpenAI-powered assistant, froze hiring, and publicly celebrated the savings ahead of its IPO. Within about a year, CEO Sebastian Siemiatkowski reversed course and began rehiring humans. His own explanation, given to Bloomberg, was that cost had become too dominant a factor in the decision, and the result was lower-quality service. Customers, it turned out, still wanted a human on the other end when things got complicated.

Klarna is not alone, and that is the important part. IBM cut around 8,000 roles, much of it in HR, and routed the work through an internal AI system called AskHR. The bot handled routine queries and documentation fine, but struggled with anything requiring empathy or judgment, and IBM ended up rehiring to fill the gaps (as reported by The Wall Street Journal). Duolingo went "AI-first" in an April 2025 memo that tied employee evaluations to AI usage, then quietly walked back the most contentious pieces after staff pushed back on using AI for its own sake. Salesforce trimmed thousands of roles on the logic that it needed fewer people with AI in the mix, then spent the following months tempering that confidence.

The pattern now has data behind it, not just anecdotes. Forrester's Predictions 2026 report found that 55% of employers regret laying off workers for AI-related reasons. Gartner projects that half of all AI-attributed layoffs will be reversed by 2027. Talent firm Robert Half has reported that close to a third of companies that cut staff for AI have already reopened those exact roles. Analysts have started calling it the "layoff boomerang," with one catch worth noting: when the roles come back, they often come back offshore or at lower pay. The correction is real, but it is not a clean reset.

Why the models still cannot just take over

Two structural limitations keep showing up, and neither is going away with the next model release.

The first is hallucination. Models still fabricate confidently, and in any workflow where being wrong has consequences, that is a hard ceiling rather than a rough edge.

The second is subtler and, for decision-makers, arguably more dangerous: sycophancy, the tendency of these systems to tell you what you want to hear. This is not a vibe; it is measured. A 2026 study in the journal Science, led by researchers at Stanford, tested 11 leading models and found their responses were nearly 50% more sycophantic than human responses, even when users described unethical or harmful behavior. OpenAI had already lived through a public version of this in April 2025, when it rolled back a GPT-4o update that had become so flattering it was agreeing with users who were plainly wrong. Anthropic has published the most work, at least publicly, on identifying and reducing the behavior, and describes deliberate "character training" to make Claude push back when it should.

Why does this matter for product work? Because the same study found that people trust and prefer the sycophantic answers. A model that validates your roadmap, your prioritization call, or your read of a user interview is pleasant to work with and quietly corrosive to good judgment. The tool that agrees with you is not the tool that makes you better.

The productivity paradox

Here is where the evidence gets genuinely uncomfortable for the "it just works" crowd.

The widely cited number is from MIT's NANDA initiative and its "GenAI Divide" report: roughly 95% of corporate generative-AI pilots delivered no measurable impact on the bottom line, while only about 5% drove real revenue acceleration. The report's authors were clear that the failures were less about model quality and more about how organizations integrated the tools; buying from specialized vendors and building partnerships succeeded around 67% of the time, while internal builds succeeded roughly a third as often. The figure went viral, and it has since drawn fair criticism on methodology, so I would treat the exact percentage as directional rather than gospel. The direction, though, is hard to dismiss.

Then there is the study that should make every "AI made me 10x faster" claim pause. In July 2025, the nonprofit METR ran a randomized controlled trial, the gold-standard design, with 16 experienced open-source developers working on 246 real tasks in repositories they knew well. The result: when AI tools were allowed, tasks took 19% longer on average. The kicker is that the same developers predicted a 24% speedup beforehand and still believed, after the fact, that they had been about 20% faster. They were measurably slower and felt faster. (METR has since been transparent that a follow-up experiment ran into selection bias, because developers who love AI increasingly refuse to work without it, so the firm now frames the productivity question as genuinely open rather than settled. That honesty is itself worth something.)

Put those two findings together, and a useful picture emerges. The wins are real but concentrated, and the gap between perceived and actual productivity is large. The 5% who win tend to do one thing well. Duolingo, after dropping the "AI-first" theatrics, reports making four to five times as much content with roughly the same headcount and raised its 2025 revenue projection past a billion dollars. Spotify's senior engineers have shifted toward supervising AI-generated code rather than writing syntax by hand, pairing Claude Code with an internal agent system. Tools like Slack and Miro keep adding AI that speeds up the work without pretending to replace the worker. The thread connecting all of them is augmentation, not substitution.

The cost wall nobody priced in

Anthropic genuinely cracked usable coding agents. That is not in dispute. What is in dispute is whether anyone can afford to run them at scale.

Uber is the clearest cautionary tale. According to reporting in The Information, its CTO said the company burned through its entire planned 2026 AI coding budget in four months. Individual engineers were spending between $500 and $2,000 a month on tokens. Claude Code usage inside the engineering org jumped from around a third to over 80% in months, and something like 70% of committed code now originates with AI. The tool was so effective that the constant use is precisely what broke the budget.

Microsoft hit the same wall. Its Experiences and Devices division, the group behind Windows, Office, Outlook, Teams, and Surface, pulled most internal Claude Code licenses after token billing blew past the annual budget, directing engineers to GitHub's Copilot CLI by a June 30, 2026, deadline (first reported by The Verge). It is worth being precise here, because the headline version of this story is misleading: Microsoft dropped the coding tool in one division, not Claude across the company. It still uses Claude models through Microsoft Foundry and Microsoft 365 Copilot. The shift that matters is the move from flat-rate licensing to token-based, pay-as-you-go billing, where the meter runs every time the agent works, and finance teams have no good way to forecast or cap it.

It starts to resemble manufacturing. Some shops run expensive robots because the economics justify it. Plenty still rely on cheaper manual labor, because for a lot of tasks, the robot does not pay for itself yet. AI coding agents are settling into the same uneven landscape, and "is it worth it here" is becoming a line-item question rather than a foregone conclusion.

So, where does this leave product managers?

For now, I see AI showing up in our work in three distinct shapes.

1. Productivity enhancers. Note-takers, dictation, transcription, and whiteboards that pull project data from across the organization into one place. These are the safest bet and the least glamorous. They quietly make us better at the job we already do, and the MIT data suggests this kind of targeted, well-integrated use is exactly where the 5% of winners live.

2. Compressors of time-to-market. Tools like Lovable and Bolt collapse the distance between an idea and a working prototype. The real value is not that they replace engineering; it is that they let us test more directions before committing to one. More cheap experiments, earlier, means better-informed bets.

3. Strain amplifiers. This is the one I think is underdiscussed, and it is the reason I am not fully in the doom camp or the utopia camp. If a team can ship more code with agents than it ever could with people alone, then product managers face more buildable options than we can possibly validate. The bottleneck does not disappear. It moves to us. And when the number of plausible decisions multiplies, choosing the right one gets harder, not easier.

The constraint that AI does not touch is the human on the other side. You cannot ship 20 features a day, not because you cannot build them, but because users are still flesh and blood. They need time to adapt, to learn, to onboard. Real-world testing with real customers has to come first, no matter how fast the code arrives. Speed of production and speed of adoption are two different clocks, and only one of them is accelerating.

Where I land

We are still in the "let's see what happens when the dust settles" phase. The honest position is to judge what is happening now rather than what the vendors promise, because taking AI companies at their word has been a reliable way to get surprised by reality. The reversals, the costs, and the productivity studies are not arguments that AI is fake. There are arguments that it is a tool with a price and a profile, and the teams that win are the ones who figure out where it actually fits.

For PMs specifically, the job is not disappearing. If anything, the part of it that is hardest to automate, deciding what is worth building and protecting users from the firehose, is about to matter more.

What about you? Where do you think this settles? I would love to hear it.

Sources

  • Klarna reversal: Entrepreneur / Bloomberg reporting, May 2025

  • IBM AskHR and rehiring: The Wall Street Journal, via ACS Information Age

  • Duolingo "AI-first" walk-back: andreaiorio.com; augmentation figures via CNBC (Sept 2025)

  • Layoff regret and reversal data: Forrester Predictions 2026; Gartner; Robert Half

  • AI sycophancy: Cheng et al., Science (2026); OpenAI GPT-4o rollback (May 2025); Associated Press coverage

  • "95% of pilots fail": MIT NANDA, "The GenAI Divide: State of AI in Business 2025," via Fortune

  • Developer productivity RCT: METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (July 2025), arXiv 2507.09089; METR experiment-design update (Feb 2026)

  • Uber AI budget: The Information, via The Next Web

  • Microsoft Claude Code in Experiences and Devices: The Verge, via Cybernews and Windows Central

  • Spotify engineering shift: Q4 2025 earnings commentary

Keep reading