
In 2024, Klarna announced that its AI customer service agent had replaced the work of 700 employees. The story ran everywhere. It became a reference point in board meetings, earnings calls, and investor presentations as evidence that the AI-driven workforce transformation was delivering.
The follow-up story received a fraction of the coverage.
By 2026, Klarna was quietly rehiring customer service staff. The AI had not been shut down. The company discovered it needed people alongside it: to handle situations the system flagged as unresolvable, to maintain the relationship quality that returning customers expected, and to carry the institutional context the system had never had access to.
Klarna was not alone. Fifty-five percent of leaders who made AI-driven workforce cuts now say they regret the decision, according to a 2026 Forrester survey. A Robert Half study found that 32 percent of hiring managers who eliminated a role citing AI had already rehired for the same or a similar function. More than a third of companies rehired over half of the eliminated roles, and most did so within six months of the original cuts.
The people coming back are not coming back at the same cost. Returning workers are landing pay increases of 20 to 35 percent. Gartner projects that by 2027, half of the companies that cut customer service staff for AI will rehire for similar functions, often under new job titles that obscure the reversal.
The scale of the correction is significant. The pattern that produced it is more significant.
Three Companies, Three Versions of the Same Failure
The reversals at Klarna, IBM, and Ford are distinct in sector, function, and scale. They share the same root cause.
Klarna: The 700 that came back
Klarna’s 2024 announcement pointed to measurable productivity metrics: the AI agent handled inquiries with the speed and consistency of 700 employees. The metrics were real. The problem was what they did not measure.
Customer service interactions contain two categories of work. The first is resolvable by policy: refunds within parameters, status updates, account changes. AI handles this well. The second is not resolvable by policy: a customer in genuine distress, an escalation that requires judgment about exceptions, a complaint that technically does not qualify for a resolution but where the cost of losing the customer exceeds the cost of making it right.
Klarna’s system handled the first category efficiently and flagged the second as unresolvable. What the company had not planned for was how much of its customer retention lived in the second category.
IBM: The 6 percent that exposed the other 94
IBM deployed AI across its human resources function. The system processed 94 percent of incoming requests without human intervention. That number was used to justify significant headcount reductions.
The 6 percent the system could not handle included ethical dilemmas, novel grievances that fell outside policy parameters, and situations where the outcome of an automated decision carried legal or reputational exposure. These were not edge cases in the colloquial sense: they were the cases that mattered most.
IBM is now tripling entry-level hiring across all business units. The company did not publicly frame this as a reversal. The underlying dynamic was straightforward: the organization had eliminated the people it needed to supervise, correct, and backstop the system it had built.
Ford: The 350 engineers and the metrics that passed
Ford’s quality control automation processed inspections faster and more consistently than manual review. By the metrics the system was given, it performed well.
The system was not given the right metrics.
Experienced engineers who had spent years on the floor had developed pattern recognition that extended beyond the formal checklist. They flagged anomalies that fell within tolerance but that experience told them would compound downstream. The automated system had no access to that experience and no framework for capturing it.
The defects the system passed went to market. Ford rehired and promoted more than 350 engineers. The rework cost more than the savings the automation had delivered.
The Organizational Decisions That Compounded Each Failure
The technology in each of these cases functioned as designed. The compounding failures were organizational.
The irreversibility of the headcount decision was not scoped
When a company replaces a function with automation, it makes a decision that looks reversible but is functionally irreversible in the short term. The people leave. Their institutional knowledge leaves with them. The informal networks, the undocumented exceptions, the judgment calls that never became policies: none of that transfers to the system, and none of it transfers back when the people are rehired.
The companies in these cases did not model the cost of reversal before making the decision. They modeled the cost of the roles being eliminated. Those are not the same calculation.
The performance metrics were inherited from the automation, not defined by the organization
In each case, the AI system was evaluated on the metrics it was designed to optimize. The organizations did not separately define what success looked like from a business outcome perspective and then verify whether the system’s metrics correlated with those outcomes.
IBM measured request resolution rates. Klarna measured response time and volume. Ford measured inspection throughput. In each case, the system performed well on its assigned metric, and the metric missed what the humans had been delivering.
This is a design decision, not a technology limitation. Someone in each organization could have asked: are we measuring the right things? That question requires technical judgment and business context simultaneously. In none of these cases was that person structurally positioned to be in the room before the decision was made.
The 6 percent was never mapped
Every automation deployment has a residual category: the work the system cannot resolve. In most enterprise implementations, this category is treated as a cleanup problem to be addressed at rollout, not a design requirement to be specified before the headcount decision.
The pattern across Klarna, IBM, and Ford is that nobody asked what the 6 percent looked like before the workforce decision was made. By the time the answer was clear, the people who had been handling it were gone.
What a Prepared Organization Looks Like
The companies that avoided this outcome shared a common characteristic: they made the augment-or-replace question explicit before reducing headcount, and they had someone with the right technical and organizational context involved in answering it.
In practice, that means four things.
Define the residual category before making the workforce decision. What work will the system flag as unresolvable? Who handles it? What does that function cost, and how does it change the ROI calculation?
Map what humans are actually delivering, not what their job descriptions say they deliver. The informal work, the judgment calls, the escalation paths that exist outside the documented process: these need to be made visible before the decision, not after the reversal.
Run the automation alongside existing headcount long enough to see what the 6 percent is made of. The pilot period is not complete when the system is stable. It is complete when the edge cases have surfaced and been assessed.
Treat the reversal cost as part of the initial ROI model. If the decision is wrong, what does it cost to correct? That number belongs in the analysis before the decision is approved.
None of this requires slowing down AI adoption. It requires that AI adoption decisions be made with the same rigor applied to any other irreversible capital commitment.
The Advisory Gap Behind the Pattern
The pattern that produced these reversals is not an AI pattern. It is a decision structure pattern.
The conversation that should have happened before each headcount decision involved three components: a technical assessment of what the system can and cannot do, a business assessment of what humans are actually delivering, and a risk assessment of the cost of being wrong. That conversation requires someone who can sit at the intersection of all three.
In most of the organizations now paying premium to reverse their decisions, that person was not in the room. The conversation that happened was between finance, HR, and senior leadership. The AI system’s actual capability boundary was never mapped. The residual work was never scoped. The irreversibility was never costed.
The result is now in the data. Fifty-five percent regret. Twenty to 35 percent pay premiums to recover. Rework that exceeded projected savings. A Gartner projection that half the sector will repeat this sequence by 2027.
The question worth asking before the next AI workforce decision: who is answering the second-order questions before the decision is final?
Not after the pilot surfaces problems. Before the decision is made.
Michael Snyder is a Fractional CTO working with founders and CEOs at growth-stage companies when technology decisions carry consequence. Based in Austin, TX.
If your organization is making AI-driven workforce decisions in the next twelve months, the Structural Clarity Diagnostic is a structured starting point for mapping what is actually at risk before the headcount decision is final.
