Back to Blog
ai-insightsJuly 6, 202620 min read

Firing the Verifiers

Sometime in the last year, a quiet shift happened in how we talk about artificial intelligence. The conversation moved from "what can it do" to "who can we replace." Earnings calls now reference AI as a labor strategy. Voluntary retirement programs cite AI-driven transformation in their framing. A senior AI executive recently estimated that all white-collar jobs would be replaced within eighteen months, and the framing wasn't a warning. It was a roadmap.

*How the AI economy is hollowing out the immune system of knowledge work*

Sometime in the last year, a quiet shift happened in how we talk about artificial intelligence. The conversation moved from "what can it do" to "who can we replace." Earnings calls now reference AI as a labor strategy. Voluntary retirement programs cite AI-driven transformation in their framing. A senior AI executive recently estimated that all white-collar jobs would be replaced within eighteen months, and the framing wasn't a warning. It was a roadmap.

Underneath the productivity narrative is a thesis nobody is naming clearly: that AI's value comes from removing humans, and that the humans being removed weren't doing anything that mattered. Both halves of that thesis are wrong, and they're wrong in the same way.

The error sits at the layer most discussions about AI skip: verification. AI systems generate fluent output regardless of accuracy. The quality of what gets produced depends entirely on a human who can tell good from bad. That capacity isn't a nice-to-have. It is the thing that turns AI output into something usable. Strip it out, and you get volume without value.

We are simultaneously deploying AI to people who can't catch its errors and removing the people who could. The error-catching capacity of the entire knowledge economy is being hollowed out at both ends. This essay is about what happens when that's true at three different scales (individual, organizational, and economic) and how each scale tells the same story through the same observable signal: we have started measuring tokens instead of judgment.

Act I: The expertise paradox

Spend an hour watching someone new to AI use it for the first time and you will see the same thing every time. They ask a question. The model answers fluently. They accept the answer. They move on.

Now spend an hour watching a domain expert use AI in their own field. They ask a question. The model answers fluently. They squint. They ask a follow-up. They cross-check. They reject half of what they get and edit the rest.

This is the expertise paradox: AI is safest in the hands of the people who need it least.

The paradox has a cruel internal structure. Verification of AI output requires roughly the same expertise as producing the output unaided. A novice can't catch hallucinations because they don't know what the truth looks like. An expert can catch them but mostly didn't need the answer in the first place. They got speed, not capability. The middle is the danger zone: enough knowledge to feel confident, not enough to catch errors near the edge of competence. Dunning-Kruger, accelerated by a confident-sounding interface.

Three sub-mechanics deepen the problem.

**Calibration asymmetry.** Experts know what they don't know. They can locate the edges of their competence and treat AI output more carefully there. AI confidently doesn't know what it doesn't know. It produces wrong answers in the same tone as correct ones. When an expert meets the model, the expert's calibration catches what the model's lack of calibration produces. When a novice meets the model, there is no calibration on either side. The wrong answer passes through clean.

**Skill atrophy.** The novice who most needs AI is also the one whose foundational skills will atrophy fastest, because they're outsourcing exactly the work that would build verification capacity. The tool meant to grow them entrenches the gap. A junior engineer who never debugs their own code by hand will never develop the intuition for what bugs look like. A medical student who reaches for AI on every differential will never build pattern recognition for the case that doesn't fit the pattern. We are creating a generation of practitioners whose ceiling is whatever the model can do, because their floor was never built underneath them.

**Trust portability.** Humans generalize trust badly. One good answer about something easy becomes blanket faith in the system. Smooth fluency reads as uniform competence even when it isn't. The novice who gets a correct answer to a simple question concludes the system is correct in general, and then trusts it on a question they had no business outsourcing.

The cruel summary: at the individual level, AI is most valuable to people who already have the metacognition to know what they need, articulate it clearly, and verify what they get back. That set of people is small. For everyone else, the tool that promises to be an equalizer is a multiplier of existing cognitive skill. It widens the gap rather than closing it.

You can already see the artifact this produces. Open any platform where AI-generated work circulates and the output is recognizable: confident, well-structured, broadly correct, specifically wrong in ways the author didn't notice and the reader can't verify. Verification, the one scarce resource, isn't being measured. What's being measured is volume.

This is where we run into the metric.

Act II: Firing the verifiers

Inside companies, the verification problem stops being a learning curve and starts being a budget item.

The people best equipped to verify AI output are also the most expensive people in the organization. Senior engineers. Tenured analysts. Mid-career PMs. The institutional memory of a function. They are also, structurally, the people most exposed when the productivity narrative around AI says we can do more with fewer. Voluntary retirement programs cite AI-driven transformation. Tech layoffs across the industry cite the need to "invest in AI" as financial justification. Leaders at the largest AI labs publicly state that eighteen months will eliminate white-collar work.

These are not separate decisions. They are pieces of the same logic: remove the people whose job involved telling good from bad, on the assumption that the new tool no longer needs them.

The logic falls apart on contact with reality, and it falls apart for the same reason the individual paradox falls apart. The output doesn't verify itself. Removing the verifiers doesn't make the output better. It makes the output less examined. The organization gets more of what looks like work and less of what produces value. For a quarter or two, the difference is invisible, partly because the people who would have noticed are gone, and partly because the metrics have quietly changed.

This is where token-counting enters as a workplace metric, and it deserves more attention than it's getting.

When companies measure AI productivity, they are increasingly measuring tokens consumed, tokens generated, queries per developer, lines of AI-assisted code shipped. These metrics share a property: they measure the volume of AI-produced output, which is the dimension AI is reliably good at. They do not measure whether the output was correct, whether it solved the right problem, or whether it shipped something the company will regret in eighteen months. The metric confirms productivity by ignoring quality. You couldn't design a sharper symbol of the verification crisis if you tried.

Token-counting is the workplace version of the trust-portability problem from Act I. A leader who can't personally verify the output of a senior engineer's work has always relied on proxies: code review, test coverage, peer reputation. Those proxies were imperfect, but they were calibrated to the work. Tokens are not. Tokens count fluency. The leader watching the token dashboard go up has no way to know whether the team is shipping value or accumulating debt, because the people who could have told them have either been laid off, taken voluntary retirement, or learned that flagging quality concerns reads as resistance to the new strategy.

There is a second-order effect worth naming. Once an organization commits publicly to AI as a cost-cutting story, the social cost of contradicting that story becomes high. A senior person who points out that AI output requires more verification than it appears to is now arguing against the strategy that justified their colleagues' layoffs. The institutional incentive to stay quiet is strong. The verifiers who remain learn to verify less visibly, or to verify less.

What you end up with is a system that produces more output, measures more output, and has fewer people in any position to say whether the output is any good. The immune cells are gone. The metric confirms productivity. The decay is invisible until something breaks publicly, and by then the people who would have caught it earlier are no longer there.

The framing matters here. AI was sold to organizations as an augmentation story: make your people more capable. It was bought as a replacement story: do the same work with fewer people. The same product, in the hands of the same vendors, was both narratives at once, depending on who was being addressed. The customer-facing story was multiplication. The shareholder-facing story was subtraction.

Both stories can't be true. The math forces a choice. And the choice that has been made, by most large organizations operating in public, is subtraction.

Cutting the apprenticeship loop

The replacement choice has a second-order effect that does not appear on quarterly dashboards. Every organization that builds something complicated runs an apprenticeship loop, even when no one calls it that. Junior engineers ship the code that senior engineers review. Junior analysts build the models that mid-level analysts critique. Junior associates draft the briefs that partners redline. The work the junior produces is rarely valuable on its own, and most of it gets rewritten. The value is in the feedback loop. The junior learns what good looks like by having theirs corrected, repeatedly, by someone who can tell. Ten years of that cycle is what produces a senior. There is no other path to that outcome, and there never has been.

AI is now being inserted into that loop in a way that breaks it from both ends.

At the top, senior reviewers are being removed as cost-cutting, on the implicit assumption that AI can do the review. At the bottom, junior roles are not being hired, on the implicit assumption that AI can produce the draft. The work itself still gets done. AI drafts. AI reviews. Output ships. The loop closes algorithmically, and the dashboards show no change.

What disappears, quietly, is the apprenticeship the loop was producing as a byproduct.

For two or three years, this looks fine. Possibly better. The system's output volume holds steady or rises. The metric stays green. What is not on any metric is the maturation curve of the people inside the system. The juniors of 2025 who were not hired do not become the mid-levels of 2028, and the seniors who were retired or laid off in 2024 do not train the juniors who were not hired. The bench in 2030 is the second-order consequence of decisions being made now, on dashboards that do not track it.

This is a familiar failure pattern at the firm level. It is what happened inside the rating agencies and the investment banks in the decade before 2008. Once the senior people who understood mortgage-backed securities cashed out, retired, or got moved to other lines of business, the institutional knowledge of what good looked like left with them. The juniors who took their seats had never been trained to evaluate the structures they were now rating, because the people who would have trained them were gone. The verification capacity inside those institutions did not fail because of any single bad decision. It failed because the apprenticeship loop had been quietly broken for a decade, and no one was measuring it until the consequences appeared all at once.

The biological framing from earlier sharpens here. A company is not a fixed set of immune cells. It is an entity that regenerates immune cells over time, by exposing junior cells to challenge under senior supervision until they mature. Removing the seniors removes the trainers. Removing the juniors removes the candidates. A firm can do either of those for a while and still function. It cannot do both. The organism stops regenerating, and the next external challenge will meet a system that no longer has functioning immunity at any layer.

The honest objection is that some parts of every junior job were always grunt work that produced no real learning, and AI eating that grunt work is fine. That is true for a slice of it. The risk is that what looks like grunt work from the outside often turns out to be the substrate of learning. A junior who spends six months writing tickets badly and getting them rewritten by a senior is not just learning to write tickets. They are learning what shipping software actually looks like, what hidden coupling feels like in practice, and where senior reviewers tend to push back, and why. Replace that six months with AI-drafted tickets that go straight to production review, and the junior has shipped more tickets while learning less about software. The output went up. The engineer did not develop.

The verification crisis as I have described it so far is about the people being removed. The succession problem is about the people who are never being made. The same productivity dashboards confirm both decisions, and the same metric, token volume, is used to justify each one. The firm that automates its junior pipeline and retires its senior reviewers is doing both at the same time and calling the savings a strategy. The strategy has a name. It is decapitalizing the institution's capacity to verify its own work, on a timeline the current decision-makers will not be around to see.

What the productivity dashboards are not pricing in is that the verifiers of 2030 do not exist yet, because the firms that needed to be making them in 2024 and 2025 decided that making them was no longer worth the cost.

This brings us to the third scale, where the math gets harder to ignore.

Act III: The subprime knowledge economy

The financial crisis of 2008 had a specific shape. Subprime mortgages, loans to borrowers with poor credit, were bundled with higher-quality mortgages into securities that were then rated as if the diversification eliminated the risk. The assumption was that defaults across tranches would not correlate. Housing prices would not fall everywhere at once. Lower-rated mortgages dressed up as AAA were not actually AAA, and when correlation broke the model, the whole structure fell.

The clearest visualization of how that structure worked came in The Big Short, where the failure was illustrated as a Jenga tower. Each block was a tranche of mortgage debt. The bottom blocks, the riskiest and most subprime tranches, were what everything else rested on. The market told itself the tower was stable because the top blocks were safe. But once enough of the bottom blocks were pulled, the rest came down regardless of what they were rated. The blocks at the top didn't fail because they were bad. They failed because the structure they were part of had been hollowed out underneath them.

The verification failure that made this possible was not opaque. Rating agencies gave AAA ratings to mortgage-backed securities they had not meaningfully scrutinized. Investment banks structured products their own traders knew to be defective. Regulators relied on the ratings, the banks, and each other. At every layer of the system, the responsibility to verify the underlying quality of the asset was outsourced to the next layer, until no one was doing it. The system kept producing AAA-rated paper anyway, because the metric, the rating itself, confirmed that things were fine. The volume of paper went up. The amount of verification behind each unit of paper went down. The metric stayed green until the day it didn't.

The shape of what's happening to information work now is similar but inverted.

Information workers, the professionals in tech, finance, consulting, law, engineering, and media, have historically been treated by lenders as AAA-grade credit risks. Stable employment. High income. Low default rate. The entire consumer credit infrastructure of major economies depends on this population existing as a stable base. They take out mortgages. They carry credit card balances. They support the asset prices in the cities they live in. Their spending is the demand floor under the consumer-facing businesses that employ everyone else.

If AI displacement does what its loudest proponents say it will do, replacing a meaningful fraction of white-collar work, then AAA-rated borrowers begin to default not because they were misclassified, but because the underlying premise of their classification (stable employment in cognitive work) becomes unreliable through no fault of their own. That is the inversion. 2008 was bad mortgages dressed up as good. This is good mortgages becoming bad.

The mechanics that broke the 2008 model break this one too, for the same reason: correlation. The diversification assumption inside the mortgage book assumes defaults across the AAA tier are uncorrelated. AI displacement is structurally correlated. It hits cognitive work, which concentrates in specific geographies (high-cost-of-living metros), specific industries (tech, finance, consulting, professional services), and specific income brackets (the ones with the largest loan balances). The diversification doesn't hold. The geographic correlation is already loaded into the book.

The first cracks would not show in the biggest banks, which are better capitalized post-Dodd-Frank. They would show in non-bank mortgage originators, which now write most US mortgages and operate thinly capitalized. Commercial real estate is already a leading indicator from a related but earlier shock, remote work. Residential follow-on tends to lag by one to three years. None of this is prediction. Some of it is happening now.

Then there is the demand side, which is the part most arguments about AI labor displacement skip.

A factory that automates its workforce gets immediate margin expansion. The fired workers stop being a cost. They also stop being customers. For any individual firm, this is fine, because its own customer base does not consist primarily of its own workers. For the industry collectively, it is not fine.

Marriner Eccles, who chaired the Federal Reserve through the Depression and the Second World War, identified the same mechanism in his account of the 1920s. He argued that wealth concentration during that decade had created an economy capable of producing more than its consumers could afford to buy. The financial structures that emerged to bridge the gap (leverage, credit expansion, speculative asset growth) were temporary substitutes for the wage income that would have sustained demand on its own. When the substitutes failed, the underlying shortfall in purchasing power became visible all at once. Eccles was writing about a different decade and a different mechanism, but the principle generalizes: a corporate sector that fires its way to higher margins is also firing its own collective customer base. The savings show up immediately on the income statement. The demand erosion shows up later, on someone else's.

The information worker tier is unusually exposed to this dynamic because it is unusually concentrated. The marginal propensity to consume drops as wealth concentrates. Even if AI preserves total income in the economy, which is contested, concentrating that income in the owners of the AI systems while removing it from the workers using them produces less aggregate demand, not more. The Eccles problem at industrial scale.

What this implies, taken seriously, is that the productivity gains the market is currently pricing into AI may be partly fictional. They are real at the level of the individual firm and the individual earnings call. They are partly illusory at the level of the economy, because the demand they require is being eroded by the same mechanism that generates them. A sector cannot indefinitely sell into a market whose ability to buy depends on the wages it is cutting.

The 2008 parallel is uncomfortable for a reason. The mechanism is similar. The blind spot is similar. The thing the model isn't pricing in is the correlation. And the metric confirming that everything is fine is once again a fluency metric (tokens, queries, AI-assisted output volume) that has nothing to do with the underlying soundness of the system.

The counterargument worth taking seriously

The honest objection is that every previous wave of automation has redistributed labor rather than eliminated it. Looms didn't end employment; they shifted it. Tractors didn't end employment; they shifted it. Spreadsheets didn't end accountants; they made them more productive. The historical record is on the side of adaptation, not collapse.

The rebuttal has two parts.

The first is that past automation was task-specific. A loom replaces a specific motion. A tractor replaces a specific labor input. AI is general-purpose, which means the set of tasks it can absorb is much larger than the set any previous tool could absorb. Comparing AI displacement to loom displacement assumes the categories are equivalent. They aren't.

The second is that the timing math doesn't work even if the historical pattern holds. The new job categories that absorb displaced workers historically take a decade or more to emerge and stabilize. A worker can be unemployed for six months and default on their mortgage. The credit system runs on shorter time horizons than labor markets do. "Long run, it works out" is not a position the mortgage book can underwrite.

Adaptation may happen. The question is whether the financial and social infrastructure can survive the transition. The historical examples we cite as success stories were also disasters at the time. The Luddites were not wrong about what was happening to them. They were wrong about being able to stop it.

What the throughline says

Three scales, one mechanism. At the individual level, AI removes friction from generating output without removing the need for someone to verify it, and the people most likely to use it are least equipped to do that verification. At the organizational level, companies remove the verifiers as a cost-cutting move and adopt productivity metrics that confirm volume while obscuring quality. At the same time, they stop hiring the juniors who would have grown into tomorrow's verifiers, breaking the succession loop that produces verification capacity in the first place. At the economic level, the people being removed are also the AAA borrower base and the demand floor of the knowledge economy, and the diversification assumption that lets the credit system tolerate individual displacement breaks under correlation.

Token-counting is the visible signal at all three levels. At the individual level: you can't tell if the tokens are right. At the organizational level: the metric measures volume, not value. At the economic level: we have built an economy that pays for tokens instead of for the judgment behind them. The same observable phenomenon recurs as you zoom out. That is how you know it is the spine of the thing, not a detail.

What does an honest response look like?

The first move is to stop pretending verification is free. AI productivity narratives assume the output, once produced, is usable. It almost never is, in domains where being wrong is costly. The cost of verification doesn't disappear when AI generates the output. It shifts from the producer to the reviewer, and it shifts forward in time. Mature organizations will start measuring AI work the way mature organizations have always measured work: by what shipped and held up, not by how much got generated.

The second move is to invest in the verification capacity currently being treated as overhead. Senior people, institutional memory, and the long-tenured pattern-recognizers who can tell when output is subtly wrong are the assets that determine whether AI deployment creates value or destroys it. Cutting them to fund the AI investment is cutting the only thing that makes the AI investment work. The same logic applies to the apprenticeship loop. Hiring and training juniors is not deferrable for two or three years until AI matures. It is how an organization regenerates its verification capacity over time. A firm that stops doing it is choosing permanent dependence on the seniors it currently has, which is not a position any institution should pick voluntarily.

The third move is structural. Verification at scale requires more than human attention. It requires infrastructure: systems that remember what has been decided, what has been disputed, what has been verified, and what has not. Most of what AI gets wrong is wrong in the same way the same organization got wrong six months ago. Persistent context, the accumulated and queryable history of what an organization has figured out, is the closest thing we have to verification infrastructure. It is not a substitute for the verifiers. It is what allows the verifiers who remain to operate at the scale the new output volume demands.

There is a version of this story that ends well. Augmentation framing, making information workers three times more effective, is commercially and economically sustainable. Replacement framing, firing half your information workers, is not, because it eats its own demand base. The vendors and operators who internalize the first framing will outlast the ones still selling the second. The companies that preserve their verification capacity, even at the cost of slower margin expansion in the short term, will be the ones with anything left to verify in five years.

The choice between augmentation and replacement is being made right now, mostly by default, mostly without anyone naming it clearly. The pattern is hard to see because the metric is confirming that everything is fine. Tokens are up. Output is up. Productivity is up.

Productivity is a measure of value created. Value depends on someone being able to tell good from bad. Removing those people on the assumption that the new tool no longer needs them is a bet that the entire knowledge economy has been priced wrong for fifty years, that the senior judgment, institutional memory, and cross-checking that organizations have paid for were overhead all along.

That bet will lose. Some of what's coming is already on the schedule. In the next two to five years, a major company will ship an AI-generated decision (a medical recommendation, a financial product, a legal filing, a piece of infrastructure code) that costs the company materially because no one verified it. The investigation will trace the failure to roles that were eliminated as redundant. The post-mortem will describe a verification gap. The board will demand a hiring plan. The hiring plan will discover that the verifiers who would have caught it are now consulting independently at three times their previous salaries, working at a competitor that did not lay them off, or gone from the industry entirely.

That will be the moment the market reprices verification. Until then, the companies that keep their verifiers will look slower and more expensive than the ones that did not. They will be the ones still operating safely in five years.

The leaders making the call now should be honest with themselves about which side of that line their organization is on. If your AI strategy still ships through people who can verify the output, you are augmenting human judgment. If your AI strategy ships through fewer people because the output is assumed to be sound, you are replacing it. Augmentation is sustainable. Replacement consumes its own foundation: the people who could verify the output, the customer base that could afford the product, and eventually the institutional knowledge that made any of it work.

The same metric, volume going up, confirms both strategies. The metric is wrong. It has always been wrong. We have just never had a technology that could produce enough fluent output to make the wrongness this expensive.

Enjoyed this article?

Subscribe to get new posts delivered to your inbox.