The Confidence Illusion.
Exploring why AI hasn't made professionals less capable, and how it has broken the feedback loop that used to keep their confidence honest.
Elena Bosch had underwritten commercial credit for eleven years at a mid-sized bank in Rotterdam. Colleagues who worked with her described her as “unnervingly right” — not infallible, but right often enough, and wrong in ways that were legible even to her, that people trusted her instinct on files that didn’t fit the standard scoring model. Her default-rate predictions had landed within two points of actual outcomes for six consecutive years. She could tell you, unprompted, the three times in the past decade she had been badly wrong, and exactly what she had missed each time.
In early 2025, the bank rolled out an AI underwriting assistant to the commercial lending team. It read financials, cross-referenced sector risk data, and produced a recommendation with a confidence score attached. Elena adopted a habit almost immediately, without being told to: she read the file cold, wrote down her own view — approve, decline, or flag — before opening the tool, and only then checked it against the model’s output. When they agreed, she moved fast. When they diverged, she went back into the file to find out why.
Her colleague Jorik, two desks over, adopted a different habit, just as unconsciously. He opened the file and the tool at the same time. He read the model’s recommendation first, then skimmed the underlying file to confirm it looked reasonable, then approved. It was faster. It felt, if anything, more rigorous — he was, after all, reading the same financials Elena was, just in a different order.
Eight months later, a portfolio quality review turned up something that unsettled the risk committee more than any single bad loan could have. Elena’s calibration — the statistical tightness between how confident she reported being in a decision and how often that decision proved correct — hadn’t moved from her pre-AI baseline. Jorik’s had come apart. On the subset of cases where the model’s recommendation later proved wrong, Jorik’s own accuracy had fallen sharply — but his self-reported confidence in those same decisions remained exactly as high as it had always been. He had no idea anything had changed. Nobody had told him his judgement was slipping, because the correction signal that would normally tell a professional that — the felt discomfort of being wrong, attributed clearly to your own call — had never arrived. The model was right often enough, and confident enough in its wrongness, that Jorik’s errors were absorbed silently into decisions that felt, from the inside, exactly as sound as they always had.
This is the Confidence Illusion. It is not a story about AI making people less capable. Every available measure of Jorik’s underlying analytical skill — his ability to read a balance sheet, to spot the sector risk his manager wanted him to spot — remained intact. What broke was something upstream of skill: the loop that let his confidence track his own accuracy. And once that loop breaks, a professional can be simultaneously skilled and structurally unable to know it.
Why confidence used to mean something
To understand what has broken, it helps to understand what calibration actually is, and where it comes from. It is not a personality trait. It is a trained response, built through a specific, repeatable mechanism: a person forms a judgement under uncertainty, reality responds, and the gap between the prediction and the outcome — assuming the person actually notices it — recalibrates the next prediction.
This is, essentially, the mechanism behind every serious account of how expertise develops. Gary Klein’s decades of research into naturalistic decision-making — how firefighters, ICU nurses, and fighter pilots make fast, accurate calls under time pressure — found that expert intuition is built from exactly this loop, repeated thousands of times, until pattern recognition becomes near-instantaneous. Philip Tetlock’s forecasting research, most visible in the “superforecasters” work, found that the individuals who consistently outperformed professional analysts on geopolitical predictions shared one habit above all others: they tracked their own predictions against outcomes, relentlessly, and adjusted. Calibration, in Tetlock’s data, was learnable — but only through that unforgiving feedback loop. People who never checked their predictions against outcomes stayed miscalibrated indefinitely, regardless of how experienced they became.
The common thread across all of this research is that confidence was never supposed to be a free-floating feeling. It was supposed to be downstream of a track record — built, case by case, through the friction of being wrong and finding out. Professionals who worked without any decision-support tool got this feedback by default, because there was no alternative: they had to form a view, because nothing else was going to form one for them, and reality corrected them whether they liked it or not.
AI has not removed reality’s ability to correct people. It has removed the requirement that a person form an independent view before consulting a second opinion — and it turns out that requirement was doing almost all of the calibrating work.
The mechanism: how offloading breaks the loop
The distinction the research keeps drawing is between two different things people can offload to AI: memory and judgement. Offloading memory — letting a tool hold facts you’d otherwise need to recall — has a long history and a relatively mild effect on calibration; people have used notebooks, databases, and search engines for this without losing their grip on how confident they should be in their own conclusions.
Offloading judgement is different. A 2026 study presented at the CHI Conference on Human Factors in Computing Systems — “Accurate but Not Confident or Confident but Not Accurate? Cognitive Offloading Impairs Confidence Calibration in Human-AI Teams” — tested this distinction directly. Participants who worked entirely unaided showed the tightest alignment between their stated confidence and their actual accuracy of any condition in the study. Participants who offloaded judgement specifically — not memory, but the act of forming an initial view — showed the sharpest overconfidence effect measured. Combined offloading of both memory and judgement produced a different, but equally distorting, metacognitive bias. The finding was blunt: the moment a person stops forming their own view before seeing an external recommendation, their stated confidence stops being a reliable signal of anything.
This connects to a separate and older body of research on automation bias and complacency. A recent review in the journal AI & Society offers a useful, precise definition of the mechanism: complacency is what happens when trust closes the gap between “I believe this system works” and “I no longer evaluate whether it worked this time.” That second clause is the entire mechanism. It is not that people trust AI too much in the abstract. It is that, case by case, they stop running the small internal check — does this feel right to me, independent of what the tool says — that used to catch errors before they became decisions. The review notes that this collapse is especially likely to occur when a tool’s recommendation aligns with a person’s initial instinct, which creates a reinforcing loop: agreement breeds trust, trust breeds less checking, less checking means agreement is confirmed more often because disagreement is never investigated.
Complacency happens when trust closes the gap between “I believe this system works” and “I no longer evaluate whether it has worked this time.” This is where users catch errors, flag edge cases, and discover the AI’s limitations — and where automation-assisted decision-makers stop doing so.
— Automation bias in human–AI collaboration, AI & Society, 2025–2026
The mechanism, stated plainly: an unaided professional forms a view, gets it wrong sometimes, and feels that wrongness directly — this is uncomfortable, and discomfort is what recalibrates confidence. An AI-assisted professional who skips the independent view never generates the raw material discomfort needs. There’s nothing for the wrongness to attach to. It gets absorbed into a decision that, from the inside, felt entirely reasonable — because it was never really the professional’s own decision to feel wrong about.
What gets lost: two faces of the same failure
The workforce data suggests this mechanism is producing two distinct, opposite-looking outcomes — and it’s worth being precise about why they’re actually the same failure.
ManpowerGroup’s 2026 Global Talent Barometer, based on interviews with nearly 14,000 workers across 19 countries, found that regular AI use jumped 13% in a single year, to 45% of the workforce — while confidence in using the technology fell 18% over the same period, the first overall decline in worker confidence the barometer had recorded in three years. The drop was steepest among the most experienced workers: down 35% among Baby Boomers, 25% among Gen X. And yet 89% of workers still report confidence in the skills required for their current role. This is not a story about people losing faith in their own competence generally. It is a much narrower and stranger story: people are losing trust specifically in their own independent judgement, in the moments where that judgement now runs alongside a machine’s.
A separate 2026 study, reported by the American Psychological Association, adds a piece that clarifies the picture further: overreliance on AI at work does not appear to measurably reduce raw cognitive ability. What it erodes is confidence in independent reasoning, and people’s sense of ownership over their own ideas — their felt sense that a conclusion is actually theirs, rather than borrowed and merely endorsed. Researchers flagged the long-term risk plainly: not that AI use makes people less intelligent, but that some professionals become less engaged in the kind of effortful, generative thinking that produces genuinely novel judgement — because the tool has quietly become the place where the thinking happens first.
Put the two findings together and the shape of the problem sharpens. One group — the Jorik pattern — offloads the initial judgement entirely, stops noticing when they’re wrong, and drifts toward silent overconfidence that nobody, including them, can see from the inside. A second group — visible in the Manpower data, concentrated among the most experienced workers — retains the instinct to form an independent view, notices that view increasingly gets second-guessed or overridden by a fluent machine recommendation, and starts to doubt instincts that were never actually unreliable. Both groups have lost the same thing: a working relationship between what they believe and what is true. One believes too much. The other no longer knows what to believe. Neither can currently tell you, reliably, which of their own judgements to trust — which is precisely the capability the feedback loop used to provide.
Three archetypes
Three professional postures are visible in this transition, and naming them helps predict where a given person or team is heading before a performance review or an audit catches it.
The Silent Drifter has adopted the Jorik pattern without noticing. They read the AI’s output before forming their own view, or alongside it rather than before it, and their confidence has not moved even as their independent accuracy — measurable only by looking at the subset of cases where the AI later proves wrong — quietly declines. They are, by every self-report metric, doing fine. The gap between their confidence and their accuracy is invisible to them by construction: nothing in their daily experience produces the discomfort that would reveal it. This is the most structurally dangerous archetype, because it self-conceals; the only way to detect it is to deliberately audit decisions after the fact, which almost nobody does until something goes wrong at scale.
The Doubt Spiral is the mirror image, and disproportionately made up of experienced professionals — the ManpowerGroup data’s Boomer and Gen X respondents. Their independent judgement is intact, arguably as strong as it has ever been, but it now runs constantly alongside a fluent, confident machine output that occasionally disagrees with it. Each disagreement produces a small, corrosive moment of self-doubt, even when the human was right — because the AI’s confidence is expressed with the same fluent certainty whether it is correct or not, and there is no longer an obvious signal for which voice to trust. Over time, this group’s stated confidence in their own judgement erodes even though their underlying accuracy has not. They are, in a real sense, the collateral damage of a workplace that has not built any mechanism for telling people when their own instinct was the right one.
The Calibrated Checker is the smallest and most valuable group, and their habit is almost embarrassingly simple: they form an independent view before they open the tool, every time, as a fixed discipline rather than an occasional practice. When the AI agrees, they move fast, exactly as anyone would. When it disagrees, they treat the disagreement as diagnostic information rather than as an automatic override in either direction — sometimes the AI is right and they update; sometimes they are right and the AI’s confident output was simply wrong in a way that fluent language obscured. Crucially, they track the outcomes of these disagreements over time, which means their confidence keeps being disciplined by the exact mechanism that built expertise in the first place. They are not smarter than the Silent Drifters or more skilled than the Doubt Spirals. They have simply refused to let the tool remove the one step — form your own view first — that made confidence mean something.
What this means in practice
For individuals, the practical implication is almost aggressively simple, which is part of why it is so easy to skip: write down your own view, even briefly, before you open the tool. Not because your first instinct is always right — often it won’t be — but because the act of committing to a view, and then finding out whether it held up, is the only mechanism that keeps your confidence tethered to your actual accuracy. Skipping this step doesn’t just cost you a marginally worse decision today. It costs you the raw material your judgement needs to stay calibrated over the next decade.
For organisations, the implication is that AI adoption without a parallel discipline for preserving independent judgement is quietly manufacturing both of the Confidence Illusion’s failure modes at once — a population that is dangerously overconfident in eroding judgement, sitting alongside a population that has stopped trusting judgement that was never actually impaired. Neither failure shows up in standard productivity or output-quality metrics in the short term, because AI-assisted output is, on average, good. It shows up later, unpredictably, in the moments the tool is wrong and nobody in the workflow retained the instinct — or the confidence — to catch it.
The practical fix organisations are experimenting with is structural rather than motivational: build the independent-judgement step into the workflow itself, so it doesn’t depend on individual discipline. Some firms are formalising this as a two-step approval process — record a human recommendation before the AI recommendation is visible, then compare. It is a small amount of friction, deliberately reintroduced, in a system that has spent two years trying to remove friction wherever it can be found. The firms doing this are, in effect, treating the independent-judgement step the way clinical trials treat a blind assessment: not as inefficiency, but as the only mechanism that keeps the measurement honest.
For hiring and development, the shift is from screening for confidence — historically an easy, visible, interview-friendly proxy for competence — to screening for calibration, which is much harder to fake and much more diagnostic. A candidate who can describe, specifically, a time they were confidently wrong and how they found out is demonstrating the exact muscle the Confidence Illusion is quietly atrophying across the workforce. A candidate who cannot produce that story, or who describes only times they were confidently right, is a much weaker signal than it used to be — high self-reported confidence, on its own, no longer reliably predicts good judgement in an AI-assisted environment. It may increasingly predict the opposite.
The closing uncomfortable truth
There is a final, uncomfortable feature of the Confidence Illusion worth naming directly: it is nearly undetectable from the inside. A Silent Drifter does not feel like someone whose judgement is eroding. They feel exactly like someone who is right most of the time, because they are — the AI is usually right, and their approval of its usually-right output feels, moment to moment, indistinguishable from good judgement. A Doubt Spiral does not feel like someone whose instincts are being unfairly eroded. They feel like someone who is simply, reasonably, less sure than they used to be, in a world that has gotten more complicated. Neither has access to the information that would tell them which category they’re in, because the information that used to provide it — the direct, attributable, felt experience of being wrong — has been quietly rerouted around them.
The people whose confidence still means something are not, on average, more talented than everyone else. They have simply refused to let the tool remove the one habit that made confidence a real signal in the first place: forming a view of their own, before they see anyone else’s — including a machine’s — and finding out, deliberately and repeatedly, whether they were right.
Everyone else is walking around with a level of confidence that has quietly stopped being about anything at all. The unsettling part is that there is currently no way to tell, from the outside or from the inside, who is who — until the moment it matters, and the gap between what someone believed and what was actually true finally becomes visible to everyone except the person who should have seen it first.


