Priya opens the shared doc at 4:50 p.m. on a Thursday. Her colleague finished the competitive analysis an hour ago — clean headers, a tidy comparison table, five confidently stated market-share figures, a closing paragraph that reads like it was written by someone who has thought about this problem for a decade.
She skims it, relieved. Then she notices one figure doesn’t match what she remembers from last quarter’s board deck. She checks. It’s wrong — not wildly wrong, just wrong enough to have been believable. So she checks the next number. Also wrong. By 6:40 p.m., she has rebuilt half the document from scratch, and she still isn’t sure which parts of the other half she can trust.
Nothing about this document looked unfinished. That was the problem. It looked, by every visual and structural signal available to Priya, exactly like good work. Harvard Business Review gave this phenomenon a name in September 2025 — workslop — and the researchers at BetterUp Labs and Stanford’s Social Media Lab who coined it found Priya’s Thursday evening is now startlingly common: 40% of 1,150 U.S. full-time employees surveyed said they’d received AI-generated work like this in the past month, costing them nearly two hours of rework, on average, per instance. 53% said it left them annoyed. 42% said it made them trust the sender less. A third said they’d rather not work with that colleague again.
What’s actually being measured here isn’t an AI capability gap. It’s a mirage, and mirages are interesting precisely because the failure isn’t in what you’re looking at. It’s in your own perceptual machinery.
Where the mirage comes from
Long before anyone trained a transformer, psychologists had already mapped the terrain AI would eventually colonise. In the 1990s, researchers including Norbert Schwarz and Rolf Reber began documenting what they called processing fluency: the ease with which the brain processes a piece of information becomes, all by itself, a signal the brain uses to judge whether that information is true. Statements printed in a clearer font were rated as more truthful than the identical statements printed in a harder-to-read one. Words that were easier to pronounce were judged more trustworthy. None of this had anything to do with the content. It was entirely about the packaging.
This is not a flaw we can shame ourselves out of. It’s a load-bearing shortcut.
Evaluating every claim on its actual merits, from scratch, every time, would be cognitively unaffordable — so the brain outsources a huge portion of its trust calibration to surface cues: confidence, fluency, structure, specificity.
For most of human history, that shortcut worked reasonably well, because producing fluent, confident, well-structured, specific-sounding content was actually hard. It required expertise, rehearsal, or both. Fluency was, imperfectly but usefully, correlated with competence. Con artists and skilled salespeople have always understood how to break that correlation deliberately, but they were rare, and rarity was itself a kind of natural defence.
Generative AI didn’t invent the fluency heuristic. It did something more consequential: it made fluency free, infinite, and instant, for anyone, on any topic, at any hour, completely severing the correlation the heuristic was built to exploit. The packaging that used to be expensive to produce, and therefore a semi-reliable signal, is now the cheapest part of the entire pipeline. Only the substance remains expensive. And substance is exactly what the fluency heuristic was never designed to check.
A related finding from the same era of research makes the picture sharper still. The illusory truth effect, first documented by Lynn Hasher, David Goldstein, and Thomas Toppino in the late 1970s, showed that people rate a statement as more likely true simply because they’ve encountered it before, independent of whether it was true the first time either. Repetition alone manufactures a feeling of familiarity, and familiarity gets misread as validity. Put fluency and repetition together and you get something close to a complete account of how confident nonsense survives contact with a smart, careful reader: it sounds right, it feels familiar, and both of those sensations arrive before any actual verification has had the chance to begin.
AI-generated text is unusually good at triggering both at once: fluent by construction, and often repeating, in slightly different words, the same plausible-sounding claim across a document, which makes that claim feel more corroborated with every restatement, even though it has exactly one source: the model’s own prior sentence.
The mechanism, made visible
It’s worth being precise about why this happens, because “AI sounds confident” undersells the mechanism. A large language model is not making a judgement about whether it knows something and then choosing how to express that judgement. It is predicting, token by token, the most statistically plausible continuation of a prompt, shaped by training and reinforcement to produce answers that human raters found satisfying. Human raters, being human, are subject to the same fluency heuristic as everyone else — they tend to rate confident, well-structured, elaborate answers more highly than hedged, uncertain, terse ones, regardless of accuracy. The result is a system trained, at a deep level, to optimise for the appearance of competence, because that’s what the training signal actually rewarded.
Researchers have now measured the downstream consequence directly. In a set of controlled experiments, participants were shown AI-generated answers that varied in length and elaborateness. Longer, more detailed explanations reliably increased participants’ confidence in the answer but did not improve their ability to tell correct answers from incorrect ones. Only 26% of participants could accurately perceive when an AI response was overconfident relative to its actual reliability. Verbosity was functioning as a proxy for expertise: length and elaborateness were triggering what researchers describe as an epistemic authority heuristic, one that bypasses genuine comprehension checking almost entirely.
Across these studies, a consistent pattern holds: the more elaborate and confident an AI’s explanation, the more people trust it, and the less that trust tracks whether the explanation is actually correct.
This is the mechanism underneath every piece of workslop Priya has ever received. It’s not that the model is lying with unusual skill. It’s that the model never learned to signal uncertainty the way a careful human expert instinctively does — the pause, the hedge, the “I’d want to double-check this before you use it,” the visible discomfort of a person who knows the difference between confidence and correctness. An LLM has no discomfort to display. It presents a fabricated statistic with exactly the same tonal certainty as a verified one, because tonal certainty was never connected to truth in the first place. It was connected to what got rewarded during training.
What gets lost
The direct cost of the fluency mirage is the two hours Priya spent on Thursday. The indirect cost is larger and slower-moving, and it shows up in three places.
The first is trust between colleagues. The workslop data on this point is unambiguous. Handing someone fluent-but-hollow work doesn’t just cost them time; it changes how they see you. Forty-two percent of workslop recipients said it made them view the sender as less trustworthy; roughly half rated the sender as less capable, less creative, or less reliable than before. AI is being used, by well-intentioned people trying to move faster, in ways that are quietly corroding the interpersonal trust that actual collaboration depends on.
The second is what happens to judgement itself when it goes unpracticed. Verification is a skill, and like any skill, it atrophies without use. If the default posture toward AI output becomes acceptance rather than scrutiny — because scrutiny is effortful and the output looks so finished — then the muscle required to catch the next, more consequential error weakens exactly when the stakes are rising.
This is not hypothetical. MIT Sloan Management Review’s 2026 research on AI verification makes the trajectory explicit: as AI systems are given longer, higher-risk, more autonomous tasks, checking whether the work was actually done correctly becomes harder and slower — not easier — even as organisational reliance on that work accelerates. The gap between how much we trust AI output and how well we can verify it is widening, not narrowing, at precisely the moment the stakes attached to that output are growing. A parallel 2026 global executive survey found 86% of leadership teams now consider AI central to strategic priority-setting — meaning the quality of human oversight has quietly become a strategic variable, not just an operational nicety.
The third is subtler: the mirage doesn’t just produce bad individual decisions. It changes what “good work” is perceived to look like across an entire organisation, because the visual and structural markers of quality — polish, structure, confident phrasing, comprehensive-seeming coverage — have become disconnected from the underlying rigor those markers used to signal. Teams start optimizing, often unconsciously, for the appearance that gets rewarded rather than the substance that was supposed to produce it. That’s a much harder problem to reverse than a single bad quarter, because it’s not a mistake — it’s a slowly shifting incentive structure that nobody explicitly chose.
Four faces of the mirage
The fluency mirage doesn’t show up as one uniform failure. In practice, it tends to take one of four recognisable shapes.
The Instant Expert. An AI-generated answer arrives with the tone and structure of someone who has spent a career in the domain — assured phrasing, domain vocabulary used correctly, a confident framing of trade-offs. Nobody asks how the model would know this, because it doesn’t sound like a guess. It sounds like expertise. The tell isn’t in what’s said; it’s in the complete absence of the specific, hard-won caveats a genuine expert would volunteer unprompted.
The Verbose Alibi. Length gets mistaken for rigour. A three-paragraph answer feels more trustworthy than a three-sentence one, even when the extra length is padding, restatement, or elaboration on an unverified premise. This is the mechanism the verbosity research captured precisely: elaborateness functions as a proxy for diligence, whether or not any additional diligence actually occurred.
The Certainty Engine. The response never hedges, never flags uncertainty, never says “I’m not fully sure about this specific figure.” Because the model’s tone doesn’t shift between confident-and-correct and confident-and-fabricated, the human reading it loses the single cue — hesitation — that would normally prompt a second look.
The Downstream Debtor. This is workslop’s actual delivery mechanism. The polished, fluent, unverified output gets passed along — to a colleague, a manager, a client — and the obligation to notice the gap between fluency and substance transfers with it, usually without anyone naming that a transfer has occurred. The person who generated it experiences speed. The person who receives it experiences the bill.
The compounding problem
All of this was true when the worst an AI system could do was write one wrong paragraph. It matters more now, because the unit of AI-produced work is no longer a paragraph — it’s a multi-step task, executed with growing autonomy across research, drafting, coding, and decision-support, each step building on the fluent-sounding output of the step before it. MIT Sloan’s 2026 research on AI verification names this directly: as systems take on longer, higher-risk, more autonomous tasks, checking whether the work was actually done correctly becomes harder and slower, not easier, because there are more intermediate steps, each one a fresh opportunity for a plausible-sounding error to get quietly folded into the next stage before anyone looks closely. A wrong number in a single answer is a two-hour problem, à la Priya. A wrong assumption embedded at step two of a fourteen-step agentic workflow is a problem nobody notices until it surfaces three departments away, dressed in the confident, finished-looking output of everything built on top of it. The mirage doesn’t stay contained to the document it started in. It propagates, fluently, at every layer it touches — and each layer adds its own varnish of polish, making the original error progressively harder to spot the further downstream it travels.
This is also why the comfortable assumption — that organisations will simply get better at catching this over time, the way they got better at spotting spam email or phishing attempts — doesn’t hold up well under scrutiny. Phishing emails get caught because their fluency is usually imperfect: a slightly wrong logo, an odd turn of phrase, a domain that’s almost right. The fluency mirage has no equivalent tell. It is, definitionally, the case where the packaging is flawless. There is no typo to notice, no broken image, no uncanny phrasing — just a confident, well-structured, entirely plausible claim that happens to be wrong, sitting in a document that looks exactly like every other document that was actually right.
What actually holds up
None of this is an argument against using AI for real work — it’s an argument for pricing verification back into how that work gets planned, evaluated, and rewarded, deliberately, rather than assuming it will happen for free because it always used to.
For individuals, that starts with treating fluency as a neutral signal rather than a positive one. A polished, confident answer earns exactly zero epistemic credit until it’s been checked; the smoother it is, the more consciously worth checking it becomes, precisely because smoothness is the cheapest thing about it now. One useful ritual: before passing along any AI-assisted output, name out loud — even just to yourself — the single number, claim, or assumption in it you are least sure about, and check that one thing before you check anything else. It’s a small habit, but it directly counters the illusory-truth trap, because it forces a specific claim to be evaluated on its own merits rather than absorbed into the general glow of a well-formatted document. Building this instinct is uncomfortable early on — it means being the person who says “let me verify that” in a meeting full of people nodding along, which can feel like friction for its own sake. It isn’t. It’s the only mechanism in the room actually distinguishing fluent from true.
For hiring managers, this reframes what a strong candidate looks like. The ability to produce fluent, well-structured work is no longer a differentiator — it’s table stakes, available to anyone with an API key. What’s scarce, and getting scarcer, is the ability to interrogate fluent output and articulate specifically why it might be wrong: to name the missing caveat, spot the plausible-but-unverified number, ask the question the confident tone was designed to make you forget to ask. That capability shows up far more reliably in a structured conversation — “here’s a polished answer, tell me what’s wrong with it” — than anywhere in a résumé, and it’s worth building directly into how candidates are evaluated rather than assuming it will surface on its own.
For leaders, the implication is structural. If output volume is being measured without measuring downstream verification cost, the organisation is tracking a number that workslop is subsidising — quietly, and usually with the time of the most conscientious people on the team, the ones who can’t bring themselves to pass along something they haven’t checked. Making verification visible — treating it as legitimate, budgeted work rather than an invisible tax someone absorbs after the fact — is not a productivity drag. It’s the correction that lets the productivity number mean anything at all. That might mean building a verification checkpoint into a workflow the same way a code review checkpoint already exists; it might mean explicitly rewarding the person who caught the error rather than only the person who moved fast. Either way, the goal is the same: make the invisible cost visible before it becomes someone’s Thursday evening.
Priya finishes rebuilding the document a little after 7. Nobody will ever see the two hours it took her — not in the deliverable, which now looks exactly as polished as it did before she started, and not in any metric her organisation tracks. That invisibility is the entire mechanism of the fluency mirage in miniature: the correction disappears into the same seamless-looking output that necessitated it, and next week, someone hands her something else that looks finished.
The uncomfortable truth isn’t that AI produces bad work. It’s that AI produces work indistinguishable, on the surface, from good work, and the surface is the only thing most of us have time to check. Fluency was never supposed to be the whole test. For most of history, it was a reasonable proxy, because faking it convincingly was expensive. That’s no longer true, for anyone, about anything. The proxy broke. What’s left, the only thing left, is the slower, less comfortable, entirely human work of actually checking — and organisations are only beginning to notice that they stopped budgeting for it right around the time they started needing it most.


