The interview that was too clean
Priya has screened engineering candidates for eleven years, long enough to develop the thing every senior recruiter eventually develops: a private, unscientific but rarely-wrong sense for when an interview is going too well. The candidate in front of her that Tuesday in November 2025 was going too well. Thirty-five minutes into a live system-design round, he had produced a clean, well-structured answer to a deliberately ambiguous scaling problem, complete with the kind of trade-off language, “it depends on your consistency requirements here”, that usually takes people years to acquire, not months. He paused for exactly the right amount of time before each point, the way someone pauses when they’re reading rather than recalling. When Priya asked one small follow-up, off-script, about a detail he hadn’t mentioned, there was a half-second delay that hadn’t been there before, and then a perfectly composed answer to a question she hadn’t actually thought was answerable that cleanly.
She didn’t flag him in the moment. She wasn’t sure enough, and being wrong about this kind of accusation carries its own cost. She let the interview finish, gave him a strong internal score, and only two weeks later, reading a report her company’s newly deployed AI-detection tool had generated on the recorded session, did she learn what her instinct had been circling: micro-patterns in his response timing, consistent with reading rather than generating in real time, flagged at 89% confidence. He had, in all likelihood, been fed the answer through a hidden overlay, the kind marketed today as “invisible to screen-share”: Cluely, Final Round AI, LockedIn AI, and a dozen less famous competitors, all built for exactly this moment, all doing a brisk business.
He didn’t get the job. But Priya’s story isn’t really about him. It’s about the fact that she almost missed it, that her company only caught it because they’d bought expensive detection software specifically built for this exact scenario, and that, by the numbers now coming out of the recruiting industry, for every candidate like him who gets caught, a comparable number don’t. Across 19,368 live technical interviews conducted between July 2025 and January 2026, 38.5% of candidates were flagged for AI-assisted cheating behaviour, with the flag rate for technical roles alone climbing to 48%. And of the candidates flagged, 61% scored above the passing threshold and received an offer anyway. The detection worked. The consequences, for most of them, did not follow.
How we got here: the short history of a proxy
The job interview has never been a perfect measure of anything. It has always been a proxy: a stand-in, chosen because it was practical, for a thing employers actually want to know and can’t directly observe: can this person do the work, reliably, under real conditions, for years. Different eras picked different proxies and believed in them with roughly the same misplaced confidence.
The whiteboard-algorithm era, which peaked in tech hiring through the 2010s, bet that the ability to invert a binary tree under time pressure correlated with the ability to ship production software. It didn’t correlate especially well: a body of research eventually made that case convincingly enough that companies started phasing whiteboard rounds out even before AI entered the picture, partly because AI coding assistants had already made memorising algorithms less relevant to the actual job. The behavioural-interview era that layered on top of it bet that a well-told story about a past conflict predicted how someone would handle a future one. That proxy had its own well-documented flaws (coachability, rehearsal, the simple fact that articulate storytellers aren’t always the strongest performers) but it survived for decades because, flawed as it was, nobody had a cheaper or faster substitute.
What’s different now isn’t that the proxy is imperfect. Proxies have always been imperfect. What’s different is that the proxy has become gameable at scale, cheaply, invisibly, by anyone with an internet connection and twenty dollars a month, and gameable in a way that produces a result which is, to a human evaluator watching in real time, functionally indistinguishable from the real thing. Previous generations of interview gaming (a coached answer, a rehearsed story, a friend’s practice questions) still required the candidate to actually absorb and reproduce the material themselves. The AI-overlay era removes that last requirement entirely. The candidate doesn’t need to know the answer. They need to be willing to read it, fluently, off a screen nobody else can see, in real time, while maintaining eye contact and a confident tone. That is a categorically easier skill to acquire than the thing the interview was originally built to measure, and it produces results that score, on every visible metric, just as well.
There’s a useful parallel in how academic testing handled a similar shock a few years earlier. When large language models first made take-home essays trivially fakeable, some institutions responded by banning AI outright and trying to detect its use after the fact: an approach that produced a steady stream of false accusations against students who simply write in a clean, structured style, and that degraded almost as fast as the detection tools it depended on. Others responded by redesigning the assessment itself: oral defences, in-class components, work that had to be produced and explained live, in the room, with no opportunity to outsource the thinking beforehand. The second approach held up considerably better, for a simple reason: it stopped trying to detect the presence of AI and started making AI assistance structurally useless for the specific thing being measured. Hiring is now living through the identical decision point, several years behind education and with far higher stakes attached to getting it wrong.
“The candidates most equipped to beat this system aren’t the most competent ones. They’re the ones most comfortable performing competence they don’t have — which is, if you sit with it for a moment, close to the exact opposite of what a hiring process is supposed to select for.”
The mechanism: an arms race with no ceiling
The tools themselves have moved fast enough that describing them risks going stale within a quarter, but the shape of the arms race is stable enough to describe with some confidence. On the candidate side: browser overlays that sit invisibly on top of a video-call window, capturing the interviewer’s question via audio or screen-share, running it through a language model, and displaying a generated answer in a window that doesn’t appear in anything the interviewer can see, because it’s rendered outside the shared application, or because it exploits how screen-share capture works at a technical level. Some candidates pair this with a Bluetooth earpiece and real-time audio coaching instead of a visual overlay, reading a script fed to them word for word. A smaller, more aggressive tier uses deepfake video tools to swap out who’s actually appearing on camera entirely, letting a more experienced person sit the interview on someone else’s behalf while their face is rendered convincingly onto the candidate’s video feed. Tools originally built for the wave of smart-glasses cheating that got them banned from the SAT are, unsurprisingly, being repurposed for the exact same function in a Zoom call.
On the employer side: the response has been to fight AI with AI. 61% of companies surveyed in Greenhouse’s 2026 AI in Hiring Report now use dedicated detection software during live interviews: analysing response-timing patterns, eye-gaze consistency, behavioural biometrics, and the subtle statistical fingerprints that distinguish someone reading generated text from someone generating an answer in real time as they speak. Another meaningful share of companies are actively evaluating similar tools. It is, in the most literal sense, an arms race: an entire specialised market has emerged on both sides of the interview table simultaneously, each side selling tools designed to defeat the other side’s tools, with candidates and recruiters as the actual customers on either end.
Nobody involved in this arms race describes it as stable or winning. 62% of hiring professionals surveyed openly admit that candidates are getting better at faking their way through AI-assisted deception faster than recruiters are getting better at catching it. Human accuracy at spotting deepfaked video in an interview setting was measured at just 55.5%, barely better than a coin flip, which is a genuinely remarkable thing to discover about a task hiring managers have spent their careers believing they were good at. And the deception isn’t limited to live interviews. 91% of recruiters in the same Greenhouse survey report encountering some form of candidate deception in the broader hiring process this year: 63% report resume exaggeration, 48% report fabricated references, and a smaller but rapidly growing share report prompt injections: hidden text embedded in a résumé file, invisible to a human reader, engineered specifically to manipulate an AI resume-screening tool into ranking the candidate higher regardless of their actual qualifications, before a human recruiter ever opens the file.
What gets lost
The most obvious cost of all this is the one that gets the headlines: a company might hire someone who genuinely cannot do the job they were hired for, discovering it only once real stakes are attached and the scaffolding is gone. That cost is real, but it’s not actually the largest one, and treating it as the whole story misses what’s happening underneath it.
The first hidden cost is speed and cost, distributed across every candidate rather than concentrated on the dishonest ones. 68% of hiring managers report being more personally involved in hiring than they were a year ago: more manual screening, more live problem-solving exercises, more in-person final rounds specifically reintroduced to strip away the conditions that make AI-assisted cheating possible. Google, Cisco, and McKinsey have all shifted at least their final interview rounds back to in-person formats for exactly this reason. That’s a rational institutional response to a real problem. It’s also a tax on everyone in the pipeline, honest and dishonest alike, in the form of longer processes, more logistical friction, and more expensive hiring per head: a cost that didn’t exist five years ago and now exists because the interview itself can no longer be trusted at face value.
The second hidden cost is a trust collapse that runs in both directions simultaneously, and is arguably more corrosive than either side’s individual behaviour. 70% of hiring managers say they now trust AI tools to make faster, better hiring decisions than they could make unassisted. Only 8% of job seekers describe AI-driven hiring processes as fair. Neither number is really about AI’s actual accuracy. Both numbers are about a hiring relationship where each side has become convinced, with substantial justification, that the other side is gaming the system: hiring managers assuming candidates are cheating, candidates assuming AI screening tools are unfairly filtering them out before a human ever looks at their application. Trust, once it degrades this broadly across an entire labour-market relationship, doesn’t come back because one detection tool gets slightly more accurate. It comes back, if it comes back at all, only once both sides believe the underlying process is worth trusting again, and right now, neither side does.
The third and least visible cost falls specifically on honest candidates, and it’s the one that should worry anyone who cares about fairness in hiring more than the headline fraud statistics do. The candidate who spent genuine hours preparing, who answers a beat slower because they’re actually generating the answer live rather than reading it, who occasionally says “let me think about that for a second” because they are, in fact, thinking, that candidate is now being scored, in real time, against a fluency benchmark partly set by people who aren’t actually thinking at all. Some of those honest candidates are starting to reach for the same tools themselves, not because they want to deceive anyone, but because refusing to use an AI overlay increasingly feels less like integrity and more like showing up unarmed to a competition where the rules have already changed for everyone else.
Four faces of the interview theatre
The people inside this system aren’t uniform, and treating “AI cheating in interviews” as one behaviour obscures more than it reveals. In practice, four recognisable archetypes show up on both sides of the table.
The Augmented Performer. Uses an AI overlay not because they lack any relevant knowledge, but because the tool removes risk from an already-stressful, high-stakes performance: reading a slightly better-structured version of an answer they could have produced themselves, just slower and less cleanly. This is the hardest case morally and the hardest case to detect, because the underlying competence often genuinely exists; what’s fabricated is the polish, not the substance.
The Panicked Improviser. Has little relevant knowledge for the specific role, discovered the tool out of desperation in a brutal hiring market, and is functionally borrowing an entire skill set for the length of one video call. This is the case the fraud statistics are mostly measuring, and the one most likely to produce a painful collapse once the job actually starts and the overlay is gone.
The Honest Straggler. Refuses to use any AI assistance during interviews, on principle or simply because they haven’t thought to, and is now competing directly against a fluency bar partly set by candidates in the first two categories: absorbing, without realising it, most of the real cost this entire dynamic imposes on the labour market.
The Detector’s Blind Spot. The hiring-side archetype: a recruiter or hiring manager who trusts their own in-person read of a candidate more than they trust detection software, unaware that human accuracy at spotting this specific category of deception sits close to chance. Not naive, exactly, just operating on an instinct that used to be reliable and now, quietly, often isn’t.
Most real interviews contain some blend of these on both sides of the table simultaneously, and the blend is precisely what makes any single detection method (human intuition, statistical timing analysis, in-person format) insufficient on its own. A company that solves for the Panicked Improviser with better timing-analysis software will still walk directly into the Augmented Performer, because the statistical fingerprint of “reading a slightly better version of an answer I could have produced myself” is far subtler than the fingerprint of “reading an answer to a question I have never encountered before.” Solving for one archetype at a time is how a detection strategy ends up looking effective in a vendor’s own case studies while quietly missing most of the actual problem.
What actually holds up
None of this is an argument that hiring is doomed, or that every candidate using AI assistance is acting in bad faith, or that detection software is a solved problem. It’s an argument for being honest about what an interview can and can’t currently tell you, and adjusting accordingly on both sides of the table.
For candidates: the signal that used to differentiate you, sounding sharp and fluent under pressure, is now available to anyone with a hidden browser tab, which means it differentiates you less than it used to, whether or not you personally use any AI assistance. What still differentiates you is depth that survives unscripted follow-up: a body of real, specific, defensible work you can discuss for fifteen minutes past the point any overlay would have run dry, in detail no generated script would have anticipated. Build toward being interrogated, not toward being fluent for exactly as long as a scripted answer lasts.
For hiring teams: stop optimising interview scoring for fluency, because fluency is now the cheapest thing in the entire process to fake. Score the follow-up chain instead: five or six specific, unscripted “why” and “what if” questions that force a candidate to reason live, off any script, about their own stated answer. Overlay tools can sustain a first answer convincingly. They struggle badly, and increasingly visibly, past the third layer of genuine improvisation, because genuine improvisation is precisely the thing they were never built to generate under real-time constraint.
For leaders: if you can’t currently describe, with some confidence, how your last ten hires actually got hired (not how fast, not how cheaply, but on what evidence you now trust) you’re running an organisation built partly on signal you can no longer verify. That gap doesn’t show up in an onboarding week. It shows up the first time someone on your team faces a genuinely novel problem with no script, no earpiece, and no panel watching their confidence instead of their actual output, and by then, the cost of finding out has moved from a bad interview to a bad quarter.
The closing thought
Priya still can’t fully explain, months later, what tipped her off about that candidate before the detection report confirmed it. She describes it as a feeling that the interview was “too finished”, that every answer arrived complete, in a way real thinking rarely does, without the small hesitations and self-corrections and half-formed detours that usually mark someone actually working through a problem live. She has started listening for those detours deliberately now, in every interview she runs, treating the mess of real thinking as a signal worth scoring rather than a flaw worth penalising.
That’s a small, human-scale adjustment to an arms race that neither candidates nor employers are close to winning outright. It won’t stop the next overlay tool, and it won’t restore the interview to whatever imperfect authority it briefly had. But it points at the thing actually worth protecting underneath all of this: not the appearance of fluency, which has become nearly worthless as a signal, but the visible, occasionally clumsy evidence of someone actually thinking, live, with nowhere to hide it and no script to hide behind. That evidence is still real. It’s just gotten a great deal harder to find in a room full of people who’ve learned exactly how to look like they have it.


