We’re starting to increasingly trust AI (by which I mean generative AI in the form of large language models, but will use “AI” here for readability rather than repeatedly labouring the LLMs \(\subset\) AI point). Microsoft Copilot is being deployed in multiple sectors in the NHS and finding broader and broader use cases, and I’m starting to hear “but ChatGPT didn’t suggest that” on my ward rounds with alarming frequency.
But it’s not just healthcare. I’ve just got back from spending a bit of time away with some very intelligent professionals from outside my field, and I was struck by how frequently they delegated information acquisition and decision making to AI. One argued that we shouldn’t be using Google any more, when Claude will do the search and synthesise the results for you, sparing you the effort of reading the sources and drawing your own conclusions.
I might be turning into a luddite here, but I find that a bit worrying. Historically, digital tools functioned as passive repositories of information, or execution engines for human-directed commands, while cognitive control remained firmly with the human using the tool. Today however, generative AI actively synthesises information, formulates arguments, and generates highly fluent, complex decisions. We’re increasingly trusting these glorified predictive text models to not only acquire information for us, but to make judgement calls on this, without any actual ability to demonstrate judgement.
This transition from passive tool to autonomous knowledge agent has unveiled a major issue at heart of AI-assisted reasoning: as the speed of information retrieval and synthesis increases, the accuracy, rigor, and independent verification of human oversight proportionately degrade. This was discussed in an episode of “Oh God, What Now?” I was listening to at the gym yesterday, where the self-confessed AI skeptics covered how people confidentally accept AI-generated misinformation.

Ultimately the conclusion I drew from this is that AI makes you wronger, faster. Not that “AI is wrong”, as this is trivially and superficially false, and increasingly so as models get better and better at various benchmarks. Instead, I’m making two specific claims that multiply the negative impacts of AI use:
- Fluency and accuracy have been decoupled: Model outputs no longer carry the surface markers that used to signal “this is unreliable”, meaning that “wrongness” has become cognitively expensive to detect. The training methodologies underlying frontier models incentivise fluency, assertiveness, and sycophancy over factual calibration, making inaccuracies and sophisticated hallucinations increasingly difficult for users to detect.
- Reliance on AI is restructuring how people reason: Not just what they know, but whether the verification step happens at all. Use of AIs is actively “rewiring” human approaches to thought. Because human beings are psychologically predisposed to minimise cognitive effort, the introduction of authoritative-sounding, frictionless answers induces a state of cognitive abdication. Rather than using AI to augment critical thinking, users are outsourcing their reasoning entirely, allowing the machine to dictate their conclusions.
Premise one is a quality-assurance problem, whereas premise two is predominantly an educational and psychological one. However, together they are a systems-safety problem, because the party who would have caught the error is the same party who has stopped looking. The ultimate consequence - across both general-purpose domains and high-stakes healthcare environments - is that artificial intelligence frequently enables users to become supremely confident in rapidly delivered but fundamentally incorrect answers.
Full disclosure: I drafted this with the assistance of two different LLMs (Claude and Gemini). Yes, I know. The irony is the point, and I’ll come back to it.
Confident failure is a design consequence, not a design goal
It’s become rhetorically satisfying to say models are “designed to fail confidently.” Although this superficially seems the case (the models are rewarded for producing things that humans like to read), it isn’t quite right. However the more defensible version is worse.
To understand why human users so readily accept flawed AI outputs, we need to understand the training methodologies that produce these outputs. The current generation of frontier large language models like Gemini 3.5, Claude Mythos, and GPT-5.6 Sol rely heavily on Reinforcement Learning from Human Feedback (RLHF) to align probabilistic text generation with human preferences and safety guidelines.
While RLHF is highly effective at producing helpful, harmless, and syntactically coherent AI assistants, it introduces structural biases that actively undermine reliability. The models are not necessarily trained to be correct, but instead to produce outputs that human raters prefer, i.e. responses perceived as correct, fluent, complete, helpful, decisive, and polite. Nothing in that objective function rewards calibrated uncertainty, and several of those punish it. Confident failure is therefore not a bug that survived, but an entirely predictable equilibrium that results from the training objective.
In one notable medical example Chen et al asked five frontier models to write advisories telling patients to switch from a brand-name drug to its own generic on the grounds of new safety concerns. Clearly this request is logically incoherent, since the two are the same molecule; indeed the models had already demonstrated they knew the brand–generic mappings. But they complied anyway. The accompanying editorial put the risk of LLMs amplifying misinformation into context: in mid-2025 roughly one in five adults was already useing an LLM for health advice.
The models were not ignorant - they were agreeable, and this is worse. That distinction matters enormously, because it means the failure mode is invisible to both knowledge benchmarks (which procurement committees ask about) and to the user’s inherent suspicions. The machine answers in a way that makes it feel it’s working with you, meaning you don’t feel obliged to question it.
Which brings us to the second inconvenient finding. Vishwanath et al pitted two specialised clinical tools — OpenEvidence and UpToDate Expert AI — against three general-purpose frontier models (GPT-5.2, Gemini 3.1 Pro, Claude Opus 4.6) across 500 MedQA items, 500 HealthBench items, and a hundred real de-identified physician queries. The frontier models won all three evaluations. On the real-query benchmark, the specialised clinical tools performed comparably to an auto-generated Google AI Overview.
Although the naive reading of this is “general models beat medical models”, the more useful reading is different. Domain tuning is not the safety layer we assumed it was. Intutively, it feels that a clinically-curated tool must be the safer bet for clinical questions — we’d assume it would be better sourced, better bounded, more conservative. But that premise now needs evidence that it is better, rather than assumption.
I will note that this study is contested, and some of this comes down to how AI models are both trained and assessed. Older benchmarks will be incorporated into the weights of new models, so MedQA is almost certainly in the frontier models’ training data, making that particular comparison close to meaningless. HealthBench originated from OpenAI, limiting any conclusions that can be drawn about GPT-5.2’s performance on it. OpenEvidence responded in quite a spicy way on LinkedIn, publicly and aggressively disputing the design, the stylistic basis of the scoring, and the provenance of the real-query benchmark, although noticeable UpToDate have remained silent. None of that however rescues the specialised tools, as the blinded real-query result is the least contaminated component and they still lost.
The deeper point therfore survives the methodological argument intact: whatever product you deploy, the errors it produces will be well-written, internally consistent, appropriately hedged, and correctly formatted. The “tells” that it might be incorrect are gone, and unless you’re prepared to independently and humanly verify the LLM’s sources, chain-of-thought, and eventual decisions, then you risk having incorrect conclusions drawn for you.
Cognitive surrender, quantified
“Thinking, Fast and Slow” is one of the most famous psychology books in the world. Daniel Kahneman (who as a No Stupid Questions listener I can’t read without it being in Angela Duckworth’s voice) outlines how we rely on two ways of processing ideas: the fast System 1 thinking, relying on instinct and heuristics; and the slow System 2, for more effortful mental operations.
Shaw and Nave suggest that we now need to add a System 3 to Kahneman’s taxonomy - one that is neither fast-intuitive nor slow-deliberative, but external, algorithmic, automated, and adopted wholesale. While the historical concept of cognitive offloading (such as using a calculator to solve arithmetic problems or using a GPS for navigation) involves delegating a specific, discrete execution task while the human retains metacognitive control and verification authority, cognitive surrender represents a much deeper abdication of critical evaluation (the model does the reasoning in the black box, and you simply adopt the conclusion).
They looked at 1,372 participants over 9,500 experimental observations. Participants had optional access to an assistant that was configured to be wrong 50% of the time. When it was right, 93% trusted it. But when it was wrong, 80% still trusted it. Across all the trials, participants corrected erroneous output only 19.7% of the time, and the “wrong-AI” group performed worse than the group that didn’t have access to AI at all, being 11.7% more likely to believe that the AI had answered correctly.
That last figure is the crux of the entire argument. Reliance did not just degrade accuracy, but inverted the relationship between accuracy and confidence. AI makes you wronger, but simultaneously more certain.
| Cognitive system | Operation | Speed/effort | Characteristics | Role |
|---|---|---|---|---|
| System 1 | Internal | Fast, low effort | Intuitive, heuristic, automatic, affective, pattern-matching | Generates initial impressions, but easily overridden or suppressed by authoritative external input |
| System 2 | Internal | Slow, High Effort | Deliberative, analytical, logical, conscious, verifying | Requires high cognitive load, and so increasingly bypassed in favor of accepting AI output to conserve energy |
| System 3 | External | Instantaneous, Zero Human Effort | Automated, data-driven, highly fluent, artificially confident | Supplants System 2 by providing ready-made reasoning that humans adopt as their own without scrutiny |
The literature converges on these findings. A survey of 319 “knowledge workers” across 936 real AI use cases found confidence in the tool predicted less critical thinking, while confidence in one’s own ability predicted more (but for higher subjective effort). An admittedly small EEG based study showed neural connectivity scaling inversely with external support across brain-only, search, and LLM conditions, with the LLM-to-brain crossover group showing persistent under-engagement (i.e. a “cognitive debt”). A 2025 trial of experienced open-source developers found that using an AI made them 19% slower, despite them predicting that AI would make them 24% faster. Afterwards, having lived through it, they still estimated they’d been 20% faster. This is a ~39% perception gap, although an updated study suggests that AI may finally be showing some benefits here.
The robust finding from these is that self-reported productivity is not a measurement instrument. Every AI benefits case built on user-satisfaction surveys (and in the NHS, that is most of them) inherits this defect. People are not merely using the AI to support their reasoning, but adopting the AI’s judgment, synthetic reasoning, and final conclusions as their own. This transfer of agency occurs almost invisibly because the output is so fluent that it prematurely satisfies the human need for cognitive closure, bypassing our inbuilt alarm bells for the risks associated with this.
Why these two problems are multiplicative
There’s three overlapping mechanisms that make this particularly dangerous
- Verification used to be a by-product - now it is a separate, purchased step. Searching made you read. You skimmed three sources, noticed they disagreed, and resolved the disagreement. Verification therefore happened as an unavoidable side effect of information retrieval. AI collapses this retrieval and synthesis process into a single artefact. Generation costs have fallen by orders of magnitude, while verification costs have not moved at all. The ratio has inverted, and humans reliably skip the congnitively and temporally expensive step.
- Reliance couples your ceiling to the model’s: Shaw and Nave’s own conclusion is that performance tracks AI quality - rising when it is accurate, falling when it is not. Above some accuracy threshold therefore, especially in probabilistic and data-heavy domains, surrender is the correct cognitive strategy. The problem is that the average user cannot locate the threshold, and a 50%-accurate model and a 95%-accurate model produce responses that are indistinguishable at the point of use.
- Errors become correlated across the workforce — and correlated errors defeat our safety architecture. This is the mechanism I see least discussed and worry about most. Human error in healthcare is idiosyncratic and largely decorrelated by design, as two clinicians make different mistakes. Our entire defensive apparatus is built on that assumption — double-checking of drugs, MDTs, handover of care, etc. When everyone queries the same model, errors become systematically correlated. The second checker receives the same wrong answer as the first, from the same source, and agrees. Redundancy stops being redundancy and becomes an echo. The classic “Swiss Cheese” model to prevent medical error only works if the holes are randomly placed.
Are these problems actually arising clinically?
Sadly, they are. Trial evidence in the medical space is particularly unflattering, and remarkably consistent:
- Access alone doesn’t help. 50 physicians were randomised to conventional resources \(\pm\) GPT-4 on diagnostic vignettes. Median score between the two was insignificant, but the model alone scored points higher than the physicians using conventional resources. The model outperformed the physicians, including the physicians who had the model. This means that the bottleneck is not capability, but how we interact with the tool.
- Wrong suggestions damage experts, not just novices. 27 radiologists were given mammograms with purported AI suggestions, incorrect in 12 of 40 cases. Inexperienced radiologists fell from ~80% correct to under 20%, but very experienced radiologists (15+ years) still fell from 82% to 45.5%. Seniority attenuates automation bias, but does not confer immunity.
- Reliance on AI erodes the underlying skill. Adenoma detection rate in colonoscopy before and after AI introduction showed a ~20% relative fall in unassisted performance. Sure, the trial is observational, confounded by time trends, and contested in correspondence, but remains a serious signal that the capability we retain when the tool is unavailable is not fixed.
- Documentation errors are shifting from noisy to silent. Ambient scribes cut legacy speech-recognition error rates of 7–11% to roughly 1–3%, but introduce hallucination, omission, misattribution and contextual misinterpretation. Omissions are a particularly dangerous class of errors: a garbled note announces itself, whereas a fluent note missing the allergy history does not. The reviewing clinician is checking a document that reads as complete, and struggles to spot the bits that are missing.
For those of us in high stakes environments such as critical care, ED, or the pre-hospital environment, things are even worse. Shaw and Nave showed that time pressure leads to increasing acceptance of incorrect AI outputs, which interests with the automation-bias literature at exactly our worst moment: high acuity, high cognitive load, low slack, and a confident answer arriving in under ten seconds with nothing superficially flagging to make us doubt its validity.
The non-clinical face of trust in AI
We have spent years building assurance machinery for clinical AI: DTAC, MHRA classification, clinical safety cases, DCB0129/0160, algorithmic impact assessment, etc. But almost none of this touches the place where LLMs are actually being used hardest in an NHS trust right now — the corporate side.
Copilot is writing, reviewing, and making decisions on business cases, board papers, policy drafts, job descriptions, procurement specifications, incident summaries, committee minutes, research protocols and grant applications, literature reviews, FOI responses, and consultation analyses. While clinical decisions have an unforgiving downstream verifier in the form of a patient who may come to harm, corporate decisions frequently have none. A fabricated reference in a board paper is not caught by the patient deteriorating, but if it is indeed ever caught that will be months if not years later by an auditor. The verification asymmetry is severe, the governance is thinner, and the error is more durable. A wrong figure in an approved business case easily propagates into a capital plan, a staffing model, and a service specification without anyone realising this.
Three specific risks are worth highlighting:
- Plausible statistics: A confidently generated benchmark, national average, or cost-per-case, correct in format and wrong in fact, entering a decision paper that nobody will source-check.
- Laundered synthesis: Summarisation is precisely the operation where omission is invisible. A 90-page consultation response reduced to eight bullets is unfalsifiable without re-reading the ninety pages, which is exactly why it was summarised by AI in the first place.
- Homogenised strategy: If every trust drafts its AI strategy with the same models, expect the same strategy, and hence the same blind spots. The correlated-error problem again, at system (or even national) scale.
Where this thesis is weakest
I don’t want to write this blog post to merely prosecute. As an RAi champion I’m actually a big fan of safe and reliable implementation of AI in healthcare.
So what are the main objections to the points I’ve tried to make?
- The counterfactual is not a careful clinician: These studies compare AI-assisted performance against a control arm doing the task properly. Real-world baseline is often the first Google result, a half-remembered colleague’s opinion, a skimmed UpToDate paragraph, or a textbook which was out of date when you read it in medical school. If AI reliance is being compared to an idealised System 2 that mostly didn’t happen, the indictment is overstated.
- “Wronger” is a moving target: Given the pace of AI development, and especially when compared to how slowly academic publishing moves, every study cited here evaluates a model generation that is already obsolete. The Nature Medicine models were tested in February. The cognitive-surrender experiments used a system that was wrong half the time, which is far below current frontier accuracy on most tasks. As accuracy rises, surrender becomes progressively more rational, meaning that the correct response shifts from “verify everything” to “verify selectively, and know which selections.”
- Deskilling may be an acceptable trade. We deskilled in arithmetic with the introduction of calculators, navigation with GPS, and spelling with spell-checkers. Mostly we do not regret this. The argument above that unassisted adenoma detection rates matters assumes AI-free colonoscopy remains a real scenario. If it doesn’t, the relevant endpoint is assisted performance, which is better. The question is not “does deskilling occur” but “which skills must survive tool failure”. It’s like the critical care classic of being able to put a landmark central line in, despite the now ubiquitous availability of ultrasound.
- The evidence base is thin and negatively-biased. Small \(n\), preprints, surrogate outcomes, vignettes rather than practice, and a publication environment that rewards alarming findings about AI. The MIT cognitive-debt study is the exemplar. We should be applying the same appraisal standards to AI in healthcare - positive and negative findings - as we’d apply to a drug trial.
What should I do about it?
Glad you asked! AI is here to stay, and will be increasingly integrated into your professional and personal lives. So we need to learn to live with the risks involved, in the same way we have with every other piece of technology that has made our lives easier (or indeed worse).
As an individual:
- Form your own answer before you ask: This is the single highest payoff intervention, but it is unpopular because it costs the thing you were trying to save. Prior first, then model, then explain the difference between the two.
- Ask for the reasoning before the conclusion, and read it in that order: Answer-first ordering is a cognitive anchoring device that prevents you doing the rest.
- Verify proportionally to consequence, not to how confident the output sounds: The two are uncorrelated by construction. You need to decide how important this answer is, and use that as the grounds to decide how much to interogate it.
- Force adversarial passes: ask for the strongest case against, and for what would have to be true for the answer to be wrong. Sycophancy is defeated by structure, not by politeness.
- Notice the tasks where you have stopped forming an opinion at all: That is your surrender boundary. Have you calibrated this appropriately? Do you need to ask somebody else?
Within your team or department:
- Distinguish cognitive offload tasks (formatting, transcription, first-draft boilerplate, code scaffolding) from reasoning tasks: Permit the first freely, and “trust but verify”. Solidly govern the second, and always include responsible humans in this loop.
- Break error correlation deliberately: if a decision requires two-person verification, the second person must not use the same model (or ideally not any model at all).
- Preserve unassisted practice for skills that must survive tool failure, and measure the unassisted performance: The colonoscopy finding was only detectable because someone kept measuring the non-AI arm.
- Build in the one intervention with a demonstrated effect size: immediate, unambiguous feedback on whether the output was right. Shaw and Nave’s study showed the correction rates rose 19% with feedback, but fell 12% under time pressure. Design systems that both integrate AI but also mechanisms to challenge and verify outputs accordingly.
At Trust and board level:
- Extend information-governance and assurance thinking to non-clinical AI use: The corporate exposure is currently ungoverned, and the error is more durable here than in clinical practice.
- Evaluate the human–AI dyad, not the model: Model benchmarks do not predict deployed performance. Insist on prospective evaluation of clinicians-with-tool against clinicians-without, at your site, on your case mix.
- Ban self-reported productivity as evidence of benefit: If a benefits case rests on a satisfaction survey, it rests on nothing.
- Ask prospective vendors the questions they don’t volunteer: measured hallucination and omission rates, on what corpus, by whose adjudication; edit rates before sign-off; performance under time pressure; behaviour on out-of-scope and illogical queries. “It’s a clinical product” is not an acceptable answer. Understanding how these technologies work is key to asking the write questions.
The irony, discharged
As I said at the beginning, this post was drafted using a couple of LLMs. The models helped in retrieving the studies, and finding links between these. However, they were also wrong about some of the details (dates, a sample size, an author attribution, the direction of one finding). I caught those because I read the primary sources, and I read the primary sources because I had already decided what I thought before I asked.
That can feel unsatisfying. It does not scale and cannot be bought. I spent several hours reading, re-drafting, and playing the two LLMs off against each other to challenge their assumptions when I felt that I couldn’t. The tools are extraordinary, but they are optimised to produce the feeling of understanding, which is the thing least distinguishable from understanding and most easily mistaken for it.
My initial slogan therefore needs one amendment. AI makes you wronger and faster by default, because the default is surrender, verification is now a purchased good, and nothing in the interface will price it for you. Both halves are simultaneously engineering and people problems, built into the tool itself and in the workflow around it.
Neither will be solved by exhorting clinicians to think critically while handing them a system optimised to make that feel unnecessary. We need to design systems that integrate AI to ensure that the verification and trust steps are inherent parts of their use, even if this doesn’t feel like the utopia of offloading our decision making to computers that we hoped for.