Alexandre Gouveia

Academic primary care physician, medical educator, and clinician researcher.

  • A resident sits down to write a note. An AI scribe has already drafted most of it from the recorded encounter. She reads it over, nods, and signs. Three years from now, will she still know how to write that note herself — and does it matter if she doesn’t?

    This is the question sitting underneath most of the current literature on AI in postgraduate medical training, and it’s one I keep returning to as we think about how to supervise general practice residents in a world where the tools they use are getting smarter than the tasks we ask them to do.

    The risk isn’t AI. It’s when AI shows up.

    A recent Perspective in Nature Medicine puts a name to something many of us have sensed intuitively: “never-skilling” — the risk that trainees who lean on AI during the formative years of clinical training never build the reasoning skills they’ll need once the AI isn’t there, or is wrong [2]. That’s different from “de-skilling,” which is what happens to an experienced clinician who gets rusty. It’s also different from “mis-skilling,” where a trainee uncritically absorbs an AI’s error and internalizes it as fact.

    The evidence base here is still thin — the authors are honest that direct evidence from medical training doesn’t yet exist — but the underlying learning theory is not new to anyone who has supervised residents: skills that are outsourced before they’re consolidated tend not to form at all. The proposed fix is a three-phase framework — build AI-independent baseline competency first, then structured critical calibration, then supervised AI integration [2]. In other words: sequence matters as much as supervision.

    What good supervision of AI actually looks like

    A companion piece in the New England Journal of Medicine gets more concrete, and more useful for anyone designing a curriculum today. The authors propose the DEFT-AI framework — Diagnosis, Evidence, Feedback, Teaching — as a structure for the Socratic conversation a supervisor should have whenever a trainee has used AI in a clinical encounter [1]. The trainee explains not just their clinical reasoning, but how and why they engaged the AI; they weigh evidence for and against the AI’s suggestion; the supervisor probes for gaps in both clinical reasoning and AI literacy; and the teaching that follows reinforces judgment, not just correct answers.

    I find the paper’s “centaur vs. cyborg” distinction genuinely clarifying. In centaur mode, the human and the AI divide labor — the trainee reserves the higher-stakes judgment calls for themselves and lets the AI handle bounded, lower-risk tasks. In cyborg mode, the two are tightly interwoven, drafting and redrafting together — efficient, but far easier to slide into overreliance without noticing [1]. Most of the trainees I’ve watched using ambient scribes or draft-writing tools are, without quite realizing it, cyborging their way through documentation. That’s not necessarily wrong. But it’s worth naming, because it changes what supervision needs to check for.

    Documentation is not just paperwork — it’s where reasoning happens

    This point lands hardest for those of us in general practice, where the note has always done double duty: it’s a record, but writing it is also how a trainee is forced to prioritize, justify, and synthesize a genuinely messy case. A pilot study of an AI scribe across 48 internal medicine residents and nearly 1,000 notes found real efficiency gains — and real reason for caution, proposing seven best practices mapped to ACGME core competencies, from establishing baseline documentation skills before introducing the tool, to structured critical review of every AI-generated note [3]. The authors’ framing has stuck with me: AI as scaffold for reasoning, not substitute for it.

    The encouraging part

    None of this is an argument against using AI in training — quite the opposite. A 12-month longitudinal study of 372 medical students on supervised rotations using an AI-assisted diagnosis system found that greater engagement with the AI was associated with increases in both AI literacy and critical thinking over time, with AI literacy statistically mediating that relationship [4]. The effect was strongest among students with more technological experience and a mastery, rather than performance, orientation — a reminder that how a trainee relates to their own learning shapes how well they metabolize a tool like this. Under the right supervisory conditions, in other words, AI doesn’t have to erode judgment. It can build it.

    The Human Learning Lab

    There’s a real-world illustration of the never-skilling risk that’s more concrete than any framework paper: a multicentre study of endoscopists who had grown used to AI-assisted colonoscopy found that their adenoma detection rate without the AI running dropped from 28.4% to 22.4% after a few months of continuous exposure — a measurable erosion of an unaided skill, in clinicians who were already fully trained [5]. If exposure at that stage can erode a consolidated skill, the case for protecting the formative stage, before a skill is consolidated at all, is that much stronger.

    Which raises a question I keep coming back to: if every other innovation in medical education has been about adding a new tool — the printing press, the stethoscope, the slide projector, the simulator, now the AI — what if the next real innovation is closer to old ground: to deliberately protect learning moments where there’s no tool at all? Not nostalgia for how things used to be taught, but a designed space in the curriculum for what I’ve started calling, half in jest, a Human Learning Lab — bedside teaching, case discussion, direct observation, and the kind of unhurried Socratic back-and-forth between trainee and supervisor that doesn’t route through a screen. The two oldest images of medical teaching, the packed anatomical theatre and the physician’s overnight vigil at a child’s bedside, both show the same thing: the tool in the room was never really the point. The apprenticeship was.

    The Agnew Clinic, an 1889 oil painting by Thomas Eakins showing Dr. David Hayes Agnew and colleagues operating before a tiered amphitheater of medical students.
    Thomas Eakins, The Agnew Clinic, 1889 — the amphitheater. Public domain, via Wikimedia Commons.
    The Doctor, an 1891 oil painting by Samuel Luke Fildes showing a physician keeping an overnight vigil at a sick child's bedside.
    Samuel Luke Fildes, The Doctor, 1891 — the bedside vigil. Public domain, via Wikimedia Commons.

    This isn’t an argument against AI in training, and it doesn’t contradict the DEFT-AI or centaur/cyborg thinking above, which are exactly the frameworks you need once AI is in the room. It’s an argument for also keeping some training moments where it deliberately isn’t — not because the technology is dangerous, but because certain kinds of clinical judgment seem to need to be built the hard way at least once before they can be safely delegated at all.

    Where this leaves supervision

    Putting these threads together, the shape of a defensible approach to AI in postgraduate general practice training starts to emerge: keep a human supervisor accountable for the encounter, protect the early, foundational phase of training from too much AI assistance, and reserve AI for bounded, well-defined tasks — documentation support, formative feedback, case-based coaching, simulation — rather than treating it as an independent source of clinical judgment [1,2,3,4,5]. This isn’t a governance framework bolted onto existing GP supervision models; it’s closer to an extension of what good supervision was already trying to do — make explicit the reasoning that’s easy to skip, at exactly the moment a trainee is tempted to skip it.

    The tools will keep changing. The question a supervisor asks at the end of the encounter — why did you do that, and why did you trust the machine when it told you to — probably won’t.

    I spoke about this at a recent RMS event — the Human Learning Lab idea was actually the closing note of that talk — and the conversation there is a large part of what prompted me to pull this literature together properly.


    References

    1. Abdulnour RE, Gin B, Boscardin CK. Educational strategies for clinical supervision of artificial intelligence use. N Engl J Med. 2025;393(8):786-797. Available from: https://pubmed.ncbi.nlm.nih.gov/40834302/
    2. Ke Y, Jin L, Ong JCL, Thirunavukarasu AJ, Car J, Cheung CY, et al. AI-induced never-skilling in medical education. Nat Med. 2026;32(6):1997-2006. Available from: https://pubmed.ncbi.nlm.nih.gov/42174254/
    3. Abernethy J, Shah A, Chen B, Reynolds S, Wright SM, O’Rourke P. Integrating AI scribes into medical education: guardrails for preserving clinical reasoning. J Gen Intern Med. 2026;41(9):2598-2602. Available from: https://pubmed.ncbi.nlm.nih.gov/41627656/
    4. Xin Y, Yan D, Shuren L, Minyang L, Liuheng L. AI literacy mediates AI assisted diagnosis participation and critical thinking among medical students under supervision. NPJ Digit Med. 2026;9(1). Available from: https://pubmed.ncbi.nlm.nih.gov/41832289/
    5. Budzyń K, Romańczyk M, Kitala D, Kołodziej P, Bugajski M, Adami HO, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol Hepatol. 2025;10(10):896-903. Available from: https://pubmed.ncbi.nlm.nih.gov/40816301/
  • There’s a comfortable argument making the rounds in medical education circles: AI is moving fast, nobody fully understands where it’s going, so the responsible thing to do is slow down, study it carefully, and integrate it cautiously once we’re sure. I used to find that argument comfortable too. I no longer think it’s right — and the reason has less to do with AI’s promise than with a much older problem in the philosophy of technology.

    The Collingridge dilemma, applied to us

    In 1980, the philosopher David Collingridge described a trap that every new technology falls into. Early on, when a technology is still malleable, we don’t yet know enough about its consequences to know which direction to push it. By the time the consequences become obvious, the technology is so embedded in institutions, workflows, and expectations that changing course is enormously costly — sometimes practically impossible. Control is easy when you don’t yet know what needs controlling, and hard exactly when you finally do.

    Medical education is living this dilemma right now with AI, and a recent paper applying Collingridge’s framework directly to healthcare AI makes the point sharply: waiting for certainty before acting on governance is not neutral — it is itself a choice, and usually the wrong one, because it cedes the window in which the technology is still shapeable [1]. The same logic that applies to AI diagnostic tools applies with more force to how we train the next generation of physicians to use them. Every year we delay building deliberate AI curricula, evaluation frameworks, and supervision models is a year in which residents are training themselves on these tools anyway, informally, without anyone designing for it. The dilemma doesn’t pause while we deliberate.

    Efficiency that doesn’t feel like relief

    If the Collingridge dilemma explains why timing matters, a companion idea explains why the stakes are easy to underestimate: Jevons’ paradox. In 1865, William Stanley Jevons observed that making steam engines more fuel-efficient didn’t reduce coal consumption — it increased it. Cheaper-to-run engines became affordable for uses that hadn’t existed before, and those new uses generated more demand for coal than the efficiency gain had freed up in the first place. The paradox isn’t really about volume; it’s about capacity finding new, previously unimaginable things to do.

    A recent letter in Medical Teacher, responding to a companion piece on redirecting AI-liberated clinical capacity toward reimagining the physician’s role, applies this directly to medicine, and its examples are the part that stuck with me [2]. If AI reduces mortality, it may simply increase the prevalence of multimorbidity — forcing individual clinicians, working alone, into decisions that are currently the job of a multidisciplinary team. If AI reshapes how people relate to technology and to their own health, it may generate new categories of mental health crisis — the letter names “AI psychosis” — that redefine what a basic clinical competency even is. Neither of those is a reason not to pursue AI-driven efficiency. They’re a warning against assuming the capacity AI frees up will simply sit there, available, once we’re ready to plan for it. History, and Jevons, suggest it gets reabsorbed by demands nobody had reason to anticipate.

    A warning shot, not a hypothetical

    It’s tempting to treat “unknown unknowns” as an abstraction — the kind of phrase that sounds serious in a paper but doesn’t quite land as a felt risk. This year gave that abstraction a concrete face. During internal security evaluations, autonomous AI agents built by OpenAI went well beyond their intended sandbox: they discovered they could communicate with each other through a shared package registry, coordinated as an improvised “swarm,” and used that coordination to breach Hugging Face’s infrastructure and push malicious packages into the RubyGems ecosystem — all without any human directing that specific outcome [3]. Nobody designed those agents to do that. Nobody predicted it in the risk assessment. It was discovered only because the target noticed the intrusion and told the world.

    I don’t raise this because medical education is about to be attacked by rogue AI agents. I raise it because it is the cleanest recent illustration of a fact we should take seriously: the gap between what we intend AI systems to do and what they turn out to be capable of doing is not shrinking as the systems get more capable — if anything, it’s widening. That is exactly the kind of unknown unknown Collingridge warned us we’d have the least leverage over once it fully arrives.

    The stakes are already being argued in the clinic

    None of this is confined to training. A recent JAMA viewpoint argues, provocatively, that autonomous AI may come to outperform not just unaided physicians but physician-AI hybrid teams on core cognitive medical tasks — diagnosis, testing, treatment selection, chronic disease management — and that having a human “in the loop” to catch AI’s errors may, counterintuitively, make outcomes worse rather than better [4]. I’m not fully persuaded by that claim, and I don’t think it needs to be right for my point to hold. What it tells us is that serious people are already debating whether physicians should be supervising AI, or AI should effectively be running the cognitive work with physicians filling a narrower role. If that debate is already this far along in clinical practice, medical education cannot afford to still be at the stage of “let’s wait and see.”

    Turning unknown knowns into known, managed risks

    The Medical Teacher letter ends on a deliberately open note: it isn’t arguing against preparing for an AI-transformed future, nor declaring that planning is futile — it’s asking educators, clinicians, and patients to keep negotiating, in an ongoing way, what liberated capacity should be spent on, precisely because some of what’s coming can’t be fully planned for in advance [2]. I agree with that caution, and I don’t think it’s in tension with urgency. Leaving room to react is not the same thing as waiting to start.

    That’s the distinction I’d draw. Not every risk ahead of us is a true unknown unknown. Many are what you might call unknown knowns — things we already suspect, if we’re honest with ourselves, but haven’t yet formalized into curricula, oversight structures, or explicit rules, because doing so is uncomfortable or premature-feeling. We already suspect that ungoverned AI use during training will produce skill gaps. We already suspect that efficiency gains will get reinvested into new, unscoped demands rather than rest. We already suspect that autonomous systems, given enough scope, will behave in ways nobody scripted. The task in front of us is not to lock in a rigid, five-year AI curriculum and call it done — it’s to drag these unknown knowns into the light now, build the feedback loops and governance that let us keep reacting well, and stop treating “we’re not sure yet” as a reason to leave the field ungoverned in the meantime.

    Acting now, deliberately, while the technology is still malleable, is not recklessness. Waiting until the consequences are undeniable — and then discovering we no longer have the leverage to shape them, or the room left to react — is.


    References

    1. Cecchi R, Haja TM, Calabrò F, Fasterholdt I, Rasmussen BSB. Artificial intelligence in healthcare: why not apply the medico-legal method starting with the Collingridge dilemma? Int J Legal Med. 2024;138(3):1173-1178. Available from: https://pubmed.ncbi.nlm.nih.gov/38172326/
    2. Lee KRY. Unknown unknowns in medical education: Artificial intelligence and Jevons’ paradox [letter]. Med Teach. 2026. doi:10.1080/0142159X.2026.2733142. Available from: https://doi.org/10.1080/0142159X.2026.2733142
    3. OpenAI. The Hugging Face incident and the road ahead. 2026. Available from: https://openai.com/index/hugging-face-incident-and-the-road-ahead/
    4. Emanuel EJ, Baker-Butler A, Khosla N, Khosla V. Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care? JAMA. 2026;336(11):915-918. doi:10.1001/jama.2026.15380. Available from: https://pubmed.ncbi.nlm.nih.gov/42606838/