Alexandre Gouveia

Academic primary care physician, medical educator, and clinician researcher.

What Should AI Supervise, and Who Should Supervise AI?

A resident sits down to write a note. An AI scribe has already drafted most of it from the recorded encounter. She reads it over, nods, and signs. Three years from now, will she still know how to write that note herself — and does it matter if she doesn’t?

This is the question sitting underneath most of the current literature on AI in postgraduate medical training, and it’s one I keep returning to as we think about how to supervise general practice residents in a world where the tools they use are getting smarter than the tasks we ask them to do.

The risk isn’t AI. It’s when AI shows up.

A recent Perspective in Nature Medicine puts a name to something many of us have sensed intuitively: “never-skilling” — the risk that trainees who lean on AI during the formative years of clinical training never build the reasoning skills they’ll need once the AI isn’t there, or is wrong [2]. That’s different from “de-skilling,” which is what happens to an experienced clinician who gets rusty. It’s also different from “mis-skilling,” where a trainee uncritically absorbs an AI’s error and internalizes it as fact.

The evidence base here is still thin — the authors are honest that direct evidence from medical training doesn’t yet exist — but the underlying learning theory is not new to anyone who has supervised residents: skills that are outsourced before they’re consolidated tend not to form at all. The proposed fix is a three-phase framework — build AI-independent baseline competency first, then structured critical calibration, then supervised AI integration [2]. In other words: sequence matters as much as supervision.

What good supervision of AI actually looks like

A companion piece in the New England Journal of Medicine gets more concrete, and more useful for anyone designing a curriculum today. The authors propose the DEFT-AI framework — Diagnosis, Evidence, Feedback, Teaching — as a structure for the Socratic conversation a supervisor should have whenever a trainee has used AI in a clinical encounter [1]. The trainee explains not just their clinical reasoning, but how and why they engaged the AI; they weigh evidence for and against the AI’s suggestion; the supervisor probes for gaps in both clinical reasoning and AI literacy; and the teaching that follows reinforces judgment, not just correct answers.

I find the paper’s “centaur vs. cyborg” distinction genuinely clarifying. In centaur mode, the human and the AI divide labor — the trainee reserves the higher-stakes judgment calls for themselves and lets the AI handle bounded, lower-risk tasks. In cyborg mode, the two are tightly interwoven, drafting and redrafting together — efficient, but far easier to slide into overreliance without noticing [1]. Most of the trainees I’ve watched using ambient scribes or draft-writing tools are, without quite realizing it, cyborging their way through documentation. That’s not necessarily wrong. But it’s worth naming, because it changes what supervision needs to check for.

Documentation is not just paperwork — it’s where reasoning happens

This point lands hardest for those of us in general practice, where the note has always done double duty: it’s a record, but writing it is also how a trainee is forced to prioritize, justify, and synthesize a genuinely messy case. A pilot study of an AI scribe across 48 internal medicine residents and nearly 1,000 notes found real efficiency gains — and real reason for caution, proposing seven best practices mapped to ACGME core competencies, from establishing baseline documentation skills before introducing the tool, to structured critical review of every AI-generated note [3]. The authors’ framing has stuck with me: AI as scaffold for reasoning, not substitute for it.

The encouraging part

None of this is an argument against using AI in training — quite the opposite. A 12-month longitudinal study of 372 medical students on supervised rotations using an AI-assisted diagnosis system found that greater engagement with the AI was associated with increases in both AI literacy and critical thinking over time, with AI literacy statistically mediating that relationship [4]. The effect was strongest among students with more technological experience and a mastery, rather than performance, orientation — a reminder that how a trainee relates to their own learning shapes how well they metabolize a tool like this. Under the right supervisory conditions, in other words, AI doesn’t have to erode judgment. It can build it.

The Human Learning Lab

There’s a real-world illustration of the never-skilling risk that’s more concrete than any framework paper: a multicentre study of endoscopists who had grown used to AI-assisted colonoscopy found that their adenoma detection rate without the AI running dropped from 28.4% to 22.4% after a few months of continuous exposure — a measurable erosion of an unaided skill, in clinicians who were already fully trained [5]. If exposure at that stage can erode a consolidated skill, the case for protecting the formative stage, before a skill is consolidated at all, is that much stronger.

Which raises a question I keep coming back to: if every other innovation in medical education has been about adding a new tool — the printing press, the stethoscope, the slide projector, the simulator, now the AI — what if the next real innovation is closer to old ground: to deliberately protect learning moments where there’s no tool at all? Not nostalgia for how things used to be taught, but a designed space in the curriculum for what I’ve started calling, half in jest, a Human Learning Lab — bedside teaching, case discussion, direct observation, and the kind of unhurried Socratic back-and-forth between trainee and supervisor that doesn’t route through a screen. The two oldest images of medical teaching, the packed anatomical theatre and the physician’s overnight vigil at a child’s bedside, both show the same thing: the tool in the room was never really the point. The apprenticeship was.

The Agnew Clinic, an 1889 oil painting by Thomas Eakins showing Dr. David Hayes Agnew and colleagues operating before a tiered amphitheater of medical students.
Thomas Eakins, The Agnew Clinic, 1889 — the amphitheater. Public domain, via Wikimedia Commons.
The Doctor, an 1891 oil painting by Samuel Luke Fildes showing a physician keeping an overnight vigil at a sick child's bedside.
Samuel Luke Fildes, The Doctor, 1891 — the bedside vigil. Public domain, via Wikimedia Commons.

This isn’t an argument against AI in training, and it doesn’t contradict the DEFT-AI or centaur/cyborg thinking above, which are exactly the frameworks you need once AI is in the room. It’s an argument for also keeping some training moments where it deliberately isn’t — not because the technology is dangerous, but because certain kinds of clinical judgment seem to need to be built the hard way at least once before they can be safely delegated at all.

Where this leaves supervision

Putting these threads together, the shape of a defensible approach to AI in postgraduate general practice training starts to emerge: keep a human supervisor accountable for the encounter, protect the early, foundational phase of training from too much AI assistance, and reserve AI for bounded, well-defined tasks — documentation support, formative feedback, case-based coaching, simulation — rather than treating it as an independent source of clinical judgment [1,2,3,4,5]. This isn’t a governance framework bolted onto existing GP supervision models; it’s closer to an extension of what good supervision was already trying to do — make explicit the reasoning that’s easy to skip, at exactly the moment a trainee is tempted to skip it.

The tools will keep changing. The question a supervisor asks at the end of the encounter — why did you do that, and why did you trust the machine when it told you to — probably won’t.

I spoke about this at a recent RMS event — the Human Learning Lab idea was actually the closing note of that talk — and the conversation there is a large part of what prompted me to pull this literature together properly.


References

  1. Abdulnour RE, Gin B, Boscardin CK. Educational strategies for clinical supervision of artificial intelligence use. N Engl J Med. 2025;393(8):786-797. Available from: https://pubmed.ncbi.nlm.nih.gov/40834302/
  2. Ke Y, Jin L, Ong JCL, Thirunavukarasu AJ, Car J, Cheung CY, et al. AI-induced never-skilling in medical education. Nat Med. 2026;32(6):1997-2006. Available from: https://pubmed.ncbi.nlm.nih.gov/42174254/
  3. Abernethy J, Shah A, Chen B, Reynolds S, Wright SM, O’Rourke P. Integrating AI scribes into medical education: guardrails for preserving clinical reasoning. J Gen Intern Med. 2026;41(9):2598-2602. Available from: https://pubmed.ncbi.nlm.nih.gov/41627656/
  4. Xin Y, Yan D, Shuren L, Minyang L, Liuheng L. AI literacy mediates AI assisted diagnosis participation and critical thinking among medical students under supervision. NPJ Digit Med. 2026;9(1). Available from: https://pubmed.ncbi.nlm.nih.gov/41832289/
  5. Budzyń K, Romańczyk M, Kitala D, Kołodziej P, Bugajski M, Adami HO, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol Hepatol. 2025;10(10):896-903. Available from: https://pubmed.ncbi.nlm.nih.gov/40816301/