Large language models (LLMs) are increasingly being integrated into routine clinical workflows from operating theatres to hospital wards. Surgeons and other clinicians are now using tools such as ChatGPT not only to transcribe notes but to draft discharge summaries and operation reports, summarise case notes and even generate literature reviews in seconds. Trained on vast corpora of text, LLMs generate rapid and detailed responses that offer clear efficiency advantages in time-constrained environments. While reliance on technology to reduce workload is not new, the conversational and context-aware nature of LLMs allows them to generate outputs that appear complete and reasoned. Yet there is an unspoken risk: when we permit algorithms to speak on our behalf, are we also granting them the ability to think for us?

In plastic surgery, this risk may be subtle because the output can seem clinically sound before the underlying reasoning has been independently constructed. A registrar can use an LLM to draft a ward-round plan, operative report or follow-up letter that appears coherent and complete, and a consultant can sign it off quickly. Yet the cognitive work that normally builds judgement and detects case-to-case variation may not have occurred. The real risk is not that LLMs write well, but that they may tempt us to stop thinking things through and start approving what ‘seems reasonable’. This shift, toward reduced scrutiny, aligns with established concerns about automation bias and unrecognised skill decay that occur when technological assistance is routinely trusted.1,2

A recent research paper by Kosmyna and colleagues at the Massachusetts Institute of Technology, alongside other electroencephalographic work, suggests that this risk might not be just a hypothetical concern.3,4 In a controlled experiment, participants who relied on ChatGPT to write essays showed reduced neural engagement and poorer memory for what they produced.3 The authors call this ‘cognitive debt’, a mental quieting that can persist even after the machine is turned off.3 This is mechanistic, preprint-level evidence, so should be read as a warning rather than proof of harm.3 Nonetheless, the signal warrants attention. When the LLM generates coherent output, the clinician’s role can shift from constructing the reasoning to editing it, and people may later struggle to explain or reproduce what was actually ‘theirs’.3 In a busy plastic surgery practice, that kind of drift could easily go unnoticed.

Plastic surgery is particularly exposed to this risk because much of its decision-making occurs in the surgeon’s mind in real time, not on a checklist. Plastic surgeons are trained to picture the reconstruction several steps ahead and to keep backup plans ready as the case unfolds. If venous outflow looks poor, they revise the anastomosis. If perfusion is borderline, they redesign or stage the reconstruction. An LLM can describe these steps neatly on the page, but it cannot supply the lived judgement that decides which step to take at the operating table.

The potential for cognitive offloading is pronounced in daily plastic surgery practice. Operation reports, consent summaries and perioperative notes can now be generated almost instantly from bullet points or voice prompts. When such practices become routine, surgeons risk displacing their own recall and clinical judgement, reviewing decisions they did not actively formulate and bypassing the reflection and variance detection, which Crebbin and colleagues identify as hallmarks of surgical expertise decision-making.5 What begins as documentation support risks evolving into cognitive substitution.

In plastic surgery, where small technical details can have significant downstream consequences, linguistic fluency can obscure clinically relevant omissions, including laterality, implant type and size, ischaemia time, recipient vessels, drain plans, anticoagulation instructions or the specific intraoperative issue that should trigger early review.6 The danger is not that the LLM makes a mistake, but that a plausible narrative reduces the impulse to question what is missing or uncertain.2

The risk is amplified in training. This is because performance can improve while learning stalls. If an LLM generates the plan’s structure, the explanation or the differential, the trainee’s role shifts from construction to selection and editing. This is consistent with the concern that artificial intelligence (AI) assistance may hinder skill development or accelerate skill decay without the LLM user’s awareness.2 Surgical education already flags over-reliance on AI as a threat to independent decision-making. This matters most when the case deviates from expectations, which is precisely when judgement is required.7

Even among experienced surgeons, the ‘illusion of oversight’ is persuasive because the narrative sounds clinically plausible. Unlike dictation or templates that largely reflect the surgeon’s own thinking, LLMs can introduce confident detail that was never verified. Existing studies on generative AI in plastic surgery have already highlighted ethical and safety concerns, including the risk that AI recommendations can lead surgeons to overlook their own clinical judgement.6 Over-reliance on LLMs is also linked with uncritical acceptance of incorrect or hallucinated outputs, which has implications for clinical documentation and patient-facing information.8

The solution is not to avoid LLMs but to use them deliberately, in ways that protect the cognitive steps that matter (Figure 1). A simple rule is that the surgeon (or trainee) must first write the indication and operative plan in their own words, identify the key risks and state a realistic Plan B. Only then should the model be used to improve clarity, structure or patient-appropriate language. This preserves the cognitive ‘reps’ that underpin safe judgement and variance detection, rather than outsourcing them to a fluent draft.2,5

Fig. 1
Fig. 1.Cognitive debt risk map for large language model (LLM) use in plastic surgery. The map stratifies common examples of LLM use by clinical stakes and degree of cognitive substitution. The examples illustrate where LLM use primarily improves expression and organisation versus where it risks replacing cognitive work that underpins safe plastic surgery judgement.

Used well, LLMs can reduce administrative overhead and improve communication. Their best role in plastic surgery is to polish and organise what the clinician had already decided, not generate the reasoning that the clinician then merely signs off on.6 If the model can produce a plan faster than a trainee can explain it, that should be treated as a prompt to rebuild the mental model, not a reason to move on.2 The cognitive process that safeguards patients cannot be delegated. In plastic surgery, passive fluency is never sufficient. AI must support, not replace, independent surgical judgement.6


Conflict of interest

The authors have no conflicts of interest to disclose.

Funding declaration

The authors received no financial support for the research, authorship and/or publication of this article.