Learning sometimes happens, despite our best efforts.

A disorienting moment

  • AI passes nursing and medical licensing exams in the 90th percentile (Gilson et al., 2023)
  • Produces care plans, reflective portfolios, and clinical case analyses that satisfy standard assessment criteria (Gilart et al., 2025)
  • Blinded clinicians rated AI responses to real patient questions as more empathetic than physicians' (Ayers et al., 2023)
  • 24/7 access to AI tutors with adaptive responses, giving more consistent feedback than most human assessors (Jurenka et al., 2024)
  • 94% of AI-written submissions entered blind into a BSc Psych programme were undetected — and outscored real students (Scarfe et al., 2024)

Sanctuary strategies

  • Denial. "AI will never replace clinical judgement": probably true; doesn't resolve the problem that AI can already complete almost all assessments
  • Retreat. "Focus on what only humans can do": the boundary keeps moving as models improve, creating a shrinking perimeter
  • Restriction. "Prevent students from using it": no policy or detector has changed this and none are likely to (see "the internet")
  • Resignation. "Our roles will diminish": performance is not zero-sum

None of these strategies identifies what we want to protect.

Four premises

  1. Learning results from what the student does and thinks, and only from what the student does and thinks (Simon, in Ambrose, 2010)
  2. The aim of teaching is to create conditions that influence what students do and think (Ramsden, 2003)
  3. Assessment aims to infer learning from observable patterns in what students do and think
  4. Certification attempts to make judgements about students' future performance

AI has not changed any of these premises.

Becoming a nurse

  • This is the language of the profession: care, fitness to practise, professional identity, clinical judgement (formation concepts, not assessment outcomes)
  • Becoming is transformational; the person who completes the programme is different from the person who began it (Barnett, 2009)
  • But the infrastructure we've built to support this aspiration is designed around artifact production, not identity formation

The system has always been measuring a proxy for the developmental process it actually cares about

We cannot observe learning directly

  • Knowledge is built through effortful retrieval, not exposure; neural connections strengthen through repeated, effortful activation; conditions that feel harder produce better long-term outcomes (Bjork & Bjork, 2009; Willingham, 2009)
  • Student engagement is therefore a reasonable proxy for improved learning outcomes (Kuh, 2007)
  • This engagement produces artifacts that serve as evidence of the engagement (Carless, 2007)
  • The work of learning (the "desirable difficulty") and the artifact it produced were reliably connected, so the proxy held

You already know what this feels like.

The swampy lowland of assessment

  • Assessment has always had architectural flaws, accepted as manageable imprecision (e.g. arbitrary cut-scores, inter-rater variability, order-effects; extraneous factors)
  • Most feedback doesn't translate to improved performance (Hattie et al., 2007)
  • What we call "cheating" is defined relative to what we expect students to do themselves, rather than an inherent property of a task
  • Students were already optimising for grades over learning

This is diagnostic information about the system, not about students.

The proxy became a target

  • When a measure becomes a target, it ceases to be a good measure (Goodhart, 1984)
  • We confused extraneous load ("What is the submission date?") and germane load ("What does good look like?") (Sweller, 1988)
  • Providing model answers, micro-criteria rubrics, word counts per section, and detailed assessment briefs, removes cognitive struggle
  • "In the varied topography of professional practice... there are those who choose the swampy lowlands" (Schön, 1983)

Rational students should focus on the artifact production.

The proxy has failed

  • A large majority of students report using AI for submitted and assessed work (HEPI, 2026; Eaton et al., 2024)
  • A student can now produce a polished, well-referenced output without reading, engaging, or understanding (Dawson, Bearman & Dollinger, 2024; Corbin, Dawson & Liu, 2025)
  • Fluency now has low information value; it is noise in the system
  • By optimising the system to value the artifact, we've created a structural problem that incentivises misaligned behaviour

The conditions that made the proxy work no longer hold.

The link between the work of becoming a nurse and the evidence of that becoming is permanently broken.

Two responses so far

  • Discursive responses: updated policies, AI declarations, and traffic light systems are attempts to change the language around assessment
    • Defensive positions that leave the foundational assumptions intact and maintain the status quo (Corbin, Dawson & Liu, 2025)
    • The logic of artifact-protection leads to arms-race dynamics
  • Structural responses: change the underlying conditions that determine what students think and do

Discursive changes cannot address structural problems.

A more accurate AI detector won't help

  • A perfect detector tells you whether AI produced the artifact, not whether the student was cognitively engaged
    • Outsource your admin to AI (probably fine)
    • Outsource your thinking to AI (less fine)
  • Those are different questions, and detection answers the wrong one
  • AI can generate either avoidance or genuine struggle; the tool is the same, the output may even be the same, but the formation is not

"Did they use AI?" -> "What did they use AI for?"

What do educators contribute?

  • Stateless: no memory across sessions, no accumulated understanding of this patient, this ward, this moment
  • Static: knowledge frozen at training; it cannot update from the unfolding encounter (see "continuous learning")
  • No context: no embodied experience of practice, no stakes, no accountability
  • Lacks the sense of meaning that comes from being present, responsible, and changed by what happens

"What can only humans do?" -> "What can humans contribute?"

AI agents raise the ceiling

  • Agents take actions: search, synthesise, draft, critique, iterate; they do real work in collaboration with the student, not just in response to them
  • Directing an agent well creates a scaffold that surfaces harder problems, challenges reasoning with more depth, and raises the engage-able ceiling
  • The question is not whether students will have agents (they already do) but whether we are designing experiences that support effective learning
  • Access is unevenly distributed: assuming every student arrives with capable agents builds existing advantage into the design — the equity question is structural, not a footnote

Sending someone to the gym to do your workouts is absurd.

Conditions supporting cognitive struggle

  • Problem-driven inquiry: creates situations of genuine uncertainty from the beginning (no clean solutions)
  • Collaborative construction: makes engagement visible and challengeable (defend your position)
  • Facilitation over instruction: holds the process without providing the answers (struggle stays with the student)
  • Metacognitive reflection: makes developing judgement visible to the student and to the assessor

All of these conditions require students to make a choice.

What is "the work"?

  • "The work" is not what you produce; it is what happens to you while you engage with the struggle that builds new understanding
  • The student who is engaged — directing, evaluating, being challenged, being changed — is doing the work
  • Our task is not to prevent AI from doing the work; it is to design conditions where students choose not to outsource the struggle

The misalignment was always there. AI has simply made it impossible to keep optimising the proxy.

A deliberately light opening, but it carries the whole argument in miniature: the system we have built is not the reason learning happens; quite often, learning happens in spite of it. Hold that thought — by the end I want to have explained why that is, and what AI has to do with it.

With the frame in place, look at what current systems can do. Generative AI passes licensing examinations at the 90th percentile. It writes care plans that are structurally and clinically coherent. It produces reflective portfolios that meet standard assessment criteria. It gives feedback on clinical reasoning at any hour. None of this is speculative; these are current capabilities, available to any student with an internet connection. The usual response is to list what AI cannot do — it cannot be present in the clinical encounter, cannot build a therapeutic relationship. All true. But it sidesteps the point: those are not the things our assessment instruments measure. What our instruments measure, AI can now produce. In alignment terms: the proxy is now trivially satisfiable without the thing it was a proxy for.

The responses to that disorientation cluster into four types. All understandable; all insufficient. Denial: AI cannot replicate the essentially human. Probably true in places — but the problem is not that AI replicates clinical judgement, it's that AI replicates the artifacts we use as evidence of clinical judgement. Retreat: rebuild assessment around what humans uniquely offer. The trouble is the boundary keeps shrinking. "At least AI can't write genuine reflections." Then it can. Defending a shrinking perimeter is unwinnable. Restriction: prohibit and detect. The data consistently shows this doesn't work. High AI-use figures are not evidence that students are dishonest — they are a rational response to what the system rewards. Resignation: cede ground to AI. Unsatisfying, and it abandons the question. Underneath all four is the question we keep avoiding: what are we actually trying to develop in nursing students — and are we designing for that?

These four premises describe the project of nursing education, and they hold regardless of whether AI is in the room. Premise 1 is neurological and constructivist: no one can learn on behalf of the student. Formation requires the student to do the cognitive work. Premise 2 contains a challenge. If learning can happen without a teacher — and it can — then the teacher's presence is not automatically valuable. Teaching must add something otherwise inaccessible. AI is now moving into this space too, which makes the question of the teacher's unique contribution more pressing, not less. Premise 3 is the epistemological problem: we cannot observe learning directly. We observe behaviour and infer understanding. Premise 4 is the regulatory claim: registration and certification are forward-looking claims about fitness to practise, based on backward-looking observations. AI did not change any of these. It changed what we were using as evidence for premise 3. We were observing products rather than doing-and-thinking, and AI can produce the products without the thinking. That is the learning alignment problem stated precisely.

The goal is not that students know about nursing. It is that they become nurses. The claim that the system is designed around artifact production rather than formation is contentious, and it's one of the weaker claims I make — but I'm going to defend it across the next few slides. The short version: the thing we say we value (becoming) is precisely the thing we never built the infrastructure to measure, so we measured the proxy instead.

The proxy was not a failure of intelligence. It was a rational solution to a real problem. We cannot observe learning directly; we observe behaviour and infer understanding. Artifacts were meant to be a window onto the engagement that produced them. To write a good essay, a student had to read widely, organise their thinking, and commit to a position. The engagement was bundled into the economics of production, so nobody had to ask what "the work" was — the work and the artifact were reliably connected. And the science of learning is unambiguous: memory is built through effortful retrieval and the struggle to make sense of something that doesn't yet make sense — Bjork's "desirable difficulties". Conditions that feel easier — re-reading, worked examples, fluency — tend to produce the illusion of learning, not the thing. Hold onto "fluency", because it returns.

Take a moment to think about your own professional development. Not what you studied — what changed you. There will be a moment — probably more than one — where your clinical understanding genuinely shifted. A patient you didn't know how to help. A decision you got wrong and had to sit with. A colleague who challenged your reasoning in a way you couldn't dismiss. Those moments were probably uncomfortable, and almost certainly not assessable. Now picture a formal assessed task: a submission brief, criteria, a word count, a deadline. The texture is completely different. That gap is not evidence that educators have failed. It is information about what we have been asking assessment to do — to stand in for experiences that are genuinely hard to structure, observe, and measure. The model for what we are actually trying to produce has always been available. We just never built the system around it.

AI has not broken assessment. It has made it impossible to ignore what was already broken. The pass mark, inter-rater variability, grade aggregation across qualitatively different tasks — known, documented flaws, accepted as manageable because the alternative seemed worse. Students were already optimising for grades over learning, already producing artifacts that met criteria without reliably doing the intellectual work those criteria were meant to evidence. AI did not introduce that behaviour; it made it more efficient and more visible. And the line I most want to land here: this is diagnostic information about the system, not about students. Every time we talk about students as though they are all cheating, we cause harm and we misread the situation. "Cheating" is not a fixed standard — it is relative to expectations we wrote for a world where producing the artifact required the engagement. When the model changes, what we ask students to do themselves can change with it.

This is the learning alignment problem stated as a law. Goodhart: when a measure becomes a target, it ceases to be a good measure. It is the exact same failure mode AI safety calls specification gaming — optimise the proxy, lose the goal. Over time we stopped treating the artifact as a proxy and started treating it as the thing itself. The drift is visible everywhere. Students spend most of their pre-submission questions on word count, spacing, referencing — and we read this as a communication problem ("be clearer!") when it is really a signal about what we taught them to care about. Cognitive science distinguishes extraneous load (navigation friction — remove it, good pedagogy) from germane load (the effort of working out what matters and what good looks like — that effort IS the learning). When we specified word counts per section, released model answers, and broke rubrics into micro-criteria, we reduced germane load. We thought we were helping. We were removing the mechanism — and handing AI a complete specification it can now follow with no student learning involved. We perfected the map, and the territory got bypassed because the map was so accurate a machine could follow it.

Generative AI has not created a cheating problem. It has severed the inferential chain between the document and the person. A student can produce a polished, well-referenced, criterion-satisfying output without having read a source, grappled with an idea, or developed any understanding. The proxy has collapsed, and the assessment system built on it is exposed as measuring something other than what it claimed to measure. This is not a claim about the prevalence of dishonesty; it is structural. The construct validity the instrument depended on is gone. The correlation between artifact quality and intellectual engagement — the assumption the whole system rested on — no longer holds. Fluency, once a reasonable signal of thinking, is now noise. [NOTE for MR: figures are UK/Canada (HEPI; Eaton). For the international GRNEN audience you might add a one-line caveat that the proportion varies by country but the direction is universal — happy to put that on the slide if you want it explicit.]

Most of what's being produced — policies, frameworks, AI-use declarations — shares one purpose: to restore faith in the artifact as valid evidence; to reconnect the broken link. Dawson and colleagues distinguish discursive responses (change the language: declarations, traffic-light frameworks specifying what students may delegate) from structural responses (change the underlying conditions that determine what students do and why). Most institutional effort is discursive. It represents real, thoughtful work — but it leaves the foundational assumption intact: the artifact is still what's being protected, with tighter lines drawn around it. If the goal is to preserve the artifact, and AI keeps improving at producing artifacts, the trajectory escalates: more restriction, more detection, more circumvention, more enforcement. The dominant question becomes "did you use AI for this?", and the relationship turns adversarial. At the end of that road is an arms race no one wins — and the erosion of the very trust that nursing education, a discipline premised on human care, depends on. In alignment terms, this is trying to patch a misaligned proxy instead of changing the objective.

This reframes the question. The worry that students are "using AI to do the work" is right if the work is the artifact, and wrong if the work is the cognitive engagement. AI can generate the conditions for genuine struggle — harder problems, faster feedback, access to complexity beyond a student's unaided reach. The artifact may look similar; the formation is not. A single use of AI does not contaminate an assignment, and there are countless ways of thinking while using it. So the question shifts from "did you use AI?" — about tool use — to "what did you use it for, and were you genuinely grappling?" — about process. Harder to answer, but the right question, and the one that changes what assessment needs to be.

What AI lacks is not capability — it is the conditions under which capability becomes meaningful. AI is stateless: no memory of prior encounters, no accumulated understanding of how a patient's condition evolved. It is static: knowledge frozen at training. Most importantly, it has no professional context — it doesn't know what it cost to break bad news that morning, or what clinical intuition says when a patient's story doesn't cohere. So the question shifts, for educators as well as students: not what only humans can do, but what humans contribute within a system where AI is also participating. The answer is context — and the educator is the one who holds it. The educator who knows this student's trajectory, this cohort's dynamics, this profession's pressures, has something AI cannot reproduce. That is the contribution.

The shift from AI-as-tool to AI-as-agent matters. A chatbot answers; an agent acts — it searches, synthesises, critiques, iterates, and persists across a student's learning. Students already using agents are working with a collaborator that has context, memory, and the capacity to do substantial cognitive work. This will accelerate, and the question is not how to prevent it but how to design for it well. I want to be honest about the equity dimension, because this audience spans very different contexts — across the US, Canada, and Africa, access to capable agents, reliable connectivity, and the time to learn to direct them is profoundly uneven. A design that silently assumes every student arrives equipped builds existing advantage into the foundations. That is not a reason to refuse the model — inequity predates AI and runs through supervision, library access, and time to think — but it is a reason to design deliberately for it rather than pretend it away. The open stance toward agents has to be matched by an honest one about access.

Nobody would pay someone to go to the gym on their behalf. The reason that's obviously absurd is that we understand the point of the gym is the change it produces in you — not the certificate of attendance. We have somehow lost that clarity about learning. The artifact is the certificate of attendance. The workout is the work.

If formation is the development of practical wisdom, what structures make it possible? Problem-based learning — which nursing and medical education know well — was designed around the answer before AI arrived. Its structural features are the same conditions under which AI integration becomes productive rather than substitutive. The alignment is structural, not retrospective. Problem-driven inquiry puts students in the swampy lowland from the start. Collaborative construction makes engagement visible and challengeable. Facilitation holds the process without handing over answers. Metacognitive reflection makes developing judgement visible — to the student and to the assessor. And these conditions realign the system: they make the proxy hard to game, because what's assessed is the engagement itself, not a detachable artifact. AI then raises the ceiling — giving access to wicked problems previously beyond reach — rather than removing the struggle.

So the questions are not "how do we update our AI policy?" or "how do we detect more reliably?" The work was always the cognitive struggle. A good essay was work because of the grappling, not the document. Clinical placement was work because of the situations navigated, not the competency log completed. The artifact was always downstream; we confused it for the thing. Formation is not content accumulation — it is transformation. The student who has genuinely grappled, been wrong, and revised their thinking is a different person; the artifact was a trace of that, never the process itself. And this brings us back to where we started. This was an alignment problem all along. We optimised a proxy — the artifact — because we could not specify or observe the thing we actually valued. AI didn't break that system; it removed the friction that let us pretend the proxy was the goal. The work — the becoming — is exactly what we were never able to write down. Our job now is to design the conditions where that work cannot be outsourced, and to stop mistaking the evidence for the thing it was always only standing in for.