A fluent draft no longer tells you whether the thinking happened
Supervision has always run on an inference: if the writing is the student’s, the work behind it probably is too. That doesn’t hold any more, and detection tools don’t repair it, because they’re still asking about the document rather than the person. What’s left is asking students to account for the decisions in their work, which only really happens in conversation.
Most of what gets written about AI and doctoral supervision is addressed to supervisors but I think it’s worth starting somewhere else, when we consider what our students are dealing with at the moment, several times a day and mostly on their own.
Consider the following. A doctoral student has a chapter due in a week and they have a tool that will produce a competent version of it in the time it takes to make coffee, and it’ll be better than what they’d produce on their own. Nobody has given them a usable version of when accepting that offer costs them something, and the guidance they do have mostly stops at “use AI to support your work, not replace it”. The problem is that this sounds right but doesn’t help much at the decision point, because it never says where “support” ends.
The draft you receive is fine. The argument holds together, the sources are real, and the prose is cleaner than last time. Then you ask the student something small and specific, why this author’s definition of the construct rather than the other one, and they don’t have a response. Nothing was faked, the reading may well have happened, but something supervisors have relied on for years has stopped working, and I think it’s worth being precise about what that is.
This is the central concern that Benita Olivier and I spent most of the last few months thinking about while we wrapped our book, Still Yours: A Doctoral Researcher’s Guide to AI, which Springer Nature expects to publish in November.
The proxy doctoral supervision runs on
You can’t observe someone thinking, so you read what they produce and reason backwards: this synthesis is coherent, so the person who wrote it probably understood the papers. That was never a perfect inference but it was cheap and mostly right, because producing fluent academic prose used to require having done the reading. Sarah Elaine Eaton calls where we’ve arrived postplagiarism; the values underneath academic integrity haven’t changed, but the test we’ve been using has stopped tracking them.
The instinct is to repair the inference by policing it, with detection tools and declarations and tighter rules about what can be used, where, and when. Let’s set aside how unreliable AI detectors are, because the deeper problem isn’t that they don’t work but that they’re aimed at the wrong problem. If a detector comes back clean then you’ve only learned that the text was probably typed by a human, but you still don’t know whether the student can explain their choices. If it flags the chapter, you’re now having a conversation about the tool rather than the work. Either way you’re back where you started, reading a document and making guesses about the person.
What you’re actually calibrating
This is the part I think is the core of the position. Doctoral supervision is a calibration exercise: over months you build a working model of where this researcher is, what they can do unaided, where their judgement is thin, which struggles are productive and which are just stuck, and you tune what you give back to them accordingly, challenge for one and scaffolding for another. Almost all that engagement is informed by the evidence they hand you.
So when the drafts have been smoothed by AI beyond what the student can actually defend, the calibration fails. Your guidance is then correctly aimed at the researcher on the page but lands on a different person sitting in front of you, and both of you can take months to notice, because everything looks fine. That seems to me a more serious problem than the integrity question, and I haven’t yet seen a policy that addresses it.
Ask them to defend it
A replacement for “did you write this?” is “can you defend the decisions that shaped this?” In practice it’s unremarkable; you pick a decision and ask why it went that way rather than the obvious alternative. Why this inclusion criterion. Why this author’s definition. What would have to be true for the opposite finding. You’re not testing recall, you’re testing whether there’s reasoning under the sentence, and reasoning can’t be produced after the fact by someone who doesn’t have it. Bearman, et al. call this evaluative judgement, and it’s a capacity that gets more valuable as producing plausible work with AI gets cheaper.
There’s a corollary that’s harder to align with the existing system. If a piece of assessed work can be completed in full by a model, an honest response is to look carefully at the artefact before looking hard at the student. Some of what we set was only ever a proxy that happened to be convenient to mark, which is the argument I made about the thesis as a proxy for the person, and a standalone literature summary was worth setting when producing one meant reading fifteen papers. A protocol defence, a design critique of a paper in the student’s own field, a viva-style exchange about why a method was chosen; these ask for the reasoning directly and they hold up much better.
Which brings me back to the question I can’t seem to move past. What are we actually adding once the machine does more of the work, and does it better than we do? If the answer is that we check its output, that won’t survive many more model releases, for the reasons I describe in the verification trap. If it’s that we hold the judgement that directs the work, then judgement is what supervision is for, and it gets built by asking people to account for what they decided.
That’s the through-line of the book we wrote, and I’d be interested to know how you’re handling it with your own students.