The last newsletter was almost entirely a list of what I’d posted on the site, partly because I wanted to highlight what kinds of things you might find there, and partly because I wasn’t sure what shape the updated newsletter should take. Moving forward it will mostly be things I’ve found useful and think you might too: a couple of reads, some quotes, a debate worth following, and a single piece of my own, with a link to everything else.

Most of my time this month was spent working on one idea: the research harness — a way of setting out explicitly what an AI agent is for in a piece of doctoral work, and how it’s allowed to operate. What I didn’t expect is how much the idea has started to change my own practice. Over the past few weeks I’ve begun building an agentic harness around my own scholarship and learning: a working setup that defines what I ‘should’ be spending my time on, what I let AI do, what stays my decision, and how it all gets recorded. It’s still mostly scaffolding, but it’s already changing how I choose what to work on. I’m interested to know if anyone reading here has tried to make their own working rules with AI explicit, rather than leaving them ad hoc? If so, I’d love to know what that looks like for you.

One thing from me

If you look at one thing I’ve posted this month, I’d suggest the one-page guide to the research harness for doctoral researchers and supervisors. It condenses the framework (which started out as an essay) into a reference card, including the seven components of the harness, what each one does, and how to start building one with a student without bolting on a layer of process. If you supervise doctoral work, it’s meant to be usable in a single conversation.

Everything else from May is collected on the new Recently added page.

Worth reading

Two pieces from this month, both about AI and learning:

  • Keren, West & Desai: Promoting clinical expertise in the age of AI: no struggle, no mastery (JAMA). An argument that leaning on AI for the cognitively demanding parts of clinical work removes the productive struggle that expertise is built from.
  • Dickinson & Marshall: Trained to stop learning (Wonkhe). A large UK survey (1,055 students) with a finding educators can actually act on: it is assessment design, not AI use itself, that decides whether AI ends up supporting or replacing learning.

Quotes that resonated with me

Students in this research are not passively working through confusing policy. Many are actively constructing personal ethical frameworks – theories of practice about the relationship between tools, effort, learning, and identity – that are often more considered than the institutional guidance they receive. They are doing this work largely alone, with little support and no recognition.

Dickinson & Marshall, Trained to stop learning

When production is cheap, opportunity cost becomes the real cost. You can’t build everything, and whatever you pick comes at the cost of everything else.

Maggie Appleton, Collaborative AI Engineering

People seek Claude’s guidance across many different areas of their life, but over three-quarters of conversations (76%) were concentrated in just four domains: health and wellness (27%), professional and career (26%), relationships (12%), and personal finance (11%)

Anthropic, How people ask Claude for personal guidance

A debate on traffic light systems

A disagreement over AI assessment scales that played out on LinkedIn early in the month. I’m quite sceptical of assessment scales and traffic light systems but found both sides worth paying attention to.

Blinded by the (traffic) lights. Mark Bassett argues that assessment scales fail where it matters most: the middle bands let one piece of student work satisfy and violate adjacent levels at once; the question of when an assessment “begins” has no answer; and the AI tools themselves draw no line between what one level permits and the next forbids. His charge is that a label ends up communicating an institutional preference about conditions that institutions can’t control.

On labels, scales, and what the AIAS actually claims to do. Mike Perkins, a lead author of the AI Assessment Scale paper, agrees with Mark’s position more than you might expect; labelling a task and walking away is useless, detection doesn’t work, and the first version of the AIAS leaned too hard on enforcement. But he argues v2 was reframed, and as a result the scale isn’t an enforcement mechanism, but is rather a tool for redesigning assessment and communicating intent. As someone who is working on several frameworks, I was reminded that “a framework develops legitimately when its central claim becomes clearer under criticism.”

Both are worth reading in full. Wherever you land in the debate, it’s nice to see these kinds of disagreements in honest engagement at a time when so much of the space is just noise.

Something for the commute

On the reg podcast. AI-powered research workflows (Inger Mewburn and Jason Downs). A practical conversation about actually building AI into a research practice.