
Opinion
The sector’s response to generative AI has been, overwhelmingly, a policing operation. We’ve bought detection tools, redesigned tasks to be “AI-proof,” rewritten integrity policies, and spent enormous energy on a single question: did the student use AI? It’s the wrong question, and our fixation on it is letting a much larger problem walk past unexamined.
In my field, finance, the prized skill for decades was the ability to produce — build the model, write the commentary, assemble the analysis. That skill is being commoditised in front of us. An AI tool will now build a three-statement model from a folder of filings in minutes. What it can’t reliably do is know where that model is wrong. The skill that survives, and grows more valuable, is review: directing the work, interrogating it, and owning the result.
So the uncomfortable question for anyone who designs assessment isn’t whether students used AI. It’s whether we’re still assessing the skill the machine has made cheap — production — instead of the one that now matters — judgement and verification. Most of our assessment still rewards production. We ask students to generate the essay, the model, the report, and we grade the artefact. AI generates a competent artefact on demand. We are, increasingly, certifying a skill the labour market has stopped paying for.
The fix isn’t detection; it’s changing what we put a grade on. If production is the part the machine now does, production can’t remain the thing we certify. What deserves assessment is the judgement that sits on top of it: whether a student can take an analysis — their own or a machine’s — find where it breaks, name the assumption holding it up, and decide whether to trust the result. That is verification, and it is a discipline in its own right, not a by-product of being good at producing.
And here is what makes it urgent: the better these tools get, the worse we become at checking them. When output is usually right and always polished, vigilance decays and we wave it through. A weak model keeps you sharp; a strong one lulls you. So the more capable the AI, the more deliberately verification has to be taught — the exact opposite of the comforting assumption that better tools need less human skill.
In practice, the high-value task in my own classes is no longer “produce the analysis.” It is: here is an AI-generated analysis — find what’s wrong, show how you checked, defend your corrections. That is verification taught deliberately rather than absorbed by accident, and it’s far harder to fake than a polished essay, because it demands the judgement the polish can’t supply.
I’d be selling the sector something dishonest if I stopped there, because this has a real limit. You cannot review what you’ve never been able to produce. The instinct that a depreciation schedule looks off is compiled from years of having built them. If we automate production out of the curriculum entirely, we don’t graduate reviewers; we graduate people who rubber-stamp confident-looking output. Verification sits on top of production competence — it doesn’t replace it. The implication is awkward: we may still need students to learn the production work we’re automating, precisely so they can judge it — like a pilot trained on manual flight before trusting the autopilot.
None of this is solved by a better detector. It’s solved by a curriculum conversation we’ve been avoiding while we argue about cheating. The integrity panic treats AI as a threat to academic standards. The deeper risk is quieter: that we keep certifying graduates in a skill that’s being automated, while never deliberately teaching the one that’s becoming the whole job. We have built, by accident, an apprenticeship that assumes production comes first and judgement comes later. That sequence is now inverted, and our assessment hasn’t caught up.
Stop asking whether they used the tool. Start asking whether we’re still testing the right thing.
Pedram Nourani — Assistant Professor of Finance, S P Jain School of Global Management, Sydney