Our research shows how AI is getting better at exams. Here’s how unis can respond
When ChatGPT arrived at the end of 2022, universities scrambled to determine whether generative AI was good enough to pass assessments and make it easy for students to cheat.
There were some big claims by AI companies at the time, including that tools showed “human-level performance on various professional and academic benchmarks”. Headlines boasted about AI passing US bar exams.
But early evidence suggested generative AI did not always live up to the hype. Our previous research showed models performed poorly on tasks requiring the essential skill of critical legal analysis. In other words, the AI models would really struggle to compete with a real law student in legal analysis.
But as AI models get more sophisticated, we wanted to know whether that was still true.
Our research
Our 2023 experiment tested generative AI on a real Australian criminal law exam. It found) “generative AI isn’t close to replacing humans in intellectually demanding tasks such as [an undergraduate] law exam”. Our new study revisited that experiment.
This time, we tested nine models from five AI providers across two compulsory law subjects at the University of Wollongong. Most of the AI tools are easily accessible via subscription.
Once the students’ criminal law and tort law (which deals with compensation) final exam questions had been prepared, we used each model to generate an answer to the same exam papers, producing 18 answers in total.
The models received the exam questions and prompts, but no subject-specific lecture notes, textbooks or curated legal materials. They did have internet access, and some had an “enhanced reasoning” feature. This gives the AI more time to plan, consider different approaches and check its reasoning before answering.
Six of the 18 papers were marked by subject coordinators (the authors here), and the remaining 12 were mixed among genuine student papers and blind-marked by tutors who did not know anything about the involvement of AI models.
What we found
Compared to 2023, the improvement was significant.
In criminal law, AI papers averaged 76.3%, outperforming 82.5% of students. In torts, they averaged 66%, outperforming 61% of students. Seven of the 18 AI papers ranked at or above the 90th percentile of students.
This is a shift from the 2023 results on criminal law, when AI papers averaged 52.5% and performed around the 22nd percentile.
More important than the grades was the models’ performance on hypothetical legal scenarios, which were one component of the overall assessment.
In 2023, AI models demonstrated a clear weakness when it came to critical analysis of these scenarios.
In our new study, that had largely disappeared. In criminal law, AI performed much better, with an average of 76% compared to 59.6% for students. In torts, AI models and students were at roughly the same level, while AI models performed dramatically better on the essay question.
What does this mean?
Overall, the study suggests AI models have become capable enough that students could use them to cheat and calls into question the integrity of results from non-supervised assessments.
This does not mean, however, that AI has become a reliable legal expert.
Performance in our study varied substantially between models and subjects. Indeed, we found AI models to have jagged capabilities. Some models were excellent in one subject and weak in the other.
Strong analysis could sit alongside poor citations, weak source selection or fabricated authorities. For some models we identified lower rates of hallucination, however we realised they also cited fewer sources.
What this study cannot tell us
We only tested two law subjects at one Australian university, using 18 AI-generated papers. Results may differ in other law subjects (for example, contracts) or other disciplines.
Prompting also matters. We tried to standardise prompts, but we had to make some changes to keep answers realistic in length.
Further, the study did not assess whether the models’ performance would remain consistent across different prompts, versions or assessment formats.
More importantly, despite all the guidelines provided to tutors, marking exam papers is impacted by subjective judgements.
How can universities respond?
Today’s job market requires graduates to work with AI responsibly and effectively. But students still need enough foundational knowledge and skills to detect errors, sharpen weak points and improve its output. If they cannot add value beyond what AI can produce alone, their work may be replaceable.
No single assessment format can achieve both aims. We propose three complementary approaches.
First, some assessment should remain AI-free and monitored in-person. This means on campus exams, oral assessments or supervised problem-solving. This ensures students have enough independent knowledge and reasoning ability to recognise when AI is wrong. For example, in first-year law, we suggest these tasks should make up the bulk of assessment.
Second, some assessments should deliberately require collaboration with AI. But universities should assess the process, not just the polished final product. This means we should evaluate students’ ability to collaborate with AI models, such as, what they prompted, which sources they checked, what errors they identified, what they accepted or rejected, and how they improved the reasoning.
Third, universities can use a “relay” model that mirrors professional practice. This approach has two versions, and each has two phases.
In a human-first version, students submit a draft in a supervised environment and then collaborate with AI to refine and resubmit it.
In an AI-first version, students collaborate with AI and submit their draft. They must then use their own knowledge and critical analysis in an supervised environment to critique, verify and improve it without access to AI, and resubmit their work.
Markers should grade each submission separately. This evaluates both students’ legal knowledge and skills, and their ability to collaborate with AI.
The goal is not to pretend students can be kept away from AI indefinitely, nor to outsource learning to the technology. Universities need assessments that still show what students can do independently, while teaching them how to supervise, challenge and improve increasingly capable AI systems.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.