
According to a report published today (the 24th) by Phys.org, a new study by Cardiff University and the University of Melbourne has found that generative artificial intelligence (GenAI) is not yet able to assess students' written work as reliably as human teachers.
Researchers used two versions of ChatGPT to grade 50 undergraduate biology papers, testing the large language models' ability to evaluate student work. ChatGPT was required to grade the papers according to 7 criteria. The researchers also used 4 different prompting methods, then compared the models' average scores and score variation with those of human graders.

Dr. William Kay of Cardiff University's School of Biosciences said: “The results show that the scores given by the models varied considerably and could not accurately predict human grading. There were significant differences when generative AI and humans graded the same set of papers. Looking only at the total paper scores, the results were relatively close; once we looked at each scoring criterion and each individual paper, the gap between the models and human graders became quite substantial.”
Except in one case, AI models generally gave higher average scores than human graders. The largest difference between AI and human average scores was 16.1 points; for an individual paper, the largest difference reached 40 points.
William Kay also found that AI tended to lower the scores of high-scoring papers while raising the scores of low-scoring assignments, systematically concentrating scores toward the middle.
The results show that, at least for now, even after extensive training, AI cannot reliably assign scores to written work with a high degree of subjectivity that are comparable to those given by humans. “The large language models we tested are currently not suitable for replacing teachers or predicting students' assignment grades.”
Researchers noted that higher education is indeed interested in whether large language models can use pattern-recognition capabilities to make assignment grading more objective while improving grading efficiency and reducing pressure on teaching staff. However, this study shows that they are not currently suitable for this purpose.
In addition, submitting assignments to AI tools without students' explicit consent raises ethical issues in itself; large language models cannot and should not be relied upon to grade students' lengthy written assignments. “As large language models continue to develop, their ability to imitate human judgment may improve in the future. But based on this study, achieving genuine consistency between AI-generated scores and human grading will likely be difficult.”
