Abstract
The rapid advancement of large language models (LLMs) in tackling quantitative problems presents both challenges and opportunities for science education. Here, we benchmark GPT-4.1 and GPT-5 models on a corpus of 601 undergraduate physics exam problems from the University of Bath over the past five years, marking answers against official numerical solutions. GPT-5 answered 82.9% of valid questions correctly, surpassing GPT-4.1 (65.7%) and showing roughly half the rate of question misinterpretation. Accuracy remains stable from first- to fourth-year problems and across subject areas, indicating robust performance throughout the physics undergraduate curriculum. These results position state-of-the-art LLMs at or above a first-class level, underscoring their transformative potential for assessment design in higher education.
| Original language | English |
|---|---|
| Article number | 033007 |
| Number of pages | 6 |
| Journal | Physics Education |
| Volume | 61 |
| Issue number | 3 |
| Early online date | 1 Jun 2026 |
| DOIs |
|
| Publication status | Published - 1 Jun 2026 |
Data Availability Statement
All data that support the findings of this study areincluded within the article (and any supplementary files).
Supplementary Material available at: https://
doi.org/10.1088/1361-6552/ae68b3/data1.
Funding
This work was financially supported by the University of Bath’s Teaching Development Fund ‘Evaluating AI-Driven Learning in Physics Education.’
| Funders |
|---|
| University of Bath |
Keywords
- Artificial Intelligence
- ChatGPT
- physics
ASJC Scopus subject areas
- Education
- General Physics and Astronomy
Fingerprint
Dive into the research topics of 'First-class AI? Assessing LLMs on undergraduate physics exams'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS