Skip to main navigation Skip to search Skip to main content

First-class AI? Assessing LLMs on undergraduate physics exams

  • Harvard University

Research output: Contribution to journalComment/debatepeer-review

Abstract

The rapid advancement of large language models (LLMs) in tackling quantitative problems presents both challenges and opportunities for science education. Here, we benchmark GPT-4.1 and GPT-5 models on a corpus of 601 undergraduate physics exam problems from the University of Bath over the past five years, marking answers against official numerical solutions. GPT-5 answered 82.9% of valid questions correctly, surpassing GPT-4.1 (65.7%) and showing roughly half the rate of question misinterpretation. Accuracy remains stable from first- to fourth-year problems and across subject areas, indicating robust performance throughout the physics undergraduate curriculum. These results position state-of-the-art LLMs at or above a first-class level, underscoring their transformative potential for assessment design in higher education.

Original languageEnglish
Article number033007
Number of pages6
JournalPhysics Education
Volume61
Issue number3
Early online date1 Jun 2026
DOIs
Publication statusPublished - 1 Jun 2026

Data Availability Statement

All data that support the findings of this study are
included within the article (and any supplementary files).
Supplementary Material available at: https://
doi.org/10.1088/1361-6552/ae68b3/data1.

Funding

This work was financially supported by the University of Bath’s Teaching Development Fund ‘Evaluating AI-Driven Learning in Physics Education.’

Funders
University of Bath

    Keywords

    • Artificial Intelligence
    • ChatGPT
    • physics

    ASJC Scopus subject areas

    • Education
    • General Physics and Astronomy

    Fingerprint

    Dive into the research topics of 'First-class AI? Assessing LLMs on undergraduate physics exams'. Together they form a unique fingerprint.

    Cite this