arrow
Return

Java academic benchmark: Exam-based evaluation of LLMs on object-oriented programming

delete2026-07-01
delete0
PRE
AI
A
Alessio Cocchieri
L
Luca Ragazzi *
G
Gianluca Aguzzi
G
Giacomo Frisoni
L
Lorenzo Molfetta
G
Gianluca Moro
M
Mirko Viroli
DOI:10.1016/j.jss.2026.113033delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Current code generation benchmarks largely overlook object-oriented programming (OOP) skills, leaving open whether large language models (LLMs) can effectively apply OOP principles to structured programming tasks. The dominance and permissiveness of Python in most benchmarks further obscure model weaknesses in core OOP principles. We introduce the Java Academic Benchmark (JAB), based on 103 authentic Java exams collected over a decade at a major university, together with 506 expert-written JUnit tests, to rigorously evaluate advanced Java programming competence with a strong focus on OOP. To complement execution-based scoring, we propose KODE, an LLM-as-a-Judge framework that assesses OOP adherence across four pedagogical dimensions. We evaluate 27 LLMs under two resolution strategies: single-attempt (one-shot) and agentic (iterative refinement with compiler and test feedback). Results reveal that, under our evaluation protocol, larger closed models match or surpass bachelor-level OOP students-especially under agentic resolution-while smaller open models lag behind. JAB's class-level design enables fine-grained error analysis, exposing recurring misconceptions. Editor's note: Open Science material was validated by the Journal of Systems and Software Open Science Board.
Keywords:
Large language models
Java
Object-oriented programming
Coding benchmark

Journal

Journal of Systems and Software cover
Journal of Systems and Software
IF:
4.1
Papers:
5.4K
Citations:
8.4K

Organization

U
university of bologna
Scholars:
735
Papers: 273
Citations: 0
Cited Papers

Cited Papers

No cited papers available