Return
Java academic benchmark: Exam-based evaluation of LLMs on object-oriented programming
DOI:10.1016/j.jss.2026.113033.png)
Abstract
En 中文
Current code generation benchmarks largely overlook object-oriented programming (OOP) skills, leaving open whether large language models (LLMs) can effectively apply OOP principles to structured programming tasks. The dominance and permissiveness of Python in most benchmarks further obscure model weaknesses in core OOP principles. We introduce the Java Academic Benchmark (JAB), based on 103 authentic Java exams collected over a decade at a major university, together with 506 expert-written JUnit tests, to rigorously evaluate advanced Java programming competence with a strong focus on OOP. To complement execution-based scoring, we propose KODE, an LLM-as-a-Judge framework that assesses OOP adherence across four pedagogical dimensions. We evaluate 27 LLMs under two resolution strategies: single-attempt (one-shot) and agentic (iterative refinement with compiler and test feedback). Results reveal that, under our evaluation protocol, larger closed models match or surpass bachelor-level OOP students-especially under agentic resolution-while smaller open models lag behind. JAB's class-level design enables fine-grained error analysis, exposing recurring misconceptions. Editor's note: Open Science material was validated by the Journal of Systems and Software Open Science Board.
Keywords:
Large language models
Java
Object-oriented programming
Coding benchmark
Journal
IF:
4.1
Papers:
5.4K
Citations:
8.4K
Organization
Cited Papers
No cited papers available

