Return
From Code Variability to Theme Convergence: AI–Human Alignment in Thematic Analysis With Claude Code
DOI:10.1177/16094069261462093.png)
Abstract
En 中文
Purpose
This study examines AI–human alignment in thematic analysis from a multi-level perspective, asking whether agreement at the level of individual codes is necessary for convergence in higher-order themes. While prior research often evaluates overall agreement, this study distinguishes between code-level variability and theme-level stability to provide a more nuanced assessment of AI-assisted qualitative analysis.
Methods
Using qualitative interview data from a doctoral study on e-commerce, the original human thematic analysis (2012, NVivo) was compared with four independent AI-assisted analyses conducted in 2026 using Claude Code. Two prompting strategies were tested: general prompts and structured multi-phase prompts. Alignment was assessed at both code and theme levels using the F1 Score, while inter-session consistency was evaluated through bidirectional mapping across session pairs.
Findings
At the code level, AI outputs showed considerable variability, with F1 Scores ranging from 55.9% to 87.6% and clear differences between prompting approaches (structured: 83.1% average vs. general: 59.2%). In contrast, theme-level alignment remained consistently high across all sessions (90.9%–100%), with strong inter-session consistency (average 86.1%). These findings indicate that although AI-generated codes may differ across sessions and from human coding, the resulting thematic structures converge reliably.
Originality
This study introduces a multi-level alignment framework and provides one of the first empirical evaluations of Claude Code as an agentic AI tool for thematic analysis. The 14-year gap between the original human analysis and AI reanalysis offers a distinctive test of AI engagement with historical qualitative data. The study identifies a pattern of hierarchical convergence, where code-level divergence coexists with theme-level stability.
Implications
The findings suggest that strict code-level agreement may not be necessary for reliable thematic conclusions. AI-assisted analysis can support theme development when structured prompting and human oversight are maintained, offering methodological guidance for integrating AI into qualitative research while preserving analytical rigor.
Journal
I
IF:
3.8
Papers:
2.2K
Citations:
1.2W

