1
Return

Self-referential consistency in stateless language models: a behavioral perspective

delete2026-08-12
delete0
PRE
AI
J
José Augusto de Lima Prestes *
DOI:10.1007/s00146-026-03292-3delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Large language models (LLMs) increasingly generate outputs that resemble introspection, including self-reference, epistemic modulation, and claims about their internal states. This study investigates whether such behaviors reflect stable underlying patterns or merely surface-level generative artifacts. We evaluated five open-weight, stateless LLMs using a structured battery of 21 introspective prompts. The main corpus comprised 1050 completions collected under a baseline decoding condition (temperature = 0.7), supplemented by 2100 additional completions generated under matched temperature conditions (temperature = 0.2 and 1.0), for a total of 3150 completions. Outputs were analyzed across four behavioral dimensions: surface-level similarity (token overlap via SequenceMatcher), semantic coherence (Sentence-BERT embeddings), inferential consistency (Natural Language Inference with a RoBERTa-large model), and diachronic continuity (stability across prompt repetitions). Construct validity was further examined through a human-evaluation layer in which 10 annotators rated 80 selected response pairs drawn from the same prompt battery on a 5-point consistency scale. Inter-rater agreement was moderate by Krippendorff’s $$\alpha$$ for ordinal ratings ( $$\alpha = 0.553$$ ), while reliability was moderate at the single-rater level and strong for aggregated ratings by intraclass correlation (ICC(2,1) = 0.564; ICC(2,k) = 0.928). The human-evaluation layer showed that lexical overlap and embedding-based semantic similarity were weak proxies for perceived self-referential consistency, whereas NLI-based indicators tracked mean human ratings much more closely. Across the matched temperature conditions, lower temperature generally increased semantic and diachronic stability, whereas higher temperature tended to increase drift and reduce coherence, though the pattern was not perfectly monotonic across all models or metrics. We therefore interpret apparent self-referential stability in the open-weight stateless LLMs evaluated here as conditional and fragile rather than robustly stable across generation regimes. Following recent behavioral frameworks, we heuristically adopt the term pseudo-consciousness to describe structured yet non-experiential self-referential output in LLMs. This use reflects a functionalist stance that avoids ontological commitments, focusing instead on behavioral regularities interpretable through Dennett’s intentional stance. This study contributes a reproducible behavioral framework, complemented by human validation and a matched decoding-temperature sensitivity analysis, for evaluating simulated introspection in LLMs. Our findings carry implications for interpretability, alignment, and user perception, highlighting the need for caution when attributing mental states to stateless generative systems based on linguistic fluency alone.
Keywords:
Large language models
Introspective simulation
Pseudo-consciousness
Self-reference
Behavioral evaluation
AI alignment

Journal

A
AI & Society
IF:
4.7
Papers:
405
Citations:
0

Organization

No organization information available
Cited Papers

Cited Papers

Citing Papers

Citing Papers