|
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
|
| Volume 187 - Issue 126 |
| Published: July 2026 |
| Authors: Keval Barvaliya |
10.5120/ijca4f210741271e
|
Keval Barvaliya . The Uncharted Interior: Empirical Evidence for Spontaneous Latent World Model Formation in Large Language Models. International Journal of Computer Applications. 187, 126 (July 2026), 26-38. DOI=10.5120/ijca4f210741271e
@article{ 10.5120/ijca4f210741271e,
author = { Keval Barvaliya },
title = { The Uncharted Interior: Empirical Evidence for Spontaneous Latent World Model Formation in Large Language Models },
journal = { International Journal of Computer Applications },
year = { 2026 },
volume = { 187 },
number = { 126 },
pages = { 26-38 },
doi = { 10.5120/ijca4f210741271e },
publisher = { Foundation of Computer Science (FCS), NY, USA }
}
%0 Journal Article
%D 2026
%A Keval Barvaliya
%T The Uncharted Interior: Empirical Evidence for Spontaneous Latent World Model Formation in Large Language Models%T
%J International Journal of Computer Applications
%V 187
%N 126
%P 26-38
%R 10.5120/ijca4f210741271e
%I Foundation of Computer Science (FCS), NY, USA
Large Language Models (LLMs) have advanced reasoning, adaptation, and generative abilities despite their limited training on next-token prediction, which is the only way they are trained in world simulation. There are recent developments in transformer architectures and representation learning that indicate that such systems can learn structured internal representations, but it is not clear how much LLMs learn a coherent latent world model. Much existing work has focused on benchmark performance and emergent behaviors, and less on the internal representational dynamics involved in inference. This study proposes the Latent World Model Emergence (LWME) framework to explore the emergence of latent world representations in transformer-based LLMs. The framework is based on linear probing, activation patching, geometric manifold analysis, and causal ablation to examine representational organization for various knowledge domains. The study proposes a World Model Fidelity Score (WMFS) that is multidimensional and measures structural coherence, causal consistency, and inferential reliability to quantify the quality of representation. Results from experiments show that latent world-model representations are consistently observed for different transformer architectures and become more powerful when the parameter scales are larger and the training data is more varied. The causal ablation experiments demonstrate that ablation of mid-layer attention circuits is essential for inferential coherence and that the targeted ablations have a significant impact on reasoning performance. Additionally, there is also a problem that the accuracy of the benchmarks is not aligned with the actual reasoning abilities of the WMFS, which means that traditional assessment of actual reasoning ability may be misleading and merge memorization with structured inference. The study offers a systematic interpretability framework and empirical evidence of the ability of advanced LLMs to form organized internal representations that have implications for the fields of AI interpretability, alignment, and safety.