Research Article

Quantifying Representational Complexity in Transformer Models via Residual Stream Spectral Entropy

by  Keval Barvaliya
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 126
Published: July 2026
Authors: Keval Barvaliya
10.5120/ijca54dc2360b436
PDF

Keval Barvaliya . Quantifying Representational Complexity in Transformer Models via Residual Stream Spectral Entropy. International Journal of Computer Applications. 187, 126 (July 2026), 9-25. DOI=10.5120/ijca54dc2360b436

                        @article{ 10.5120/ijca54dc2360b436,
                        author  = { Keval Barvaliya },
                        title   = { Quantifying Representational Complexity in Transformer Models via Residual Stream Spectral Entropy },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 126 },
                        pages   = { 9-25 },
                        doi     = { 10.5120/ijca54dc2360b436 },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Keval Barvaliya
                        %T Quantifying Representational Complexity in Transformer Models via Residual Stream Spectral Entropy%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 126
                        %P 9-25
                        %R 10.5120/ijca54dc2360b436
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Despite remarkable advances in the reasoning and generalization capabilities of large language models, the internal representational mechanisms that underlie these behaviors remain poorly understood. Current frameworks of evaluation are mainly based on the assumption of benchmark performance as a proxy for model capability and provide limited information about the organization and transformation of information in the network during inference. One of the main open questions in current interpretability work is the disconnect between the behavioral assessment and the understanding of the mechanisms. In this work, a principled metric, Residual Stream Spectral Entropy (RSSE), is proposed for measuring the representational complexity of transformer-based language models by analyzing their residual stream activations. The RSSE represents the entropy of the singular value distribution of the layer-wise residual representations, and hence a normalized score of the distribution of the representational energy of a model across latent computational subspaces. RSSE works directly on the activation geometry of the internal representations of the task, without relying on labels or probabilities of output. RSSE is measured across a variety of transformer model families with a wide range of scales, architectures, and training goals, using a variety of reasoning and language understanding benchmarks. The results presented in this paper indicate that higher-order reasoning tasks consistently demonstrate wider distributions across the spectrum than do factual retrieval and low-complexity classification tasks and that the RSSE grows monotonically as both increase in scale and complexity of the task. Interestingly, models with similar reasoning accuracy tend to produce similar entropy profiles despite their different architectures, indicating that entropy might be a feature of the computation of transformers in general, and not be solely a function of any individual model choice. These results provide a novel approach for investigating emergent capabilities, representation scaling, and the boundaries of interpretability for large neural systems.

References
  • Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., . . . Amodei, D. (2020). Language Models are Few-Shot Learners. ArXiv. https://arxiv.org/abs/2005.14165
  • Zhou, K., Ji, J., Tianyi, T., Hou, Y., Zhao, W. X., Li, J., … Wen, J.-R. (2023). A survey of large language models. arXiv preprint arXiv:2303.18223. https://arxiv.org/abs/2303.18223
  • Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., … Fedus, W. (2022). Emergent Abilities of Large Language Models. Transactions on Machine Learning Research, 2022-August.
  • Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., … Zhang, Y. (2023). Sparks of artificial general intelligence: Early experiments with GPT-4. arXiv preprint arXiv:2303.12712. https://arxiv.org/abs/2303.12712
  • Kästner, L., & Crook, B. (2024). Explaining AI through mechanistic interpretability. European Journal for Philosophy of Science, 14(4). https://doi.org/10.1007/s13194-024-00614-4
  • Bereska, L., & Gavves, E. (2024). Mechanistic Interpretability for AI Safety: A Review. Transactions on Machine Learning Research, 2024.
  • Nanda, N., Chan, L., Lieberum, T., Smith, J., & Steinhardt, J. (2023). Progress measures for grokking via mechanistic interpretability. In Proceedings of the International Conference on Learning Representations (ICLR 2023). https://arxiv.org/abs/2301.05217
  • Golgoon, A., Filom, K., & Ravi Kannan, A. (2024). Mechanistic interpretability of large language models with applications to the financial services industry. In ICAIF 2024 - 5th ACM International Conference on AI in Finance (pp. 660–668). Association for Computing Machinery, Inc. https://doi.org/10.1145/3677052.3698612
  • Bengio, Y., Courville, A., & Vincent, P. (2013). Representation learning: A review and new perspectives. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(8), 1798–1828. https://doi.org/10.1109/TPAMI.2013.50
  • Geiger, A., Ibeling, D., Zur, A., Chaudhary, M., Chauhan, S., Huang, J., … Icard, T. (2025). Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability. Journal of Machine Learning Research, 26.
  • Sharkey, L., Chughtai, B., Batson, J., Lindsey, J., Wu, J., Bushnaq, L., … McGrath, T. (2025). Open Problems in Mechanistic Interpretability. Transactions on Machine Learning Research, 2025-September, 1–89.
  • Elhady, A., Agirre, E., & Artetxe, M. (2025). Emergent Abilities of Large Language Models under Continued Pre-training for Language Adaptation. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 32174–32186). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.acl-long.1547
  • Ichien, N., Stamenković, D., & Holyoak, K. J. (2024). Large Language Model Displays Emergent Ability to Interpret Novel Literary Metaphors. Metaphor and Symbol, 39(4), 296–309. https://doi.org/10.1080/10926488.2024.2380348
  • Schaeffer, R., Miranda, B., & Koyejo, S. (2023). Are Emergent Abilities of Large Language Models a Mirage? In Advances in Neural Information Processing Systems (Vol. 36). Neural Information Processing Systems Foundation.
  • Lu, S., Bigoulaeva, I., Sachdeva, R., Madabushi, H. T., & Gurevych, I. (2024). Are Emergent Abilities in Large Language Models Just In-Context Learning? In Proceedings of the Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 5098–5139). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2024.acl-long.279
  • Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., … Olah, C. (2021). A mathematical framework for transformer circuits. Transformer Circuits Thread. https://transformer-circuits.pub/2021/framework/index.html
  • Chen, X., Ding, M., Wang, X., Xin, Y., Mo, S., Wang, Y., … Wang, J. (2024). Context Autoencoder for Self-supervised Representation Learning. International Journal of Computer Vision, 132(1), 208–223. https://doi.org/10.1007/s11263-023-01852-4
  • Le-Khac, P. H., Healy, G., & Smeaton, A. F. (2020). Contrastive Representation Learning: A Framework and Review. IEEE Access, 8, 193907–193934. https://doi.org/10.1109/ACCESS.2020.3031549
  • Uelwer, T., Robine, J., Wagner, S. S., Höftmann, M., Upschulte, E., Konietzny, S., … Harmeling, S. (2025, April 1). A survey on self-supervised methods for visual representation learning. Machine Learning. Springer. https://doi.org/10.1007/s10994-024-06708-7
  • Zhang, D., Yin, J., Zhu, X., & Zhang, C. (2020, March 1). Network Representation Learning: A Survey. IEEE Transactions on Big Data. Institute of Electrical and Electronics Engineers Inc. https://doi.org/10.1109/TBDATA.2018.2850013
  • Guo, W., Wang, J., & Wang, S. (2019). Deep Multimodal Representation Learning: A Survey. IEEE Access, 7, 63373–63394. https://doi.org/10.1109/ACCESS.2019.2916887
  • Ju, W., Fang, Z., Gu, Y., Liu, Z., Long, Q., Qiao, Z., … Zhang, M. (2024, May 1). A Comprehensive Survey on Deep Graph Representation Learning. Neural Networks. Elsevier Ltd. https://doi.org/10.1016/j.neunet.2024.106207
  • Khoshraftar, S., & Aijun, A. N. (2024). A Survey on Graph Representation Learning Methods. ACM Transactions on Intelligent Systems and Technology, 15(1). https://doi.org/10.1145/3633518
  • Liu, P., Liu, Z., Gao, Z. F., Gao, D., Zhao, W. X., Li, Y., … Wen, J. R. (2024). Do Emergent Abilities Exist in Quantized Large Language Models: An Empirical Study. In 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings (pp. 5174–5190). European Language Resources Association (ELRA).
  • Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., … Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361. https://arxiv.org/
  • Pasupuleti, M. K. (2022). Scaling Laws for Transformers on Low-Dimensional Data: A Statistical and Approximation Theory Perspective. International Journal of Academic and Industrial Research Innovations(IJAIRI), 02(01), 191–202. https://doi.org/10.62311/nesx/rp-5-01-2022
  • Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., … Sifre, L. (2022). Training compute-optimal large language models. In Advances in Neural Information Processing Systems (Vol. 35, pp. 30016–30030). https://arxiv.org/abs/2203.15556
  • Yin, Y., Zhao, Y., Zheng, M., Lin, K., Ou, J., Chen, R., … Gai, K. (2025). Towards Precise Scaling Laws for Video Diffusion Transformers. In Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition (pp. 18155–18165). IEEE Computer Society. https://doi.org/10.1109/CVPR52734.2025.01692
  • Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners. OpenAI Blog, 1(8). Available at: https://openai.com/blog/better-language-models
  • Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., … Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288. https://arxiv.org/abs/2307.09288
  • Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., … Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1–67. http://jmlr.org/papers/v21/20-1307.html
  • Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., … Schulman, J. (2021). Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168. https://arxiv.org/abs/2110.14168
  • Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., & Choi, Y. (2019). HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4791–4800). https://doi.org/10.18653/v1/P19-1472
  • Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., & Tafjord, O. (2018). Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv preprint arXiv:1803.05457. https://arxiv.org/abs/1803.05457
  • Joshi, M., Choi, E., Weld, D. S., & Zettlemoyer, L. (2017). TriviaQA: A reading comprehension dataset over Wikipedia and the web. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (pp. 1601–1611). https://doi.org/10.18653/v1/P17-1147
  • Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., … Petrov, S. (2019). Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7, 452–466. https://doi.org/10.1162/tacl_a_00276
  • Roy, O., & Vetterli, M. (2007). The effective rank: A measure of effective dimensionality. In Proceedings of the 15th European Signal Processing Conference (EUSIPCO 2007) (pp. 606–610).
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Mechanistic Interpretability Representation Learning Emergent Abilities in Large Language Models Transformer Scaling Laws Residual Stream Analysis

Powered by PhDFocusTM