Research Article

Policy-Governed Self-Healing Loops for Observability Pipelines in Distributed Cloud Systems

by  Sneha Gullapalli
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 127
Published: July 2026
Authors: Sneha Gullapalli
10.5120/ijca0959b18e3a0a
PDF

Sneha Gullapalli . Policy-Governed Self-Healing Loops for Observability Pipelines in Distributed Cloud Systems. International Journal of Computer Applications. 187, 127 (July 2026), 25-32. DOI=10.5120/ijca0959b18e3a0a

                        @article{ 10.5120/ijca0959b18e3a0a,
                        author  = { Sneha Gullapalli },
                        title   = { Policy-Governed Self-Healing Loops for Observability Pipelines in Distributed Cloud Systems },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 127 },
                        pages   = { 25-32 },
                        doi     = { 10.5120/ijca0959b18e3a0a },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Sneha Gullapalli
                        %T Policy-Governed Self-Healing Loops for Observability Pipelines in Distributed Cloud Systems%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 127
                        %P 25-32
                        %R 10.5120/ijca0959b18e3a0a
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Observability pipelines have become a critical component of distributed cloud systems, collecting and processing metrics, logs, traces, and events that support monitoring, diagnosis, and automated operations. Failures within these pipelines, including queue saturation, exporter outages, schema drift, timestamp delays, memory pressure, and backend unavailability, can significantly reduce system visibility and affect operational decision-making. This paper presents SHIELD-OP, a policy-governed self-healing framework designed to improve the resilience of observability pipelines. The framework extends the Monitor–Analyze–Plan–Execute (MAPE) loop by incorporating policy constraints, verification mechanisms, rollback procedures, cooldown controls, and risk-aware remediation actions. A reproducible simulation study involving 50 random seeds and 120 failure episodes per seed compares SHIELD-OP with static alerting, retry-based recovery, ungoverned automation, and MAPE-K without verification. Results indicate that SHIELD-OP reduces recovery time and telemetry loss while maintaining a lower rate of unsafe actions, demonstrating the value of governance-aware self-healing for cloud observability infrastructure.

References
  • Kephart, J. O., and Chess, D. M. 2003. The Vision of Autonomic Computing. Computer 36, 1, 41–50.
  • Salehie, M., and Tahvildari, L. 2009. Self-Adaptive Software: Landscape and Research Challenges. ACM Transactions on Autonomous and Adaptive Systems 4, 2.
  • Weyns, D., Malek, S., and Andersson, J. 2012. FORMS: Unifying Reference Model for Formal Specification of Distributed Self-Adaptive Systems. ACM Transactions on Autonomous and Adaptive Systems 7, 1.
  • Huebscher, M. C., and McCann, J. A. 2008. A Survey of Autonomic Computing: Degrees, Models, and Applications. ACM Computing Surveys 40, 3.
  • Arcaini, P., Riccobene, E., and Scandurra, P. 2015. Modeling and Analyzing MAPE-K Feedback Loops for Self-Adaptation. In IEEE/ACM SEAMS.
  • de Lemos, R., et al. 2013. Software Engineering for Self-Adaptive Systems: A Second Research Roadmap. Lecture Notes in Computer Science, vol. 7475, Springer.
  • Filieri, A., Tamburrelli, G., and Ghezzi, C. 2016. Supporting Self-Adaptation via Quantitative Verification and Sensitivity Analysis at Run Time. IEEE Transactions on Software Engineering.
  • Gheibi, O., Weyns, D., and Quin, F. 2021. Applying Machine Learning in Self-Adaptive Systems: A Systematic Literature Review. ACM Transactions on Autonomous and Adaptive Systems.
  • Chen, Y., et al. 2025. AIOpsLab: A Holistic Framework to Evaluate AI Agents for Enabling Autonomous Clouds. arXiv:2501.06706.
  • Shetty, M., et al. 2024. Building AI Agents for Autonomous Clouds: Challenges and Design Principles. In ACM Symposium on Cloud Computing.
  • Wang, T., and Qi, G. 2024. A Comprehensive Survey on Root Cause Analysis in Microservices: Methodologies, Challenges, and Trends. arXiv:2408.00803.
  • Soldani, J., and Brogi, A. 2022. Anomaly Detection and Failure Root Cause Analysis in Microservice-Based Cloud Applications: A Survey. ACM Computing Surveys.
  • Pham, L., Ha, H., and Zhang, H. 2025. RCAEval: A Benchmark for Root Cause Analysis of Microservice Systems. In Proceedings of the ACM Web Conference.
  • OpenTelemetry. 2026. OpenTelemetry Collector. Documentation.
  • OpenTelemetry. 2026. Collector Resiliency. Documentation.
  • Cloud Native Computing Foundation. 2025. OpenTelemetry Specification.
  • Kubernetes. 2026. Configure Liveness, Readiness and Startup Probes. Documentation.
  • Kubernetes. 2026. Pod Lifecycle. Documentation.
  • National Institute of Standards and Technology. 2023. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1.
  • Yazdanparast, Z. 2024. A Survey on Self-Healing Software Systems. arXiv:2403.00455.
  • International Journal of Computer Applications. 2026. IJCA Scope and Topics. Foundation of Computer Science.
  • Gullapalli, S. 2026. Telemetry Quality Indexing for AI-Ready Observability in Multi-Node Cloud Systems: A Multi-Dimensional Framework for Distributed Infrastructure Diagnosis. International Journal of Engineering Development and Research (IJEDR) 14, 2, 878–. DOI: 10.56975/ijedr.v14i2.308419.
  • Divi, V. R. 2026. Multi-Agent Debate for Software Architecture Governance: A Framework for ADR Generation, Risk Review, and Deployment Readiness. International Journal of Engineering Development and Research (IJEDR) 14, 2, 862–. DOI: 10.56975/ijedr.v14i2.308420.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Self-healing observability AIOps OpenTelemetry policy governance MAPE-K telemetry pipelines automated remediation distributed cloud systems

Powered by PhDFocusTM