|
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
|
| Volume 187 - Issue 127 |
| Published: July 2026 |
| Authors: Sneha Gullapalli |
10.5120/ijca0959b18e3a0a
|
Sneha Gullapalli . Policy-Governed Self-Healing Loops for Observability Pipelines in Distributed Cloud Systems. International Journal of Computer Applications. 187, 127 (July 2026), 25-32. DOI=10.5120/ijca0959b18e3a0a
@article{ 10.5120/ijca0959b18e3a0a,
author = { Sneha Gullapalli },
title = { Policy-Governed Self-Healing Loops for Observability Pipelines in Distributed Cloud Systems },
journal = { International Journal of Computer Applications },
year = { 2026 },
volume = { 187 },
number = { 127 },
pages = { 25-32 },
doi = { 10.5120/ijca0959b18e3a0a },
publisher = { Foundation of Computer Science (FCS), NY, USA }
}
%0 Journal Article
%D 2026
%A Sneha Gullapalli
%T Policy-Governed Self-Healing Loops for Observability Pipelines in Distributed Cloud Systems%T
%J International Journal of Computer Applications
%V 187
%N 127
%P 25-32
%R 10.5120/ijca0959b18e3a0a
%I Foundation of Computer Science (FCS), NY, USA
Observability pipelines have become a critical component of distributed cloud systems, collecting and processing metrics, logs, traces, and events that support monitoring, diagnosis, and automated operations. Failures within these pipelines, including queue saturation, exporter outages, schema drift, timestamp delays, memory pressure, and backend unavailability, can significantly reduce system visibility and affect operational decision-making. This paper presents SHIELD-OP, a policy-governed self-healing framework designed to improve the resilience of observability pipelines. The framework extends the Monitor–Analyze–Plan–Execute (MAPE) loop by incorporating policy constraints, verification mechanisms, rollback procedures, cooldown controls, and risk-aware remediation actions. A reproducible simulation study involving 50 random seeds and 120 failure episodes per seed compares SHIELD-OP with static alerting, retry-based recovery, ungoverned automation, and MAPE-K without verification. Results indicate that SHIELD-OP reduces recovery time and telemetry loss while maintaining a lower rate of unsafe actions, demonstrating the value of governance-aware self-healing for cloud observability infrastructure.