SAFE-HealCloud: Safety-Aware, Agentic Self-Healing for Cloud Infrastructure
Prudvi Saisaran Ponduru,
Pavani Priya Vyshnavi Nandanavanam and
Sai Kesav Kumar Ponduru
International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 2026, vol. 12, issue 4, 163-188
Abstract:
Cloud infrastructure failures are increasingly difficult to detect, diagnose, and remediate because production environments combine microservices, Kubernetes control loops, service meshes, serverless workloads, infrastructure-as-code, continuous delivery, and heterogeneous telemetry. Reactive monitoring and manual incident response remain necessary, but they do not scale to the volume, velocity, and causal complexity of modern cloud operations. This paper provides a structured synthesis of scholarly, industry, and standards-based work on AI-driven self-healing for cloud infrastructure, with emphasis on AIOps, AgentOps, LLM-based cloud operations, anomaly detection, causal root cause analysis, graph learning, reinforcement learning, automated remediation, Kubernetes self-healing, observability, chaos engineering, and self-healing infrastructure-as-code. We propose SAFE-HealCloud, a safety-aware, agentic, feedback-driven framework that integrates telemetry ingestion, multimodal observability, anomaly detection, failure prediction, causal RCA, retrieval-augmented LLM reasoning, policy guardrails, risk-scored remediation planning, controlled execution, verification, rollback, human approval, and continuous learning. A formal model defines cloud state, observability vectors, failure states, action spaces, remediation policies, rewards, constraints, and reliability objectives. We also specify reproducible Kubernetes-based evaluation designs, metrics, algorithms, risk controls, and operational use cases. The central finding is that near-term practical value lies in graduated autonomy: low-risk reversible actions can be automated, medium-risk actions should be canaried and policy-gated, and high-risk changes should remain human-approved. Designed in this way, AI-driven self-healing can reduce detection and recovery times, preserve error budgets, improve operator productivity, and strengthen digital resilience without sacrificing safety, auditability, or governance.
Keywords: AIOps; self-healing cloud infrastructure; autonomous remediation; observability; Kubernetes; LLM agents; root cause analysis; chaos engineering; SRE; resilience engineering; AgentOps; graph neural networks; causal inference; reinforcement learning; infrastructure as code; policy as code (search for similar items in EconPapers)
Date: 2026
Note: Article URL: https://ijsrcseit.com/home/article/view/CSEIT26124220
References: Add references at CitEc
Citations:
Downloads: (external link)
https://ijsrcseit.com/home/article/view/CSEIT26124220 Article URL (text/html)
https://ijsrcseit.com/home/article/download/CSEIT26124220/CSEIT26124220 Full text (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:jbh:ijsrcs:v12:y2026:i4:id:2121
DOI: 10.32628/CSEIT26124220
Access Statistics for this article
More articles in International Journal of Scientific Research in Computer Science, Engineering and Information Technology from International Journal of Scientific Research in Computer Science, Engineering and Information Technology
Bibliographic data for series maintained by Pankaj Sharma (USA) ().