Self-Healing Datacenter Network Fabrics: Autonomous Remediation in Hyperscale and Distributed Cloud Infrastructure
Vijaya Bhaskar Methuku
International Journal of Scientific Research in Computer Science, Engineering and Information Technology, 2026, vol. 12, issue 2, 616-632
Abstract:
Hyperscale data centers use switching fabrics to keep large numbers of computers continuously connected. However, with the growing demands of distributed AI workloads and edge deployments, relying only on passive fault tolerance is no longer enough to ensure network resilience at scale. Traditional redundancy designs handle isolated failures well, but they break down when gradual degradations accumulate silently across thousands of interconnected paths before any single alarm is triggered. This gap in visibility and response is a key challenge addressed by this paper. In contrast, a self-healing Network fabric continuously performs anomaly detection, correlates telemetry signals across multiple layers into clear diagnoses, and autonomously executes pre-validated mitigations. A controlled evaluation using an 8-leaf, 4-spine Clos simulation environment characterized detection latency, mitigation time, and convergence behavior across ten distinct failure classes and compared outcomes against a human-operator baseline; detection latency averaged 847 milliseconds, while automated mitigation time of 2.3 seconds represented a 19.6× mitigation time improvement over the human baseline with zero packet loss across 42 automated remediations during a 7-day simulation period. A four-layer governance model integrates blast-radius containment directly into the remediation pipeline, distinguishing the framework from intent-based and AIOps-driven approaches that treat governance separately. Incident knowledge repositories enable pattern recognition and policy retrieval at machine speed, while rollback mechanisms and full audit provenance ensure governance obligations are satisfied on every automated action. These results demonstrate that autonomous remediation can meet the operational demands of AI-scale infrastructure without compromising safety governance.
Keywords: Self-Healing Networks; Clos Fabric Resilience; Autonomous Network Remediation; AIOps-Driven Orchestration; Telemetry-Based Anomaly Detection (search for similar items in EconPapers)
Date: 2026
Note: Article URL: https://ijsrcseit.com/home/article/view/CSEIT26121392
References: Add references at CitEc
Citations:
Downloads: (external link)
https://ijsrcseit.com/home/article/view/CSEIT26121392 Article URL (text/html)
https://ijsrcseit.com/home/article/download/CSEIT26121392/CSEIT26121392 Full text (application/pdf)
Related works:
This item may be available elsewhere in EconPapers: Search for items with the same title.
Export reference: BibTeX
RIS (EndNote, ProCite, RefMan)
HTML/Text
Persistent link: https://EconPapers.repec.org/RePEc:jbh:ijsrcs:v12:y2026:i2:id:1964
DOI: 10.32628/CSEIT26121392
Access Statistics for this article
More articles in International Journal of Scientific Research in Computer Science, Engineering and Information Technology from International Journal of Scientific Research in Computer Science, Engineering and Information Technology
Bibliographic data for series maintained by Pankaj Sharma ().