Muhammad Bilal, Jon Crowcroft, Ruizhi Wang, Xiaolong Xu, S. Dustdar
A survey paper systematically organizing architectures, evaluation, and safety of LLM-based network operations and AIOps agents.
Increasing adoption of LLM agents in network and IT operations raises critical reliability and safety concerns due to the immediate impact of operational decisions. Existing evaluations focus on static QA, failing to capture real workflow complexity.
Organizes literature along four axes: autonomy hierarchy, tool scope, evidence traces, and assurance contracts. Identifies consistent patterns across tasks like telemetry query recommendation, diagnosis, root-cause analysis, configuration synthesis, change planning, and limited self-healing. Argues that operational reliability depends on surrounding mechanisms rather than the model itself, and proposes workflow-centered evaluation (trace quality, bounded tool use, safe proposal generation, sandboxed replay, canary trials).
Concludes that progress in intelligent NetOps/AIOps depends on treating autonomy as a constrained operational control problem and ensuring output reliability, auditability, and safe deployment. Includes analysis of security, privacy, and governance risks.