You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit f6b62b5
Browse filesBrowse the repository at this point in the historyBrowse files
NetOpsBench is an open benchmark arena for agentic network troubleshooting — run reproducible fault scenarios on live SONiC-VS / Containerlab topologies, plug in any troubleshooting agent, and score it across quality and efficiency dimensions.
22
22
23
-
-**2026-05**: 🎉 **Initial Release** - NetOpsBench is now available as an open-source benchmark for closed-loop agentic DCN troubleshooting.
24
-
- Public SDK with `run_scenario()` and `run_suite()` APIs
25
-
- Custom diagnostic agent framework and examples
26
-
- Scenario generation tooling for XS, Small, Medium, and Large topologies
27
-
- Complete benchmark documentation and methodology
28
-
- Cross-model evaluation results and scoring framework
23
+
## Overview
29
24
30
-
NetOpsBench evaluates agentic RCA inside a real reproducible data-center network loop: build a SONiC-VS / Containerlab fabric, inject a controlled fault, observe Pingmesh and telemetry symptoms, call an agent, and score the diagnosis against ground truth.
25
+
NetOpsBench provides: (1) an interactive and realistic environment mimicking production networks, with common tracing and telemetry tooling; (2) comprehensive and reproducible benchmarks covering a wide range of faults and failures; (3) an extensible architecture with an open SDK to readily integrate with various agent paradigms and observability tools, allowing users to try out their own agentic workflows.
31
26
32
27
It is built for researchers and engineers who want to compare LLM-backed, symbolic, heuristic, or hybrid troubleshooting strategies on the same operational benchmark, not just on static logs or hand-written prompts.
-**2026-05**: 🎉 **Initial Release** - NetOpsBench is now available as an open-source benchmark for agentic network troubleshooting.
34
+
- Public SDK with `run_scenario()` and `run_suite()` APIs
35
+
- Custom troubleshooting agent framework and examples
36
+
- Scenario generation tooling for XS, Small, Medium, and Large topologies
37
+
- Complete benchmark documentation and methodology
38
+
- Cross-model evaluation results and scoring framework
39
+
36
40
## Why It Matters
37
41
38
42
| Capability | What NetOpsBench gives you |
39
43
|---|---|
40
-
| Run the network, not just logs | Faults execute against live SONiC/Containerlab topologies with trafficand observability. |
41
-
| Bring any agentic strategy| Pass a Python object with `diagnose(context)` into the public SDK. |
44
+
| Run an interactive environment, not just logs | Faults execute against live SONiC/Containerlab topologies with traffic, tracing, and telemetry tooling. |
45
+
| Bring any agentic workflow| Pass a Python object with `diagnose(context)` into the public SDK. |
42
46
| Score localization, not only detection | Measure fault type, device, interface, runtime, tool calls, and token cost. |
43
47
| Scale with isolated worker pools | Run suites across independent labs and merge results into one `BenchmarkReport`. |
44
48
45
49
## Quick Start
46
50
47
-
NetOpsBench runtime execution requires Linux because Containerlab depends on Linux networking primitives.
51
+
> NetOpsBench runtime execution requires Linux because Containerlab depends on Linux networking primitives.
The first successful run should produce a `BenchmarkReport` with case-level scores, timing, and artifact paths. For Docker, Containerlab, and runtime setup details, read [Quickstart](docs/content/docs/getting-started/index.mdx).
68
+
The first successful run produces a `BenchmarkReport` with case-level scores, timing, and artifact paths. For Docker, Containerlab, and runtime setup details, read [Quickstart](docs/content/docs/getting-started/index.mdx).
63
69
64
-
## Run a Scenario YAML
70
+
###Run a scenario from Python
65
71
66
72
```python
67
73
from examples.agents import MinimalDeepAgent
@@ -77,11 +83,11 @@ with NetOpsBench(workspace=".") as bench:
77
83
print(report.summary)
78
84
```
79
85
80
-
Scenario YAML files define the benchmark case: topology scale, traffic profile, fault type, target device, and interface-level ground truth when applicable. Use [Python API Guide](docs/content/docs/api/quickstart.mdx) for `run_scenario(...)`, `run_suite(...)`, and `workers=N`; use[Custom Diagnostic Agents](docs/content/docs/build-your-agent/custom-agents.mdx) when you are ready to replace `MinimalDeepAgent` with your own RCA strategy.
86
+
Scenario YAML files define the benchmark case: topology scale, traffic profile, fault type, target device, and interface-level ground truth when applicable. Use the [Python API Guide](docs/content/docs/api/quickstart.mdx) for `run_scenario(...)`, `run_suite(...)`, and `workers=N`; see[Custom Troubleshooting Agents](docs/content/docs/build-your-agent/custom-agents.mdx) when you are ready to replace `MinimalDeepAgent` with your own strategy.
81
87
82
-
## Benchmark What Matters
88
+
## Benchmark Results
83
89
84
-
NetOpsBench reports detection, fault type, device/interface localization, runtime, tool calls, and token usage so agent quality and operational cost can be compared together.
90
+
NetOpsBench reports detection, fault type, device/interface localization, runtime, tool calls, and token usage so troubleshooting quality and operational cost can be compared together.
Copy file name to clipboardExpand all lines: docs/content/docs/architecture/system-overview.mdx
+3-3Lines changed: 3 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -3,7 +3,7 @@ title: System Overview
3
3
description: The benchmark architecture from fault injection to scored diagnosis reports.
4
4
---
5
5
6
-
NetOpsBench evaluates diagnosis agents through a closed loop: inject a controlled data-center network fault, observe the fabric through active and passive signals, pass a structured diagnosis context to an agent, and score the result against ground truth.
6
+
NetOpsBench evaluates troubleshooting agents through an interactive benchmark loop: inject a controlled data-center network fault, observe the fabric through tracing and telemetry evidence, pass a structured diagnosis context to an agent, and score the result against ground truth.
@@ -15,7 +15,7 @@ NetOpsBench evaluates diagnosis agents through a closed loop: inject a controlle
15
15
| Runtime provisioning | The SDK provisions a Linux/Containerlab runtime, starts observability, and prepares Pingmesh workers. | Users can run a benchmark without hand-managing the lab. |
16
16
| Fault episode | NetOpsBench injects the selected fault and runs the observation window. | The agent sees symptoms produced by an actual running network, not a static prompt. |
17
17
| Diagnosis context | The platform assembles topology, symptoms, Pingmesh summaries, runtime metadata, and tool access into `DiagnosticContext`. | Agents receive a consistent machine-readable input surface. |
18
-
| Agent reasoning | A custom agent investigates evidence and returns `DiagnosisResult`. | Different RCA strategies can plug into the same benchmark loop. |
18
+
| Agent reasoning | A custom agent investigates evidence and returns `DiagnosisResult`. | Different troubleshooting strategies can plug into the same benchmark loop. |
19
19
| Scoring and reporting | The evaluator compares verdict, fault type, device, and interface against ground truth, then emits reports and artifacts. | Results are comparable across agents and runs. |
Copy file name to clipboardExpand all lines: docs/content/docs/benchmark/methodology.mdx
+4-4Lines changed: 4 additions & 4 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -21,7 +21,7 @@ The agent returns a `DiagnosisResult` containing a verdict, location, confidence
21
21
22
22
### Topology scales
23
23
24
-
All topologies use a spine-leaf fabric running SONiC-VS switches. Scale increases both the number of network devices and the density of clients, making diagnosis progressively harder as the fault signal is diluted across a larger Pingmesh matrix.
24
+
All topologies use a spine-leaf fabric running SONiC-VS switches. Scale increases both the number of network devices and the density of clients, making diagnosis progressively harder as the fault evidence is distributed across a larger Pingmesh matrix.
NetOpsBench evaluates whether diagnosis agents can turn telemetry and topology evidence into actionable root-cause localization. The benchmark does not just ask whether an agent notices that something is wrong; it asks whether the agent identifies the correct device, interface, and fault domain as topology scale increases.
6
+
NetOpsBench evaluates whether troubleshooting agents can turn tracing, telemetry, and topology evidence into actionable root-cause localization. The benchmark does not just ask whether an agent notices that something is wrong; it asks whether the agent identifies the correct device, interface, and fault domain as topology scale increases.
7
7
8
8
## What these results mean
9
9
@@ -48,13 +48,13 @@ Diagnosis quality is a funnel. An agent must detect the fault, then localize it
Most models maintain high detection. OpenAI GPT-5.4 and DeepSeek V4 Pro hold 100% from XS through Medium and stay near 80% on Large. MiniMax M2.7 climbs from 25% on XS to 94% on Large. Detection is a starting point, not a final RCA outcome.
51
+
Most models maintain high detection. OpenAI GPT-5.4 and DeepSeek V4 Pro hold 100% from XS through Medium and stay near 80% on Large. MiniMax M2.7 climbs from 25% on XS to 94% on Large. Detection is a starting point, not a final troubleshooting outcome.
Device localization is the first real bottleneck. Kimi K2.6 reaches 91.7% on Small but drops to 33.3% on Large. DeepSeek V4 Pro peaks at 95.8% on Medium then falls to 37.5%. OpenAI GPT-5.4 is the strongest on Large at 45.8%, still below a production-grade RCA target.
57
+
Device localization is the first real bottleneck. Kimi K2.6 reaches 91.7% on Small but drops to 33.3% on Large. DeepSeek V4 Pro peaks at 95.8% on Medium then falls to 37.5%. OpenAI GPT-5.4 is the strongest on Large at 45.8%, still below a production-grade troubleshooting target.
58
58
59
59
### Interface localization
60
60
@@ -92,14 +92,14 @@ The best choice depends on whether the priority is localization quality, fast tr
92
92
| Model | Strongest signal | Main risk | Best fit |
93
93
|---|---|---|---|
94
94
| Kimi K2.6 | Strong Small-scale composite, early device localization | Large quality drops while time and token cost rise sharply | Deep-diagnosis runs with generous reasoning budget |
95
-
| DeepSeek V4 Pro | Balanced quality on XS, Small, Medium | Large localization still degrades materially | Quality-oriented RCA evaluation |
95
+
| DeepSeek V4 Pro | Balanced quality on XS, Small, Medium | Large localization still degrades materially | Quality-oriented troubleshooting evaluation |
96
96
| OpenAI GPT-5.4 | Fast, low tool calls, strong detection | Interface localization collapses at Large | Fast triage, latency-sensitive flows |
97
97
| MiniMax M2.7 | Lower runtime/token footprint, improving detection | Lower absolute localization at Large | Budget-sensitive baselines and broad triage |
98
98
99
99
## How to read these results
100
100
101
101
- For triage, prioritize detection rate and average time.
102
-
- For RCA automation, prioritize device localization, interface localization, and composite score.
102
+
- For troubleshooting automation, prioritize device localization, interface localization, and composite score.
103
103
- For cost governance, compare localization quality against tool calls and tokens.
104
104
- For deployment readiness, judge by Large absolute values rather than XS-to-Large deltas alone.
Copy file name to clipboardExpand all lines: docs/content/docs/build-your-agent/custom-agents.mdx
+3-3Lines changed: 3 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -1,9 +1,9 @@
1
1
---
2
-
title: Custom Diagnostic Agents
3
-
description: Implement an RCA agent and evaluate it with the NetOpsBench benchmark loop.
2
+
title: Custom Troubleshooting Agents
3
+
description: Implement a troubleshooting agent and evaluate it with the NetOpsBench benchmark loop.
4
4
---
5
5
6
-
This is the main public integration path for NetOpsBench. After [Quickstart](/docs/getting-started) works, replace the example agent with your own RCA strategy and run it against the same generated network-fault scenarios.
6
+
This is the main public integration path for NetOpsBench. After [Quickstart](/docs/getting-started) works, replace the example agent with your own troubleshooting strategy and run it against the same generated network-fault scenarios.
7
7
8
8
Your agent can be LLM-backed, heuristic, symbolic, retrieval-augmented, or hybrid. The benchmark contract stays the same: NetOpsBench calls `diagnose(context)`, the agent receives `DiagnosticContext`, and the agent returns `DiagnosisResult`.
Copy file name to clipboardExpand all lines: docs/content/docs/contribute/custom-faults.mdx
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -4,7 +4,7 @@ description: Add project-local or external faults through the supported SDK exte
4
4
---
5
5
This guide explains the **public** fault extension path for NetOpsBench.
6
6
7
-
Custom fault packs belong in the contributor area because they extend the benchmark fault surface. If your goal is to evaluate an RCA strategy, start with [Custom Diagnostic Agents](/docs/build-your-agent/custom-agents) instead.
7
+
Custom fault packs belong in the contributor area because they extend the benchmark fault surface. If your goal is to evaluate a troubleshooting strategy, start with [Custom Troubleshooting Agents](/docs/build-your-agent/custom-agents) instead.
Copy file name to clipboardExpand all lines: docs/content/docs/examples/llm-fault-type-judge.mdx
+2-2Lines changed: 2 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -52,10 +52,10 @@ The detailed report includes a `fault_type_judgment` field. Its `mode` is `deter
52
52
53
53
## When to use it
54
54
55
-
Use LLM-as-judge when comparing agents that produce useful free-form RCA labels. Leave it disabled when you need strict taxonomy compliance or fully offline deterministic scoring.
55
+
Use LLM-as-judge when comparing agents that produce useful free-form troubleshooting labels. Leave it disabled when you need strict taxonomy compliance or fully offline deterministic scoring.
0 commit comments