Skip to content

Commit f6b62b5

Browse files
committed
Update docs and pages
1 parent 84a5527 commit f6b62b5

14 files changed

Lines changed: 222 additions & 142 deletions

File tree

‎README.md‎

Lines changed: 26 additions & 20 deletions
Original file line numberDiff line numberDiff line change
@@ -1,7 +1,7 @@
1-
# NetOpsBench
1+
# NetOpsBench: Open Arena for NetOps in AI Infrastructure
22

33
<p align="center">
4-
<strong>An Interactive Arena for AI Infrastructure Diagnosis</strong>
4+
<strong>Fair, reproducible benchmarks for agentic network troubleshooting.</strong>
55
</p>
66

77
<p align="center">
@@ -18,33 +18,39 @@
1818
<a href="docs/content/docs/benchmark/results.mdx">Benchmark Results</a>
1919
</p>
2020

21-
## News
21+
NetOpsBench is an open benchmark arena for agentic network troubleshooting — run reproducible fault scenarios on live SONiC-VS / Containerlab topologies, plug in any troubleshooting agent, and score it across quality and efficiency dimensions.
2222

23-
- **2026-05**: 🎉 **Initial Release** - NetOpsBench is now available as an open-source benchmark for closed-loop agentic DCN troubleshooting.
24-
- Public SDK with `run_scenario()` and `run_suite()` APIs
25-
- Custom diagnostic agent framework and examples
26-
- Scenario generation tooling for XS, Small, Medium, and Large topologies
27-
- Complete benchmark documentation and methodology
28-
- Cross-model evaluation results and scoring framework
23+
## Overview
2924

30-
NetOpsBench evaluates agentic RCA inside a real reproducible data-center network loop: build a SONiC-VS / Containerlab fabric, inject a controlled fault, observe Pingmesh and telemetry symptoms, call an agent, and score the diagnosis against ground truth.
25+
NetOpsBench provides: (1) an interactive and realistic environment mimicking production networks, with common tracing and telemetry tooling; (2) comprehensive and reproducible benchmarks covering a wide range of faults and failures; (3) an extensible architecture with an open SDK to readily integrate with various agent paradigms and observability tools, allowing users to try out their own agentic workflows.
3126

3227
It is built for researchers and engineers who want to compare LLM-backed, symbolic, heuristic, or hybrid troubleshooting strategies on the same operational benchmark, not just on static logs or hand-written prompts.
3328

3429
![NetOpsBench pipeline architecture](docs/public/assets/pipeline_architecture.png)
3530

31+
## News
32+
33+
- **2026-05**: 🎉 **Initial Release** - NetOpsBench is now available as an open-source benchmark for agentic network troubleshooting.
34+
- Public SDK with `run_scenario()` and `run_suite()` APIs
35+
- Custom troubleshooting agent framework and examples
36+
- Scenario generation tooling for XS, Small, Medium, and Large topologies
37+
- Complete benchmark documentation and methodology
38+
- Cross-model evaluation results and scoring framework
39+
3640
## Why It Matters
3741

3842
| Capability | What NetOpsBench gives you |
3943
|---|---|
40-
| Run the network, not just logs | Faults execute against live SONiC/Containerlab topologies with traffic and observability. |
41-
| Bring any agentic strategy | Pass a Python object with `diagnose(context)` into the public SDK. |
44+
| Run an interactive environment, not just logs | Faults execute against live SONiC/Containerlab topologies with traffic, tracing, and telemetry tooling. |
45+
| Bring any agentic workflow | Pass a Python object with `diagnose(context)` into the public SDK. |
4246
| Score localization, not only detection | Measure fault type, device, interface, runtime, tool calls, and token cost. |
4347
| Scale with isolated worker pools | Run suites across independent labs and merge results into one `BenchmarkReport`. |
4448

4549
## Quick Start
4650

47-
NetOpsBench runtime execution requires Linux because Containerlab depends on Linux networking primitives.
51+
> NetOpsBench runtime execution requires Linux because Containerlab depends on Linux networking primitives.
52+
53+
### Install and run via CLI
4854

4955
```bash
5056
git clone https://github.com/NetX-lab/NetOpsBench.git
@@ -59,9 +65,9 @@ export OPENAI_API_KEY=...
5965
PYTHONPATH=. python examples/01_run_scenario.py --vendor openai
6066
```
6167

62-
The first successful run should produce a `BenchmarkReport` with case-level scores, timing, and artifact paths. For Docker, Containerlab, and runtime setup details, read [Quickstart](docs/content/docs/getting-started/index.mdx).
68+
The first successful run produces a `BenchmarkReport` with case-level scores, timing, and artifact paths. For Docker, Containerlab, and runtime setup details, read [Quickstart](docs/content/docs/getting-started/index.mdx).
6369

64-
## Run a Scenario YAML
70+
### Run a scenario from Python
6571

6672
```python
6773
from examples.agents import MinimalDeepAgent
@@ -77,11 +83,11 @@ with NetOpsBench(workspace=".") as bench:
7783
print(report.summary)
7884
```
7985

80-
Scenario YAML files define the benchmark case: topology scale, traffic profile, fault type, target device, and interface-level ground truth when applicable. Use [Python API Guide](docs/content/docs/api/quickstart.mdx) for `run_scenario(...)`, `run_suite(...)`, and `workers=N`; use [Custom Diagnostic Agents](docs/content/docs/build-your-agent/custom-agents.mdx) when you are ready to replace `MinimalDeepAgent` with your own RCA strategy.
86+
Scenario YAML files define the benchmark case: topology scale, traffic profile, fault type, target device, and interface-level ground truth when applicable. Use the [Python API Guide](docs/content/docs/api/quickstart.mdx) for `run_scenario(...)`, `run_suite(...)`, and `workers=N`; see [Custom Troubleshooting Agents](docs/content/docs/build-your-agent/custom-agents.mdx) when you are ready to replace `MinimalDeepAgent` with your own strategy.
8187

82-
## Benchmark What Matters
88+
## Benchmark Results
8389

84-
NetOpsBench reports detection, fault type, device/interface localization, runtime, tool calls, and token usage so agent quality and operational cost can be compared together.
90+
NetOpsBench reports detection, fault type, device/interface localization, runtime, tool calls, and token usage so troubleshooting quality and operational cost can be compared together.
8591

8692
![Composite benchmark score](docs/public/assets/benchmark/fig_avg_score.png)
8793

@@ -95,7 +101,7 @@ Read [Benchmark Methodology](docs/content/docs/benchmark/methodology.mdx) for sc
95101
| Understand the benchmark loop | [System Overview](docs/content/docs/architecture/system-overview.mdx) |
96102
| Understand parallel execution | [Worker Pool Execution](docs/content/docs/architecture/worker-pool-execution.mdx) |
97103
| Use NetOpsBench from Python | [Python API Guide](docs/content/docs/api/quickstart.mdx) |
98-
| Plug in your own RCA agent | [Custom Diagnostic Agents](docs/content/docs/build-your-agent/custom-agents.mdx) |
104+
| Plug in your own troubleshooting agent | [Custom Troubleshooting Agents](docs/content/docs/build-your-agent/custom-agents.mdx) |
99105
| Interpret benchmark scores | [Benchmark Methodology](docs/content/docs/benchmark/methodology.mdx) |
100106
| Deploy or inspect observability | [Operations](docs/content/docs/operations/deployment.mdx) |
101107

@@ -114,7 +120,7 @@ If you use NetOpsBench in your research, please cite:
114120
```bibtex
115121
@software{netopsbench2026,
116122
author = {Yang, Yitao and Xu, Hong},
117-
title = {{NetOpsBench}: An Interactive Arena for Agentic RCA in AI infrastructure},
123+
title = {{NetOpsBench}: Open Arena for NetOps in AI Infrastructure},
118124
year = {2026},
119125
url = {https://github.com/netx-lab/NetOpsBench},
120126
}

‎docs/content/docs/api/reference.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -111,5 +111,5 @@ python -m netopsbench.platform.topology.metadata_generator <clab-yaml> [output.j
111111

112112
- [Python API Guide](/docs/api/quickstart)
113113
- [CLI Reference](/docs/api/cli)
114-
- [Custom Diagnostic Agents](/docs/build-your-agent/custom-agents)
114+
- [Custom Troubleshooting Agents](/docs/build-your-agent/custom-agents)
115115
- [Architecture](/docs/architecture/system-overview)

‎docs/content/docs/architecture/system-overview.mdx‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -3,7 +3,7 @@ title: System Overview
33
description: The benchmark architecture from fault injection to scored diagnosis reports.
44
---
55

6-
NetOpsBench evaluates diagnosis agents through a closed loop: inject a controlled data-center network fault, observe the fabric through active and passive signals, pass a structured diagnosis context to an agent, and score the result against ground truth.
6+
NetOpsBench evaluates troubleshooting agents through an interactive benchmark loop: inject a controlled data-center network fault, observe the fabric through tracing and telemetry evidence, pass a structured diagnosis context to an agent, and score the result against ground truth.
77

88
![NetOpsBench pipeline architecture](/assets/pipeline_architecture.png)
99

@@ -15,7 +15,7 @@ NetOpsBench evaluates diagnosis agents through a closed loop: inject a controlle
1515
| Runtime provisioning | The SDK provisions a Linux/Containerlab runtime, starts observability, and prepares Pingmesh workers. | Users can run a benchmark without hand-managing the lab. |
1616
| Fault episode | NetOpsBench injects the selected fault and runs the observation window. | The agent sees symptoms produced by an actual running network, not a static prompt. |
1717
| Diagnosis context | The platform assembles topology, symptoms, Pingmesh summaries, runtime metadata, and tool access into `DiagnosticContext`. | Agents receive a consistent machine-readable input surface. |
18-
| Agent reasoning | A custom agent investigates evidence and returns `DiagnosisResult`. | Different RCA strategies can plug into the same benchmark loop. |
18+
| Agent reasoning | A custom agent investigates evidence and returns `DiagnosisResult`. | Different troubleshooting strategies can plug into the same benchmark loop. |
1919
| Scoring and reporting | The evaluator compares verdict, fault type, device, and interface against ground truth, then emits reports and artifacts. | Results are comparable across agents and runs. |
2020

2121
## Alert-driven diagnosis
@@ -57,5 +57,5 @@ In shorthand: **scenario -> agent -> session -> report**.
5757
- [Worker Pool Execution](/docs/architecture/worker-pool-execution) explains parallel execution, worker isolation, and report aggregation.
5858
- [Examples](/docs/examples/run-scenario-vs-suite) shows runnable scripts for the main paths.
5959
- [Benchmark Methodology](/docs/benchmark/methodology) explains how diagnosis outputs are scored.
60-
- [Custom Diagnostic Agents](/docs/build-your-agent/custom-agents) explains how external RCA strategies plug into the loop.
60+
- [Custom Troubleshooting Agents](/docs/build-your-agent/custom-agents) explains how external troubleshooting strategies plug into the loop.
6161
- [Python API Guide](/docs/api/quickstart) shows how to run NetOpsBench from Python.

‎docs/content/docs/benchmark/methodology.mdx‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,7 @@ The agent returns a `DiagnosisResult` containing a verdict, location, confidence
2121

2222
### Topology scales
2323

24-
All topologies use a spine-leaf fabric running SONiC-VS switches. Scale increases both the number of network devices and the density of clients, making diagnosis progressively harder as the fault signal is diluted across a larger Pingmesh matrix.
24+
All topologies use a spine-leaf fabric running SONiC-VS switches. Scale increases both the number of network devices and the density of clients, making diagnosis progressively harder as the fault evidence is distributed across a larger Pingmesh matrix.
2525

2626
| Scale | Spines | Leafs | Clients | Network devices |
2727
|---|---:|---:|---:|---:|
@@ -111,9 +111,9 @@ A full benchmark report aggregates case-level outcomes into summary metrics:
111111
| Metric | Interpretation |
112112
|---|---|
113113
| `detection_accuracy` | Fault/healthy verdict quality. |
114-
| `device_localization_rate` | Device-level RCA precision. |
115-
| `interface_localization_rate` | Interface-level RCA precision for applicable cases. |
116-
| `localization_composite_score` | Combined localization signal used for ranking. |
114+
| `device_localization_rate` | Device-level troubleshooting precision. |
115+
| `interface_localization_rate` | Interface-level troubleshooting precision for applicable cases. |
116+
| `localization_composite_score` | Combined localization metric used for ranking. |
117117
| `average_score` | Mean case score over the run. |
118118
| `avg_time_seconds` | Mean diagnosis latency per case. |
119119
| `avg_tool_calls` | Mean number of tool calls per case. |

‎docs/content/docs/benchmark/results.mdx‎

Lines changed: 5 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -3,7 +3,7 @@ title: Benchmark Results
33
description: "Cross-model NetOpsBench results: localization quality, scaling behavior, and runtime cost."
44
---
55

6-
NetOpsBench evaluates whether diagnosis agents can turn telemetry and topology evidence into actionable root-cause localization. The benchmark does not just ask whether an agent notices that something is wrong; it asks whether the agent identifies the correct device, interface, and fault domain as topology scale increases.
6+
NetOpsBench evaluates whether troubleshooting agents can turn tracing, telemetry, and topology evidence into actionable root-cause localization. The benchmark does not just ask whether an agent notices that something is wrong; it asks whether the agent identifies the correct device, interface, and fault domain as topology scale increases.
77

88
## What these results mean
99

@@ -48,13 +48,13 @@ Diagnosis quality is a funnel. An agent must detect the fault, then localize it
4848

4949
![Detection accuracy](/assets/benchmark/fig_detection.png)
5050

51-
Most models maintain high detection. OpenAI GPT-5.4 and DeepSeek V4 Pro hold 100% from XS through Medium and stay near 80% on Large. MiniMax M2.7 climbs from 25% on XS to 94% on Large. Detection is a starting point, not a final RCA outcome.
51+
Most models maintain high detection. OpenAI GPT-5.4 and DeepSeek V4 Pro hold 100% from XS through Medium and stay near 80% on Large. MiniMax M2.7 climbs from 25% on XS to 94% on Large. Detection is a starting point, not a final troubleshooting outcome.
5252

5353
### Device localization
5454

5555
![Device localization](/assets/benchmark/fig_device_loc.png)
5656

57-
Device localization is the first real bottleneck. Kimi K2.6 reaches 91.7% on Small but drops to 33.3% on Large. DeepSeek V4 Pro peaks at 95.8% on Medium then falls to 37.5%. OpenAI GPT-5.4 is the strongest on Large at 45.8%, still below a production-grade RCA target.
57+
Device localization is the first real bottleneck. Kimi K2.6 reaches 91.7% on Small but drops to 33.3% on Large. DeepSeek V4 Pro peaks at 95.8% on Medium then falls to 37.5%. OpenAI GPT-5.4 is the strongest on Large at 45.8%, still below a production-grade troubleshooting target.
5858

5959
### Interface localization
6060

@@ -92,14 +92,14 @@ The best choice depends on whether the priority is localization quality, fast tr
9292
| Model | Strongest signal | Main risk | Best fit |
9393
|---|---|---|---|
9494
| Kimi K2.6 | Strong Small-scale composite, early device localization | Large quality drops while time and token cost rise sharply | Deep-diagnosis runs with generous reasoning budget |
95-
| DeepSeek V4 Pro | Balanced quality on XS, Small, Medium | Large localization still degrades materially | Quality-oriented RCA evaluation |
95+
| DeepSeek V4 Pro | Balanced quality on XS, Small, Medium | Large localization still degrades materially | Quality-oriented troubleshooting evaluation |
9696
| OpenAI GPT-5.4 | Fast, low tool calls, strong detection | Interface localization collapses at Large | Fast triage, latency-sensitive flows |
9797
| MiniMax M2.7 | Lower runtime/token footprint, improving detection | Lower absolute localization at Large | Budget-sensitive baselines and broad triage |
9898

9999
## How to read these results
100100

101101
- For triage, prioritize detection rate and average time.
102-
- For RCA automation, prioritize device localization, interface localization, and composite score.
102+
- For troubleshooting automation, prioritize device localization, interface localization, and composite score.
103103
- For cost governance, compare localization quality against tool calls and tokens.
104104
- For deployment readiness, judge by Large absolute values rather than XS-to-Large deltas alone.
105105

‎docs/content/docs/build-your-agent/custom-agents.mdx‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,9 +1,9 @@
11
---
2-
title: Custom Diagnostic Agents
3-
description: Implement an RCA agent and evaluate it with the NetOpsBench benchmark loop.
2+
title: Custom Troubleshooting Agents
3+
description: Implement a troubleshooting agent and evaluate it with the NetOpsBench benchmark loop.
44
---
55

6-
This is the main public integration path for NetOpsBench. After [Quickstart](/docs/getting-started) works, replace the example agent with your own RCA strategy and run it against the same generated network-fault scenarios.
6+
This is the main public integration path for NetOpsBench. After [Quickstart](/docs/getting-started) works, replace the example agent with your own troubleshooting strategy and run it against the same generated network-fault scenarios.
77

88
Your agent can be LLM-backed, heuristic, symbolic, retrieval-augmented, or hybrid. The benchmark contract stays the same: NetOpsBench calls `diagnose(context)`, the agent receives `DiagnosticContext`, and the agent returns `DiagnosisResult`.
99

‎docs/content/docs/contribute/custom-faults.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@ description: Add project-local or external faults through the supported SDK exte
44
---
55
This guide explains the **public** fault extension path for NetOpsBench.
66

7-
Custom fault packs belong in the contributor area because they extend the benchmark fault surface. If your goal is to evaluate an RCA strategy, start with [Custom Diagnostic Agents](/docs/build-your-agent/custom-agents) instead.
7+
Custom fault packs belong in the contributor area because they extend the benchmark fault surface. If your goal is to evaluate a troubleshooting strategy, start with [Custom Troubleshooting Agents](/docs/build-your-agent/custom-agents) instead.
88

99
## Recommended starting point
1010

‎docs/content/docs/contribute/repository-layout.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -35,7 +35,7 @@ Key sub-packages:
3535
- `platform/observability/` — Telegraf config generation, InfluxDB bucket management, observability validation
3636
- `platform/scenario/` — scenario parsing, validation, generation, and execution
3737
- `platform/toolkit/` — internal toolkit core; `toolkit/mcp/` adapts it for MCP protocol tools
38-
- `platform/pingmesh/` — pingmesh generation, probing, and detection
38+
- `platform/pingmesh/` — pingmesh generation, path checks, and detection
3939

4040
## Scripts
4141

‎docs/content/docs/examples/llm-fault-type-judge.mdx‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -52,10 +52,10 @@ The detailed report includes a `fault_type_judgment` field. Its `mode` is `deter
5252

5353
## When to use it
5454

55-
Use LLM-as-judge when comparing agents that produce useful free-form RCA labels. Leave it disabled when you need strict taxonomy compliance or fully offline deterministic scoring.
55+
Use LLM-as-judge when comparing agents that produce useful free-form troubleshooting labels. Leave it disabled when you need strict taxonomy compliance or fully offline deterministic scoring.
5656

5757
## Related docs
5858

5959
- [Benchmark Methodology](/docs/benchmark/methodology)
6060
- [Run Scenario vs Run Suite](/docs/examples/run-scenario-vs-suite)
61-
- [Custom Diagnostic Agents](/docs/build-your-agent/custom-agents)
61+
- [Custom Troubleshooting Agents](/docs/build-your-agent/custom-agents)

‎docs/content/docs/getting-started/index.mdx‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -105,7 +105,7 @@ For more command-line asset generation, result inspection, and runtime cleanup c
105105

106106
## Next steps
107107

108-
- [Build Your Agent](/docs/build-your-agent/custom-agents) explains how to plug in your own RCA strategy.
108+
- [Build Your Agent](/docs/build-your-agent/custom-agents) explains how to plug in your own troubleshooting strategy.
109109
- [Benchmark Methodology](/docs/benchmark/methodology) defines detection, localization, and efficiency scoring.
110110
- [Benchmark Results](/docs/benchmark/results) shows a completed cross-model run.
111111
- [Examples](/docs/examples/run-scenario-vs-suite) compares single-scenario, suite, scale, and batch scripts.

0 commit comments

Comments
 (0)