Skip to content

Commit 7a02972

Browse files
committed
docs: refine release results presentation
1 parent b5dbb23 commit 7a02972

10 files changed

Lines changed: 6233 additions & 132 deletions

‎README.md‎

Lines changed: 11 additions & 16 deletions
Original file line numberDiff line numberDiff line change
@@ -24,18 +24,11 @@ NetOpsBench is an open benchmark arena for agentic network troubleshooting — r
2424

2525
## Why NetOpsBench
2626

27-
Developing and evaluating agentic root cause analysis methods for network troubleshooting remains challenging, with three core bottlenecks hindering further advancement:
28-
29-
![NetOpsBench motivation](docs/public/assets/Motivation.png)
30-
31-
| Gap | The problem | How NetOpsBench closes it |
32-
|---|---|---|
33-
| **No fair comparison** | Varied network topologies, fault sets, observability tools, and evaluation metrics hinder the comparison of agentic troubleshooting strategies across the research community. | NetOpsBench unifies fault scenarios, observability access and scoring rules to support agent comparison on a shared benchmark. |
34-
| **Non-reproducible faults** | Real network incidents cannot be reliably reproduced or labeled with consistent ground truth, slowing iterative improvement and evaluation of troubleshooting agents. | Containerlab + SONiC-VS inject controlled, reproducible faults with stable labels, so every run is an identical, repeatable episode. |
35-
| **Non-Interactive Environment** | Static topology snapshots and logs cannot provide live probing and telemetry signals required by agents for diagnostic work. | NetOpsBench offers an interactive environment for agents to operate within live networks, capturing real-time Pingmesh data, gNMI telemetry and switch CLI evidence during every episode. |
36-
37-
27+
Troubleshooting agents are difficult to compare when the network, incident, and evidence change from run to run. NetOpsBench turns those variables into a controlled live benchmark:
3828

29+
- **Reproducible incidents** — labeled faults run against repeatable SONiC-VS and Containerlab topologies.
30+
- **Interactive evidence** — agents inspect live Pingmesh, BGP, gNMI, syslog, and switch state instead of static logs.
31+
- **Comparable outcomes** — one evaluator measures detection, localization, efficiency, and tool use across agent strategies.
3932

4033
## Overview
4134

@@ -111,7 +104,13 @@ NetOpsBench reports detection, fault type, device/interface localization, runtim
111104
| Fat-tree K=8 | Fat-tree | 80 | 128 | 70 |
112105
| Fat-tree K=12 | Fat-tree | 180 | 144 | 70 |
113106

114-
![NetOpsBench v0.2.0 DeepSeek validation](docs/public/assets/benchmark/fig_deepseek_v02_quality.png)
107+
**Diagnosis score** is the mean end-to-end case score: healthy cases require the correct verdict, while fault cases receive localization credit only after the fault is detected. **Fault detection F1** measures the fault-versus-healthy decision independently.
108+
109+
![NetOpsBench v0.2.0 quality across seven scales](docs/public/assets/benchmark/fig_deepseek_v02_overview.svg)
110+
111+
The largest validated Fat-tree profile provides a compact case-level view. Each square below is one K=12 case; detailed cross-topology observability analysis remains in the full results.
112+
113+
![Fat-tree K=12 case outcomes grouped by fault family](docs/public/assets/benchmark/fig_deepseek_v02_k12_cases.svg)
115114

116115
Read the [v0.2.0 release notes](docs/content/docs/releases/v0.2.0.mdx), [Benchmark Methodology](docs/content/docs/run-benchmarks/methodology.mdx), and [Benchmark Results](docs/content/docs/run-benchmarks/results.mdx) for the full validation snapshot and scoring definitions.
117116

@@ -134,10 +133,6 @@ The public [NetOpsBench Trace Dataset](https://huggingface.co/datasets/yyyyyt/ne
134133
- Global community: [NetOpsBench Slack](https://join.slack.com/t/netopsbench/shared_invite/zt-3zhhfangj-2U4dU_NSfCy1rcOM1dmuvQ)
135134
- Chinese-language community: [NetOpsBench Feishu group](https://applink.feishu.cn/client/chat/chatter/add_by_link?link_token=595v4390-2a51-4db0-baa4-811821b47448)
136135

137-
## Contributing
138-
139-
Contributions are welcome for benchmark scenarios, fault types, SDK ergonomics, documentation, and evaluation workflows.
140-
141136
## License
142137

143138
NetOpsBench is released under the MIT License. See [LICENSE](LICENSE).

‎docs/content/docs/run-benchmarks/methodology.mdx‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -144,11 +144,11 @@ The judge affects only fault-type classification. Verdict, device localization,
144144
| Metric | Interpretation |
145145
|---|---|
146146
| `detection_accuracy` | Fault/healthy verdict accuracy. |
147-
| `detection_f1` | Fault/healthy verdict F1-score used in the results plots. |
147+
| `detection_f1` | Fault/healthy verdict F1-score, shown publicly as **Fault detection F1**. |
148148
| `device_localization_rate` | Device-level troubleshooting precision. |
149149
| `interface_localization_rate` | Interface-level precision for applicable cases. |
150150
| `localization_composite_score` | Combined localization metric used for comparison. |
151-
| `average_score` | Mean case score over the run. |
151+
| `average_score` | Mean verdict-gated case score, shown publicly as **Diagnosis score**. Healthy cases require the correct verdict; fault cases receive localization credit only after the fault is detected. |
152152
| `avg_time_seconds` | Mean diagnosis latency per case. |
153153
| `avg_tool_calls` | Mean tool calls per case. |
154154

‎docs/content/docs/run-benchmarks/results.mdx‎

Lines changed: 10 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -9,18 +9,20 @@ This page records versioned benchmark snapshots. Treat each snapshot as a refere
99

1010
This 2026-08 release snapshot used `minimal-deepagent` with `deepseek-v4-pro`, thinking disabled, temperature 0, and two runtime workers for the XS–Large campaigns. All 109 cases completed under the NetOpsBench v0.2.0 runtime, Rust client-agent, traffic, and Pingmesh contract.
1111

12-
| Scale | Cases | Primary reward (%) | Detection F1 (%) | Device loc. (%) | Interface loc. (%) | Fault type (%) | Avg time (s) | Avg tools | Input tokens (K) |
12+
| Scale | Cases | Diagnosis score (%) | Fault detection F1 (%) | Device loc. (%) | Interface loc. (%) | Fault type (%) | Avg time (s) | Avg tools | Input tokens (K) |
1313
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|
1414
| XS | 14 | 75.0 | 95.7 | 75.0 | 57.1 | 41.7 | 61.3 | 22.6 | 174.3 |
1515
| Small | 15 | 83.3 | 95.7 | 83.3 | 57.1 | 58.3 | 65.1 | 20.9 | 221.4 |
1616
| Medium | 28 | 71.4 | 85.7 | 70.8 | 42.9 | 37.5 | 79.6 | 23.5 | 298.4 |
1717
| Large | 52 | 70.2 | 84.3 | 68.8 | 42.9 | 45.8 | 84.3 | 25.2 | 405.7 |
1818

19-
![DeepSeek 0.2 release quality by scale](/assets/benchmark/fig_deepseek_v02_quality.png)
19+
![DeepSeek 0.2 release quality across seven scales](/assets/benchmark/fig_deepseek_v02_overview.svg)
2020

21-
Across all scales:
21+
The overview spans all 319 v0.2 Agent-scored cases. The table and aggregate summary in this subsection focus on the 109 XS–Large cases; Xlarge and Fat-tree details follow below.
2222

23-
- Primary reward was **72.9%** (`79.5 / 109`).
23+
Across XS–Large:
24+
25+
- Diagnosis score was **72.9%** (`79.5 / 109`). This is a verdict-gated case score, not an RL reward: healthy cases require the correct verdict, while fault cases receive credit for correct device/interface localization only after detecting the fault.
2426
- Verdict accuracy was **80.7%** (`88 / 109`), including **13 / 13** healthy cases and no false positive.
2527
- The Agent detected **75 / 96** positive cases (**78.1%**).
2628
- Device localization was **69 / 96** (**71.9%**) over all positive cases and **69 / 75** (**92.0%**) conditional on the Agent reporting a fault.
@@ -59,13 +61,15 @@ The machine-readable [v0.2.0 result snapshot](/assets/benchmark/deepseek_v02_rel
5961

6062
The large-topology campaigns use composite summaries because infrastructure failures were fixed and only the affected cases were rerun. Selection is by scenario ID: a valid retry replaces an invalid infrastructure/API attempt, while successful original attempts remain unchanged.
6163

62-
| Topology | Catalog coverage | Agent-scored | Snapshot | Primary reward (%) | Detection F1 (%) | Device loc. (%) | Device loc. given detection (%) | Interface loc. (%) | Fault type (%) | Avg time (s) | Input tokens (K) |
64+
| Topology | Catalog coverage | Agent-scored | Snapshot | Diagnosis score (%) | Fault detection F1 (%) | Device loc. (%) | Device loc. given detection (%) | Interface loc. (%) | Fault type (%) | Avg time (s) | Input tokens (K) |
6365
|---|---:|---:|---|---:|---:|---:|---:|---:|---:|---:|---:|
6466
| Xlarge CLOS | 70 / 70 | 70 | Retry-adjusted 0.2 composite | 56.4 | 81.1 | 56.1 | 82.2 | 30.0 | 43.9 | 97.5 | 727.8 |
6567
| Fat-tree K=8 | 70 / 70 | 70 | Resumed 0.2 composite* | 63.6 | 80.0 | 63.6 | 95.5 | 45.0 | 34.8 | 102.9 | 594.7 |
6668
| Fat-tree K=12 | 70 / 70 | 70 | Retry-adjusted 0.2 composite | 55.0 | 73.1 | 56.1 | 97.4 | 27.5 | 30.3 | 102.9 | 784.3 |
6769

68-
![DeepSeek large-topology results](/assets/benchmark/fig_deepseek_v02_large_topologies.png)
70+
![Fat-tree K=12 case outcomes grouped by fault family](/assets/benchmark/fig_deepseek_v02_k12_cases.svg)
71+
72+
The K=12 figure is a representative case-level deep dive for the largest validated Fat-tree profile. Each square is one of its 70 Agent-scored cases; the outcome distribution must not be assumed to apply unchanged to the other topologies.
6973

7074
\* The K=8 composite combines 16 complete results from the initial run with 54 missing-scenario results from a resumed run. Two orphan traces produced during the interrupted process were excluded because they did not complete post-recovery validation.
7175

0 commit comments

Comments
 (0)