You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 7a02972
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: README.md
+11-16Lines changed: 11 additions & 16 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -24,18 +24,11 @@ NetOpsBench is an open benchmark arena for agentic network troubleshooting — r
24
24
25
25
## Why NetOpsBench
26
26
27
-
Developing and evaluating agentic root cause analysis methods for network troubleshooting remains challenging, with three core bottlenecks hindering further advancement:
|**No fair comparison**| Varied network topologies, fault sets, observability tools, and evaluation metrics hinder the comparison of agentic troubleshooting strategies across the research community. | NetOpsBench unifies fault scenarios, observability access and scoring rules to support agent comparison on a shared benchmark. |
34
-
|**Non-reproducible faults**| Real network incidents cannot be reliably reproduced or labeled with consistent ground truth, slowing iterative improvement and evaluation of troubleshooting agents. | Containerlab + SONiC-VS inject controlled, reproducible faults with stable labels, so every run is an identical, repeatable episode. |
35
-
|**Non-Interactive Environment**| Static topology snapshots and logs cannot provide live probing and telemetry signals required by agents for diagnostic work. | NetOpsBench offers an interactive environment for agents to operate within live networks, capturing real-time Pingmesh data, gNMI telemetry and switch CLI evidence during every episode. |
36
-
37
-
27
+
Troubleshooting agents are difficult to compare when the network, incident, and evidence change from run to run. NetOpsBench turns those variables into a controlled live benchmark:
38
28
29
+
-**Reproducible incidents** — labeled faults run against repeatable SONiC-VS and Containerlab topologies.
30
+
-**Interactive evidence** — agents inspect live Pingmesh, BGP, gNMI, syslog, and switch state instead of static logs.
31
+
-**Comparable outcomes** — one evaluator measures detection, localization, efficiency, and tool use across agent strategies.
**Diagnosis score** is the mean end-to-end case score: healthy cases require the correct verdict, while fault cases receive localization credit only after the fault is detected. **Fault detection F1** measures the fault-versus-healthy decision independently.
108
+
109
+

110
+
111
+
The largest validated Fat-tree profile provides a compact case-level view. Each square below is one K=12 case; detailed cross-topology observability analysis remains in the full results.
112
+
113
+

115
114
116
115
Read the [v0.2.0 release notes](docs/content/docs/releases/v0.2.0.mdx), [Benchmark Methodology](docs/content/docs/run-benchmarks/methodology.mdx), and [Benchmark Results](docs/content/docs/run-benchmarks/results.mdx) for the full validation snapshot and scoring definitions.
117
116
@@ -134,10 +133,6 @@ The public [NetOpsBench Trace Dataset](https://huggingface.co/datasets/yyyyyt/ne
134
133
- Global community: [NetOpsBench Slack](https://join.slack.com/t/netopsbench/shared_invite/zt-3zhhfangj-2U4dU_NSfCy1rcOM1dmuvQ)
|`interface_localization_rate`| Interface-level precision for applicable cases. |
150
150
|`localization_composite_score`| Combined localization metric used for comparison. |
151
-
|`average_score`| Mean case score over the run. |
151
+
|`average_score`| Mean verdict-gated case score, shown publicly as **Diagnosis score**. Healthy cases require the correct verdict; fault cases receive localization credit only after the fault is detected. |
152
152
|`avg_time_seconds`| Mean diagnosis latency per case. |
Copy file name to clipboardExpand all lines: docs/content/docs/run-benchmarks/results.mdx
+10-6Lines changed: 10 additions & 6 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -9,18 +9,20 @@ This page records versioned benchmark snapshots. Treat each snapshot as a refere
9
9
10
10
This 2026-08 release snapshot used `minimal-deepagent` with `deepseek-v4-pro`, thinking disabled, temperature 0, and two runtime workers for the XS–Large campaigns. All 109 cases completed under the NetOpsBench v0.2.0 runtime, Rust client-agent, traffic, and Pingmesh contract.
11
11
12
-
| Scale | Cases |Primary reward (%) |Detection F1 (%) | Device loc. (%) | Interface loc. (%) | Fault type (%) | Avg time (s) | Avg tools | Input tokens (K) |
12
+
| Scale | Cases |Diagnosis score (%) |Fault detection F1 (%) | Device loc. (%) | Interface loc. (%) | Fault type (%) | Avg time (s) | Avg tools | Input tokens (K) |

19
+

20
20
21
-
Across all scales:
21
+
The overview spans all 319 v0.2 Agent-scored cases. The table and aggregate summary in this subsection focus on the 109 XS–Large cases; Xlarge and Fat-tree details follow below.
22
22
23
-
- Primary reward was **72.9%** (`79.5 / 109`).
23
+
Across XS–Large:
24
+
25
+
- Diagnosis score was **72.9%** (`79.5 / 109`). This is a verdict-gated case score, not an RL reward: healthy cases require the correct verdict, while fault cases receive credit for correct device/interface localization only after detecting the fault.
24
26
- Verdict accuracy was **80.7%** (`88 / 109`), including **13 / 13** healthy cases and no false positive.
25
27
- The Agent detected **75 / 96** positive cases (**78.1%**).
26
28
- Device localization was **69 / 96** (**71.9%**) over all positive cases and **69 / 75** (**92.0%**) conditional on the Agent reporting a fault.
@@ -59,13 +61,15 @@ The machine-readable [v0.2.0 result snapshot](/assets/benchmark/deepseek_v02_rel
59
61
60
62
The large-topology campaigns use composite summaries because infrastructure failures were fixed and only the affected cases were rerun. Selection is by scenario ID: a valid retry replaces an invalid infrastructure/API attempt, while successful original attempts remain unchanged.
61
63
62
-
| Topology | Catalog coverage | Agent-scored | Snapshot |Primary reward (%) |Detection F1 (%) | Device loc. (%) | Device loc. given detection (%) | Interface loc. (%) | Fault type (%) | Avg time (s) | Input tokens (K) |
64
+
| Topology | Catalog coverage | Agent-scored | Snapshot |Diagnosis score (%) |Fault detection F1 (%) | Device loc. (%) | Device loc. given detection (%) | Interface loc. (%) | Fault type (%) | Avg time (s) | Input tokens (K) |

71
+
72
+
The K=12 figure is a representative case-level deep dive for the largest validated Fat-tree profile. Each square is one of its 70 Agent-scored cases; the outcome distribution must not be assumed to apply unchanged to the other topologies.
69
73
70
74
\* The K=8 composite combines 16 complete results from the initial run with 54 missing-scenario results from a resumed run. Two orphan traces produced during the interrupted process were excluded because they did not complete post-recovery validation.
0 commit comments