A CPU benchmark designed to measure performance on workloads characteristic of AI agent execution.
AI agents (like coding assistants, autonomous research tools, and tool-using LLM systems) stress CPUs differently than traditional benchmarks measure. Research shows that tool processing on CPUs accounts for 50-90% of total latency in agentic workloads (Georgia Tech/Intel, 2025), making the CPU the actual bottleneck — not the GPU.
Agent workloads involve rapid context switching, heavy JSON serialization, large-text manipulation, tree-search planning, concurrent tool dispatch, subprocess spawning, diff computation, HTML parsing, schema validation, streaming response parsing, code edit application, and memory-intensive context management — often all interleaved.
AgentiCPUMark targets these specific workload patterns to give a realistic picture of how a CPU performs when running AI agent infrastructure.
| Benchmark | What it measures |
|---|---|
| JSON Processing | Serialization & deserialization of tool schemas, function calls, and conversations |
| Text Processing | Tokenizer-like chunking, regex parsing, BPE-like frequency counting |
| Tree Search | Monte Carlo tree search and beam search simulating agent planning/reasoning |
| Diff/Patch | Line-by-line comparison, unified diff generation, and patch application |
| HTML Parsing | DOM tree construction, CSS-like queries, and HTML-to-markdown conversion |
| Schema Validation | JSON Schema compilation, validation, and malformed JSON repair |
| Streaming Parse | SSE stream parsing with fragmented tool-call JSON accumulation |
| Code Edit Apply | Exact substring search, uniqueness verification, and edit application |
| Agentic Loop | Full Glob->Grep->Read->Edit->Verify cycle with growing conversation state |
| Memory Pressure | Large working-set manipulation, burst alloc/dealloc, context window management |
| Benchmark | What it measures |
|---|---|
| Context Switching | Rapid coroutine/thread switching simulating agent reasoning loops |
| Concurrent Dispatch | Thread-pool throughput for parallel tool execution with shared state |
| Subprocess Spawning | Process fork/exec, pipe management, and output collection |
pip install .Run the full suite:
agenticpumarkRun a specific benchmark:
agenticpumark --benchmark json_processingOptions:
--benchmark NAME Run only the named benchmark
--iterations N Timed iterations per benchmark (default: 5)
--threads N Max threads for concurrent benchmarks (default: CPU count)
--json Output results as JSON
--verbose Show detailed per-iteration results with statistics
Each benchmark runs 1 untimed warm-up iteration before timed runs begin. This stabilizes CPU caches and Python's internal optimizations.
If the coefficient of variation (CV = stddev/mean) exceeds 5% after the initial iterations, additional runs are automatically added (up to 4 extra) to improve statistical confidence. This follows the Phoronix Test Suite approach.
Each benchmark reports:
- Median time (primary metric, robust to outliers)
- Mean, StdDev, and CV (coefficient of variation)
- 95% Confidence Interval (using t-distribution)
- Min/Max times across all iterations
Each benchmark's raw time is normalized against a reference system:
Score = 1000 * (reference_time / actual_time)
A score of 1000 = reference performance (AMD Ryzen 9 7950X baseline). Higher is better.
Three composite scores are computed using weighted geometric means:
| Score | Benchmarks included | What it tells you |
|---|---|---|
| Single-Agent Speed | All sequential benchmarks | How fast a single agent executes |
| Multi-Agent Throughput | All concurrent benchmarks | How well the CPU scales with parallel agents |
| Overall Composite | All benchmarks | Combined agentic CPU performance |
The geometric mean prevents any single benchmark from dominating through outlier performance (following SPEC CPU methodology).
| Benchmark | Weight |
|---|---|
| Agentic Loop | 13% |
| JSON Processing | 10% |
| Subprocess Spawning | 9% |
| Text Processing | 8% |
| Tree Search | 8% |
| Context Switching | 7% |
| Concurrent Dispatch | 7% |
| Diff/Patch | 7% |
| Streaming Parse | 7% |
| Code Edit Apply | 7% |
| HTML Parsing | 6% |
| Schema Validation | 6% |
| Memory Pressure | 5% |
Weights reflect the relative frequency and CPU cost of each workload pattern in real agent systems, informed by profiling data from the AgentCgroup and CPU-Centric Agentic AI research.
===========================================================================
AgentiCPUMark v0.1.0
CPU Benchmark for AI Agentic Workloads
===========================================================================
CPU: AMD Ryzen 9 7950X 16-Core Processor
Cores: 32
Platform: Linux-6.8.0-generic-x86_64-with-glibc2.39
Python: 3.13.0
Arch: x86_64
---------------------------------------------------------------------------
Benchmark Time (s) StdDev CV Ops Score
---------------------------------------------------------------------------
context_switching 0.048 0.002 4.2% 30,000 1042
json_processing 0.051 0.001 2.0% 4,300 980
...
---------------------------------------------------------------------------
Single-Agent Speed 1015
Multi-Agent Throughput 1030
OVERALL COMPOSITE 1020
(Reference: 1000 = AMD Ryzen 9 7950X | Higher is better)
===========================================================================
The initial benchmark suite (v0.1.0–v0.2.0) was designed from published research:
- A CPU-Centric Perspective on Agentic AI (Georgia Tech/Intel) — profiled five agentic workloads and found CPU tool processing accounts for 50-90% of total latency
- AgentCgroup — measured OS-level resource usage of AI agents, revealing 15.4x peak/avg memory ratios and burst-silence CPU utilization patterns
Two additional benchmarks (Streaming Parse and Code Edit Apply) were added after live introspection of Claude Code (Anthropic's AI coding agent). We ran Claude Code on the AgentiCPUMark repo itself and traced its behavior:
-
Streaming Parse: Claude Code receives LLM responses as Server-Sent Event (SSE) streams. Tool call arguments arrive as
input_json_deltafragments — partial JSON strings that the client must accumulate chunk-by-chunk, attempting aJSON.parse()on every fragment to detect completion. Profiling showed this is the agent's main hot loop, running continuously during every API response. Multiple sub-agents can stream in parallel, multiplexing fragment accumulation across independent sessions. This pattern is universal across agent clients (Cursor, Copilot, etc.) and was not covered by our existing JSON Processing benchmark, which tests complete serialize/deserialize cycles rather than incremental fragment assembly. -
Code Edit Apply: Claude Code's Edit tool applies changes via exact substring matching, not traditional diff/patch. It loads the entire file, searches for an
old_string, verifies the match is unique (requiring a second full scan), snapshots the original for undo, and applies the replacement. This is fundamentally different from the LCS-based diff computation in our Diff/Patch benchmark. For large files, the uniqueness scan is expensive. When the target string isn't unique, the agent must extend context by including surrounding lines and re-scan — a pattern of iteratively widening substring searches. Agents typically chain 5-20 such edits per task.
Both patterns were confirmed by measuring Claude Code's client-side overhead across multi-tool tasks (34ms/turn client processing, 7-22 turns per task).
The Agentic Loop benchmark was added after a deeper study of Claude Code's behavior when writing code. We ran five different coding tasks (small feature, test file creation, multi-file refactor, bug fix, complex multi-file feature) and traced every tool call, turn count, token flow, and timing:
Task Type | Turns | Wall(s) | Input Tokens | Output Tokens
Small feature | 9 | 32 | 132K | 1.7K
Test file (new) | 23 | 155 | 559K | 6.2K
Multi-file refact | 10 | 24 | 101K | 1.1K
Bug find+fix | 3 | 16 | 52K | 0.5K
Complex feature | 30 | 206 | 801K | 8.4K
Every task followed the same computational cycle: Glob (discover project structure) -> Grep (find symbols) -> Read (load file) -> Edit (substring search + uniqueness check + replace) -> Bash (verify with tests) -> Serialize (append results to growing conversation history) -> loop. The critical insight is that conversation state grows monotonically — turn 1 serializes ~50K tokens to JSON; turn 30 serializes ~800K tokens. The entire history is re-serialized on every turn for API calls and session persistence, making later turns progressively more expensive. No other benchmark captures this growing-context serialization pressure combined with real file I/O, regex search, and edit application in a single integrated workload.
- Zero external dependencies: Pure Python stdlib. No pip install needed beyond the package itself.
- Deterministic: Fixed random seeds where applicable for reproducibility.
- Real workload patterns: Every benchmark targets an actual computational pattern observed in production agent systems, not synthetic micro-benchmarks.
- Statistical rigor: Warm-up runs, adaptive iteration counts, and full statistical reporting following established benchmarking best practices.
- A CPU-Centric Perspective on Agentic AI — Georgia Tech / Intel (2025)
- AgentCgroup: OS Resources of AI Agents — profiling data on agent resource usage patterns
- SPEC CPU 2017 Methodology — scoring and composite methodology
- AI Agents That Matter — Princeton (2024), agent benchmarking methodology
Contributions welcome! Please open an issue to discuss new benchmark ideas before submitting a PR.
MIT