Skip to content

Token-usage benchmark: HEY CLI vs hey-axi output (synthetic, reproducible) - #6

Closed
tangentus wants to merge 1 commit into
mainfrom
feat/token-benchmark
Closed

tangentus wants to merge 1 commit into
mainfrom
feat/token-benchmark

Conversation

@tangentus

Copy link
Copy Markdown
Contributor

What

A reproducible token-usage benchmark comparing what an agent reads from raw HEY CLI output vs hey-axi output for 9 representative commands: box list, box view imbox, a long thread read, label list, screener list, event week, todo list, a paged search, and an error response.

  • npm run bench:tokens counts tokens with js-tiktoken o200k_base and cl100k_base for:

    • (a) hey … --json
    • (b) hey … --styled (human output)
    • (c) hey-axi default TOON
    • (d) hey-axi --json

    hey-axi runs for real against a replay of the captured HEY response. --write refreshes docs/benchmarks.md; --real uses your own captures.

  • npm run bench:capture runs the real HEY CLI v1.7.0 binary against bench/lib/mock-hey-api.mjs, a local mock API that answers GET only and returns 405 for anything else. It uses HEY_BASE_URL, a dummy HEY_TOKEN and a throwaway HOME, so both outputs come from HEY's real formatting code. --real instead runs the read-only allowlist (bench/commands.mjs) against your logged-in account and writes to the gitignored bench/captures/real/.

  • npm run bench:fixtures regenerates the synthetic API responses from basecamp/hey-sdk openapi.json @ go/v0.31.1 (the SDK hey 1.7.0 is built on). Shapes come from the spec, values are invented, and per-endpoint shapers produce the structure the CLI interprets. Everything is checked in and marked synthetic: true.

  • docs/benchmarks.md covers results, takeaways, methodology, limitations, and how to rerun on your own mailbox. Linked from the README.

  • test/bench.test.js: the commands are on the read-only allowlist, the mock refuses non-GET, the benchmark runs on the synthetic captures, and the docs table matches a fresh run.

Results (synthetic data, measured)

Measured on synthetic captures from hey 1.7.0, tokenizers from js-tiktoken. Negative % = hey-axi output is larger.

o200k_base (GPT-4o / GPT-4.1 / o-series)

Command hey --json hey --styled hey-axi (TOON) hey-axi --json TOON vs hey --json TOON vs --styled
box list 768 62 503 588 34.5% -711.3%
box view imbox 20,598 857 15,591 15,310 24.3% -1719.3%
thread read (long thread) 2,811 1,506 2,409 2,308 14.3% -60.0%
label list 379 58 214 246 43.5% -269.0%
screener list 784 162 465 553 40.7% -187.0%
event week 5,573 471 4,629 4,234 16.9% -882.8%
todo list 1,684 136 1,373 1,226 18.5% -909.6%
search (paged) 6,634 864 5,024 4,959 24.3% -481.5%
error (thread not found) 25 6 21 21 16.0% -250.0%
Total 39,256 4,122 30,229 29,445 23.0% -633.4%

cl100k_base (GPT-4 / GPT-3.5)

Command hey --json hey --styled hey-axi (TOON) hey-axi --json TOON vs hey --json TOON vs --styled
box list 761 62 497 580 34.7% -701.6%
box view imbox 20,563 862 15,565 15,192 24.3% -1705.7%
thread read (long thread) 2,852 1,559 2,454 2,327 14.0% -57.4%
label list 380 58 214 245 43.7% -269.0%
screener list 787 163 468 553 40.5% -187.1%
event week 5,552 471 4,633 4,167 16.6% -883.7%
todo list 1,685 136 1,373 1,217 18.5% -909.6%
search (paged) 6,647 869 5,031 4,929 24.3% -478.9%
error (thread not found) 25 6 21 21 16.0% -250.0%
Total 39,252 4,186 30,256 29,231 22.9% -622.8%

Takeaways:

  • TOON saves about 23% vs hey --json, and 35–44% on flat lists.
  • On nested payloads, TOON is roughly the size of compact JSON, sometimes slightly larger.
  • HEY's --styled output is 5–24× smaller than either JSON form, because it shows only a few columns.
  • The next big win would be a field-selection / compact mode in hey-axi, not a different encoding.

Limitations

  • Synthetic payloads fill every documented field, so absolute sizes may run higher than real data.
  • There is no reliable offline Claude tokenizer: @anthropic-ai/tokenizer is Claude 2-era. OpenAI encodings are used as a proxy.

How tested

  • npm test: 85/85 pass (81 existing + 4 new). Uses js-tiktoken, a new devDependency pinned in the lockfile.
  • I ran the capture locally with the official hey 1.7.0 linux binary against the mock. No login was used, no real account or mailbox was touched, and nothing was sent.

Needs real-HEY testing

On your Mac: npm ci && npm run bench:capture -- --real && npm run bench:tokens -- --real. It's read-only, the output is gitignored, and BENCH_THREAD_ID picks a long thread.

Targets main. Not for auto-merge.

- bench/generate-api-fixtures.mjs: SYNTHETIC HEY API responses generated
  from basecamp/hey-sdk openapi.json (go/v0.31.1, the SDK behind hey 1.7.0)
- bench/capture.mjs: runs the real hey binary against a local read-only
  mock API (or, with --real, read-only commands against your own account;
  output gitignored) and records --json and --styled output
- bench/tokens.mjs (npm run bench:tokens): counts tokens with js-tiktoken
  o200k_base and cl100k_base for hey --json, hey --styled, hey-axi TOON
  and hey-axi --json; --write refreshes docs/benchmarks.md
- docs/benchmarks.md: results, methodology, limitations, rerun guide
- test/bench.test.js: allowlist, mock is GET-only, bench runs, docs in sync
@tangentus

Copy link
Copy Markdown
Contributor Author

Landed in main via #8 (squash commit 421fd57), which included this benchmark.

@tangentus
tangentus deleted the feat/token-benchmark branch October 6, 2026 03:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant