Skip to content

[P2] Per-request cent truncation erases cumulative cost for short calls #78

Description

@EvanProgramming

Summary

Cost is converted to integer cents inside every _track_cost call. Any individual request costing less than one cent is truncated to zero before aggregation, so large numbers of ordinary short requests can be reported as free.

Environment

  • OpenKyrozen main at 6b2eaa2980c5f7e1ea313b5aa378818db8f7c349
  • macOS, Python 3.12.13
  • Deterministic runtime reproduction; no credential required

Reproduction

reset_cost_tracker()
for _ in range(1000):
    _track_cost("deepseek", {"prompt_tokens": 0, "completion_tokens": 100})
print(get_cost_summary())

Observed: deepseek: 0K in / 100K out ~0c.

Even using the repository's own legacy output rate of $1.10 per million tokens, 100,000 output tokens should total $0.11. The loss occurs because each 100-token request is truncated independently with int(...).

Expected behavior

Accumulate tokens or sub-cent precision first, and round only when formatting the aggregate for display.

Impact

Cost reports systematically undercount workloads made of short requests, including agents that perform many small tool-planning calls.

Acceptance criteria

  • Repeated sub-cent requests accumulate to the same total as one request with the combined token count.
  • Internal accounting retains sufficient precision; display rounding happens only in get_cost_summary().
  • Tests cover 1,000 requests of 100 completion tokens and mixed prompt/completion usage.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

P2area: costToken usage, pricing, accounting, and spend reportingbugSomething isn't working

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions