Skip to content

captures: persist session-attribution fields in the req envelope (device_id, entrypoint, workload) #98

Description

@jsboige

Problem

Every traffic diagnosis (leak hunts, consumption analyses, split-brain checks) has to reconstruct WHO drove a request from heuristics over the body: workspace by most-cited path root, cron-vs-interactive by a per-request header stamp that misses tool-loop continuations, session identity not at all. Measured 2026-09-13: ~20-30 ad-hoc re-derivations across past analyses, each re-discovering the same traps (see #97's script header for the current list).

The proxy already SEES the ground truth at capture time but does not persist it.

Proposal

In the request-capture path, parse what is already in the request and write it into the envelope (top level, next to machine/model):

  • device_id8 — first 8 hex of metadata.user_id's device_id (client installation identity)
  • entrypoint / workload — from the x-anthropic-billing-header system[0] block (cc_entrypoint, cc_workload)
  • workspace root — from the capture-side path conventions, or left to scripts if not derivable server-side

Values only, never secrets; absent fields simply omitted. This makes native-consumption.py and every future attribution question exact instead of heuristic, at ~zero runtime cost.

Non-goals

Session-id synthesis (needs a client-side change); body rewriting; retroactive migration of archives.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions