Skip to content

analyze_traffic can overflow the policy agent's context window on high-volume users #36

Description

@martinmain93

Summary

The policy-builder agent's analyze_traffic tool summarises every distinct
endpoint pattern for a user and returns them in a single tool result. For a
high-volume bot, that result can be large enough to overflow the agent model's
context window on the following turn. The tool appears to run (progress ticks
along), then the turn fails with a generic "network error" and no policy is
produced.

Repro

  1. Point CrabTrap at an audit log for a busy agent — in our case ~955k requests
    over 7 days for a single bot user.
  2. Open the policy generator for that user and use the default "analyze traffic
    and build a policy" flow.
  3. analyze_traffic enumerates the distinct endpoint patterns, summarises each
    via the fast model, concatenates them all into one tool result, and the next
    agent call exceeds the context window → "network error".

Expected vs. actual

Expected: the agent analyzes a high-volume user's traffic and produces a policy.

Actual: for a high-volume user, analyze_traffic runs (progress advances),
then the turn fails with a generic "network error" and no policy is produced —
because the single tool result overflows the agent model's context window.

Root cause

toolAnalyzeTraffic calls AggregatePathGroups and then summarises and returns
all groups in one string. Two compounding problems:

  • The number of distinct (method, path) patterns is unbounded. In our data it
    was ~15k groups.
  • Each group also triggers a fast-model summarisation call, so the fan-out is
    unbounded too.

There is no cap, no pagination, and no graceful handling when the resulting
context is too large.

Why "just cap to the top N by count" is not the right fix

We measured the obvious mitigation (keep the top-N groups by request count) and
it's actively wrong for a security tool. In our data the 15,759 patterns
collapse to only 201 distinct hosts. The top 200 patterns by volume:

  • cover 93% of request volume, but
  • only 16 of 201 hosts — i.e. they silently drop 185 destinations.

For an egress policy the destination is the unit that matters, and the rare,
low-volume host in the long tail (an unexpected bucket, a one-off POST) is
exactly what you want a policy author to see. Ranking by volume hides it.

Proposed direction

Give the agent bounded, explorable views instead of one unbounded dump:

  • group_by="host" (default): per-host request + distinct-endpoint counts —
    cheap breadth, no fast-model calls.
  • group_by="endpoint": individual patterns, with optional host/path_prefix
    filters and opt-in per-endpoint summaries.
  • limit (default 50, hard-capped) + offset pagination; every result reports
    totals and how many rows remain.

Summarisation then runs only over the returned page. As a backstop, the agent
loop can detect provider context-length errors, shrink the largest tool result,
and retry / degrade gracefully rather than dying.

I opened a PR here to address this #35 which
worked well on our data in our testing

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions