Summary
The policy-builder agent's analyze_traffic tool summarises every distinct
endpoint pattern for a user and returns them in a single tool result. For a
high-volume bot, that result can be large enough to overflow the agent model's
context window on the following turn. The tool appears to run (progress ticks
along), then the turn fails with a generic "network error" and no policy is
produced.
Repro
- Point CrabTrap at an audit log for a busy agent — in our case ~955k requests
over 7 days for a single bot user.
- Open the policy generator for that user and use the default "analyze traffic
and build a policy" flow.
analyze_traffic enumerates the distinct endpoint patterns, summarises each
via the fast model, concatenates them all into one tool result, and the next
agent call exceeds the context window → "network error".
Expected vs. actual
Expected: the agent analyzes a high-volume user's traffic and produces a policy.
Actual: for a high-volume user, analyze_traffic runs (progress advances),
then the turn fails with a generic "network error" and no policy is produced —
because the single tool result overflows the agent model's context window.
Root cause
toolAnalyzeTraffic calls AggregatePathGroups and then summarises and returns
all groups in one string. Two compounding problems:
- The number of distinct
(method, path) patterns is unbounded. In our data it
was ~15k groups.
- Each group also triggers a fast-model summarisation call, so the fan-out is
unbounded too.
There is no cap, no pagination, and no graceful handling when the resulting
context is too large.
Why "just cap to the top N by count" is not the right fix
We measured the obvious mitigation (keep the top-N groups by request count) and
it's actively wrong for a security tool. In our data the 15,759 patterns
collapse to only 201 distinct hosts. The top 200 patterns by volume:
- cover 93% of request volume, but
- only 16 of 201 hosts — i.e. they silently drop 185 destinations.
For an egress policy the destination is the unit that matters, and the rare,
low-volume host in the long tail (an unexpected bucket, a one-off POST) is
exactly what you want a policy author to see. Ranking by volume hides it.
Proposed direction
Give the agent bounded, explorable views instead of one unbounded dump:
group_by="host" (default): per-host request + distinct-endpoint counts —
cheap breadth, no fast-model calls.
group_by="endpoint": individual patterns, with optional host/path_prefix
filters and opt-in per-endpoint summaries.
limit (default 50, hard-capped) + offset pagination; every result reports
totals and how many rows remain.
Summarisation then runs only over the returned page. As a backstop, the agent
loop can detect provider context-length errors, shrink the largest tool result,
and retry / degrade gracefully rather than dying.
I opened a PR here to address this #35 which
worked well on our data in our testing
Summary
The policy-builder agent's
analyze_traffictool summarises every distinctendpoint pattern for a user and returns them in a single tool result. For a
high-volume bot, that result can be large enough to overflow the agent model's
context window on the following turn. The tool appears to run (progress ticks
along), then the turn fails with a generic "network error" and no policy is
produced.
Repro
over 7 days for a single bot user.
and build a policy" flow.
analyze_trafficenumerates the distinct endpoint patterns, summarises eachvia the fast model, concatenates them all into one tool result, and the next
agent call exceeds the context window → "network error".
Expected vs. actual
Expected: the agent analyzes a high-volume user's traffic and produces a policy.
Actual: for a high-volume user,
analyze_trafficruns (progress advances),then the turn fails with a generic "network error" and no policy is produced —
because the single tool result overflows the agent model's context window.
Root cause
toolAnalyzeTrafficcallsAggregatePathGroupsand then summarises and returnsall groups in one string. Two compounding problems:
(method, path)patterns is unbounded. In our data itwas ~15k groups.
unbounded too.
There is no cap, no pagination, and no graceful handling when the resulting
context is too large.
Why "just cap to the top N by count" is not the right fix
We measured the obvious mitigation (keep the top-N groups by request count) and
it's actively wrong for a security tool. In our data the 15,759 patterns
collapse to only 201 distinct hosts. The top 200 patterns by volume:
For an egress policy the destination is the unit that matters, and the rare,
low-volume host in the long tail (an unexpected bucket, a one-off POST) is
exactly what you want a policy author to see. Ranking by volume hides it.
Proposed direction
Give the agent bounded, explorable views instead of one unbounded dump:
group_by="host"(default): per-host request + distinct-endpoint counts —cheap breadth, no fast-model calls.
group_by="endpoint": individual patterns, with optionalhost/path_prefixfilters and opt-in per-endpoint summaries.
limit(default 50, hard-capped) +offsetpagination; every result reportstotals and how many rows remain.
Summarisation then runs only over the returned page. As a backstop, the agent
loop can detect provider context-length errors, shrink the largest tool result,
and retry / degrade gracefully rather than dying.
I opened a PR here to address this #35 which
worked well on our data in our testing