Repository navigation
fix(builder): paginate analyze_traffic and recover from context overflow - #35
Conversation
The policy-builder agent's analyze_traffic tool summarised and returned every distinct endpoint pattern in a single tool result. For a high-volume user that could be many thousands of patterns, producing a tool result large enough to overflow the agent model's context window on the next turn — surfacing to the user as a generic "network error" after the tool appeared to run. Redesign the tool to give the agent bounded, explorable views instead of one unbounded dump: - group_by="host" (default): per-host request and distinct-endpoint counts — cheap breadth of egress destinations, no fast-model calls. - group_by="endpoint": individual (method, path) patterns, optionally with host/path_prefix filters and opt-in per-endpoint summaries. - limit (default 50, hard-capped 100) + offset pagination; limit=0 returns counts only. Every result reports totals and how many rows remain. Summarisation now runs only over the returned page, so the fast-model fan-out is bounded too. The existing URL normalizer is unchanged; granularity is the agent's choice via parameters. As a backstop, the agent loop detects context-length errors from the provider (Bedrock / Anthropic / OpenAI), shrinks the largest tool result in the conversation and retries, and degrades to an actionable message instead of a fatal error if it still cannot fit. The admin UI surfaces each analyze_traffic call's parameters and result count so repeated drill-down calls are legible rather than looking redundant. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
| Filename | Overview |
|---|---|
| internal/builder/agent.go | Core redesign: paginated analyze_traffic (host/endpoint modes, limit/offset), context-recovery loop, appendSummary upsert, and supporting helpers — all well-structured with clear invariants. |
| internal/llm/errors.go | Introduces ErrContextLength sentinel and bodyIndicatesContextLength; the "context window" substring is slightly broader than the other three patterns. |
| internal/llm/bedrock.go | Adds isBedrockContextLengthError gated on ValidationException type; double-wrapping via fmt.Errorf("%w: %w") is valid Go 1.20+ and tested. |
| internal/llm/anthropic.go | ErrContextLength wrapped only on HTTP 400, correctly excluding 429 rate-limit errors. |
| internal/llm/openai.go | Same pattern as Anthropic adapter; ErrContextLength restricted to HTTP 400 responses. |
| internal/admin/pg_audit_reader.go | Adds hostFilter and pathPrefix parameters to AggregatePathGroups SQL; correctly uses starts_with instead of LIKE for literal prefix matching. |
| internal/builder/agent_test.go | Comprehensive new tests covering host/endpoint modes, pagination, limit clamping, drill-down, summarization bounding, context-error recovery, and ErrContextLength classification. |
| web/src/components/PolicyDetail.tsx | Surfaces analyze_traffic scope/paging params, first-line result summary, and context_recovery notice in the agent event stream UI. |
| internal/admin/policies_integration_test.go | Updated stub signatures and test fixtures to include full-URL path patterns and new group_by/summarize parameters. |
Reviews (6): Last reviewed commit: "fix(llm): classify context overflow by s..." | Re-trigger Greptile
Address PR review feedback: - AggregatePathGroups filtered path_prefix with `LIKE $5 || '%'`, so `_` and `%` in the prefix acted as wildcards — e.g. path_prefix "/api/v1/funds_" matched "/api/v1/funds/..." (725k rows vs 0 on real data). Use starts_with() for a literal prefix match. - Narrow isContextLengthError's "too long" check to "is too long" so unrelated errors are less likely to be misclassified as context overflow. - Have the test stub honor path_prefix (literal, mirroring starts_with) and add a test covering literal path_prefix filtering. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
| // request because it exceeded the model's context window. It matches on message | ||
| // text to stay provider-agnostic (Bedrock, direct Anthropic, and OpenAI all | ||
| // surface this differently), tolerating the surrounding wrapping each adds. | ||
| func isContextLengthError(err error) bool { |
There was a problem hiding this comment.
I think we should use http status code to distinguish the actual errors so we are not unnecessarily categorizing rate-limiting issues along with this, see below
HTTP status codes for context-overflow vs rate-limit
| Provider | Context too long | Rate / token throttle | Quota exhausted |
|---|---|---|---|
| Anthropic | 400 invalid_request_error — message "prompt is too long" (errors, rate-limits) |
429 rate_limit_error (errors, rate-limits) |
n/a — billing handled separately (errors) |
| OpenAI | 400 invalid_request_error with code: "context_length_exceeded" (cookbook, error-codes) |
429 rate_limit_exceeded — "Rate limit reached for requests" (error-codes, cookbook) |
429 insufficient_quota — "You exceeded your current quota..." |
| (error-codes, cookbook) | |||
| Bedrock | HTTP 200 with stopReason: "model_context_window_exceeded"; 400 ValidationException if the request is malformed (Converse API) |
429 ThrottlingException (Converse API, troubleshooting) |
rolled into ThrottlingException (AWS account quotas) |
| (troubleshooting) |
Addresses PR review (@bjhaid): the agent loop's context-overflow detector matched on message substrings, and its "token"+"exceed" catch-all would also match 429 rate-limit errors (e.g. OpenAI "rate_limit_exceeded ... tokens per minute"). Those got swallowed as context overflow — the user saw misleading "narrow your request" guidance and all retries were burned against an already throttled endpoint. Classify in the adapters, where the HTTP status / exception type is available, and wrap a shared sentinel (llm.ErrContextLength): - Anthropic / OpenAI: only HTTP 400 with a context-length message/code. - Bedrock: only a ValidationException with a too-long message (never ThrottlingException). Status/type is the gate (never 429); the message is the discriminator within it, since only OpenAI emits a distinct context_length_exceeded code. builder now uses errors.Is(err, llm.ErrContextLength) instead of string matching. Adds adapter tests asserting 429s are not classified as context overflow. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
Thanks! |
The policy-builder agent's analyze_traffic tool summarised and returned every distinct endpoint pattern in a single tool result. For a high-volume user that could be many thousands of patterns, producing a tool result large enough to overflow the agent model's context window on the next turn — surfacing to the user as a generic "network error" after the tool appeared to run.
Redesign the tool to give the agent bounded, explorable views instead of one unbounded dump:
Summarisation now runs only over the returned page, so the fast-model fan-out is bounded too. The existing URL normalizer is unchanged; granularity is the agent's choice via parameters.
As a backstop, the agent loop detects context-length errors from the provider (Bedrock / Anthropic / OpenAI), shrinks the largest tool result in the conversation and retries, and degrades to an actionable message instead of a fatal error if it still cannot fit.
The admin UI surfaces each analyze_traffic call's parameters and result count so repeated drill-down calls are legible rather than looking redundant.