Skip to content

fix(builder): paginate analyze_traffic and recover from context overflow - #35

Merged
bjhaid merged 3 commits into
brexhq:mainfrom
martinmain93:fix-analyze-traffic-context-overflow
Jun 25, 2026
Merged

bjhaid merged 3 commits into
brexhq:mainfrom
martinmain93:fix-analyze-traffic-context-overflow

Conversation

@martinmain93

Copy link
Copy Markdown
Contributor

The policy-builder agent's analyze_traffic tool summarised and returned every distinct endpoint pattern in a single tool result. For a high-volume user that could be many thousands of patterns, producing a tool result large enough to overflow the agent model's context window on the next turn — surfacing to the user as a generic "network error" after the tool appeared to run.

Redesign the tool to give the agent bounded, explorable views instead of one unbounded dump:

  • group_by="host" (default): per-host request and distinct-endpoint counts — cheap breadth of egress destinations, no fast-model calls.
  • group_by="endpoint": individual (method, path) patterns, optionally with host/path_prefix filters and opt-in per-endpoint summaries.
  • limit (default 50, hard-capped 100) + offset pagination; limit=0 returns counts only. Every result reports totals and how many rows remain.

Summarisation now runs only over the returned page, so the fast-model fan-out is bounded too. The existing URL normalizer is unchanged; granularity is the agent's choice via parameters.

As a backstop, the agent loop detects context-length errors from the provider (Bedrock / Anthropic / OpenAI), shrinks the largest tool result in the conversation and retries, and degrades to an actionable message instead of a fatal error if it still cannot fit.

The admin UI surfaces each analyze_traffic call's parameters and result count so repeated drill-down calls are legible rather than looking redundant.

The policy-builder agent's analyze_traffic tool summarised and returned every
distinct endpoint pattern in a single tool result. For a high-volume user that
could be many thousands of patterns, producing a tool result large enough to
overflow the agent model's context window on the next turn — surfacing to the
user as a generic "network error" after the tool appeared to run.

Redesign the tool to give the agent bounded, explorable views instead of one
unbounded dump:

- group_by="host" (default): per-host request and distinct-endpoint counts —
  cheap breadth of egress destinations, no fast-model calls.
- group_by="endpoint": individual (method, path) patterns, optionally with
  host/path_prefix filters and opt-in per-endpoint summaries.
- limit (default 50, hard-capped 100) + offset pagination; limit=0 returns
  counts only. Every result reports totals and how many rows remain.

Summarisation now runs only over the returned page, so the fast-model fan-out
is bounded too. The existing URL normalizer is unchanged; granularity is the
agent's choice via parameters.

As a backstop, the agent loop detects context-length errors from the provider
(Bedrock / Anthropic / OpenAI), shrinks the largest tool result in the
conversation and retries, and degrades to an actionable message instead of a
fatal error if it still cannot fit.

The admin UI surfaces each analyze_traffic call's parameters and result count
so repeated drill-down calls are legible rather than looking redundant.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@greptile-apps

greptile-apps Bot commented Jun 24, 2026 •

Copy link
Copy Markdown

Greptile Summary

This PR fixes context-window overflow in the policy-builder agent by redesigning analyze_traffic into a paginated, bounded tool (host/endpoint modes, limit/offset, hard cap of 100 rows per call) and adding a completeWithContextRecovery backstop that shrinks the largest tool result and retries on context-length errors from all three providers.

  • analyze_traffic now exposes group_by, host, path_prefix, summarize, limit, and offset; summarization runs only over the returned page, bounding the fast-model fan-out to maxAnalyzeLimit calls.
  • ErrContextLength sentinel added in internal/llm/errors.go; each adapter wraps it only after a status-code/exception-type gate (HTTP 400 or Bedrock ValidationException) to prevent rate-limit errors from being misclassified.
  • The admin UI now renders each analyze_traffic call\u2019s scope/paging params and first-line result count, making repeated drill-down calls legible.

Confidence Score: 5/5

Safe to merge; the change is well-scoped, additive, and covered by thorough unit and integration tests.

No defects were found on the changed paths. The context-recovery loop, pagination logic, SQL filter expressions, and provider error classification are all correct and validated by the new test suite. The two observations flagged are minor hardening suggestions that do not affect current behaviour.

internal/llm/errors.go — the "context window" substring match warrants a second look if broader Bedrock error vocabularies are a concern.

Important Files Changed

Filename Overview
internal/builder/agent.go Core redesign: paginated analyze_traffic (host/endpoint modes, limit/offset), context-recovery loop, appendSummary upsert, and supporting helpers — all well-structured with clear invariants.
internal/llm/errors.go Introduces ErrContextLength sentinel and bodyIndicatesContextLength; the "context window" substring is slightly broader than the other three patterns.
internal/llm/bedrock.go Adds isBedrockContextLengthError gated on ValidationException type; double-wrapping via fmt.Errorf("%w: %w") is valid Go 1.20+ and tested.
internal/llm/anthropic.go ErrContextLength wrapped only on HTTP 400, correctly excluding 429 rate-limit errors.
internal/llm/openai.go Same pattern as Anthropic adapter; ErrContextLength restricted to HTTP 400 responses.
internal/admin/pg_audit_reader.go Adds hostFilter and pathPrefix parameters to AggregatePathGroups SQL; correctly uses starts_with instead of LIKE for literal prefix matching.
internal/builder/agent_test.go Comprehensive new tests covering host/endpoint modes, pagination, limit clamping, drill-down, summarization bounding, context-error recovery, and ErrContextLength classification.
web/src/components/PolicyDetail.tsx Surfaces analyze_traffic scope/paging params, first-line result summary, and context_recovery notice in the agent event stream UI.
internal/admin/policies_integration_test.go Updated stub signatures and test fixtures to include full-URL path patterns and new group_by/summarize parameters.

Reviews (6): Last reviewed commit: "fix(llm): classify context overflow by s..." | Re-trigger Greptile

Comment thread internal/admin/pg_audit_reader.go Outdated
Address PR review feedback:

- AggregatePathGroups filtered path_prefix with `LIKE $5 || '%'`, so `_` and `%`
  in the prefix acted as wildcards — e.g. path_prefix "/api/v1/funds_" matched
  "/api/v1/funds/..." (725k rows vs 0 on real data). Use starts_with() for a
  literal prefix match.
- Narrow isContextLengthError's "too long" check to "is too long" so unrelated
  errors are less likely to be misclassified as context overflow.
- Have the test stub honor path_prefix (literal, mirroring starts_with) and add
  a test covering literal path_prefix filtering.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Comment thread internal/builder/agent.go
// request because it exceeded the model's context window. It matches on message
// text to stay provider-agnostic (Bedrock, direct Anthropic, and OpenAI all
// surface this differently), tolerating the surrounding wrapping each adds.
func isContextLengthError(err error) bool {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should use http status code to distinguish the actual errors so we are not unnecessarily categorizing rate-limiting issues along with this, see below

HTTP status codes for context-overflow vs rate-limit

Provider Context too long Rate / token throttle Quota exhausted
Anthropic 400 invalid_request_error — message "prompt is too long" (errors, rate-limits) 429 rate_limit_error (errors, rate-limits) n/a — billing handled separately (errors)
OpenAI 400 invalid_request_error with code: "context_length_exceeded" (cookbook, error-codes) 429 rate_limit_exceeded — "Rate limit reached for requests" (error-codes, cookbook) 429 insufficient_quota — "You exceeded your current quota..."
(error-codes, cookbook)
Bedrock HTTP 200 with stopReason: "model_context_window_exceeded"; 400 ValidationException if the request is malformed (Converse API) 429 ThrottlingException (Converse API, troubleshooting) rolled into ThrottlingException (AWS account quotas)
(troubleshooting)

Addresses PR review (@bjhaid): the agent loop's context-overflow detector
matched on message substrings, and its "token"+"exceed" catch-all would also
match 429 rate-limit errors (e.g. OpenAI "rate_limit_exceeded ... tokens per
minute"). Those got swallowed as context overflow — the user saw misleading
"narrow your request" guidance and all retries were burned against an already
throttled endpoint.

Classify in the adapters, where the HTTP status / exception type is available,
and wrap a shared sentinel (llm.ErrContextLength):

- Anthropic / OpenAI: only HTTP 400 with a context-length message/code.
- Bedrock: only a ValidationException with a too-long message (never
  ThrottlingException).

Status/type is the gate (never 429); the message is the discriminator within it,
since only OpenAI emits a distinct context_length_exceeded code. builder now uses
errors.Is(err, llm.ErrContextLength) instead of string matching. Adds adapter
tests asserting 429s are not classified as context overflow.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@bjhaid
bjhaid merged commit d55d1ab into brexhq:main Jun 25, 2026
4 checks passed
@bjhaid

bjhaid commented Jun 25, 2026

Copy link
Copy Markdown
Contributor

Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants