You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
LCORE-1675: Documentation for conversation compaction
Document the conversation compaction feature across OpenAPI spec,
configuration guide, architecture overview, and query endpoint docs.
Add context_status field ("full"/"summarized") to QueryResponse and
StreamingQueryResponse documentation. Create comprehensive user guide
at docs/user_doc/conversation_compaction.md with configuration
examples, behavior details, and FAQ.
Copy file name to clipboardExpand all lines: docs/design/conversation-compaction/conversation-compaction-spike.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -146,7 +146,7 @@ The "buffer zone" is the most recent turns kept verbatim (not summarized).
146
146
147
147
Anthropic's compaction uses token-based thresholds throughout — the buffer is implicit (whatever fits after the compaction block).
148
148
149
-
**Recommendation**: **Hybrid with degrading guard**. Start with the last 4 turns. If their token count exceeds the available budget, degrade to 3, then 2, then 1, then 0. This handles pathological cases where a few large turns (e.g., with tool results) would overflow the context even after summarizing everything else.
149
+
**Recommendation**: **Hybrid with degrading guard**. Start with the last 4 turns. If their token count exceeds the available budget, degrade to 3, then 2, then 1, then 0. The token estimate only counts message text — non-message items (tool calls, tool results) are not included — so the guard may underestimate turns that carry large tool-result payloads.
Copy file name to clipboardExpand all lines: docs/design/conversation-compaction/conversation-compaction.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -177,7 +177,7 @@ When triggered, split conversation into:
177
177
- **Summary zone**: Oldest turns that will be summarized.
178
178
- **Buffer zone**: Most recent turns kept verbatim.
179
179
180
-
Buffer zone uses a degrading guard: start with N turns (default 4), estimate their token count. If they exceed the available budget (context window minus summary minus new query), reduce to N-1 turns and re-estimate. Continue degrading (4→3→2→1→0) until the buffer fits. This handles pathological cases where a few large turns (e.g., with tool results) consume most of the context.
180
+
Buffer zone uses a degrading guard: start with N turns (default 4), estimate their token count. If they exceed the available budget (`context_window * buffer_max_ratio`), reduce to N-1 turns and re-estimate. Continue degrading (4→3→2→1→0) until the buffer fits. The token estimate only counts message text — non-message items (tool calls, tool results) are not included — so the guard may underestimate turns that carry large tool-result payloads.
**Purpose:** Automatically summarize older conversation turns when the conversation history approaches the LLM's context window limit, preventing HTTP 413 failures and enabling arbitrarily long conversations.
376
+
377
+
**Design Philosophy (Option A):** Once compaction triggers, LCore takes ownership of the context sent to the LLM. The `conversation` parameter is dropped from the OGX call (`omit_conversation=True`), and LCore constructs the input explicitly from summaries + recent turns + new query. The full original history remains in OGX for auditing.
-`partition_conversation()` — Splits conversation items into old and recent chunks using a *degrading guard*: starts with the configured `buffer_turns` and shrinks one pair at a time until the recent chunk fits the token budget
385
+
-`summarize_chunk()` — Single LLM call to produce a `ConversationSummary` from older turns
386
+
-`recursively_resummarize()` — Folds multiple accumulated summaries into one when they approach the context limit
387
+
388
+
2.**Runtime Integration Layer** (`utils/conversation_compaction.py`) — Manages side effects:
389
+
- Per-conversation locking (serializes concurrent requests on the same conversation)
390
+
- Compaction state loading (cache-preferred with marker fallback)
391
+
- Marker persistence (`[lightspeed:compaction-summary]` sentinel in conversation items)
392
+
-`CompactionStartedEvent` emission for streaming progress indicators
393
+
-`apply_compaction()` (async generator) — Main entry point used by all endpoints
394
+
-`store_compacted_turn()` — Appends user query + LLM output when in compacted mode
-`buffer_max_ratio` (default: `0.3`) — Max fraction of window for the buffer
451
+
452
+
Models must have context windows registered via `inference.context_windows` (a map of model ID to token count).
453
+
454
+
**Concurrency:** A per-conversation lock dictionary serializes concurrent compaction requests on the same conversation. Lock entries are reference-counted and cleaned up when the last waiter exits.
455
+
456
+
---
457
+
373
458
## 3. Request Processing Pipeline
374
459
375
460
This section illustrates how requests flow through LCore from initial receipt to final response.
@@ -400,11 +485,12 @@ Here's how a real query flows through the system:
400
485
5.**Model Selection** - Use configured default model (e.g., `meta-llama/Llama-3.1-8B-Instruct`)
401
486
6.**Context Building** - Retrieve conversation history, query RAG vector stores for relevant docs, determine available MCP tools
402
487
7.**Shield moderation** - LCore-owned direct-run moderation (and agent capabilities where applicable) using shields configured in LCORE config
403
-
8.**OGX / agent call** - Send request with system prompt, RAG context, and MCP tools
10.**Post-Processing** - Generate conversation summary if new
406
-
11.**Store Results** - Save to Cache DB, User DB, consume quota, update metrics
407
-
12.**Return Response** - Complete LLM response with referenced documents, token usage, and remaining quota
488
+
8.**Conversation compaction** - If enabled and estimated tokens exceed the threshold, summarize older turns and rebuild the context (see [Section 2.11](#211-conversation-compaction-utilscompactionpy-utilsconversation_compactionpy))
489
+
9.**OGX / agent call** - Send request with system prompt, RAG context, and MCP tools
Copy file name to clipboardExpand all lines: docs/devel_doc/openapi.md
+4-2Lines changed: 4 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -2738,7 +2738,7 @@ user's query to a selected OGX LLM and returning the generated response.
2738
2738
- mcp_headers: Headers that should be passed to MCP servers.
2739
2739
2740
2740
### Returns:
2741
-
- QueryResponse: Contains the conversation ID and the LLM-generated response.
2741
+
- QueryResponse: Contains the conversation ID, the LLM-generated response, and a `context_status` field indicating whether the conversation context is `"full"` or `"summarized"`.
2742
2742
2743
2743
### Raises:
2744
2744
- HTTPException:
@@ -3021,7 +3021,7 @@ content type text/event-stream.
3021
3021
- mcp_headers: Headers that should be passed to MCP servers.
3022
3022
3023
3023
### Returns:
3024
-
- SSE-formatted events for the query lifecycle.
3024
+
- SSE-formatted events for the query lifecycle. Includes a `context_status` field (`"full"` or `"summarized"`) in the `end` event payload indicating whether conversation compaction was applied. When compaction is triggered, a `compaction` SSE event is emitted before inference begins.
3025
3025
3026
3026
### Raises:
3027
3027
- HTTPException:
@@ -7996,6 +7996,7 @@ Attributes:
7996
7996
input_tokens: Number of tokens sent to LLM.
7997
7997
output_tokens: Number of tokens received from LLM.
7998
7998
available_quotas: Quota available as measured by all configured quota limiters.
7999
+
context_status: Indicates whether the conversation context sent to the LLM is "full" (complete history) or "summarized" (older turns were summarized). Present in QueryResponse and in the end event payload (EndEventData) of the streaming query response; omitted from /v1/responses (OpenAI-compatible) and /a2a responses.
7999
8000
8000
8001
8001
8002
| Field | Type | Description |
@@ -8010,6 +8011,7 @@ Attributes:
8010
8011
| available_quotas | object | Quota available as measured by all configured quota limiters |
8011
8012
| tool_calls | array | List of tool calls made during response generation |
8012
8013
| tool_results | array | List of tool results |
8014
+
| context_status | string | Indicates whether the conversation context sent to the LLM is `"full"` (complete history) or `"summarized"` (older turns were summarized). Present in QueryResponse and in the `end` event payload (EndEventData) of the streaming query response; omitted from `/v1/responses` (OpenAI-compatible) and `/a2a` responses. |
|`context_status`| string |`"full"`| Whether the conversation context is `"full"` (complete history) or `"summarized"` (older turns were summarized via conversation compaction) |
148
149
149
150
**`referenced_documents` items:**
150
151
@@ -242,7 +243,7 @@ Emitted when the full response is assembled.
242
243
243
244
#### 7. `end`
244
245
245
-
Emitted last on success. Contains metadata.
246
+
Emitted last on success. Contains metadata including `context_status` (`"full"` or `"summarized"`).
246
247
247
248
```json
248
249
{
@@ -251,7 +252,8 @@ Emitted last on success. Contains metadata.
251
252
"referenced_documents": [],
252
253
"truncated": null,
253
254
"input_tokens": 11,
254
-
"output_tokens": 19
255
+
"output_tokens": 19,
256
+
"context_status": "full"
255
257
},
256
258
"available_quotas": {"UserQuotaLimiter": 998911}
257
259
}
@@ -327,9 +329,9 @@ Both endpoints share the same pre-processing pipeline:
327
329
11. Prepare Responses API parameters (model, system prompt, tools, MCP headers)
328
330
12. Extract image attachments separately for multimodal input construction
**`/v1/query` then:** applies conversation compaction (blocking), calls the LLM, generates topic summary, consumes tokens, stores results, returns JSON. When compaction is applied, the response includes `context_status: "summarized"`; otherwise `context_status: "full"`.
331
333
332
-
**`/v1/streaming_query` then:** generates a `request_id`, starts the SSE stream, emits events as the LLM generates tokens, performs post-stream cleanup (topic summary, token consumption, persistence).
334
+
**`/v1/streaming_query` then:** generates a `request_id`, starts the SSE stream, applies compaction if needed (emitting a `compaction` SSE event), emits events as the LLM generates tokens, performs post-stream cleanup (topic summary, token consumption, persistence). The `end` event includes `context_status` indicating whether compaction was applied.
333
335
334
336
---
335
337
@@ -409,7 +411,8 @@ curl -X POST http://localhost:8090/v1/query \
Copy file name to clipboardExpand all lines: docs/user_doc/config.md
+45Lines changed: 45 additions & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -214,6 +214,51 @@ Attributes:
214
214
| buffer_turns | integer | Number of recent turns to keep verbatim. |
215
215
| buffer_max_ratio | number | Maximum fraction of context window the buffer zone can occupy, regardless of buffer_turns. |
216
216
217
+
### How to enable conversation compaction
218
+
219
+
Compaction is disabled by default. To enable it, add a `compaction` section to your `lightspeed-stack.yaml` and set `enabled: true`. You must also register context window sizes for the models you use via the `inference.context_windows` map so the compaction trigger can calculate when older turns should be summarized.
220
+
221
+
**Minimal configuration:**
222
+
223
+
```yaml
224
+
inference:
225
+
default_provider: openai
226
+
default_model: gpt-4o-mini
227
+
context_windows:
228
+
openai/gpt-4o-mini: 128000
229
+
230
+
compaction:
231
+
enabled: true
232
+
```
233
+
234
+
**Full configuration with all options:**
235
+
236
+
```yaml
237
+
inference:
238
+
default_provider: openai
239
+
default_model: gpt-4o-mini
240
+
context_windows:
241
+
openai/gpt-4o-mini: 128000
242
+
openai/gpt-4o: 128000
243
+
244
+
compaction:
245
+
enabled: true
246
+
threshold_ratio: 0.7# trigger at 70% of context window (default)
247
+
token_floor: 4096# minimum tokens before compaction can fire (default)
buffer_max_ratio: 0.3# buffer may use at most 30% of the window (default)
250
+
```
251
+
252
+
**Key considerations:**
253
+
254
+
- `context_windows` is required. Models absent from this map have no registered window and compaction will not trigger for them.
255
+
- `threshold_ratio`controls how aggressively compaction fires. Lower values compact sooner; higher values wait longer (closer to the window limit).
256
+
- `buffer_turns`sets how many recent user/assistant turn pairs are kept in full. A degrading guard automatically reduces this if the buffer itself would exceed `buffer_max_ratio` of the window.
257
+
- `token_floor`prevents compaction from triggering on very short conversations.
258
+
- When compaction is disabled (the default), requests that exceed the context window surface as HTTP 413.
259
+
260
+
For a comprehensive explanation of the feature, see the [Conversation Compaction Guide](conversation_compaction.md).
0 commit comments