Skip to content

Commit 991baca

Browse files
committed
LCORE-1675: Documentation for conversation compaction
Document the conversation compaction feature across OpenAPI spec, configuration guide, architecture overview, and query endpoint docs. Add context_status field ("full"/"summarized") to QueryResponse and StreamingQueryResponse documentation. Create comprehensive user guide at docs/user_doc/conversation_compaction.md with configuration examples, behavior details, and FAQ.
1 parent d8e4151 commit 991baca

10 files changed

Lines changed: 333 additions & 20 deletions

File tree

‎docs/README.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -29,6 +29,8 @@ See the full documentation at [`../README.md`](../README.md) or browse sub-pages
2929

3030
[OKP guide](https://lightspeed-core.github.io/lightspeed-stack/user_doc/okp_guide.html)
3131

32+
[Conversation compaction](https://lightspeed-core.github.io/lightspeed-stack/user_doc/conversation_compaction.html)
33+
3234
[Authentication and Authorization](https://lightspeed-core.github.io/lightspeed-stack/user_doc/auth.html)
3335

3436
[User data collection](https://lightspeed-core.github.io/lightspeed-stack/user_doc/user_data_collection.html)

‎docs/design/conversation-compaction/conversation-compaction-spike.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -146,7 +146,7 @@ The "buffer zone" is the most recent turns kept verbatim (not summarized).
146146

147147
Anthropic's compaction uses token-based thresholds throughout — the buffer is implicit (whatever fits after the compaction block).
148148

149-
**Recommendation**: **Hybrid with degrading guard**. Start with the last 4 turns. If their token count exceeds the available budget, degrade to 3, then 2, then 1, then 0. This handles pathological cases where a few large turns (e.g., with tool results) would overflow the context even after summarizing everything else.
149+
**Recommendation**: **Hybrid with degrading guard**. Start with the last 4 turns. If their token count exceeds the available budget, degrade to 3, then 2, then 1, then 0. The token estimate only counts message text — non-message items (tool calls, tool results) are not included — so the guard may underestimate turns that carry large tool-result payloads.
150150

151151
## Decision 10: Concurrency during compaction
152152

‎docs/design/conversation-compaction/conversation-compaction.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -177,7 +177,7 @@ When triggered, split conversation into:
177177
- **Summary zone**: Oldest turns that will be summarized.
178178
- **Buffer zone**: Most recent turns kept verbatim.
179179
180-
Buffer zone uses a degrading guard: start with N turns (default 4), estimate their token count. If they exceed the available budget (context window minus summary minus new query), reduce to N-1 turns and re-estimate. Continue degrading (4→3→2→1→0) until the buffer fits. This handles pathological cases where a few large turns (e.g., with tool results) consume most of the context.
180+
Buffer zone uses a degrading guard: start with N turns (default 4), estimate their token count. If they exceed the available budget (`context_window * buffer_max_ratio`), reduce to N-1 turns and re-estimate. Continue degrading (4→3→2→1→0) until the buffer fits. The token estimate only counts message text — non-message items (tool calls, tool results) are not included — so the guard may underestimate turns that carry large tool-result payloads.
181181

182182
## Additive summarization
183183

‎docs/devel_doc/ARCHITECTURE.md‎

Lines changed: 92 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -35,7 +35,7 @@ To keep requests on-topic and protect sensitive data, LCore applies **safety shi
3535
- **Multi-Provider Support**: Works with multiple LLM providers (Ollama, OpenAI, Watsonx, etc.)
3636
- **Enterprise Security**: Authentication, authorization (RBAC), and secure credential management
3737
- **Resource Management**: Token-based quota limits and usage tracking
38-
- **Conversation Management**: Multi-turn conversations with history and caching
38+
- **Conversation Management**: Multi-turn conversations with history, caching, and automatic compaction
3939
- **RAG Integration**: Retrieval-Augmented Generation for context-aware responses
4040
- **Tool Orchestration**: Model Context Protocol (MCP) server integration
4141
- **Observability**: Prometheus metrics, structured logging, and health checks
@@ -370,6 +370,91 @@ External A2A requests go through LCore's standard authentication system (K8s, RH
370370

371371
---
372372

373+
### 2.11 Conversation Compaction (`utils/compaction.py`, `utils/conversation_compaction.py`)
374+
375+
**Purpose:** Automatically summarize older conversation turns when the conversation history approaches the LLM's context window limit, preventing HTTP 413 failures and enabling arbitrarily long conversations.
376+
377+
**Design Philosophy (Option A):** Once compaction triggers, LCore takes ownership of the context sent to the LLM. The `conversation` parameter is dropped from the OGX call (`omit_conversation=True`), and LCore constructs the input explicitly from summaries + recent turns + new query. The full original history remains in OGX for auditing.
378+
379+
**Architecture:**
380+
381+
The compaction system is split into two layers:
382+
383+
1. **Pure Logic Layer** (`utils/compaction.py`) — Side-effect-free functions:
384+
- `partition_conversation()` — Splits conversation items into old and recent chunks using a *degrading guard*: starts with the configured `buffer_turns` and shrinks one pair at a time until the recent chunk fits the token budget
385+
- `summarize_chunk()` — Single LLM call to produce a `ConversationSummary` from older turns
386+
- `recursively_resummarize()` — Folds multiple accumulated summaries into one when they approach the context limit
387+
388+
2. **Runtime Integration Layer** (`utils/conversation_compaction.py`) — Manages side effects:
389+
- Per-conversation locking (serializes concurrent requests on the same conversation)
390+
- Compaction state loading (cache-preferred with marker fallback)
391+
- Marker persistence (`[lightspeed:compaction-summary]` sentinel in conversation items)
392+
- `CompactionStartedEvent` emission for streaming progress indicators
393+
- `apply_compaction()` (async generator) — Main entry point used by all endpoints
394+
- `store_compacted_turn()` — Appends user query + LLM output when in compacted mode
395+
396+
**Data Flow:**
397+
398+
```text
399+
User Query → Enabled + Context Window Registered?
400+
│
401+
No │ Yes
402+
↓ │ ↓
403+
Pass-through Acquire Lock
404+
↓
405+
Fetch Conversation Items
406+
↓
407+
Load Compaction State
408+
(cache → marker fallback)
409+
↓
410+
Estimate Tokens
411+
(instructions + summaries +
412+
recent items + new query)
413+
↓
414+
Exceeds Threshold?
415+
│
416+
No │ Yes
417+
↓ │ ↓
418+
Skip Partition (old | recent)
419+
│ ↓
420+
│ Summarize Old Chunk (LLM call)
421+
│ ↓
422+
│ Write Marker + Cache Summary
423+
│ ↓
424+
├─→ Recursive Fold (if needed)
425+
↓
426+
Build Explicit Input:
427+
[summaries + recent + query]
428+
↓
429+
Set omit_conversation=True
430+
↓
431+
Release Lock → Continue to LLM
432+
```
433+
434+
**Endpoint Integration:**
435+
436+
| Endpoint | Mode | Cache | `context_status` |
437+
|---|---|---|---|
438+
| `/v1/query` | Blocking (`apply_compaction_blocking()`) | Yes | Yes (`"full"` / `"summarized"`) |
439+
| `/v1/streaming_query` | Streaming (`apply_compaction()` generator) | Yes | Yes (in `end` event) |
440+
| `/v1/responses` | Blocking | Yes | No (OpenAI-compatible, silent) |
441+
| `/a2a` | Blocking, marker-only (no cache) | No | No (A2A protocol scope) |
442+
443+
**Configuration:**
444+
445+
Compaction is controlled by `CompactionConfiguration` in `lightspeed-stack.yaml`:
446+
- `enabled` (default: `false`) — Master switch
447+
- `threshold_ratio` (default: `0.7`) — Fraction of context window that triggers compaction
448+
- `token_floor` (default: `4096`) — Minimum token count before compaction can fire
449+
- `buffer_turns` (default: `4`) — Recent turns kept verbatim
450+
- `buffer_max_ratio` (default: `0.3`) — Max fraction of window for the buffer
451+
452+
Models must have context windows registered via `inference.context_windows` (a map of model ID to token count).
453+
454+
**Concurrency:** A per-conversation lock dictionary serializes concurrent compaction requests on the same conversation. Lock entries are reference-counted and cleaned up when the last waiter exits.
455+
456+
---
457+
373458
## 3. Request Processing Pipeline
374459

375460
This section illustrates how requests flow through LCore from initial receipt to final response.
@@ -400,11 +485,12 @@ Here's how a real query flows through the system:
400485
5. **Model Selection** - Use configured default model (e.g., `meta-llama/Llama-3.1-8B-Instruct`)
401486
6. **Context Building** - Retrieve conversation history, query RAG vector stores for relevant docs, determine available MCP tools
402487
7. **Shield moderation** - LCore-owned direct-run moderation (and agent capabilities where applicable) using shields configured in LCORE config
403-
8. **OGX / agent call** - Send request with system prompt, RAG context, and MCP tools
404-
9. **LLM Processing** - Stack / agent generates response, may invoke MCP tools, returns token counts
405-
10. **Post-Processing** - Generate conversation summary if new
406-
11. **Store Results** - Save to Cache DB, User DB, consume quota, update metrics
407-
12. **Return Response** - Complete LLM response with referenced documents, token usage, and remaining quota
488+
8. **Conversation compaction** - If enabled and estimated tokens exceed the threshold, summarize older turns and rebuild the context (see [Section 2.11](#211-conversation-compaction-utilscompactionpy-utilsconversation_compactionpy))
489+
9. **OGX / agent call** - Send request with system prompt, RAG context, and MCP tools
490+
10. **LLM Processing** - Stack / agent generates response, may invoke MCP tools, returns token counts
491+
11. **Post-Processing** - Generate conversation summary if new
492+
12. **Store Results** - Save to Cache DB, User DB, consume quota, update metrics
493+
13. **Return Response** - Complete LLM response with referenced documents, token usage, and remaining quota
408494

409495
**Key Takeaways:**
410496
- RAG enhances responses with relevant documentation

‎docs/devel_doc/openapi.md‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -2738,7 +2738,7 @@ user's query to a selected OGX LLM and returning the generated response.
27382738
- mcp_headers: Headers that should be passed to MCP servers.
27392739

27402740
### Returns:
2741-
- QueryResponse: Contains the conversation ID and the LLM-generated response.
2741+
- QueryResponse: Contains the conversation ID, the LLM-generated response, and a `context_status` field indicating whether the conversation context is `"full"` or `"summarized"`.
27422742

27432743
### Raises:
27442744
- HTTPException:
@@ -3021,7 +3021,7 @@ content type text/event-stream.
30213021
- mcp_headers: Headers that should be passed to MCP servers.
30223022

30233023
### Returns:
3024-
- SSE-formatted events for the query lifecycle.
3024+
- SSE-formatted events for the query lifecycle. Includes a `context_status` field (`"full"` or `"summarized"`) in the `end` event payload indicating whether conversation compaction was applied. When compaction is triggered, a `compaction` SSE event is emitted before inference begins.
30253025

30263026
### Raises:
30273027
- HTTPException:
@@ -7996,6 +7996,7 @@ Attributes:
79967996
input_tokens: Number of tokens sent to LLM.
79977997
output_tokens: Number of tokens received from LLM.
79987998
available_quotas: Quota available as measured by all configured quota limiters.
7999+
context_status: Indicates whether the conversation context sent to the LLM is "full" (complete history) or "summarized" (older turns were summarized). Present in QueryResponse and in the end event payload (EndEventData) of the streaming query response; omitted from /v1/responses (OpenAI-compatible) and /a2a responses.
79998000

80008001

80018002
| Field | Type | Description |
@@ -8010,6 +8011,7 @@ Attributes:
80108011
| available_quotas | object | Quota available as measured by all configured quota limiters |
80118012
| tool_calls | array | List of tool calls made during response generation |
80128013
| tool_results | array | List of tool results |
8014+
| context_status | string | Indicates whether the conversation context sent to the LLM is `"full"` (complete history) or `"summarized"` (older turns were summarized). Present in QueryResponse and in the `end` event payload (EndEventData) of the streaming query response; omitted from `/v1/responses` (OpenAI-compatible) and `/a2a` responses. |
80138015

80148016

80158017
## QuotaExceededResponse

‎docs/devel_doc/query_endpoint.md‎

Lines changed: 9 additions & 6 deletions
Original file line numberDiff line numberDiff line change
@@ -145,6 +145,7 @@ The optional `solr` field configures Solr inline RAG behavior:
145145
| `tool_results` | array[object] | `[]` | Tool call results |
146146
| `rag_chunks` | array[object] | `[]` | *(Deprecated)* RAG chunks used |
147147
| `truncated` | boolean | `false` | *(Deprecated)* Always `false` |
148+
| `context_status` | string | `"full"` | Whether the conversation context is `"full"` (complete history) or `"summarized"` (older turns were summarized via conversation compaction) |
148149

149150
**`referenced_documents` items:**
150151

@@ -242,7 +243,7 @@ Emitted when the full response is assembled.
242243

243244
#### 7. `end`
244245

245-
Emitted last on success. Contains metadata.
246+
Emitted last on success. Contains metadata including `context_status` (`"full"` or `"summarized"`).
246247

247248
```json
248249
{
@@ -251,7 +252,8 @@ Emitted last on success. Contains metadata.
251252
"referenced_documents": [],
252253
"truncated": null,
253254
"input_tokens": 11,
254-
"output_tokens": 19
255+
"output_tokens": 19,
256+
"context_status": "full"
255257
},
256258
"available_quotas": {"UserQuotaLimiter": 998911}
257259
}
@@ -327,9 +329,9 @@ Both endpoints share the same pre-processing pipeline:
327329
11. Prepare Responses API parameters (model, system prompt, tools, MCP headers)
328330
12. Extract image attachments separately for multimodal input construction
329331

330-
**`/v1/query` then:** applies conversation compaction (blocking), calls the LLM, generates topic summary, consumes tokens, stores results, returns JSON.
332+
**`/v1/query` then:** applies conversation compaction (blocking), calls the LLM, generates topic summary, consumes tokens, stores results, returns JSON. When compaction is applied, the response includes `context_status: "summarized"`; otherwise `context_status: "full"`.
331333

332-
**`/v1/streaming_query` then:** generates a `request_id`, starts the SSE stream, emits events as the LLM generates tokens, performs post-stream cleanup (topic summary, token consumption, persistence).
334+
**`/v1/streaming_query` then:** generates a `request_id`, starts the SSE stream, applies compaction if needed (emitting a `compaction` SSE event), emits events as the LLM generates tokens, performs post-stream cleanup (topic summary, token consumption, persistence). The `end` event includes `context_status` indicating whether compaction was applied.
333335

334336
---
335337

@@ -409,7 +411,8 @@ curl -X POST http://localhost:8090/v1/query \
409411
"tool_calls": [],
410412
"tool_results": [],
411413
"rag_chunks": [],
412-
"truncated": false
414+
"truncated": false,
415+
"context_status": "full"
413416
}
414417
```
415418

@@ -500,7 +503,7 @@ data: {"event": "token", "data": {"id": 2, "token": " an"}}
500503
501504
data: {"event": "turn_complete", "data": {"id": 50, "token": "Kubernetes is an open-source..."}}
502505
503-
data: {"event": "end", "data": {"referenced_documents": [], "truncated": null, "input_tokens": 11, "output_tokens": 50}, "available_quotas": {"UserQuotaLimiter": 998950}}
506+
data: {"event": "end", "data": {"referenced_documents": [], "truncated": null, "input_tokens": 11, "output_tokens": 50, "context_status": "full"}, "available_quotas": {"UserQuotaLimiter": 998950}}
504507
```
505508

506509
### Streaming Query Interrupt

‎docs/index.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -34,6 +34,8 @@ product questions using backend LLM services, agents, and RAG databases.
3434

3535
[OKP guide](https://lightspeed-core.github.io/lightspeed-stack/user_doc/okp_guide.html)
3636

37+
[Conversation compaction](https://lightspeed-core.github.io/lightspeed-stack/user_doc/conversation_compaction.html)
38+
3739
[Authentication and Authorization](https://lightspeed-core.github.io/lightspeed-stack/user_doc/auth.html)
3840

3941
[User data collection](https://lightspeed-core.github.io/lightspeed-stack/user_doc/user_data_collection.html)

‎docs/user_doc/config.md‎

Lines changed: 45 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -214,6 +214,51 @@ Attributes:
214214
| buffer_turns | integer | Number of recent turns to keep verbatim. |
215215
| buffer_max_ratio | number | Maximum fraction of context window the buffer zone can occupy, regardless of buffer_turns. |
216216

217+
### How to enable conversation compaction
218+
219+
Compaction is disabled by default. To enable it, add a `compaction` section to your `lightspeed-stack.yaml` and set `enabled: true`. You must also register context window sizes for the models you use via the `inference.context_windows` map so the compaction trigger can calculate when older turns should be summarized.
220+
221+
**Minimal configuration:**
222+
223+
```yaml
224+
inference:
225+
default_provider: openai
226+
default_model: gpt-4o-mini
227+
context_windows:
228+
openai/gpt-4o-mini: 128000
229+
230+
compaction:
231+
enabled: true
232+
```
233+
234+
**Full configuration with all options:**
235+
236+
```yaml
237+
inference:
238+
default_provider: openai
239+
default_model: gpt-4o-mini
240+
context_windows:
241+
openai/gpt-4o-mini: 128000
242+
openai/gpt-4o: 128000
243+
244+
compaction:
245+
enabled: true
246+
threshold_ratio: 0.7 # trigger at 70% of context window (default)
247+
token_floor: 4096 # minimum tokens before compaction can fire (default)
248+
buffer_turns: 4 # recent turns kept verbatim (default)
249+
buffer_max_ratio: 0.3 # buffer may use at most 30% of the window (default)
250+
```
251+
252+
**Key considerations:**
253+
254+
- `context_windows` is required. Models absent from this map have no registered window and compaction will not trigger for them.
255+
- `threshold_ratio` controls how aggressively compaction fires. Lower values compact sooner; higher values wait longer (closer to the window limit).
256+
- `buffer_turns` sets how many recent user/assistant turn pairs are kept in full. A degrading guard automatically reduces this if the buffer itself would exceed `buffer_max_ratio` of the window.
257+
- `token_floor` prevents compaction from triggering on very short conversations.
258+
- When compaction is disabled (the default), requests that exceed the context window surface as HTTP 413.
259+
260+
For a comprehensive explanation of the feature, see the [Conversation Compaction Guide](conversation_compaction.md).
261+
217262

218263
## Configuration
219264

0 commit comments

Comments
 (0)