You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Sep 23, 2026. It is now read-only.
When MCP servers are configured, all tool schemas (names, descriptions, input JSON schemas) are injected into the LLM context at the start of every session. With multiple MCP servers this can consume thousands of tokens of context budget before the user sends a single message.
This leaves less room for conversation history and actual work, and hurts quality on long tasks — especially in sessions that never end up using most of the configured MCP tools.
Lazy-load MCP tool schemas into the LLM context only when tools are actually needed, rather than injecting all schemas at session start. A few approaches:
On-demand discovery: inject a lightweight list_tools stub first; only fetch and inject full schemas when the model determines a tool domain is relevant to the current task
Task-scoped injection: analyse the user's first message and only inject schemas for servers likely to be relevant
Deferred full load: keep schemas out of the initial context and inject them progressively as tools are invoked
Prior art
Claude Code uses a ToolSearch deferred-tool primitive — MCP tool schemas are not loaded into context until explicitly fetched, keeping the base context lean
Users with many MCP servers configured (10+ servers, 50+ tools) can lose a significant fraction of their context window before any work begins. Lazy loading would make Kimi Code CLI viable for heavily MCP-configured environments without sacrificing context budget.
Problem
When MCP servers are configured, all tool schemas (names, descriptions, input JSON schemas) are injected into the LLM context at the start of every session. With multiple MCP servers this can consume thousands of tokens of context budget before the user sends a single message.
This leaves less room for conversation history and actual work, and hurts quality on long tasks — especially in sessions that never end up using most of the configured MCP tools.
Related issues
Proposed solution
Lazy-load MCP tool schemas into the LLM context only when tools are actually needed, rather than injecting all schemas at session start. A few approaches:
list_toolsstub first; only fetch and inject full schemas when the model determines a tool domain is relevant to the current taskPrior art
ToolSearchdeferred-tool primitive — MCP tool schemas are not loaded into context until explicitly fetched, keeping the base context leanImpact
Users with many MCP servers configured (10+ servers, 50+ tools) can lose a significant fraction of their context window before any work begins. Lazy loading would make Kimi Code CLI viable for heavily MCP-configured environments without sacrificing context budget.