An AI-enabled router with persistent memory. Register any A2A-compatible agent in one YAML file and MBM routes requests to the right one — automatically.
MBM is a master AI agent that:
- Routes incoming requests to the best-fit registered agent using LLM-based intent matching.
- Self-serves when no external agents are configured — answering directly with its own LLM and your local knowledge base.
- Remembers — maintains a local markdown-based knowledge base that grows with your interactions.
You build your AI agents. You add them to MBM's config. MBM handles the routing.
flowchart LR
User([User Request])
Config[/Config: routing rules/]
subgraph MBM
Master[MBM Master / Master AI Agent\nRouting + Retrieval + Fallback]
end
subgraph Workers[Worker Agents - user supplied]
PA[Personalized Agents]
EA[Extended Agents]
end
subgraph Stores[Context Stores]
Raw[(User Context Raw DB)]
Vec[(Common Vector Storage)]
Graph[(Graph Database)]
end
Embedder[Embedder / Embedding Model]
Classifier[Classifier\nperiodic run]
%% Request + routing
User --> Master
Config --> Master
%% Retrieval (read path - the fix)
Vec -- semantic context --> Master
Graph -- relational context --> Master
%% Routing to workers, with fallback
Master -- "if worker matches (context passed)" --> PA
Master -- "if worker matches (context passed)" --> EA
PA -- result --> Master
EA -- result --> Master
Master -- "if NO worker: serve directly" --> User
Master -- response --> User
%% Ingestion (write path)
Master -- raw context --> Raw
Raw --> Embedder
Embedder -- create embeddings --> Vec
Raw --> Classifier
Classifier -- entities + relationships --> Graph
%% Optional: persist outputs back as new context
PA -. persist output .-> Raw
EA -. persist output .-> Raw
| Component | Role |
|---|---|
| MBM Master (Master AI Agent) | Entry point. Owns routing, config, context retrieval, and Worker invocation. Falls back to serving the request itself when no Worker matches. |
| Config | User-provided. Declares which Worker agent handles which kind of request. Drives the router. |
| Worker Agents | AI agents the user brings. Split (in the original drawing) into Personalized Agents and Extended Agents. See open question 2. |
| User Context Raw (DB) | Source of truth for raw, un-embedded context. |
| Embedder (Embedding Model) | Turns raw context into vectors. |
| Common Vector Storage | Semantic recall. Queried at retrieval time. |
| Classifier (periodic run) | Reads raw/vector context on a schedule, extracts entities and relationships. |
| Graph Database | Stores the structured entities/relationships the Classifier produces. Queried at retrieval time for relational context. |
git clone https://github.com/yourusername/mbm
cd mbm
go build -o mbm ./cmd/mbm
sudo mv mbm /usr/local/bin/ # optionalmbm initCreates ~/.mbm/config.yaml with sensible defaults.
export OPENAI_API_KEY=sk-...
# or set llm.api_key in ~/.mbm/config.yamlMBM works with any OpenAI-compatible API — OpenAI, Anthropic (via proxy), Ollama, LiteLLM.
mbm agentsMBM A2A Multi-Agent System
══════════════════════════════════════════
Master → http://localhost:8090
Memory → http://localhost:8091
Chat → http://localhost:8092
══════════════════════════════════════════
curl -s -X POST http://localhost:8090/ \
-H "Content-Type: application/json" \
-d '{
"jsonrpc": "2.0",
"id": "1",
"method": "tasks/send",
"params": {
"id": "task-01",
"message": {
"role": "user",
"parts": [{"type": "text", "text": "What do I know about Go concurrency?"}]
}
}
}'Edit ~/.mbm/config.yaml and add entries under agent_routes:
agent_routes:
- name: coding-assistant
url: http://localhost:9001
description: >
Expert at code review, debugging, refactoring, and architecture decisions.
Handles questions about any programming language.
tags: [code, review, debug, refactor, architecture]
enabled: true
auto_discover: true # fetches /.well-known/agent.json at startup to enrich capabilities
- name: research-agent
url: https://research.mycompany.internal
description: Searches internal docs, summarises papers, answers research questions.
tags: [research, docs, papers, analysis]
enabled: true
auth:
type: bearer
token: ${RESEARCH_AGENT_TOKEN} # resolved from environment at runtimeRestart MBM after editing the config. At startup the master:
- Loads all enabled routes.
- Optionally fetches each agent's card (
auto_discover: true) to enrich description and tags automatically. - Logs each registered agent.
Every incoming request is then routed to the best-fit agent via an LLM routing call.
MBM supports three authentication schemes for external agents:
| Type | Config keys | Header sent |
|---|---|---|
bearer |
token |
Authorization: Bearer <token> |
api_key |
key, header (optional, default X-API-Key) |
X-API-Key: <key> |
basic |
username, password |
Authorization: Basic <b64> |
All credential values support ${ENV_VAR} expansion — secrets never live in plaintext config.
auth:
type: api_key
header: X-Custom-Auth
key: ${MY_SECRET_KEY}User message
│
▼
LLM routing call (lightweight — returns JSON {"agent": "name"})
│
├─ Agent matched → forward task via A2A (tasks/send or tasks/sendSubscribe)
│
└─ No match / empty routes → self-serve
│
├─ Search local memory for relevant context
└─ Stream LLM reply with memory context injected
The routing call is fast — it asks the LLM to return a single JSON object. If the response is malformed, MBM retries up to 2 times then falls back to self-serve.
When agent_routes is empty or no configured agent matches the request:
- Searches the local knowledge base (keyword match across your
.mdfiles). - Injects relevant excerpts as context.
- Streams the LLM reply.
Build your knowledge base:
mbm add golang "Goroutines are lightweight threads managed by the Go runtime..."
mbm add rust "Ownership rules: each value has exactly one owner..."
mbm search "concurrency"# ── LLM ──────────────────────────────────────────────────────────────────────
llm:
provider: openai # openai | ollama | local
model: gpt-4-turbo
api_key: "" # or set OPENAI_API_KEY env var
base_url: "" # leave empty for OpenAI; set for Ollama / LiteLLM
temperature: 0.7
max_tokens: 2000
# ── Memory (local knowledge base) ────────────────────────────────────────────
memory:
folder: ~/.mbm/memory/topics
watch: true
auto_tag: true
# ── Agent routes ──────────────────────────────────────────────────────────────
# Register external A2A agents here.
# Leave empty → MBM self-serves all requests.
agent_routes:
- name: my-agent
url: http://localhost:9001
description: "What this agent is good at (used by the routing LLM)"
tags: [tag1, tag2]
enabled: true
auto_discover: true # fetch /.well-known/agent.json on startup
auth:
type: bearer # bearer | api_key | basic
token: ${MY_TOKEN}MBM implements the Agent-to-Agent (A2A) protocol. Any A2A-compliant service can be registered as a route.
| Method | Path | Description |
|---|---|---|
GET |
/.well-known/agent.json |
Agent card — name, description, skills, capabilities |
POST |
/ |
JSON-RPC 2.0 task endpoint |
| Method | Description |
|---|---|
tasks/send |
Send a task, wait for completed result |
tasks/sendSubscribe |
Send a task, receive SSE stream of chunks |
{
"jsonrpc": "2.0",
"id": "req-1",
"method": "tasks/send",
"params": {
"id": "task-id",
"message": {
"role": "user",
"parts": [
{"type": "text", "text": "Your question or request"}
]
}
}
}Use method: "tasks/sendSubscribe". The response is Content-Type: text/event-stream with events:
TaskStatusUpdateEvent— state transitions (working,completed,failed)TaskArtifactUpdateEvent— response content chunks (stream these to the user)
mbm init # Initialise ~/.mbm/ config and memory directory
mbm agents # Start master + built-in sub-agents (A2A mode)
mbm server # Start legacy REST API (port 8542)
mbm chat # Interactive TUI chat
# Memory management
mbm add <topic> <content> # Add a memory topic
mbm get <topic> # Retrieve a topic
mbm list # List all topics
mbm search <query> # Keyword search
mbm delete <topic> # Delete a topic
mbm context # Export all memory as markdown
mbm config # Show current configuration# Port overrides
mbm agents --master-port 8090 --memory-port 8091 --chat-port 8092Any HTTP service that speaks A2A can be registered. The minimum required surface:
GET /.well-known/agent.json
POST / (JSON-RPC 2.0)
tasks/send
tasks/sendSubscribe (optional, for streaming)
Agent card shape:
{
"name": "my-agent",
"description": "What this agent does",
"url": "http://localhost:9001",
"version": "1.0.0",
"capabilities": {"streaming": true},
"skills": [
{"id": "skill.id", "name": "Skill name", "description": "...", "tags": ["tag"]}
]
}See internal/agents/memory/agent.go and internal/agents/chat/agent.go for reference Go implementations.
cmd/mbm/
main.go CLI entrypoint (Cobra)
commands/
agents.go mbm agents — starts all servers
chat.go mbm chat — Bubbletea TUI
server.go mbm server — REST API (legacy)
add|get|list|... Memory management
internal/
a2a/
types.go A2A protocol types
server.go Generic A2A HTTP server
client.go A2A HTTP client + auth
registry.go In-process agent registry
agents/
master/agent.go ★ Config-driven router + self-serve fallback
memory/agent.go Built-in memory CRUD agent
chat/agent.go Built-in LLM chat agent
application/service/
memory_service.go Markdown CRUD + search
infrastructure/
llm/openai.go OpenAI-compatible client
llm/tools.go ChatWithTools (function calling)
server/server.go Legacy REST API
ui/ Bubbletea TUI
pkg/
config/config.go Viper-based config (includes agent_routes)
logger/logger.go
| Component | Technology |
|---|---|
| Language | Go 1.25+ |
| TUI | Bubbletea + Lip Gloss |
| Markdown | Goldmark |
| Config | Viper |
| CLI | Cobra |
| Protocol | A2A (JSON-RPC 2.0 + SSE) |
- Vector search (Qdrant / pgvector)
- Hierarchical memory tiers (STM → MTM → LTM)
- Graph-based entity relationships between topics
- Agent health-check and circuit breaker
-
mbm agents add <url>— live registration without restart - Web UI overhaul
- Shell completions (bash, zsh, fish)
- Fork and clone.
- Create a feature branch.
go build ./... && go vet ./...- Open a PR with a description of what and why.
Issues and discussions welcome.