Skip to content

feat: Model routing / tiered dispatch for AgentLauncher #382

Description

@jafreck

Summary

Add a ModelRouter interface and per-invocation model selection to AgentLauncher, enabling cost-aware model escalation based on task complexity, retry history, and error patterns.

Motivation

The framework's AgentLauncher uses a single model from BackendRuntimeConfig for all invocations. In practice, agent workloads vary dramatically in complexity — a simple formatting fix can use a cheap/fast model, while a complex architectural refactoring needs the most capable model available.

AAMF implements a tiered routing system:

  • Normal tier — default model for standard tasks
  • Heavy tier — upgraded model for complex tasks or first-retry escalation
  • Critical tier — most capable model for tasks that have failed multiple times

Routing decisions consider: symbol count, dependency depth, file count, estimated complexity, and retry history. This reduces costs significantly — most tasks route to the normal tier while only 10-15% escalate.

The framework should provide the routing mechanism without opinionating on the scoring heuristics, so consumers can plug in domain-specific logic.

Proposed API

ModelRouter Interface

/** Selects a model for a given invocation based on task characteristics and retry history. */
interface ModelRouter {
  /**
   * Select the model to use for this invocation.
   * 
   * @returns The model identifier string (e.g., 'claude-sonnet-4.6', 'gpt-5.1-codex').
   *          Return undefined to use the default model from config.
   */
  selectModel(context: ModelRoutingContext): string | undefined;
}

interface ModelRoutingContext {
  /** The agent being invoked. */
  agent: string;
  /** Which attempt this is (1 = first try). */
  attempt: number;
  /** Maximum attempts configured. */
  maxAttempts: number;
  /** Error message from the previous attempt, if this is a retry. */
  previousError?: string;
  /** Infrastructure error classification from the previous attempt, if available. */
  previousErrorClass?: string;
  /** Arbitrary task metadata provided by the consumer (complexity scores, file counts, etc.). */
  taskMetadata?: Record<string, unknown>;
  /** The default model from config (for reference in routing decisions). */
  defaultModel: string;
}

/** Result of a routing decision, for observability. */
interface ModelRoutingDecision {
  selectedModel: string;
  tier: string;
  reason: string;
  /** Numeric score that drove the decision (for metrics). */
  score?: number;
}

AgentLauncher Integration

interface AgentLauncherConfig {
  // ... existing BackendRuntimeConfig fields ...

  /** Optional model router for per-invocation model selection. */
  modelRouter?: ModelRouter;
}

class AgentLauncher {
  constructor(config: AgentLauncherConfig, logger: BackendLoggerLike);
  
  /**
   * Launch an agent, using the model router (if configured) to select
   * the model for this specific invocation.
   */
  launchAgent(invocation: AgentInvocation, worktreePath: string): Promise<AgentResult>;
}

AgentInvocation Extension

interface AgentInvocation {
  // ... existing fields ...

  /** Model override for this specific invocation. Takes precedence over router. */
  modelOverride?: string;

  /** Arbitrary metadata passed to the model router for routing decisions. */
  routingMetadata?: Record<string, unknown>;
}

AgentResult Extension

interface AgentResult {
  // ... existing fields ...

  /** The model routing decision made for this invocation, if a router was used. */
  routingDecision?: ModelRoutingDecision;
}

Event

interface ModelRoutingEvent {
  type: 'model-routing-decision';
  agent: string;
  taskId?: string;
  selectedModel: string;
  tier: string;
  reason: string;
  score?: number;
}

Built-in Routers (Optional)

Consider providing a simple built-in router as a starting point:

/** Simple tier-based router that escalates on retry. */
export function createEscalationRouter(tiers: {
  normal: string;
  heavy: string;
  critical: string;
  escalateAfterAttempt?: number;  // default: 1 (escalate on first retry)
}): ModelRouter;

Implementation Notes

  • Model selection precedence: invocation.modelOverride > modelRouter.selectModel() > config.agent.model
  • The router is called before the backend's invoke() — the selected model is passed through to the backend
  • routingDecision on AgentResult enables post-run analytics (cost per tier, escalation rate, etc.)
  • The router is synchronous by design — routing decisions should be fast heuristics, not async operations
  • Consumers who don't need routing simply don't set modelRouter — zero overhead for the common case

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions