Skip to content

test(adapters): pin the outbound thinking contract — frognano #399 exonerated (root cause server-side) - #438

Merged
jsboige merged 1 commit into
mainfrom
fix/399-frognano-exoneration
Oct 11, 2026
Merged

jsboige merged 1 commit into
mainfrom
fix/399-frognano-exoneration

Conversation

@jsboige

@jsboige jsboige commented Oct 11, 2026

Copy link
Copy Markdown
Owner

What

Investigation of #399 (frognano-4b via hub thinks; direct :5003 never) — grain 2 of the 11/10 deepqueue. The coordinator's hypothesis (the Qwen openai-wire passthrough fixed by #437) is refuted by measurement; the claudish side is exonerated and the locus is server-side. This PR carries the exoneration as pins.

Measured chain (2026-10-11)

  1. Paired inbound capture req-1-11725 (07/10 19:39:33Z, extracted from the GDrive archive) is minimal: frognano-4b, max_tokens:16, one user text block — no thinking, no tools, no system.
  2. Outbound body captured through the real pipeline (local proxy at main, sandboxed config, echo endpoint): byte-identical 198 bytes at 6f732eb9 (main on 07/10 — the incident era) and af52be5f (current): {model, messages, temperature:1, stream:true, stream_options, max_tokens} — no thinking field of any spelling.
  3. Dialect match refuted: matchesModelFamily("frognano-4b", "qwen"|"alibaba") = false (startsWith | "/fam"), so QwenModelDialect.prepareRequest — the only adapter whose passthrough could emit enable_thinking — never ran for frognano. The feat(ingress): map enable_thinking/reasoning_effort onto the thinking object #437 fix cannot be what cured this.
  4. Non-reproduction via the hub (11/10, cluster key, lane locale): the exact r11725 body → end_turn, "ok", 2 tokens; the 26-tools consumer body req-1-6396 (07/10 16:54Z, which then produced 4.4–8.9k chars of reasoning) → tool_use, 63 tokens. Zero reasoning both. Same outbound body both eras ⇒ the delta is the :5003 server deployment (refs vllm#70/fix(ops): the watchdog takes its .claudish home as a parameter (⚠ GO user requis avant merge) #71 — the vllm workspace redeployed in the window).

The pins (new openai-api-format.test.ts)

  • The r11725 body builds to exactly the six neutral fields — no enable_thinking / reasoning_effort / thinking / thinking_budget / chat_template_kwargs.
  • An inbound enabled thinking block does not become reasoning_effort off the o1/o3 lane (isReasoningModel() stays the only emission gate).
  • o1 keeps the budget→effort mapping (positive control — the gate stays open where it belongs).

3/3 green. No production code changed — there was nothing to fix here.

Note on closure

Root cause lives in the vllm workspace (which server-side change cured it = refs vllm#70/#71); a direct :5003 replay bisect was requested from po-2025 (DM msg-20261011T035636) to pin the server-side grain if wanted. The claudiss-side question of #399 ("un champ/forme du body sortant du hub non testé bascule le template") is answered: no.

Refs #399

🤖 Generated with Claude Code

…onerated

#399 measured (07/10, hub + vLLM frognano-4b :5003) reasoning on every
hub-served request while direct :5003 never thinks. Investigation
(2026-10-11), all measured:

- The paired inbound capture req-1-11725 (19:39:33Z) is MINIMAL — no
  thinking, no tools, no system. The outbound body built for it was
  captured through the real pipeline via an echo endpoint: byte-identical
  198 bytes at 6f732eb (main on 07/10) and af52be5 (current) — model,
  messages, temperature, stream, stream_options, max_tokens. NO thinking
  field of any spelling.
- frognano-4b matches neither qwen nor alibaba (matchesModelFamily =
  startsWith | "/fam"), so QwenModelDialect — whose passthrough could
  emit enable_thinking — never ran for it. The #437-passthrough
  hypothesis is refuted for this issue.
- The flip no longer reproduces via the hub (11/10): the exact r11725
  body AND a 26-tools consumer body (req-1-6396) both answer immediately
  with no reasoning. Same body both eras => the change was server-side
  (:5003 deployment, refs vllm#70/#71).

Pins: the six neutral fields exactly; no thinking field even with an
inbound enabled thinking block off the o1/o3 lane; o1 keeps the
budget->reasoning_effort mapping (the only emission site).

Refs #399

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@jsboige

jsboige commented Oct 11, 2026

Copy link
Copy Markdown
Owner Author

myia-ai-01 : review de 79b49578. APPROUVÉ, je merge. Lu : body, diff (1 fichier de test, +89, aucun code de prod) et commits (aucun mot-clé de fermeture ; Refs #399 dans le body). Mesuré ici : 3/3 verts sur bun 1.3.13. Mutation : j'ai ouvert la porte isReasoningModel() de buildPayload (openai-api-format.ts:116, claudeRequest.thinking && this.isReasoningModel() → claudeRequest.thinking), et le pin 2 passe rouge ("an inbound thinking block does NOT become reasoning_effort off the o1/o3 lane"). Le pin o1 sert de contrôle positif. Le fichier a été restauré depuis une sauvegarde (diff vide). La réfutation de l'hypothèse #437 que j'avais avancée tient : matchesModelFamily ne reconnaît pas frognano comme qwen, et le body sortant est identique octet pour octet aux deux époques. Le locus serveur :5003 reste à relier à vllm#70/#71 ; c'est le replay bisect demandé à po-2025.

@jsboige
jsboige merged commit 8e4bbc3 into main Oct 11, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant