Repository navigation
test(adapters): pin the outbound thinking contract — frognano #399 exonerated (root cause server-side) - #438
Conversation
…onerated #399 measured (07/10, hub + vLLM frognano-4b :5003) reasoning on every hub-served request while direct :5003 never thinks. Investigation (2026-10-11), all measured: - The paired inbound capture req-1-11725 (19:39:33Z) is MINIMAL — no thinking, no tools, no system. The outbound body built for it was captured through the real pipeline via an echo endpoint: byte-identical 198 bytes at 6f732eb (main on 07/10) and af52be5 (current) — model, messages, temperature, stream, stream_options, max_tokens. NO thinking field of any spelling. - frognano-4b matches neither qwen nor alibaba (matchesModelFamily = startsWith | "/fam"), so QwenModelDialect — whose passthrough could emit enable_thinking — never ran for it. The #437-passthrough hypothesis is refuted for this issue. - The flip no longer reproduces via the hub (11/10): the exact r11725 body AND a 26-tools consumer body (req-1-6396) both answer immediately with no reasoning. Same body both eras => the change was server-side (:5003 deployment, refs vllm#70/#71). Pins: the six neutral fields exactly; no thinking field even with an inbound enabled thinking block off the o1/o3 lane; o1 keeps the budget->reasoning_effort mapping (the only emission site). Refs #399 Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
|
myia-ai-01 : review de |
What
Investigation of #399 (frognano-4b via hub thinks; direct :5003 never) — grain 2 of the 11/10 deepqueue. The coordinator's hypothesis (the Qwen openai-wire passthrough fixed by #437) is refuted by measurement; the claudish side is exonerated and the locus is server-side. This PR carries the exoneration as pins.
Measured chain (2026-10-11)
req-1-11725(07/10 19:39:33Z, extracted from the GDrive archive) is minimal:frognano-4b,max_tokens:16, one user text block — no thinking, no tools, no system.main, sandboxed config, echo endpoint): byte-identical 198 bytes at6f732eb9(main on 07/10 — the incident era) andaf52be5f(current):{model, messages, temperature:1, stream:true, stream_options, max_tokens}— no thinking field of any spelling.matchesModelFamily("frognano-4b", "qwen"|"alibaba")= false (startsWith |"/fam"), soQwenModelDialect.prepareRequest— the only adapter whose passthrough could emitenable_thinking— never ran for frognano. The feat(ingress): map enable_thinking/reasoning_effort onto the thinking object #437 fix cannot be what cured this.end_turn, "ok", 2 tokens; the 26-tools consumer bodyreq-1-6396(07/10 16:54Z, which then produced 4.4–8.9k chars of reasoning) →tool_use, 63 tokens. Zero reasoning both. Same outbound body both eras ⇒ the delta is the :5003 server deployment (refs vllm#70/fix(ops): the watchdog takes its .claudish home as a parameter (⚠ GO user requis avant merge) #71 — the vllm workspace redeployed in the window).The pins (new
openai-api-format.test.ts)enable_thinking/reasoning_effort/thinking/thinking_budget/chat_template_kwargs.thinkingblock does not becomereasoning_effortoff the o1/o3 lane (isReasoningModel()stays the only emission gate).3/3 green. No production code changed — there was nothing to fix here.
Note on closure
Root cause lives in the vllm workspace (which server-side change cured it = refs vllm#70/#71); a direct :5003 replay bisect was requested from po-2025 (DM msg-20261011T035636) to pin the server-side grain if wanted. The claudiss-side question of #399 ("un champ/forme du body sortant du hub non testé bascule le template") is answered: no.
Refs #399
🤖 Generated with Claude Code