Skip to content

Commit fbb088e

Browse files
committed
fix(decide,llm): close review gaps in model detection and completion advice
Review of the #238 -> #243 stack found three accuracy gaps in behavior that 1.7.9 would ship for the first time: - `claude-mythos-preview` (and any unversioned Fable/Mythos id) did not match the model-trait pattern, so it still received `temperature` and got an HTTP 400. Fable/Mythos ids now match with or without a numeric version. - A remote `verify_completion` probability between the certainty bound (0.75) and the 0.85 completion bar was reported as a decisive `is_complete=false`. It is now `uncertain` with a null conclusion, matching the documented "uncertain completion stays null" contract. - The local completion heuristic missed Mocha/Jest-style `N passing`, treated `errors: none` as a failure, and needed `not passing`/`0 passing` handling once `passing` counts as success. Every new test fails on the previous head. Rebind the public offline fixtures to immutable v125 evidence; aggregates are unchanged from v124. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014ksPGXPZQTikJYpaGHa6R6
1 parent 1b1a0b3 commit fbb088e

15 files changed

Lines changed: 772 additions & 27 deletions

‎BENCHMARKS.md‎

Lines changed: 5 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -94,14 +94,14 @@ interpretation and do not count as additional benchmark-quality gains.
9494
### Public numeric evidence registry
9595

9696
Every exact public aggregate retained below comes from the checked-in, public-safe
97-
[`offline-fixtures-v124.json`](docs/benchmark-evidence/offline-fixtures-v124.json) artifact. Its
97+
[`offline-fixtures-v125.json`](docs/benchmark-evidence/offline-fixtures-v125.json) artifact. Its
9898
SHA-256 is
99-
`dc9e35cc4c20f164d6adb33858bd444f8001f4f9dad7ad1be618e8c567adbcf6`, also recorded in the
99+
`1f74971d6213a188b31cf58f6ff6132a487da36d22455292b028dadc202a8feb`, also recorded in the
100100
adjacent `.sha256` file. The artifact contains no raw questions, answers, prompts, customer data,
101101
or per-record content fingerprints.
102102

103103
The fixture-suite digest is
104-
`05007294ab0c835fea98d65239576f3221084213e2e022d8c84328e12014a6cc`. The artifact defines
104+
`42cb269b867a1e3210d0da77d6a040abc974cd44317b91ad10a2d483d07c1586`. The artifact defines
105105
the digest algorithm and records the SHA-256 of every suite and dataset file. Each evidence ID
106106
also binds its exact command through `sha256(UTF-8 exact command)`:
107107

@@ -123,10 +123,10 @@ Historical LoCoMo, graph, handoff, consolidation, and security figures remain pr
123123
source artifacts but are omitted from the current chart until each has a matching immutable,
124124
public-safe artifact. The chart labels coding outcomes, external datasets, and operational
125125
capacity as pending evaluation tracks rather than implying scores. Regenerate it with
126-
`python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v124.json --output docs/images/context-efficiency.svg` after selecting the report to publish.
126+
`python scripts/render_benchmark_report.py --report docs/benchmark-evidence/offline-fixtures-v125.json --output docs/images/context-efficiency.svg` after selecting the report to publish.
127127

128128
The companion examples are also generated from that artifact with
129-
`python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v124.json --output docs/images/evidence-backed-agent-examples.svg`.
129+
`python -m scripts.render_benchmark_examples --report docs/benchmark-evidence/offline-fixtures-v125.json --output docs/images/evidence-backed-agent-examples.svg`.
130130
The historical-to-executable mapping is in
131131
[`docs/BENCHMARK_CHANGE_COVERAGE.md`](docs/BENCHMARK_CHANGE_COVERAGE.md).
132132

‎CHANGELOG.md‎

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -33,7 +33,13 @@ All notable changes to Engraphis are documented here. Format loosely follows
3333
the API defaults, `.env.example`, and the provider guide with `claude-sonnet-5-5`.
3434
- Documented how to back the experimental Jev decision adapter with Claude: pin an exact model
3535
id, keep fallback disabled, and avoid sampling parameters and forced tool choice.
36-
- Refreshed the public offline fixtures and source bindings in immutable v124 evidence.
36+
- Recognized `claude-mythos-preview` and other unversioned Fable/Mythos ids, so they no longer
37+
receive `temperature` (an HTTP 400) and get the thinking-model output headroom.
38+
- A remote `verify_completion` probability between the certainty bound and the 0.85 completion
39+
bar is now `uncertain` with a null `is_complete`, instead of a decisive failure.
40+
- Local completion checks recognize `N passing` (Mocha/Jest style), `not passing`, `0 passing`
41+
and `errors: none`.
42+
- Refreshed the public offline fixtures and source bindings in immutable v125 evidence.
3743

3844
## [1.7.9] - 2026-09-29
3945

‎README.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -84,9 +84,9 @@ neither is an end-to-end question-answer score. Coding outcomes, external datase
8484
operational capacity remain separate pending evaluation tracks until their artifacts are selected.
8585

8686
These values are evidence IDs `offline-chunking` and `offline-performance` in
87-
[`offline-fixtures-v124.json`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/benchmark-evidence/offline-fixtures-v124.json),
87+
[`offline-fixtures-v125.json`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/docs/benchmark-evidence/offline-fixtures-v125.json),
8888
SHA-256
89-
`dc9e35cc4c20f164d6adb33858bd444f8001f4f9dad7ad1be618e8c567adbcf6`.
89+
`1f74971d6213a188b31cf58f6ff6132a487da36d22455292b028dadc202a8feb`.
9090
[`BENCHMARKS.md`](https://github.com/Coding-Dev-Tools/engraphis/blob/main/BENCHMARKS.md#public-numeric-evidence-registry)
9191
records the matching suite digest, exact commands, and per-command config digests. The offline
9292
fixture registry intentionally excludes external, model-dependent, consolidation, productivity,

‎docs/LLM_PROVIDERS.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -109,9 +109,9 @@ Choosing a model:
109109
- Retired ids such as `claude-3-5-sonnet-20241022` and `claude-3-5-haiku-20241022` are rejected by
110110
the API; the connection test reports them as an HTTP 404.
111111

112-
Opus 4.7 and later, Sonnet 5 and later, and Fable reject `temperature` and similar sampling
113-
parameters, so Engraphis omits them for those models. Opus 5 and later, Sonnet 5 and later, and
114-
Fable also think before answering by default. Engraphis sends `ENGRAPHIS_LLM_EFFORT` (`low`,
112+
Opus 4.7 and later, Sonnet 5 and later, Fable and Mythos (including `claude-mythos-preview`)
113+
reject `temperature` and similar sampling parameters, so Engraphis omits them for those models.
114+
Opus 5 and later, Sonnet 5 and later, Fable and Mythos also think before answering by default. Engraphis sends `ENGRAPHIS_LLM_EFFORT` (`low`,
115115
`medium`, `high`, `xhigh`, or `max`; default `medium`) for those models and keeps at least 4096
116116
output tokens available so hidden reasoning cannot crowd out the reply. Other providers and older
117117
Claude models ignore the setting.

0 commit comments

Comments
 (0)