Skip to content

fix: preserve malformed LLM responses as incomplete analysis - #796

Open
yashrajp22 wants to merge 25 commits into
mainfrom
yashraj/fix-batch-compat-parse-failures
Open

yashrajp22 wants to merge 25 commits into
mainfrom
yashraj/fix-batch-compat-parse-failures

Conversation

@yashrajp22

@yashrajp22 yashrajp22 commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

The batch compatibility parser converted malformed JSON or an invalid response schema into an empty successful result. Those failures now use the core structured-response retry and inspection-ledger path. Confidence validation catches null and non-numeric values; meta soft fields are repaired before validation while findings remain strict. Optional prose assessments do not discard otherwise valid finding verdicts.

The compatibility guard checks both keyword-only constructor parameters and forwards numeric or callable timeouts. Pool waits and key retries share a monotonic deadline; cancellation and constructor errors release slots, and shorter connection budgets remain intact.

Validation: 608 tests passed against a freshly installed wheel core, including 29 pool tests; the preceding source suite passed 579 tests. Ten added regressions reproduce failures against the original source and wheel core. Contributor tools are not packaged in the wheel, so wheel checks use the contributor checkout with the installed core. Tests use deterministic fake providers; these focused tests made no live provider calls.

Combined verification across the updated PRs: 6,243 regression tests passed against source and again against the freshly installed wheel, with seven conditional skips and four expected failures per run. All 19 source/wheel sample pairs matched. The 12-skill corpus retained its findings and risk ratings; four former hangs now finish with explicit partial-analysis results. The 93 extension tests passed. Two synthetic live NVIDIA Build checks passed on the final wheel: benign-note was complete/SAFE, and the exfiltration sample retained SSD-3 with complete semantic and meta analysis and a DO_NOT_INSTALL recommendation. All seven recorded LLM analyses succeeded. Live checks used the configured model/reasoning defaults through a test-only proxy that kept the real credential outside the scanner.

@rng1995 rng1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[SkillSpector Review]

Hi @yashrajp22, thank you for fixing the compatibility constructor and both raw-JSON parsers together, so the batch scanner's LLM analyzers now fail visibly instead of returning empty results! The new compat test runs the real run_batches_detailed/arun_batches_detailed loop with a fake model and checks the ledger outcome, which is exactly where this bug was hidden.

Value and readiness: The problem is real on main.

  • Core meta path. Main's MetaAnalyzerResult turns {"findings": "invalid"} into []. The result is one provider call, a successful empty batch, and meta reported as completed.
  • Batch-scan compat mode. Here the same bug is currently hidden by another one: _patched_base_init rejects timeout=, so all four LLM analyzers are recorded as unavailable before any call is made. If only the constructor fix is applied to main's parsers, a not JSON response gives is_complete=true, every analyzer completed and llm_call_log ok=True. Shipping the constructor fix and the parser fix together is therefore the right shape.

With this PR's logic transcribed onto main, every malformed or schema-invalid response I tried takes exactly 4 calls. The inputs were invalid JSON, {"findings": "invalid"}, a top-level list or null, prose around JSON, and empty-string findings. Each one ends as llm_structured_response_invalid, with ledger skipped, the analyzer degraded and completeness partial. That holds for sync and async, and for discovery and meta. Fenced valid JSON and {"findings": []} still take one call and complete, and a graph run's risk score is unchanged.

The core change is sound, and the results above verify it. Two required follow-through fixes remain, so I am requesting changes. The existing contrib test for _patched_base_init now fails against the timeout forwarding (finding 1). The contrib docs still say malformed responses return [], which is the opposite of the new behaviour (finding 2). Both are small. Once they are in, this is ready for final maintainer review; CI also still has to run. Findings 3 to 8 are non-blocking.

Material findings

  1. [Blocker] contrib/batch_scan/runner.py:141 (test at contrib/batch_scan/tests/test_monkeypatch_fragility.py:275): with the timeout forwarding, the existing contrib test for this function now fails. The fix is small, but it is required because this PR changes the function under test.
    • _patched_base_init now always calls _original_base_init(..., node=node, timeout=timeout).
    • test_patched_init_forwards_keyword_only_node mocks _original_base_init and asserts assert_called_once_with(instance, "prompt", "model", node="semantic_security_discovery"), which compares kwargs exactly.
    • On main the file passes 28/28. Replaying this test against a transcription of the head's function fails with Actual: ... node='semantic_security_discovery', timeout=None.
    • CI does not collect the file: pyproject.toml:117 sets testpaths = ["tests"], and make test-ci runs tests/.
    • The wider contrib suite is already red on main: test_pool_wiring.py errors at collection, and the rest gives 16 failed and 29 errors, mostly from missing API keys.
    • The production change is correct; only the test expectation is stale. The only check that the timeout is forwarded is tests/test_batch_scan_security.py:68 (analyzer._timeout == 7).
    • Expected fix:
      • Add timeout=None to the assertion.
      • Add a sibling case that passes timeout=7 and asserts it is forwarded unchanged.
      • Add a case that passes a callable deadline and asserts call_args.kwargs["timeout"] is deadline.
      • All three assertions pass against the transcription.
  2. [Blocker] contrib/batch_scan/docs/DESIGN.md:286-297: the contrib docs still say that malformed responses return []. After this PR they state the opposite of the actual behaviour.
    • Stale text. The PR changes no docs, and these places still describe the old behaviour:
      • DESIGN.md says invalid JSON and schema violations are "returned as []". Its "Error propagation" paragraph says the analyzer "returns [] (no findings for that file)".
      • contrib/batch_scan/docs/README.md:344-347 (Known Limitation 5) says the user "won't know which findings were lost".
      • contrib/batch_scan/tests/docs/TEST_DESIGN.md:98 describes the same [] behaviour.
      • The runner.py:331 comment says "silent degradation if broken".
    • What the head does.
      • runner.py:156-158 and :184-185 raise _StructuredResponseValidationError. The batch is retried and then reported as incomplete.
      • The WARNING lines now come from core: "LLM structured response validation failed for ... retrying" and "... after 4 attempts".
      • A missing model_validate would now raise AttributeError past the narrowed except. It would be recorded as llm_batch_failed, not a silent [].
    • Effect. Operators who read these docs will expect silent drops. They will not understand the new "incomplete skill(s)" counts in the batch reports.
    • Expected fix.
      • Rewrite the DESIGN.md Patch 2/3 bullets and the "Error propagation" paragraph. Say that the batch is retried up to 4 attempts, then recorded as ledger skipped / llm_structured_response_invalid, with the analyzer degraded and the skill counted as incomplete. Say that the scan continues, and give the new WARNING text.
      • Narrow README Limitation 5 rather than deleting it. GapFillAnalyzer.parse_response (contrib/batch_scan/gap_fill.py) still returns [] on malformed output until #797 lands.
      • Update the TEST_DESIGN.md row and the runner.py:331 comment.
      • BUGS_FOUND.md and REVIEW_RESPONSE.md are history and can stay as they are.
  3. [Non-blocking] src/skillspector/nodes/meta_analyzer.py:141: a non-JSON overall_assessment string now fails the whole meta batch, even though nothing reads that field.
    • The change. _parse_stringified_assessment now returns json.loads(v) for any string, where main fell back to None for unparseable strings.
    • Nothing reads the field. A grep of src/ and contrib/ finds no reader: LLMMetaAnalyzer.parse_response (meta_analyzer.py:436-447) and _patched_meta_parse read only .findings.
    • Measured. I compared main's model with the transcribed validator, using one valid verdict plus an overall_assessment of "HIGH risk: exfiltrates credentials", "LOW" or "".
      • main makes 1 call and applies the verdict.
      • The PR makes 4 calls and ends llm_structured_response_invalid. The finding keeps review outcome failed.
    • Scope. main was already strict for strings that parse but have the wrong shape, such as '"LOW"' or '42'. The PR extends that strictness to unparseable strings.
    • Effect. Providers without strict json_schema (tool/function calling, CLI providers, compat mode) can put prose here. A model that does this habitually costs 3 extra calls per batch and loses valid meta verdicts. It fails closed, but it costs coverage.
    • Expected fix.
      • Keep findings strict.
      • For overall_assessment, restore the None fallback (optionally also returning None for non-dict values), or drop the unused field.
      • Limit the invalid-string cases in tests/nodes/test_meta_analyzer.py:1098-1104 to findings.
      • Add a case where a prose overall_assessment validates and the verdict is kept.
  4. [Non-blocking] contrib/batch_scan/runner.py:156: a null or non-numeric confidence raises TypeError, so the batch is not retried.
    • Cause.
      • Both _normalize_confidence validators call float(v) in mode="before" (src/skillspector/llm_analyzer_base.py:613, src/skillspector/nodes/meta_analyzer.py:100).
      • Pydantic converts only ValueError and AssertionError into ValidationError, so "confidence": null raises TypeError on main's models.
      • runner.py:156 and :184 catch only (json.JSONDecodeError, ValidationError).
      • _invoke_batch_with_retries does not retry a TypeError, because it is not a provider error.
    • Measured (transcription).
      • confidence null or [1]: 1 call, llm_batch_failed, ledger failed. This holds for discovery and meta, sync and async.
      • "abc": 4 calls and llm_structured_response_invalid.
      • On main the null input is a silent empty success (ledger completed), so the PR still improves this case.
    • Effect. The compat prompt's "never use null" rule (runner.py:223) suggests these models do emit nulls. Such a batch is lost without a retry and is labelled a generic failure, which contradicts the claim that malformed responses follow retry.
    • Expected fix.
      • In both validators, wrap float(v) and re-raise TypeError/ValueError as ValueError, so both the core and compat paths retry. Alternatively, add TypeError to the compat except tuples.
      • Add a "confidence": null case to test_compat_parse_failures_retry_and_remain_visible.
  5. [Non-blocking] contrib/batch_scan/runner.py:183: _sanitize_meta_finding runs after validation, so it can never repair the quirks it was written for.
    • Why it never fires.
      • MetaAnalyzerFinding.impact is a Literal, and explanation/remediation are str (meta_analyzer.py:108-112).
      • model_validate at :183 therefore rejects these inputs before the sanitizer at :188 sees them: an impact of "none", null, "catastrophic" or "High", and a null remediation or explanation.
      • On valid input the sanitizer is a no-op. DESIGN.md:299-302 calls these quirks recoverable soft errors.
    • Effect.
      • On main these inputs became a silent [].
      • With the PR they cost 4 calls and leave the meta batch incomplete. One bad item fails every finding in the batch.
      • The ordering bug predates the PR; the PR changes its consequence.
    • Expected fix. Since this function is being rewritten anyway:
      • Sanitize the raw dicts in data["findings"] before model_validate, and casefold impact first so that "High" is not downgraded to "low".
      • Drop the post-validation call.
      • A follow-up PR is also fine.
  6. [Non-blocking] contrib/batch_scan/runner.py:94 (pre-existing, outside the diff): in multi-key pool mode, _pooled_get_chat_model(model=None) still rejects timeout=.
    • How it fails.
      • Core always calls get_chat_model(model=model, timeout=...) (llm_analyzer_base.py:938, :1021).
      • When create_api_key_pool_from_env() finds two or more keys, batch_scan.py:195-196 calls set_api_pool, which installs this factory.
      • With the PR's constructor transcribed, LLMAnalyzerBase, LLMMetaAnalyzer and GapFillAnalyzer all raise TypeError: _pooled_get_chat_model() got an unexpected keyword argument 'timeout'.
      • The nodes record the analyzer as unavailable (semantic_security_discovery.py:327).
    • Effect. The PR moves this failure from _patched_base_init to the factory. In pool mode, neither the restored compat analysis nor the new malformed-response handling runs. It fails closed.
    • Not fixed elsewhere. No other open PR in this batch, and not #763, changes this factory.
    • Expected fix.
      • Change the signature to def _pooled_get_chat_model(model=None, *, timeout=None).
      • Pass timeout to PooledChatModel, which already accepts it, and to the fallback.
      • Add a pooled-mode constructor test.
      • Alternatively, state in the PR body that pool mode is out of scope and track it separately.
  7. [Non-blocking] contrib/batch_scan/runner.py:311: the Patch 1 guard in _verify_patch_targets does not check the newly forwarded timeout parameter (added to the patched signature at runner.py:133).
    • What the guard checks. The guard (runner.py:304-322) checks only the positional signature, a keyword-only node and response_schema. It is byte-identical to main.
    • What it misses.
      • On main the guard passes even though Patch 1 is already broken by timeout, which is the drift this PR fixes.
      • I tested a transcribed init against hypothetical core inits that remove or rename timeout. The guard passes, and construction then fails separately for each analyzer.
    • Mitigation. The new CI test (tests/test_batch_scan_security.py:68) would catch drift within this repo. The guard matters when contrib runs against a different installed skillspector, which is its documented purpose (contrib/batch_scan/CONTRIBUTING.md:132).
    • Expected fix. Require a keyword-only timeout in the guard, as it already does for node, and add a matching TestGuardPatch1Init case.
  8. [Non-blocking] tests/nodes/test_meta_analyzer.py:1099: two of the six new parametrized cases already pass on main.
    • The two cases. overall_assessment = "42" and '"wrong type"' parse on main and already fail OverallAssessment's model_type check there, so they do not test this change.
    • The rest. The other four cases do test the change, and a revert of either validator is still caught. The six new cases plus the two updated tests at tests/nodes/test_llm_analyzer_base.py:995-1001 make eight:
      • Reverting only the findings validator makes 5 of the 8 fail.
      • Reverting only the assessment validator makes 1 of the 8 fail.
    • Expected fix (optional).
      • Use unparseable values such as "" or "{not json" for overall_assessment, or drop those cases if finding 3 restores the fallback.
      • Optionally add findings="null".
      • Optionally move the ValidationError import to module level.

PIC tradeoffs:

  • Retry cost. A persistently malformed compat batch now costs up to 4 sequential provider calls instead of 1, with 0.5/1/2 s backoff. The gain is visible incompleteness. Semantic analyzers in batch scan run under a 90 s per-skill wall clock and a 30 s per-request ceiling. A provider that habitually wraps JSON in prose could therefore turn a partial result into a whole-skill timeout. I did not measure this with a real provider.
  • Strict compat parser. The compat parser only strips code fences. Core CLI providers instead use the prose-tolerant _extract_json_object (src/skillspector/llm_utils.py:207). Providers that wrap JSON in prose now get incomplete scans instead of empty ones. Reusing the extractor would recover more of these responses.
  • Private core symbol. contrib now imports _StructuredResponseValidationError (runner.py:46), and #797 does the same. It is the only exception a custom parse_response can raise to get retry and ledger handling: a ValidationError raised directly is a ValueError, and is re-raised as misconfiguration. A public alias would be cleaner. If the symbol is renamed, the import fails loudly at load time.
  • Logs. Core's generic structured-response warning replaces the separate "invalid JSON" and "schema validation failed" warnings. Operators lose that distinction, but raw LLM output (pydantic input_value) no longer reaches the logs.
  • TP4. _TP4Analyzer has no compat parse patch. On main it failed at construction. Now it spends one provider call and then fails with NotImplementedError. The final ledger outcome is the same.

Verification and gaps:

  • Baselines on main, using main's own code.
    • The compat constructor raises TypeError on timeout=. In a graph run, all four LLM analyzers are unavailable and no LLM calls are made.
    • On the core meta path, {"findings": "invalid"} gives 1 call, a success with [], and meta completed.
    • With only the constructor fix applied, not JSON responses are reported complete.
  • PR behaviour, from transcriptions. I transcribed runner.py:133-191 and meta_analyzer.py:126-142 onto main's code and ran them in the main venv.
    • The 4-call results described above.
    • In a graph run: completeness partial, static findings kept with llm_review_outcome failed, and a risk score of 83/CRITICAL, the same as main.
    • Neither the ledger events nor the logs contain the raw response.
    • ValidationError subclasses ValueError, but it is wrapped before reaching the except (ValueError, NotImplementedError): raise. Both loops catch the wrapper (llm_analyzer_base.py:1352, :1470), so no new crash path appears.
  • Tests: which fail on main.
    • The 16-case test_compat_parse_failures_retry_and_remain_visible fails on main in every case, because of the constructor TypeError. All 16 pass under transcription, including the recover-on-retry case (2 calls).
    • The two updated tests at tests/nodes/test_llm_analyzer_base.py:995-1001 fail on main.
    • Four of the six new meta cases fail on main; see finding 8.
    • test_valid_stringified_meta_fields_are_supported passes on both and serves as a guard.
    • Two tests-pro tests that currently fail on main with the timeout TypeError should pass with this PR. That is inferred; I did not run them.
  • Lint. ruff check and ruff format --check pass on the changed src/ and tests/ files.
  • Commits. All seven author commits carry a sign-off. 7ed4487 is the automated merge of main.
  • CI: No check runs exist on the head. The CI workflow run is action_required and needs a maintainer to approve it, and it must finish green before merge.
  • Conflicts: The PR is mergeable with main (current with a0d489e). It conflicts with #797 in contrib/batch_scan/runner.py, where both rewrite the skillspector.llm_analyzer_base import, and both add tests to tests/test_batch_scan_security.py. The conflict is textual and the two PRs agree: #797 raises the same exception from gap-fill's parse_response, which covers the gap-fill [] path this PR leaves alone. Whichever PR lands second needs a rebase.
  • Head update: I reviewed 7ed4487e6b7e3363be1ddf18381c83952000c8c3. The current head d1bcf2c4e02a831b3b47b54daad2636397411694 only adds automated merges of main (#691, #577, #608) from update-pr-branches.yml; those commits touch none of this PR's files (#691 only routes structured-output binding through an overridable method in llm_analyzer_base.py), so line references to llm_analyzer_base.py and llm_utils.py below are to the current head. The PR's own diff is unchanged (same patch-id), so this review applies to the current head.
  • I did not run the PR's tests or code, per policy. Every PR-behaviour result above comes from my transcriptions run on main's code (Python 3.13, pydantic 2.13.4). CI uses Python 3.12.
  • Gaps.
    • No live provider was exercised. I inferred from the code, without checking, that real LangChain structured-output parsers surface the stricter meta errors as ValidationError.
    • How often real models emit a prose overall_assessment, impact: "none", a null remediation or a null confidence is unknown.
    • I did not run the contrib tests-pro suite or mutation_max.py against the head.
    • I did not verify the PR body's claims: 536 targeted tests, 12 pinned sample skills, and source/wheel parity.

Decision: Changes Requested (reviewed head 7ed4487e6b7e3363be1ddf18381c83952000c8c3; current head d1bcf2c4e02a831b3b47b54daad2636397411694 only adds merges of main)

Comment thread contrib/batch_scan/runner.py
Comment thread contrib/batch_scan/runner.py
Comment thread contrib/batch_scan/runner.py
Comment thread contrib/batch_scan/runner.py
Comment thread contrib/batch_scan/runner.py
Comment thread src/skillspector/nodes/meta_analyzer.py Outdated
Comment thread tests/nodes/test_meta_analyzer.py Outdated
@yashrajp22

Copy link
Copy Markdown
Collaborator Author

Thanks for the pool timeout feedback! Waiting for a slot and retrying another key now use the same remaining deadline. Cancellation and constructor errors release the slot, and shorter connection limits stay intact. The sync and async regression tests pass against the installed wheel core.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants