Skip to content

benchmark: live-mode scoring benchmark for src/scoring.js (LLM judge layer) #412

Description

@devswha

Context

tests/quality/README.md ("What it does NOT measure") notes that the deterministic quality benchmark intentionally excludes LLM-based scoring (src/scoring.js) because it is non-deterministic and adds API cost/latency, and that "a separate live-mode benchmark would be its own follow-up". This was the only follow-up note in the repo without a tracking issue — filing it so it doesn't get lost.

Scope

Non-goals

  • No CI gate: live results are diagnostics, consistent with the existing live-quality stance.

Source: tests/quality/README.md follow-up note, surfaced during the 2026-06-11 remaining-work audit.

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkCorpus, metrics, calibration, or quality gate workenhancementNew feature or requestpriority: lowTriage: longer-horizon research or large build

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions