Skip to content

research: blinded human evaluation panel — does score-low text actually read human? #159

Description

@devswha

Why it matters

Every signal in patina is auto-scored. There is no human-rated checkpoint anywhere. A score of 25 might mean "humans agree this reads human" — or "LLM-judge thinks LLM-rewritten text reads human while humans still detect it". The "score must match human intuition" north-star is unfalsified.

Acceptance criteria

  • docs/research/human-eval-panel.md design doc: panel size, blinding protocol, fixture pool, analysis plan
  • Pilot panel: 30 paragraphs × 5 raters
  • Report inter-rater agreement (Krippendorff's alpha) and correlation patina score vs human rating

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestpriority: lowTriage: longer-horizon research or large buildresearchEvaluation methodology or external comparison

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions