Skip to content

ecosystem: publish suspect-zones benchmark corpus as a HuggingFace dataset #212

Description

@devswha

Why it matters

README headlines (91% / 76% / 13–25%) are unverifiable without the corpus. Publishing as HF dataset is the NLP-research reproducibility move and the prerequisite for a citation / arXiv paper.

Acceptance criteria

  • huggingface.co/datasets/devswha/patina-suspect-zones with current tests/fixtures/suspect-zones/ contents
  • Data card: per-source licensing, language coverage, how 91/76 were computed, known biases
  • CI step (manual trigger ok for v1) re-uploads on tagged release
  • Cross-linked from README, FAQ, research notes, latest.md
  • Licensing gate: review each fixture for redistribution rights before upload

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    benchmarkCorpus, metrics, calibration, or quality gate workecosystemPackage, dataset, marketplace, or third-party surface workenhancementNew feature or requestpriority: lowTriage: longer-horizon research or large build

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions