Skip to content

Add Typed Evals: Jev-powered evaluation and calibration - #180

Merged
fatwang2 merged 3 commits into
fatwang2:mainfrom
Aaryanverma:main
Oct 9, 2026
Merged

fatwang2 merged 3 commits into
fatwang2:mainfrom
Aaryanverma:main

Conversation

@Aaryanverma

Copy link
Copy Markdown
Contributor

Description

Adding Typed Evals, an open-source Python toolkit that uses Jev to evaluate LLM, RAG, and agent outputs.

Alongside evaluation metrics and custom rubrics, it supports calibrating Jev scores against human pass/fail labels.

The repo also includes a TRIVIA+ benchmark comparing raw and calibrated Jev scores on held-out data.

I've included source paths for the Jev backend, calibration documentation, and benchmark results to help with the review.

Disclosure: I'm the maintainer of Typed Evals. This is an author submission.

Benchmark: https://github.com/TrustifAI/typed_evals/blob/main/docs/BENCHMARK.md

Added a JSON file for Typed Evals with metadata and evidence.
@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Jev review

Needs maintainer review

GitHub listing (not a review criterion): 15 stars · organization TrustifAI · 1 follower

Suggested category: sdk · category confidence: 0.31

Criterion Probability of yes Result
Concrete Jev integration 0.96 pass
Description supported by evidence 0.94 pass
Usable setup documentation 0.95 pass

Follow-up:

  • Category confidence below policy threshold
  • Proposed category differs from suggested category (sdk)
  • Full file omitted because it exceeds the evidence budget: typed_evals/evaluation/evaluator.py; maintainer review required

Evidence inspected:

PR commit: c25ca5707b8bfca1965bcfb78078bf4b1883ac7a
Model: typesafe-ai/jev via vercel · policy: 9675860e2c86 · input tokens: 13342
Review run and JSON report

Model judgments use Jev only; this comment is generated from a template. Probabilities are model judgments, not verified accuracy. A maintainer decides whether to merge.

Aaryanverma and others added 2 commits October 8, 2026 18:42

@fatwang2 fatwang2 left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maintainer review of 91939ab: approved after manual source review.

At TrustifAI/typed_evals@ab9fc8a I inspected the official TypeSafe SDK backend, evaluator (including the file omitted from automated discovery), calibration core, calibration documentation and setup/examples. The backend calls system_one with typed questions; the evaluator applies a separate fitted curve per metric, using binary pass/fail labels and reserving validation content. This supports the functional description without asserting generally calibrated Jev probabilities or a universal accuracy gain. The public MIT license and usable setup are present. research fits its evaluation/calibration purpose; I manually resolve the advisory model's low-confidence sdk suggestion.

The advisory model report covers the preceding entry commit c25ca57. The only later change is README regeneration; this review covers the current head. I verified that the README is byte-for-byte the trusted generator's output and accept that mechanical regeneration with this entry. Current-head catalog CI, local validation/build and all four combined catalog tests passed.

@fatwang2
fatwang2 merged commit ace3e96 into fatwang2:main Oct 9, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants