Repository navigation
Add Typed Evals: Jev-powered evaluation and calibration - #180
Conversation
Added a JSON file for Typed Evals with metadata and evidence.
Jev reviewNeeds maintainer review GitHub listing (not a review criterion): 15 stars · organization TrustifAI · 1 follower Suggested category: sdk · category confidence: 0.31
Follow-up:
Evidence inspected:
PR commit: Model judgments use Jev only; this comment is generated from a template. Probabilities are model judgments, not verified accuracy. A maintainer decides whether to merge. |
Corrected capitalization of 'Agent' in the description.
fatwang2
left a comment
There was a problem hiding this comment.
Maintainer review of 91939ab: approved after manual source review.
At TrustifAI/typed_evals@ab9fc8a I inspected the official TypeSafe SDK backend, evaluator (including the file omitted from automated discovery), calibration core, calibration documentation and setup/examples. The backend calls system_one with typed questions; the evaluator applies a separate fitted curve per metric, using binary pass/fail labels and reserving validation content. This supports the functional description without asserting generally calibrated Jev probabilities or a universal accuracy gain. The public MIT license and usable setup are present. research fits its evaluation/calibration purpose; I manually resolve the advisory model's low-confidence sdk suggestion.
The advisory model report covers the preceding entry commit c25ca57. The only later change is README regeneration; this review covers the current head. I verified that the README is byte-for-byte the trusted generator's output and accept that mechanical regeneration with this entry. Current-head catalog CI, local validation/build and all four combined catalog tests passed.
Description
Adding Typed Evals, an open-source Python toolkit that uses Jev to evaluate LLM, RAG, and agent outputs.
Alongside evaluation metrics and custom rubrics, it supports calibrating Jev scores against human pass/fail labels.
The repo also includes a TRIVIA+ benchmark comparing raw and calibrated Jev scores on held-out data.
I've included source paths for the Jev backend, calibration documentation, and benchmark results to help with the review.
Disclosure: I'm the maintainer of Typed Evals. This is an author submission.
Benchmark: https://github.com/TrustifAI/typed_evals/blob/main/docs/BENCHMARK.md