AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
-
Updated
Jul 26, 2026 - Python
AuditPilot: auditable enterprise AI agents for evidence-grounded workflows, governed tools, evaluation harnesses, human review, and remediation delivery.
The open-source MultiAgentOps evaluation and verification harness for any industry business workflow.
An end-to-end framework for running, sandboxing, and scoring agentic LLMs on complex data-science and econometric replication tasks.
Detecting Relational Boundary Erosion in AI systems. A framework for testing whether models maintain honest, calibrated, and appropriate boundaries.
VLA ≠ VLM. Side-by-side viewer running NVIDIA Alpamayo R1 (vision-language-action) alongside Qwen2.5-VL (vision-language) on the same 44-sec SF dashcam clip at 5 Hz. 220 paired traces. Surfaces what an action-trained model sees that a scene-trained model doesn't, and vice versa.
RAG service that treats abstention as a feature: cited answers, a faithfulness gate, and a measured coverage-vs-false-answer curve. Chunking x retriever evaluation grid vs planted ground truth shows why retrieval metrics alone mislead. From-scratch BM25, LSA + RRF hybrid, FastAPI, MLflow, 31 tests, fully offline CI.
Cited document Q&A over your PDFs. FastAPI + pgvector with hybrid retrieval, reranking, per-claim citation verification, and a published benchmark comparing five retrieval configurations on a 30-question eval set.
Wellness verification harness for companion AI. Multi-turn adversarial suites grounded in six decades of mental-health research and current clinical standards (988, VERA-MH) and law (SB 243). Point it at any chat endpoint — get an evidence-backed, reproducible report.
Does a CLAUDE.md actually change how Claude behaves? An ablation harness: run adversarial traps with the rules and without them, grade blind, and test whether the difference is real.
Autonomous financial research agent combining live market data, financial news, sentiment analysis, and private RAG with transparent execution.
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild
Enterprise RAG lab using AWS Bedrock, Snowflake, MuleSoft, Python, and an evaluation harness for regulated lending scenarios.
An LLM-powered training-evaluation platform that scores open-ended scenario responses 0 to 10 against rubrics, with an evaluation harness that benchmarks the AI scorer against human-labelled scores.
Constitutional governance platform for multi-agent AI systems — 261 personas, 17 divisions, a judiciary, RBAC, and a written constitution
AI content engine using an anxiety-indexed behavioral science KB, multi-stage LangGraph pipeline, and calibrated LLM-as-judge evaluation harness
Closed-loop LLM factory in one monorepo: data pipeline, trainer, eval-gated checkpoint promotion, serving, and an agent whose single tool is a self-extending CLI.
AIの成果が「本物か、まぐれか」を統計で切り分ける4層検証パイプライン(取得→防御→検証→運用)。LLM評価・実ログ障害トリアージに実適用済み
Prompt-evaluation toolkit: run golden-case prompts, route models, track cost, and leaderboard.
Chaos engineering for multi-agent LLM meshes — which topology survives which failure? Runs on Nebius Serverless: CPU Job sweeps against a vLLM Endpoint, per-trial Object Storage checkpoints, kill-and-recover by design.
DoE Project
Add a description, image, and links to the evaluation-harness topic page so that developers can more easily learn about it.
To associate your repository with the evaluation-harness topic, visit your repo's landing page and select "manage topics."