Skip to content

feat: implement full retrieval contract and RAGAS evaluation - #1

Merged
ACBBZ merged 23 commits into
mainfrom
feat/rag-v2-retrieval-ragas
Jul 29, 2026
Merged

ACBBZ merged 23 commits into
mainfrom
feat/rag-v2-retrieval-ragas

Conversation

@ACBBZ

@ACBBZ ACBBZ commented Jul 29, 2026 •

Copy link
Copy Markdown
Owner

Summary

Implements the retrieval and evaluation portion of the approved RAG V2 design:

  • resolves effective retrieval options from request fields and environment defaults
  • adds explicit vector, full_text, hybrid, and auto retrieval modes
  • implements PostgreSQL full-text retrieval with tenant, knowledge-base, document, and metadata filters
  • executes vector and full-text retrieval concurrently in hybrid mode
  • fuses hybrid candidates with weighted Reciprocal Rank Fusion
  • preserves candidate order during PostgreSQL hydration
  • carries per-stage retrieval methods and scores through reranking
  • applies document filters directly in Milvus
  • applies metadata filters in PostgreSQL full-text retrieval and vector-result hydration
  • commits successful request transactions and rolls back failed requests
  • adds deterministic retrieval metrics and an optional RAGAS 0.4.3 HTTP evaluation runner
  • adds an evaluation dataset template, environment configuration, documentation, and CI contract checks

API behavior

POST /v1/retrieval/search continues to support the existing boolean options and now also accepts:

{"retrieval_mode": "vector | full_text | hybrid | auto"}

Explicitly requested capabilities are no longer silently ignored. Invalid combinations and requests with no enabled retriever return validation errors. Responses include effective_options, and chunks retain retrieval_methods and per-stage scores.

Migration

Apply 0004_retrieval_v2 before enabling full-text or hybrid retrieval:

  • converts chunks.metadata to JSONB
  • adds a generated PostgreSQL search_vector
  • adds GIN indexes for full-text and metadata filtering
  • adds a scoped retrieval index

RAGAS

The optional eval dependency group pins:

  • ragas==0.4.3
  • langchain-community>=0.3,<0.4 to avoid the known RAGAS import incompatibility with langchain-community 0.4.x

The runner calls the real RAG API and reports deterministic Hit Rate@K, Precision@K, Recall@K, and MRR, with optional Faithfulness, Answer Relevancy, Context Precision, Context Recall, and Factual Correctness metrics.

Verification

  • CI: editable .[dev] installation succeeded
  • Ruff: passed
  • full pytest suite: passed
  • evaluation contract: editable .[dev,eval] installation succeeded
  • RAGAS collections API imports: passed
  • deterministic evaluation tests: passed

Known limitation

The fixed Milvus V1 schema does not contain a generic metadata JSON field. Vector metadata filtering is therefore enforced during PostgreSQL hydration after an over-fetched Milvus search. Returned results obey the filter, but very selective filters may reduce recall. True pre-ANN metadata filtering requires a future Milvus V2 collection and reindexing.

Deferred design stages

The asynchronous ingestion worker, document-version activation state machine, token-aware parser/chunker redesign, and structured citation validation are intentionally kept out of this focused PR so the retrieval contract and evaluation layer can be reviewed and deployed independently.

@ACBBZ
ACBBZ marked this pull request as ready for review July 29, 2026 02:37
Copilot AI review requested due to automatic review settings July 29, 2026 02:37
@ACBBZ
ACBBZ merged commit a31ccd6 into main Jul 29, 2026
4 checks passed

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Implements the “Retrieval V2” contract end-to-end (option resolution, vector/full-text/hybrid execution, filtering/hydration semantics, and per-stage scoring) and adds an offline evaluation layer (deterministic metrics + optional RAGAS runner), including migration, docs, and CI checks.

Changes:

  • Added explicit retrieval modes (vector, full_text, hybrid, auto) with effective option resolution and validation.
  • Implemented PostgreSQL full-text retrieval + filtered hydration, plus hybrid concurrency and weighted RRF fusion with per-stage score/method tracking.
  • Added deterministic retrieval metrics and an optional HTTP evaluation runner with a CI contract workflow and supporting documentation/config.

Reviewed changes

Copilot reviewed 23 out of 23 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
tests/test_retrieval_pipeline.py Verifies retrieval mode execution, filter propagation, and rerank/rewrite usage with effective options.
tests/test_retrieval_options.py Covers resolver behavior for defaults, overrides, hybrid normalization, and validation errors.
tests/test_lexical_retrieval.py Tests full-text SQL filtering and stable hydration ordering.
tests/test_evaluation_metrics.py Validates deterministic Hit Rate/Precision/Recall/MRR metric behavior.
tests/test_database_session.py Verifies request-scoped commit/rollback behavior for DB sessions.
rag/storage/milvus_store.py Adds document ID filtering to Milvus search and propagates stage scores/methods.
rag/schemas.py Extends schemas for retrieval_mode, filters model, effective_options, and per-stage scoring fields.
rag/retrieval/postgres_store.py Introduces PostgreSQL full-text retrieval and ordered hydration with document/metadata filters.
rag/retrieval/pipeline.py Orchestrates effective options, hybrid concurrency, fusion, hydration, rerank, and response diagnostics.
rag/retrieval/options.py Adds effective option model + resolver mapping modes/booleans/defaults into a validated configuration.
rag/retrieval/fusion.py Implements weighted RRF fusion and carries per-stage methods/scores into fused results.
rag/evaluation/runner.py Adds JSONL-driven offline runner calling the real API and optionally scoring with RAGAS.
rag/evaluation/deterministic_metrics.py Implements deterministic retrieval metrics used by the runner.
rag/evaluation/client.py Adds an HTTP client wrapper for calling the retrieval endpoint during evaluation.
rag/evaluation/init.py Declares evaluation package purpose.
pyproject.toml Adds optional eval dependency group and includes evaluation package.
migrations/versions/0004_retrieval_v2.py Adds JSONB metadata + generated search_vector and indexes for retrieval v2.
evals/datasets/smoke.jsonl Adds a template evaluation dataset case.
docs/superpowers/plans/2026-07-29-rag-v2-retrieval-ragas.md Captures implementation plan and staged task breakdown.
docs/retrieval-v2-and-ragas.md Documents retrieval modes, filters, migration, and evaluation/RAGAS usage.
app/api/dependencies.py Adds commit/rollback lifecycle to session dependency and wires retrieval pipeline to Postgres store.
.github/workflows/rag-eval.yml Adds CI contract checks for eval deps and an opt-in/scheduled evaluation workflow.
.env.example Adds evaluation and RAGAS environment variable templates.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread rag/retrieval/options.py
Comment on lines +40 to +45
if options.retrieval_mode and options.retrieval_mode != "auto":
mode = options.retrieval_mode
vector_search = mode in {"vector", "hybrid"}
full_text_search = mode in {"full_text", "hybrid"}
hybrid_search = mode == "hybrid"
else:
Comment thread app/api/dependencies.py
Comment on lines +45 to +50
try:
yield session
await session.commit()
except BaseException:
await session.rollback()
raise
Comment on lines +76 to +80
DATASET="${{ inputs.dataset || 'evals/datasets/smoke.jsonl' }}"
EXTRA_ARGS=""
if [ "${{ inputs.use_ragas || 'true' }}" = "true" ]; then
EXTRA_ARGS="--ragas"
fi
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants