Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Agentic AI eBook RAG Chatbot

A retrieval-augmented chatbot that answers questions strictly from one document: the Agentic AI eBook. Built with LangGraph, Pinecone, text embeddings, FastAPI and a Streamlit chat UI.

Every response returns the final answer, the retrieved context chunks (with page numbers and similarity scores) and a confidence score. If the eBook doesn't contain the answer, the bot says so instead of guessing.

Demo

Chat answering a question with citations and a confidence score

Sources and page numbers Out-of-scope question is declined
Expanded sources panel Refusal for a question the eBook doesn't cover

Answers include inline citations, the retrieved chunks with page numbers and similarity scores, and a confidence score. Questions the eBook can't answer are declined instead of guessed.

Features

Required

  • PDF ingestion: extract, clean, chunk, embed, store in Pinecone (idempotent, re-runnable)
  • LangGraph RAG pipeline: retrieve, generate, grounded in the PDF
  • Chat API (FastAPI) and chat UI (Streamlit)
  • Response contains answer + retrieved chunks + confidence

Extras

  • Adaptive, self-correcting graph: cheap signals (similarity scores, citation coverage) decide when to spend an LLM call on relevance grading or fact-checking; retries and corrective rewrites are bounded by a hard per-question LLM budget
  • Refuses out-of-scope questions instead of hallucinating
  • Inline citations ([1], [2]) that map to page-numbered source chunks
  • Conversation memory per session, with follow-up question rewriting
  • Token streaming over Server-Sent Events with live pipeline status in the UI
  • Confidence score blending retrieval strength and fact-check result
  • Adaptive reranking via Pinecone-hosted bge-reranker-v2-m3 (only for ambiguous results)
  • Cost and latency visibility: every response reports llm_calls, which optional steps ran, and per-node timings
  • Provider-agnostic: OpenAI / Gemini / Groq for the LLM; Pinecone / OpenAI / Google / HuggingFace embeddings
  • Feedback loop: thumbs up/down stored via POST /feedback
  • Docker + docker-compose, GitHub Actions CI, 95 unit tests (no API keys needed to run them)

Architecture (short version)

question -> rewrite_query -> retrieve --+-- weak ------------------------------------> no_answer   (0 LLM calls)
            (only if the follow-up      +-- strong ----> select_context --+
             depends on history)        +-- ambiguous -> grade_documents --+--> generate --> verify_grounding --> finalize
                                  ^                                       |       ^               |
                                  +-------- transform_query <-------------+       +-- (once) -----+--> no_answer
                                            (bounded retry)

grade_documents (LLM) and the LLM part of verify_grounding are optional: they run only when free signals (similarity score, citation coverage, lexical support) are inconclusive. See LLM usage.

Full explanation, node table, diagrams and design trade-offs: docs/ARCHITECTURE.md.

Project structure

app/
  config.py            all settings (env-driven)
  models.py            LLM + embedding factories
  vectorstore.py       Pinecone index / vector store helpers
  ingestion/           loader.py (PDF), chunker.py, pipeline.py
  rag/                 state.py, prompts.py, nodes.py, graph.py, retriever.py, scoring.py, heuristics.py
  errors.py            quota / rate-limit detection
  api/                 main.py (routes), schemas.py, streaming.py, feedback.py
ui/streamlit_app.py    chat UI
scripts/               ingest.py, run_sample_queries.py, inspect_retrieval.py
tests/                 pytest suite (fake LLM + fake retriever)
docs/                  ARCHITECTURE.md, SAMPLE_QUERIES.md, sample_queries.json, SAMPLE_OUTPUTS.md

Setup

1. Prerequisites

2. Install

git clone https://github.com/ShubhamB74/agentic-rag-chatbot.git
cd agentic-rag-chatbot

python -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate
pip install -r requirements.txt

cp .env.example .env               # then edit .env: add PINECONE_API_KEY and one LLM key

3. Ingest the eBook (once)

python -m scripts.ingest           # or: make ingest

This downloads the PDF to data/, extracts text page by page, chunks it, embeds the chunks and upserts them to Pinecone (creating the serverless index if needed). Re-running is safe: chunk IDs are deterministic. Use --reset to wipe the namespace first.

If the download is blocked (some servers reject scripts), save the PDF manually as data/Ebook-Agentic-AI.pdf and run the command again; it will use the local file.

4. Run

# terminal 1: API  ->  http://localhost:8000/docs
uvicorn app.api.main:app --reload --port 8000        # or: make api

# terminal 2: UI   ->  http://localhost:8501
streamlit run ui/streamlit_app.py                    # or: make ui

Docker alternative

docker compose run --rm api python -m scripts.ingest   # once
docker compose up --build                              # API on :8000, UI on :8501

Using the API

Method Path Purpose
POST /chat Answer + chunks + confidence in one JSON response
POST /chat/stream Same, as Server-Sent Events (status, token, reset, final, error)
GET /sessions/{id}/history Conversation so far
DELETE /sessions/{id} Forget a conversation
POST /feedback Thumbs up/down for an answer
GET /feedback/stats Feedback totals
GET /health Status + number of indexed vectors

Interactive docs: http://localhost:8000/docs.

curl -s -X POST http://localhost:8000/chat \
  -H "Content-Type: application/json" \
  -d '{"question": "What is Agentic AI?", "session_id": "demo"}'

Shape of the response (values are illustrative):

{
  "session_id": "demo",
  "message_id": "9f2c...",
  "question": "What is Agentic AI?",
  "standalone_question": "What is Agentic AI?",
  "answer": "Agentic AI refers to ... [1][2]",
  "answered": true,
  "confidence": 0.86,
  "confidence_label": "high",
  "grounding_score": 0.95,
  "stop_reason": null,
  "unsupported_claims": [],
  "retrieved_chunks": [
    {"chunk_id": "a1b2...", "text": "...", "page": 3, "score": 0.71,
     "rerank_score": null, "used": true, "citation": 1}
  ],
  "trace": ["skip_rewrite(no history)", "retrieve(attempt=1, hits=6, top=0.66, tier=strong)",
            "select_context(used=3)", "generate",
            "verify_grounding(method=heuristic, score=0.90, confident)", "finalize"],
  "stats": {
    "llm_calls": 1, "retrieval_attempts": 1, "retrieval_tier": "strong",
    "rewrite_used": false, "grader_used": false, "verifier_used": false, "reranker_used": false,
    "grounding_method": "heuristic",
    "grounding_signals": {"claims": 4, "citation_coverage": 1.0, "lexical_support": 0.86,
                          "citations_valid": true, "numbers_supported": true, "words": 74},
    "node_latency_ms": {"rewrite_query": 0, "retrieve": 310, "select_context": 0,
                        "generate": 1650, "verify_grounding": 1, "finalize": 0}
  },
  "latency_ms": 1970
}

Reuse the same session_id to continue a conversation. When the eBook can't answer, you get "answered": false, "confidence": 0, and a stop_reason such as weak_retrieval, no_relevant_context, model_declined or failed_grounding_check.

Sample queries

Six queries covering definitions, synthesis, a follow-up (memory), and an out-of-scope question are listed in docs/SAMPLE_QUERIES.md. Run them and generate a report with real outputs:

python -m scripts.run_sample_queries                 # adaptive mode -> docs/SAMPLE_OUTPUTS.md
python -m scripts.run_sample_queries --mode dev      # cheapest: about 1 LLM call per question
python -m scripts.run_sample_queries --mode full     # grader + verifier on every question
python -m scripts.run_sample_queries --only 4,5,6    # just these queries

The runner is built for small free-tier quotas: each finished answer is cached immediately in data/.sample_cache.json, a quota error stops the run cleanly (nothing is lost), and re-running the same command only spends LLM calls on queries that are not done yet. See the workflow below.

LLM usage and cost control

The pipeline is adaptive: it spends an LLM call only when free signals are inconclusive.

Situation LLM calls Why
Clear in-scope question, strong retrieval, well-cited answer 1 generate only
Follow-up that depends on chat history ("those", "what about costs?") 2 + query rewrite
Ambiguous retrieval 3 + relevance grader + fact-checker
Obviously out-of-scope (very low similarity) 0 declined from the scores alone
Verifier rejects the draft up to the budget one corrective rewrite, then decline

Hard cap: MAX_LLM_CALLS (default 4) per question; optional steps are skipped once it is spent. Every response includes a stats object (llm_calls, retrieval_tier, grader_used, verifier_used, grounding_method, node_latency_ms, ...) and the Streamlit "Pipeline details" popover shows it.

Mode Settings Use it for
adaptive (default) ADAPTIVE_MODE=true normal use and demos
dev ENABLE_GRADER=false ENABLE_VERIFIER=false day-to-day development, about 1 call per question
full ADAPTIVE_MODE=false targeted evaluation: grader on every retrieval, verifier on every answer

The free grounding heuristic never rejects an answer by itself; only an LLM verdict can. Heuristic-only answers get a slightly lower confidence ceiling than LLM-verified ones.

Working within a free-tier quota

  1. Calibrate for free. python -m scripts.inspect_retrieval runs retrieval only (no LLM calls) and shows the similarity scores and tier for each sample query, then suggests MIN_RELEVANCE_SCORE / STRONG_SCORE. The defaults are estimates, so do this first.
  2. Develop in dev mode, and switch to --mode full only for a few targeted queries.
  3. Resume instead of restart. If the quota runs out, switch API key or model and re-run the same command.
  4. Keep LLM_MAX_RETRIES=1 (the default): provider SDKs retry 429 errors with backoff, which can burn extra quota.
  5. Free-tier limits are usually per model and change over time; check your provider's rate-limit page. A lighter model (for example a "flash-lite" variant) or Groq's free tier may allow many more requests per day.

Configuration

Everything is set via environment variables (see .env.example). The ones you will most likely touch:

Variable Default Notes
LLM_PROVIDER / LLM_MODEL openai / gpt-4o-mini gemini, groq also supported
EMBEDDING_PROVIDER / EMBEDDING_MODEL pinecone / llama-text-embed-v2 changing it requires a new PINECONE_INDEX_NAME
TOP_K 6 chunks retrieved per query
CHUNK_SIZE / CHUNK_OVERLAP 1000 / 150 characters; re-ingest with --reset after changing
USE_RERANKER false Pinecone reranker; in adaptive mode it only runs for ambiguous retrieval
ADAPTIVE_MODE true false = grader + verifier on every question (evaluation)
ENABLE_GRADER / ENABLE_VERIFIER true false = never call that LLM step (dev mode)
MAX_LLM_CALLS 4 hard per-question budget
MIN_RELEVANCE_SCORE 0.25 top similarity below this is declined with 0 LLM calls
STRONG_SCORE 0.50 top similarity at or above this skips the grader
LLM_MAX_RETRIES 1 provider-level retries (each can consume quota)
MIN_GROUNDING_SCORE 0.60 answers below this are rejected

Tuning the confidence score

Cosine similarity ranges differ between embedding models. Ask a few in-scope questions, look at the score of the used chunks, then set SCORE_HIGH near a typical good match and SCORE_LOW near the score of an unrelated chunk. That keeps the retrieval half of the confidence meaningful for your model.

Development

pip install -r requirements-dev.txt
pytest          # 95 tests, no network or API keys needed
ruff check .

The tests inject a scripted fake LLM and fake retriever into build_graph(...), so the full pipeline (retry loop, refusals, regeneration, streaming, memory) is verified offline.

Troubleshooting

Symptom Fix
PINECONE_API_KEY is not set create .env from .env.example
Index ... has dimension X but embedding model produces Y you changed embedding model; use a new PINECONE_INDEX_NAME
Everything is answered "not in the eBook" check /health shows vectors > 0 (did ingestion run?); lower MIN_RELEVANCE_SCORE; inspect retrieved_chunks
"Very little text was extracted" warning the PDF is image-based and would need OCR
429 / "quota or rate limit was reached" free-tier quota spent: see Working within a free-tier quota; use dev mode or another key/model
Legitimate questions are declined with weak_retrieval MIN_RELEVANCE_SCORE is too high for your embedding model: run scripts.inspect_retrieval and lower it
UI says "API not reachable" start the API first, or fix the URL in the sidebar

Limitations

  • Only text is read from the PDF; tables and figures are not interpreted.
  • Chat memory is in-process (MemorySaver) and resets when the server restarts.
  • The confidence score is a heuristic for ranking answers, not a calibrated probability.
  • No authentication or rate limiting yet; add both before exposing the API publicly.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages