A retrieval-augmented chatbot that answers questions strictly from one document: the Agentic AI eBook. Built with LangGraph, Pinecone, text embeddings, FastAPI and a Streamlit chat UI.
Every response returns the final answer, the retrieved context chunks (with page numbers and similarity scores) and a confidence score. If the eBook doesn't contain the answer, the bot says so instead of guessing.
| Sources and page numbers | Out-of-scope question is declined |
|---|---|
![]() |
![]() |
Answers include inline citations, the retrieved chunks with page numbers and similarity scores, and a confidence score. Questions the eBook can't answer are declined instead of guessed.
Required
- PDF ingestion: extract, clean, chunk, embed, store in Pinecone (idempotent, re-runnable)
- LangGraph RAG pipeline: retrieve, generate, grounded in the PDF
- Chat API (FastAPI) and chat UI (Streamlit)
- Response contains answer + retrieved chunks + confidence
Extras
- Adaptive, self-correcting graph: cheap signals (similarity scores, citation coverage) decide when to spend an LLM call on relevance grading or fact-checking; retries and corrective rewrites are bounded by a hard per-question LLM budget
- Refuses out-of-scope questions instead of hallucinating
- Inline citations (
[1],[2]) that map to page-numbered source chunks - Conversation memory per session, with follow-up question rewriting
- Token streaming over Server-Sent Events with live pipeline status in the UI
- Confidence score blending retrieval strength and fact-check result
- Adaptive reranking via Pinecone-hosted
bge-reranker-v2-m3(only for ambiguous results) - Cost and latency visibility: every response reports
llm_calls, which optional steps ran, and per-node timings - Provider-agnostic: OpenAI / Gemini / Groq for the LLM; Pinecone / OpenAI / Google / HuggingFace embeddings
- Feedback loop: thumbs up/down stored via
POST /feedback - Docker + docker-compose, GitHub Actions CI, 95 unit tests (no API keys needed to run them)
question -> rewrite_query -> retrieve --+-- weak ------------------------------------> no_answer (0 LLM calls)
(only if the follow-up +-- strong ----> select_context --+
depends on history) +-- ambiguous -> grade_documents --+--> generate --> verify_grounding --> finalize
^ | ^ |
+-------- transform_query <-------------+ +-- (once) -----+--> no_answer
(bounded retry)
grade_documents (LLM) and the LLM part of verify_grounding are optional: they run only when free signals
(similarity score, citation coverage, lexical support) are inconclusive. See LLM usage.
Full explanation, node table, diagrams and design trade-offs: docs/ARCHITECTURE.md.
app/
config.py all settings (env-driven)
models.py LLM + embedding factories
vectorstore.py Pinecone index / vector store helpers
ingestion/ loader.py (PDF), chunker.py, pipeline.py
rag/ state.py, prompts.py, nodes.py, graph.py, retriever.py, scoring.py, heuristics.py
errors.py quota / rate-limit detection
api/ main.py (routes), schemas.py, streaming.py, feedback.py
ui/streamlit_app.py chat UI
scripts/ ingest.py, run_sample_queries.py, inspect_retrieval.py
tests/ pytest suite (fake LLM + fake retriever)
docs/ ARCHITECTURE.md, SAMPLE_QUERIES.md, sample_queries.json, SAMPLE_OUTPUTS.md
- Python 3.11+
- A Pinecone API key (free): https://app.pinecone.io
- An LLM API key: OpenAI, Google Gemini or Groq
git clone https://github.com/ShubhamB74/agentic-rag-chatbot.git
cd agentic-rag-chatbot
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env # then edit .env: add PINECONE_API_KEY and one LLM keypython -m scripts.ingest # or: make ingestThis downloads the PDF to data/, extracts text page by page, chunks it, embeds the chunks and upserts them to
Pinecone (creating the serverless index if needed). Re-running is safe: chunk IDs are deterministic.
Use --reset to wipe the namespace first.
If the download is blocked (some servers reject scripts), save the PDF manually as
data/Ebook-Agentic-AI.pdfand run the command again; it will use the local file.
# terminal 1: API -> http://localhost:8000/docs
uvicorn app.api.main:app --reload --port 8000 # or: make api
# terminal 2: UI -> http://localhost:8501
streamlit run ui/streamlit_app.py # or: make uidocker compose run --rm api python -m scripts.ingest # once
docker compose up --build # API on :8000, UI on :8501| Method | Path | Purpose |
|---|---|---|
POST |
/chat |
Answer + chunks + confidence in one JSON response |
POST |
/chat/stream |
Same, as Server-Sent Events (status, token, reset, final, error) |
GET |
/sessions/{id}/history |
Conversation so far |
DELETE |
/sessions/{id} |
Forget a conversation |
POST |
/feedback |
Thumbs up/down for an answer |
GET |
/feedback/stats |
Feedback totals |
GET |
/health |
Status + number of indexed vectors |
Interactive docs: http://localhost:8000/docs.
curl -s -X POST http://localhost:8000/chat \
-H "Content-Type: application/json" \
-d '{"question": "What is Agentic AI?", "session_id": "demo"}'Shape of the response (values are illustrative):
{
"session_id": "demo",
"message_id": "9f2c...",
"question": "What is Agentic AI?",
"standalone_question": "What is Agentic AI?",
"answer": "Agentic AI refers to ... [1][2]",
"answered": true,
"confidence": 0.86,
"confidence_label": "high",
"grounding_score": 0.95,
"stop_reason": null,
"unsupported_claims": [],
"retrieved_chunks": [
{"chunk_id": "a1b2...", "text": "...", "page": 3, "score": 0.71,
"rerank_score": null, "used": true, "citation": 1}
],
"trace": ["skip_rewrite(no history)", "retrieve(attempt=1, hits=6, top=0.66, tier=strong)",
"select_context(used=3)", "generate",
"verify_grounding(method=heuristic, score=0.90, confident)", "finalize"],
"stats": {
"llm_calls": 1, "retrieval_attempts": 1, "retrieval_tier": "strong",
"rewrite_used": false, "grader_used": false, "verifier_used": false, "reranker_used": false,
"grounding_method": "heuristic",
"grounding_signals": {"claims": 4, "citation_coverage": 1.0, "lexical_support": 0.86,
"citations_valid": true, "numbers_supported": true, "words": 74},
"node_latency_ms": {"rewrite_query": 0, "retrieve": 310, "select_context": 0,
"generate": 1650, "verify_grounding": 1, "finalize": 0}
},
"latency_ms": 1970
}Reuse the same session_id to continue a conversation. When the eBook can't answer, you get
"answered": false, "confidence": 0, and a stop_reason such as weak_retrieval, no_relevant_context,
model_declined or failed_grounding_check.
Six queries covering definitions, synthesis, a follow-up (memory), and an out-of-scope question are listed in docs/SAMPLE_QUERIES.md. Run them and generate a report with real outputs:
python -m scripts.run_sample_queries # adaptive mode -> docs/SAMPLE_OUTPUTS.md
python -m scripts.run_sample_queries --mode dev # cheapest: about 1 LLM call per question
python -m scripts.run_sample_queries --mode full # grader + verifier on every question
python -m scripts.run_sample_queries --only 4,5,6 # just these queriesThe runner is built for small free-tier quotas: each finished answer is cached immediately in
data/.sample_cache.json, a quota error stops the run cleanly (nothing is lost), and re-running the same command
only spends LLM calls on queries that are not done yet. See the workflow below.
The pipeline is adaptive: it spends an LLM call only when free signals are inconclusive.
| Situation | LLM calls | Why |
|---|---|---|
| Clear in-scope question, strong retrieval, well-cited answer | 1 | generate only |
| Follow-up that depends on chat history ("those", "what about costs?") | 2 | + query rewrite |
| Ambiguous retrieval | 3 | + relevance grader + fact-checker |
| Obviously out-of-scope (very low similarity) | 0 | declined from the scores alone |
| Verifier rejects the draft | up to the budget | one corrective rewrite, then decline |
Hard cap: MAX_LLM_CALLS (default 4) per question; optional steps are skipped once it is spent.
Every response includes a stats object (llm_calls, retrieval_tier, grader_used, verifier_used,
grounding_method, node_latency_ms, ...) and the Streamlit "Pipeline details" popover shows it.
| Mode | Settings | Use it for |
|---|---|---|
| adaptive (default) | ADAPTIVE_MODE=true |
normal use and demos |
| dev | ENABLE_GRADER=false ENABLE_VERIFIER=false |
day-to-day development, about 1 call per question |
| full | ADAPTIVE_MODE=false |
targeted evaluation: grader on every retrieval, verifier on every answer |
The free grounding heuristic never rejects an answer by itself; only an LLM verdict can. Heuristic-only answers get a slightly lower confidence ceiling than LLM-verified ones.
- Calibrate for free.
python -m scripts.inspect_retrievalruns retrieval only (no LLM calls) and shows the similarity scores and tier for each sample query, then suggestsMIN_RELEVANCE_SCORE/STRONG_SCORE. The defaults are estimates, so do this first. - Develop in dev mode, and switch to
--mode fullonly for a few targeted queries. - Resume instead of restart. If the quota runs out, switch API key or model and re-run the same command.
- Keep
LLM_MAX_RETRIES=1(the default): provider SDKs retry 429 errors with backoff, which can burn extra quota. - Free-tier limits are usually per model and change over time; check your provider's rate-limit page. A lighter model (for example a "flash-lite" variant) or Groq's free tier may allow many more requests per day.
Everything is set via environment variables (see .env.example). The ones you will most likely touch:
| Variable | Default | Notes |
|---|---|---|
LLM_PROVIDER / LLM_MODEL |
openai / gpt-4o-mini |
gemini, groq also supported |
EMBEDDING_PROVIDER / EMBEDDING_MODEL |
pinecone / llama-text-embed-v2 |
changing it requires a new PINECONE_INDEX_NAME |
TOP_K |
6 |
chunks retrieved per query |
CHUNK_SIZE / CHUNK_OVERLAP |
1000 / 150 |
characters; re-ingest with --reset after changing |
USE_RERANKER |
false |
Pinecone reranker; in adaptive mode it only runs for ambiguous retrieval |
ADAPTIVE_MODE |
true |
false = grader + verifier on every question (evaluation) |
ENABLE_GRADER / ENABLE_VERIFIER |
true |
false = never call that LLM step (dev mode) |
MAX_LLM_CALLS |
4 |
hard per-question budget |
MIN_RELEVANCE_SCORE |
0.25 |
top similarity below this is declined with 0 LLM calls |
STRONG_SCORE |
0.50 |
top similarity at or above this skips the grader |
LLM_MAX_RETRIES |
1 |
provider-level retries (each can consume quota) |
MIN_GROUNDING_SCORE |
0.60 |
answers below this are rejected |
Cosine similarity ranges differ between embedding models. Ask a few in-scope questions, look at the score of the
used chunks, then set SCORE_HIGH near a typical good match and SCORE_LOW near the score of an unrelated
chunk. That keeps the retrieval half of the confidence meaningful for your model.
pip install -r requirements-dev.txt
pytest # 95 tests, no network or API keys needed
ruff check .The tests inject a scripted fake LLM and fake retriever into build_graph(...), so the full pipeline (retry loop,
refusals, regeneration, streaming, memory) is verified offline.
| Symptom | Fix |
|---|---|
PINECONE_API_KEY is not set |
create .env from .env.example |
Index ... has dimension X but embedding model produces Y |
you changed embedding model; use a new PINECONE_INDEX_NAME |
| Everything is answered "not in the eBook" | check /health shows vectors > 0 (did ingestion run?); lower MIN_RELEVANCE_SCORE; inspect retrieved_chunks |
| "Very little text was extracted" warning | the PDF is image-based and would need OCR |
429 / "quota or rate limit was reached" |
free-tier quota spent: see Working within a free-tier quota; use dev mode or another key/model |
Legitimate questions are declined with weak_retrieval |
MIN_RELEVANCE_SCORE is too high for your embedding model: run scripts.inspect_retrieval and lower it |
| UI says "API not reachable" | start the API first, or fix the URL in the sidebar |
- Only text is read from the PDF; tables and figures are not interpreted.
- Chat memory is in-process (
MemorySaver) and resets when the server restarts. - The confidence score is a heuristic for ranking answers, not a calibrated probability.
- No authentication or rate limiting yet; add both before exposing the API publicly.


