Job: Smart multi-source information fetching
Query → Auto-detect Strategy → Execute Search → Deduplicate → Results
Three Search Modes:
Semantic Search (Vector-based)
- Converts query to 3072D vector (OpenAI
text-embedding-3-large) - Searches Pinecone for similar document chunks
- Fast: Pre-computed embeddings at upload time
Graph Search (Knowledge-based)
- Searches Neo4j for matching entities
- Traverses 2-hop relationships
- Returns structured knowledge graph data
Hybrid Search (Best of both)
- Runs both searches in parallel
- Weighted fusion (50/50 by default)
- Auto-selected based on query keywords
Query Analysis Examples:
"What is AI?"→ Semantic (keywords: "what is", "explain")"How is AI related to ML?"→ Graph (keywords: "related to", "connected")"Explain deep learning applications"→ Hybrid (both patterns)
Job: Autonomous agent deciding WHAT tools to use
Query → GPT-4 Intent Analysis → Select Tools → Build Plan → Execute Parallel
7 Specialized Tools:
vector_search- Semantic similarity in Pineconeentity_search- Find graph nodes in Neo4jgraph_traversal- Multi-hop relationship explorationrelationship_path- Shortest path between entitiescypher_query- Custom Neo4j queriessemantic_similarity- Compare concept embeddingshybrid_search- Combined vector + graph
Smart Execution:
- Parallel execution where possible
- Sequential for dependencies (e.g., entity_search → graph_traversal)
- Priority-based ordering
Job: Multi-step reasoning with self-refinement
Query → Plan → Execute → Synthesize → Check Confidence → Refine → Answer
Streaming Reasoning Steps:
- Thought: "Analyzing query..."
- Action: "Created 3-step plan (complexity: moderate)"
- Tool Execution: "vector_search found 5 chunks, entity_search found 2 entities"
- Observation: "Found relationship: ML -[INCLUDES]-> Deep Learning"
- Synthesis: GPT-4 combines all results into coherent answer
- Refinement: If confidence < 85%, ask follow-up queries (max 3 iterations)
- Final Answer: Markdown-formatted response with citations
Self-Improvement Loop:
while (confidence < 0.85 && iterations < 3) {
critique = "Answer lacks specific examples"
additionalQuery = "Find AI healthcare case studies"
newResults = await queryMore()
answer = await resynthesize()
}Redis Caching (3 places):
- Query results:
search:${userId}:${query}(30 min TTL) - Intent analysis:
query-intent:${query}(1 hour) - Reasoning chains:
reasoning:${userId}:${query}(1 hour)
Parallel Processing: Vector + Graph searches run simultaneously
Pre-computed Embeddings: All documents embedded at upload, not query time
Indexed Databases: Pinecone (vector index) + Neo4j (graph index)
Model: OpenAI text-embedding-3-large
Benefits:
- Higher semantic precision (captures nuances)
- Better similarity matching (accurate relevance scores)
- Richer context representation
- State-of-the-art performance
Trade-off: More storage, slightly slower search (worth it for accuracy)
User: "How does machine learning relate to deep learning?"
LAYER 1 - Intelligent Retrieval:
├─ Detects: "hybrid" strategy (semantic + relational keywords)
├─ Semantic: Pinecone returns 5 relevant chunks
└─ Graph: Neo4j finds ML/DL entities + [INCLUDES] relationship
LAYER 2 - Query Planner:
├─ GPT-4 analyzes intent: "relational query"
├─ Selects tools: [vector_search, entity_search, graph_traversal]
└─ Executes in parallel → 3 tool results
LAYER 3 - Reasoning Engine:
├─ Synthesizes: "Deep learning is a subset of ML that uses neural networks..."
├─ Confidence: 92%
└─ Streams answer with reasoning chain visible to user