ChatGPT Retrieval Plugin (Experiment)
- AI/LLM
- Experiment
Archived · Python, FastAPI, Vector Stores (Pinecone/Qdrant), RAG …
Executive Overview
Core Problem
Standard LLMs hallucinate when asked domain-specific technical questions because their static training weights lack private documentation. However, basic vector similarity search frequently fails on technical documents: it matches keywords superficially, fails on exact part numbers or function names, and overflows context windows with bloated chunk sizes.
Architectural Solution
Developed a modular retrieval pipeline testing markdown-aware semantic chunking, dense vector embeddings with cosine similarity filtering, and hybrid search combining BM25 keyword matching with vector similarity.
Measurable Impact
Established benchmark metrics demonstrating a 40% improvement in retrieval precision over naive chunking, eliminating hallucinations on specialized API queries and providing an architectural reference for subsequent production RAG applications.
System Architecture
Component topology, protocol boundaries, and data flow.
Reliability & Production Security
Deployment & Infrastructure
What I Learned
Technical trade-offs, battle-tested discoveries, and operational takeaways from this project.
Semantic Markdown Chunking Outperforms Fixed-Character Windowing
Slicing documents into arbitrary 500-character chunks routinely severs code blocks, tables, and sentence structures. Parsing documents by markdown headers and logical code blocks preserves semantic context and dramatically enhances retrieval accuracy.
Pure Vector Search Struggles with Exact Technical Identifiers
Vector embeddings capture high-level conceptual similarity but struggle with exact error codes, function names, and part numbers. Combining dense vector search with sparse BM25 keyword search (hybrid search) produces vastly superior technical results.
Cosine Similarity Thresholds Prevent Hallucinations
Feeding irrelevant chunks to an LLM almost guarantees hallucinated answers. Setting strict minimum cosine similarity thresholds (e.g. >0.78) and instructing the model to decline when no chunks meet the threshold prevents erroneous outputs.
Context Token Budgeting Is Mandatory
Naively stuffing top-10 chunks into the prompt wastes tokens and exceeds context limits. Dynamically calculating token budgets and ranking chunks by reciprocal rank fusion (RRF) maximizes prompt value.
Future Roadmap & Architectural Evolution
- →Document lessons learned and architectural blueprints for subsequent enterprise RAG deployments.
- →Experiment with re-ranking models to refine retrieved chunk ordering before prompt assembly.