ChatGPT Retrieval Plugin (Experiment)

  • AI/LLM
  • Experiment

Archived · Python, FastAPI, Vector Stores (Pinecone/Qdrant), RAG …

Executive Overview

An exploratory research project investigating retrieval-augmented generation (RAG) patterns, vector embeddings, chunking strategies, and hybrid semantic search architectures to ground LLM responses in proprietary documentation.
The Challenge & Bottleneck

Core Problem

Standard LLMs hallucinate when asked domain-specific technical questions because their static training weights lack private documentation. However, basic vector similarity search frequently fails on technical documents: it matches keywords superficially, fails on exact part numbers or function names, and overflows context windows with bloated chunk sizes.

Engineering Approach

Architectural Solution

Developed a modular retrieval pipeline testing markdown-aware semantic chunking, dense vector embeddings with cosine similarity filtering, and hybrid search combining BM25 keyword matching with vector similarity.

Quantified Outcomes

Measurable Impact

Established benchmark metrics demonstrating a 40% improvement in retrieval precision over naive chunking, eliminating hallucinations on specialized API queries and providing an architectural reference for subsequent production RAG applications.

System Architecture

Component topology, protocol boundaries, and data flow.

ChatGPT Retrieval Plugin (Experiment) System Topology
Architecture Flow
CLIENT CONSUMERWeb & API CallsHTTPS / REST PayloadsJSON Schema InputBOUNDARY GATEWAYNginx / Reverse ProxyTLS TerminationRate Limiting & AuthNSERVICE CORE LOGIC• Domain Services & Controllers• DTO Runtime Validation• AWS Secrets Manager Config• Health Readiness ProbesPERSISTENCEPostgreSQL / RedisACID TransactionsDocker / EKS Hosted

Reliability & Production Security

Implemented strict cosine similarity threshold gating to reject low-confidence retrievals with graceful 'I do not have enough context' fallbacks, token budgeting algorithms to prevent context window overflow, and input query sanitization.

Deployment & Infrastructure

Containerized FastAPI Python service connecting to Pinecone and local Qdrant vector databases, deployed in isolated Docker containers with automated evaluation benchmarks.
Engineering Post-Mortem & Insights

What I Learned

Technical trade-offs, battle-tested discoveries, and operational takeaways from this project.

1

Semantic Markdown Chunking Outperforms Fixed-Character Windowing

Slicing documents into arbitrary 500-character chunks routinely severs code blocks, tables, and sentence structures. Parsing documents by markdown headers and logical code blocks preserves semantic context and dramatically enhances retrieval accuracy.

2

Pure Vector Search Struggles with Exact Technical Identifiers

Vector embeddings capture high-level conceptual similarity but struggle with exact error codes, function names, and part numbers. Combining dense vector search with sparse BM25 keyword search (hybrid search) produces vastly superior technical results.

3

Cosine Similarity Thresholds Prevent Hallucinations

Feeding irrelevant chunks to an LLM almost guarantees hallucinated answers. Setting strict minimum cosine similarity thresholds (e.g. >0.78) and instructing the model to decline when no chunks meet the threshold prevents erroneous outputs.

4

Context Token Budgeting Is Mandatory

Naively stuffing top-10 chunks into the prompt wastes tokens and exceeds context limits. Dynamically calculating token budgets and ranking chunks by reciprocal rank fusion (RRF) maximizes prompt value.

Future Roadmap & Architectural Evolution

  • →Document lessons learned and architectural blueprints for subsequent enterprise RAG deployments.
  • →Experiment with re-ranking models to refine retrieved chunk ordering before prompt assembly.
ChatGPT Retrieval Plugin (Experiment) | Siddhant Ghosh