LLM-Based Subjective Answer Evaluation (DRF)
- AI/LLM
- Backend
In Progress · Django, Django REST Framework, LLM Inference (LLaMA), Celery …
Executive Overview
Core Problem
Subjective long-form answers are time-consuming to grade manually and exhibit extreme inconsistency across human interviewers. Naively querying LLMs synchronously via HTTP blocks web threads for 5-15 seconds per evaluation, invites prompt injection vulnerabilities, and suffers from non-deterministic scoring variance.
Architectural Solution
Designed an asynchronous evaluation pipeline where submissions are ingested via DRF endpoints, stored with pending status, and queued into Celery workers for batched model inference. Implemented system prompts with strict few-shot grading rubrics, Pydantic-based JSON schema enforcement on model outputs, and heuristic sanity checks that detect score drift.
Measurable Impact
Decreased evaluation turnaround time from hours to under 8 seconds per answer, achieved an 88% alignment score with senior human evaluators, and preserved 100% web API responsiveness during peak candidate submission batches.
System Architecture
Component topology, protocol boundaries, and data flow.
Reliability & Production Security
Deployment & Infrastructure
What I Learned
Technical trade-offs, battle-tested discoveries, and operational takeaways from this project.
Synchronous LLM Calls Destroy Web API Throughput
LLM inference takes seconds, not milliseconds. Decoupling submission intake from evaluation via asynchronous message queues is mandatory to keep user-facing APIs responsive.
Strict JSON Schema Enforcement Prevents Production Crashes
Raw model responses cannot be trusted blindly. Validating LLM outputs against strict schemas (with automatic retry on malformed JSON) prevents parsing exceptions in downstream databases.
Few-Shot Rubrics Significantly Reduce Score Variance
Open-ended prompts produce fluctuating evaluations. Providing 3-5 concrete benchmark examples (poor, average, exceptional) in the system prompt anchors the model and slashes score volatility.
Prompt Injection Defense Must Be Built Into Input Pre-Processing
Untrusted user submissions will inevitably attempt jailbreaks. Stripping command-like tokens and using dual-prompt validation (grading criteria separate from candidate input) prevents manipulation.
Future Roadmap & Architectural Evolution
- →Implement fine-tuned LoRA adapter checkpoints tailored to specific domain vocabularies.
- →Add an interactive reviewer dashboard for human auditors to inspect and calibrate borderline grading decisions.