01
Retrieval
Hybrid dense plus sparse retrieval over Qdrant and Elasticsearch, with Kafka feeding real-time updates, behind a QLoRA-fine-tuned Llama 3.1.
07Real-Time Retrieval · Source Verification
QLoRA-fine-tuned Llama 3.1 with hybrid dense + sparse retrieval and automated citation and fact-verification.
Hybrid dense plus sparse retrieval over Qdrant and Elasticsearch, with Kafka feeding real-time updates, behind a QLoRA-fine-tuned Llama 3.1.
vLLM streaming at 150ms P95 time-to-first-token, holding 1K+ concurrent users on 4×A10G GPUs.