Back to Research Notes Index
Stanford // 2024TOPIC: RAGTechnical Note
Lost in the Middle: How Language Models Use Long Contexts
AUTHORS: Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, Percy Liang | PUBLISHED & REVIEWED: 2025-02-02
CORE ARCHITECTURAL THESIS & FINDING
Language models exhibit U-shaped performance curves in long-context retrieval, performing best when relevant information is at the very beginning or end of the input context.
MY INTERPRETATION
Naively stuffing entire document collections into large context windows leads to retrieval blindspots. High-precision RAG requires reranking and strategic chunk positioning.
PRACTICAL IMPLEMENTATION
Built cross-encoder reranking and perimeter context ordering in SourceLatch to place high-relevance chunks at optimal attention boundaries.
RETRIEVAL_INTELLIGENCE // HYBRID_RAGVECTOR RETRIEVAL & RERANKING
STEP 01
Chunking
512 token splits
STEP 02
Embedding
Dense vectors
STEP 03
Qdrant Search
Cosine sim (k=25)
STEP 04
Cross-Encoder
Rerank top-5
STEP 05
Grounded Gen
With citations
SYSTEM LIMITATIONS, RUNTIME OVERHEAD & PRODUCTION CONSTRAINTS
- •Cross-encoder reranking introduces slight latency overhead (approx. 40ms per query).
EVIDENCE & REPRODUCIBILITY METHODOLOGY
Attention curve analysis on multi-document question answering (arXiv:2307.03172).