>_
ARSLAN VUZMAL LONEAI & Systems Engineer
Back to Research Notes Index
UC Berkeley // 2024TOPIC: ML SystemsTechnical Note

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

AUTHORS: Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V. Le, Christopher Ré, Azalia Mirhoseini | PUBLISHED & REVIEWED: 2025-01-18
CORE ARCHITECTURAL THESIS & FINDING

Coverage of correct solutions scales as a power law with repeated test-time sampling and external verifiers, rivaling models orders of magnitude larger.

MY INTERPRETATION

Allocating compute to test-time search and deterministic verification (such as unit tests and AST validators) delivers higher reliability than relying solely on larger pre-trained parameters.

PRACTICAL IMPLEMENTATION

Built automated sampling and verification iteration loops in Orchestrion where candidate code diffs are evaluated against test suites before operator presentation.

RETRIEVAL_INTELLIGENCE // HYBRID_RAGVECTOR RETRIEVAL & RERANKING
STEP 01
Chunking
512 token splits
STEP 02
Embedding
Dense vectors
STEP 03
Qdrant Search
Cosine sim (k=25)
STEP 04
Cross-Encoder
Rerank top-5
STEP 05
Grounded Gen
With citations
SYSTEM LIMITATIONS, RUNTIME OVERHEAD & PRODUCTION CONSTRAINTS
  • Inference token costs scale with sampling count N without early-exit verifiers.
  • Requires deterministic verifiers to avoid rewarding plausible-sounding hallucinations.
EVIDENCE & REPRODUCIBILITY METHODOLOGY

Empirical analysis of arXiv:2407.21787 and test-time compute scaling laws.