A trained model gives one answer per prompt — unless you spend more compute at inference. Generate many answers and keep the best, or search a tree of reasoning paths. Same frozen weights, better answers, more compute.
The simplest way to spend more compute at inference: don't settle for the first answer. Sample N of them, then keep the single best.
Best-of-N generates every candidate in full, then throws most away. Tree search is smarter — it scores partial reasoning and pours compute into the branches that look promising.
More inference compute buys more quality — but not for free, and not forever. Here's the same budget spent two ways.
Searching at inference means running the model many times — so every run has to be cheap. Next, mixed precision: storing and multiplying weights in fewer bits (FP16, BF16, FP8) to make each of those forward passes faster and smaller.