← BlogLLM Evaluation

Verifying LLM Outputs: From Consensus to Process Scores

Isaac Kargar3 min read

  • LLM
  • Verification
  • Evaluation
  • AI Safety

Verification can happen at several points in a generated solution. A system may compare final answers across samples, use an outcome reward model to rank complete solutions, or score each intermediate step. These methods answer different questions and should be compared with their benchmark, model, sample count, and scoring protocol.

Consensus voting

Majority voting samples several answers and returns the most common final answer. In Solving Quantitative Reasoning Problems with Language Models, Minerva 540B reached 33.6% on the MATH benchmark with one greedy sample. With maj1@64, which samples 64 solutions and selects the most common parsed answer, it reached 50.3%.

The MATH score is accuracy on the test set. Minerva parses the final answer and compares mathematically equivalent forms with SymPy. The 33.6% and 50.3% values are therefore a comparison between one sample and 64-sample majority voting for Minerva 540B, rather than a general consensus rate.

Best-of-N with an outcome reward model

An outcome reward model (ORM) scores a complete solution, usually from the final answer or the full sequence, and the system chooses the highest-scoring candidate. The reward model is a selector. It does not prove that the chosen reasoning is correct, and its quality depends on the data and labels used to train it.

Process reward models

Let’s Verify Step by Step compares outcome supervision with process supervision. Its process reward model (PRM) predicts whether each step is correct, then combines the step probabilities into a score for the solution. The large-scale experiment uses a fixed GPT-4-derived generator and evaluates selection on a representative subset of 500 MATH test problems. For each problem, the researchers generated 1,860 solutions and selected the best candidate according to the reward model. Correctness was checked from the final answer.

SelectorMATH problems solved at best-of-1,860
Majority voting69.6%
Outcome reward model (ORM)72.4%
Process reward model (PRM)78.2%

The table reports the paper’s best-of-1,860 selection experiment on its representative MATH test subset. A PRM estimates step quality from learned labels; it does not prove reasoning faithfulness or correctness. See the paper’s results for the exact training and scoring setup.

The PRM’s advantage in this experiment comes from finer supervision. Human labelers marked solution steps as positive, negative, or neutral in PRM800K. An ORM receives only an outcome label, so a solution that reaches the right final answer through a bad intermediate step can be treated as correct. That failure mode is one reason to inspect the process signal separately.

The authors also report active learning for process supervision. They select convincing wrong answers for labeling rather than sampling all solutions uniformly, and estimate a 2.6-times improvement in data efficiency for that collection strategy. It is a result about their labeling setup, not a universal reduction in the cost of training a reward model.

A process reward model marks individual reasoning steps before ranking complete solutions
The source figure compares outcome and process reward models across test-time search budgets. The readable values above come from the same Let’s Verify Step by Step results.

What verification can and cannot establish

Consensus reduces the chance of choosing an uncommon answer under the assumptions of the task. An ORM can rank candidates when its outcome labels are reliable. A PRM can expose a likely error earlier in a chain, but it remains a learned estimator and can be miscalibrated or trained on a biased set of steps. A verification layer should therefore be evaluated on the target task, with held-out problems and a scoring rule that checks the property the application actually needs.

Work with Nazmi

Build your AI system with Nazmi.

Tell us what you are building, what exists today, and where your team needs help.

Start a conversation or book a 20-minute call →