How Uncertainty Estimation Scales with Sampling in Reasoning Models
Maksym Del, Markus Kängsepp, Marharyta Domnich, et al.
This paper investigates how well AI reasoning models can estimate their own uncertainty by examining two methods: having models verbalize their confidence and checking if multiple attempts reach the same answer (self-consistency). Testing across math, science, and humanities tasks, the researchers found that combining both methods works best—even just two samples together improved accuracy by up to 12% better than using either method alone, with the strongest results in mathematics.
uncertainty estimationreasoning modelschain-of-thoughtself-consistency