Research & Papers

Viable Path Entropy: new metric reveals 1.5B model can outthink 3B

A paper from Tiantian Zhang introduces a novel way to measure AI's reasoning depth under reflection.

Deep Dive

In a new paper on arXiv, Tiantian Zhang introduces Mirror Horizon and Viable Path Entropy (VPE), a formal measure for evaluating an AI system's bounded reflective capability. Mirror Theory argues that intelligence should be judged not by what a model represents in a single pass, but by what coherent continuations it can sustain under repeated reflection. VPE operationalizes this by decomposing capability into two parts: the probability of reaching a viable continuation and the diversity of verified continuation modes across successful rollouts. The theory introduces concepts like local underdetermining constraint, taste as invariant-selecting pressure, and reflection as taste-guided resolution of underdetermination.

Zhang tested the theory on GSM8K math reasoning tasks using Qwen2.5-Instruct models (1.5B and 3B) with 32 sampled rollouts per problem and two reflection horizons (96 and 160 token budgets). Results showed that increasing the token budget significantly expanded verified reachability, reduced zero-reachability, increased verified-mode entropy, and improved smoothed VPE. At 160 tokens, the smaller Qwen2.5-1.5B realized the strongest mirror horizon among the tested models, outperforming the larger Qwen2.5-3B. This finding challenges the assumption that more parameters always yield better reasoning, instead highlighting that accessible verified continuation capacity under bounded reflection protocols better defines actual capability.

Key Points
  • VPE decomposes capability into probability of reaching a viable continuation and diversity of verified continuation modes
  • Increasing token budget from 96 to 160 on GSM8K boosted verified reachability and reduced zero-reachability
  • Qwen2.5-1.5B achieved a stronger mirror horizon than Qwen2.5-3B at 160 tokens, despite having fewer parameters

Why It Matters

Challenges the parameter-count arms race: AI capability may hinge on accessible reasoning paths under reflection.

📬 Get the top 10 AI stories daily