Loading…
Fetching the tree index.
One tree per problem. The model writes a first solution c1. At each of 10 verification points, 4 independent verifier calls judge the current solution, and each verdict routes one revision. Says correct asks the reviser to re-check anyway; minor fix asks for a repair; critical flaw asks for a fresh attempt. The first successful revision (branch 0 unless it failed) becomes the next spine solution, so the spine runs c1 → c10, and c10 is the final output. The other three revisions at each point are one-step leaves.
Nodes are candidate solutions, coloured by the external judge (gpt-oss-20b comparing the final \boxed{} answer with the gold answer): ✓ correct, ✗ incorrect. The letter names the candidate's final answer; the table under the tree lists them. Edges show what the verifier said about the parent and how the revision was asked for: solid for says correct → re-check, dashed for minor fix → revise, dotted for critical flaw → regenerate. Hover or focus a node or edge for details, and click a node to read the candidate, the critique that produced it, and the reasoning traces. Arrow keys move between nodes; j/k step through problems.
Overview summarises all 1,649 trees: how often each revision mode changes the judge's label, how well verdicts track the judge, and how trees differ by difficulty (how often Qwen3.6-35B solved the problem in MathArena's 12 attempts).
Fetching the tree index.