Department of Computer Science, Stanford University · *equal contribution
Test-time scaling has enabled large language models to solve increasingly complex reasoning, discovery, and optimization tasks, but existing approaches often suffer from fixed search strategies, or hypothesis-centric reasoning that under-explores informative evidence.
Inspired by the hypothetico-deductive method and dual-space models of scientific discovery, we introduce SHELDON, a scientific reasoning harness that jointly models hypothesis and observation spaces. By treating observations as first-class objects, SHELDON can identify gaps in its current understanding, uncover new hypotheses, and design investigations that better distinguish among alternatives. Across diagnosis, abstract reasoning, and optimization benchmarks, SHELDON outperforms evaluated baselines by up to 26.7 points. On system root cause analysis tasks, SHELDON discovers 29.8 percentage points more ground-truth causes and investigates 34.4 percentage points more of the discovered causes than other baselines, demonstrating that joint hypothesis-observation search leads to broader and more effective scientific reasoning at test time.
SHELDON stands for Systematic Hypothesis Exploration and Learning through Directed Observation and Navigation. Most agent harnesses keep a set of hypotheses and attach evidence to them. SHELDON also keeps the evidence as its own searchable space, so observations gathered while testing one hypothesis can seed a hypothesis nobody had proposed yet.
Candidate explanations, solutions, or design ideas. Organized along axes (method family, assumptions, free parameters) split into regions, so the harness can see which directions are untried.
Evidence and evidence-grounded insights from the case material and tools. Also organized by axes and regions (data sources, time ranges, entities, scales). New observations stay provisional until confirmed and cannot refute a hypothesis on their own.
An evidence relation E labels each hypothesis–observation pair as support, refute, neutral, or unresolved, which gives the agent a quantitative view of how well each working hypothesis is doing.
Inspects the case material, defines axes and regions for both spaces, proposes new working hypotheses in under-explored regions, and may revive set-aside ideas after a cooldown.
Investigates each hypothesis in depth: assumptions, supporting and refuting evidence, and, when a scoring function exists, an implemented and evaluated artifact.
Looks at the working set as a whole and seeks tests whose outcomes discriminate between competing hypotheses or expose shared failure modes.
Refines or repairs hypotheses based on the observations and marks each as retained or set aside, guiding the next round of exploration.
The loop ends when the budget is spent or a round produces no new hypotheses, observations, or directions; a final evidence assessment returns the best-supported solution. SHELDON works with or without an explicit scoring function: on diagnosis and root-cause tasks the observations act as an implicit score.
Charge errors begin at 20:35 (o7) and that payment returns Invalid token (o14). Neither test targeted a payment fault, but in round 2, Explore combines the two observations into a new hypothesis, paymentFailure, which turns out to be a true root cause.All methods use GPT-5.6-terra with high reasoning effort; SHELDON delegates code implementation to GPT-5.6-luna. Baselines span reasoning agents (ReAct, GoS), coding agents (Codex), recursive agents (RLM), meta-evolution (EvoX), and autonomous research agents (Arbor, CORAL, ScientistOne).
| Method | Family | DiagnosisArena | ORCA-Bench | ARC-AGI-2 | QNN topology |
|---|---|---|---|---|---|
| ReAct | Reasoning agent | 72.00 | 40.00 | 25.00 | 0.80 |
| Codex | Coding agent | 76.00 | 40.00 | 24.44 | 0.82 |
| EvoX | Meta-evolution | — | — | 47.50 | 0.85 |
| RLM | Recursive agent | 78.00 | 40.00 | 15.00 | 0.48 |
| GoS | Reasoning agent | 68.00 | 30.00 | — | — |
| SHELDON (ours) | Scientific reasoning | 84.00 | 66.67 | 59.17 | 1.00 |
EvoX requires a scoring function; GoS requires benchmark-specific sub-agent prompts. On QNN topology, SHELDON finds a ten-gate, two-qubit classifier that reaches a perfect score, above both the strongest agent and the human baseline.
| System | ADRB1 | CDK2 | PPARA | PPARG | F2 | PGR |
|---|---|---|---|---|---|---|
| Codex | 5.97 | 6.92 | 7.46 | 6.38 | 9.31 | 7.40 |
| ReAct | 8.49 | 8.32 | 8.98 | 8.40 | 7.79 | 7.67 |
| EvoX | 8.63 | 7.19 | 8.62 | 7.41 | 8.36 | 9.17 |
| CORAL | 11.01 | 9.32 | 9.64 | 9.69 | 9.10 | 10.65 |
| SHELDON (ours) | 11.26 | 10.26 | 9.85 | 9.47 | 9.85 | 9.99 |
SHELDON is best on four of six targets and exceeds the library-best reference score on CDK2, F2, and PGR.
Five SkyDiscover ADRS tasks, three runs per method on the same hardware. “Mean” averages the final score over the three runs, “Best” is the best run, and “Cost” is the model spend in USD at which the best solution was reached (lower is better). Cloudcast is a cost to minimize; the other four are scores to maximize. Bold marks the best mean or best run on each task.
| System | EPLB ↑ | PRISM ↑ | Cloudcast ↓ | Txn Scheduling ↑ | LLM_SQL ↑ | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean | Best | Cost $ | Mean | Best | Cost $ | Mean | Best | Cost $ | Mean | Best | Cost $ | Mean | Best | Cost $ | |
| Codex | 0.1441 | 0.1441 | 0.45 | 25.60 | 25.65 | 0.23 | 677.10 | 618.10 | 0.24 | 3069 | 3558 | 0.19 | 0.6877 | 0.7109 | 0.22 |
| ReAct | 0.1440 | 0.1440 | 0.28 | 26.24 | 26.26 | 0.54 | 621.11 | 618.10 | 0.49 | 3932 | 4000 | 0.30 | 0.7001 | 0.7069 | 0.64 |
| EvoX | 0.1318 | 0.1433 | 9.70 | 26.26 | 26.26 | 2.00 | 618.10 | 618.08 | 0.15 | 4009 | 4184 | 9.51 | 0.7149 | 0.7196 | 12.58 |
| ScientistOne | 0.1319 | 0.1440 | — | 26.26 | 26.26 | — | 624.10 | 618.10 | — | 2533 | 3962 | — | 0.7034 | 0.7137 | — |
| Arbor | 0.1433 | 0.1435 | 6.86 | 26.25 | 26.26 | 2.31 | 618.80 | 618.10 | 2.52 | 3668 | 3759 | 4.79 | 0.7130 | 0.7218 | 1.23 |
| CORAL | 0.1436 | 0.1436 | 12.01 | 26.26 | 26.26 | 0.30 | 619.00 | 619.00 | 0.24 | 4077 | 4201 | 10.89 | 0.7115 | 0.7163 | 1.23 |
| SHELDON (ours) | 0.1441 | 0.1441 | 0.80 | 26.26 | 26.26 | 1.28 | 618.09 | 618.07 | 6.84 | 4198 | 4274 | 9.84 | 0.7204 | 0.7245 | 14.42 |
SHELDON has the best or tied-best mean and best run on every task. The clearest margins are on transaction scheduling and LLM_SQL, where the discovered programs use qualitatively new strategies rather than tuned variants of the seed. Several baselines tie on EPLB, PRISM and Cloudcast, which suggests those tasks have little remaining headroom. ScientistOne is evaluated on its released programs, so its cost is not available. Without any seed program, SHELDON still recovers 95 to 100 percent of the seeded improvement on four of the six optimization tasks.
| Variant | DiagnosisArena | ORCA-Bench | ARC-AGI-2 | QNN topology |
|---|---|---|---|---|
| SHELDON (full) | 84.00 | 66.67 | 59.17 | 1.00 |
| without comparative analysis | 78.00 | 56.67 | 50.00 | 0.883 |
| without observation space | 70.00 | 53.33 | 50.00 | 0.867 |
| without continued exploration | 82.00 | 53.33 | 57.70 | 0.733 |
Removing the observation space costs the most consistently. On ORCA-Bench, SHELDON's hypotheses cover 59 of 77 ground-truth causes (76.6%) versus 36 (46.8%) for GoS, and it investigates 58 of its 59 (98.3%) versus 23 of 36 (63.9%). Of the 59 causes found, 40 came from Explore and 19 from refining existing hypotheses.
@article{joshi2026sheldon,
title = {{SHELDON}: Reasoning through Dual Search over Hypothesis and Observation Spaces},
author = {Joshi, Harshit and Shethia, Priyank and Lam, Monica S.},
journal = {arXiv preprint},
year = {2026},
url = {https://github.com/stanford-oval/sheldon}
}