Preprint · 2026

SHELDON: Reasoning through Dual Search over Hypothesis and Observation Spaces

Harshit Joshi*  ·  Priyank Shethia*  ·  Monica S. Lam

Department of Computer Science, Stanford University  ·  *equal contribution

Dual-space search. Hypotheses guide which evidence to seek; observations are kept as first-class objects, linked to hypotheses by evidence relations, and can seed hypotheses nobody proposed.

Abstract

Test-time scaling has enabled large language models to solve increasingly complex reasoning, discovery, and optimization tasks, but existing approaches often suffer from fixed search strategies, or hypothesis-centric reasoning that under-explores informative evidence.

Inspired by the hypothetico-deductive method and dual-space models of scientific discovery, we introduce SHELDON, a scientific reasoning harness that jointly models hypothesis and observation spaces. By treating observations as first-class objects, SHELDON can identify gaps in its current understanding, uncover new hypotheses, and design investigations that better distinguish among alternatives. Across diagnosis, abstract reasoning, and optimization benchmarks, SHELDON outperforms evaluated baselines by up to 26.7 points. On system root cause analysis tasks, SHELDON discovers 29.8 percentage points more ground-truth causes and investigates 34.4 percentage points more of the discovered causes than other baselines, demonstrating that joint hypothesis-observation search leads to broader and more effective scientific reasoning at test time.

Overview of SHELDON: a hypothesis space and an observation space linked by evidence relations, a search trajectory on LLM_SQL, and bar charts of SHELDON against baselines on four benchmarks.
(a) SHELDON searches a hypothesis space and an observation space jointly; edges in the hypothesis space are refinements. (b) New hypotheses drive improvements on a system-optimization task. (c) Scores on DiagnosisArena, ORCA-Bench, ARC-AGI-2, and QNN circuit topology.

Two spaces, one investigation

SHELDON stands for Systematic Hypothesis Exploration and Learning through Directed Observation and Navigation. Most agent harnesses keep a set of hypotheses and attach evidence to them. SHELDON also keeps the evidence as its own searchable space, so observations gathered while testing one hypothesis can seed a hypothesis nobody had proposed yet.

Hypothesis space H

Candidate explanations, solutions, or design ideas. Organized along axes (method family, assumptions, free parameters) split into regions, so the harness can see which directions are untried.

Observation space O

Evidence and evidence-grounded insights from the case material and tools. Also organized by axes and regions (data sources, time ranges, entities, scales). New observations stay provisional until confirmed and cannot refute a hypothesis on their own.

An evidence relation E labels each hypothesis–observation pair as support, refute, neutral, or unresolved, which gives the agent a quantitative view of how well each working hypothesis is doing.

Each round has four stages

Explore

Inspects the case material, defines axes and regions for both spaces, proposes new working hypotheses in under-explored regions, and may revive set-aside ideas after a cooldown.

CandidateAnalysis

Investigates each hypothesis in depth: assumptions, supporting and refuting evidence, and, when a scoring function exists, an implemented and evaluated artifact.

ComparativeAnalysis

Looks at the working set as a whole and seeks tests whose outcomes discriminate between competing hypotheses or expose shared failure modes.

Exploit

Refines or repairs hypotheses based on the observations and marks each as retained or set aside, guiding the next round of exploration.

The reasoning loop on a cache-eviction incident: working hypotheses and observations evolve round by round, and the evidence table records what each explanation accounts for.

The loop ends when the budget is spent or a round produces no new hypotheses, observations, or directions; a final evidence assessment returns the best-supported solution. SHELDON works with or without an explicit scoring function: on diagnosis and root-cause tasks the observations act as an implicit score.

An ORCA-Bench incident: in round 1 hypotheses h2 and h6 yield observations o7 (Charge errors begin at 20:35) and o14 (Invalid token). In round 2, Explore combines them into a new hypothesis h10, paymentFailure, while h13 and h15 are refinements of round-1 hypotheses.
Why the observation space matters. In this ORCA-Bench incident, two separate round-1 investigations record that Charge errors begin at 20:35 (o7) and that payment returns Invalid token (o14). Neither test targeted a payment fault, but in round 2, Explore combines the two observations into a new hypothesis, paymentFailure, which turns out to be a true root cause.

Results

All methods use GPT-5.6-terra with high reasoning effort; SHELDON delegates code implementation to GPT-5.6-luna. Baselines span reasoning agents (ReAct, GoS), coding agents (Codex), recursive agents (RLM), meta-evolution (EvoX), and autonomous research agents (Arbor, CORAL, ScientistOne).

Headline results. Highest score on all four tasks; the blue arrow marks the gain over the strongest baseline, which changes from task to task.

Diagnosis, abstract reasoning, and circuit discovery

MethodFamilyDiagnosisArenaORCA-BenchARC-AGI-2QNN topology
ReActReasoning agent72.0040.0025.000.80
CodexCoding agent76.0040.0024.440.82
EvoXMeta-evolution——47.500.85
RLMRecursive agent78.0040.0015.000.48
GoSReasoning agent68.0030.00——
SHELDON (ours)Scientific reasoning84.0066.6759.171.00

EvoX requires a scoring function; GoS requires benchmark-specific sub-agent prompts. On QNN topology, SHELDON finds a ten-gate, two-qubit classifier that reaches a perfect score, above both the strongest agent and the human baseline.

Molecular optimization (docking score, higher is better)

SystemADRB1CDK2PPARAPPARGF2PGR
Codex5.976.927.466.389.317.40
ReAct8.498.328.988.407.797.67
EvoX8.637.198.627.418.369.17
CORAL11.019.329.649.699.1010.65
SHELDON (ours)11.2610.269.859.479.859.99

SHELDON is best on four of six targets and exceeds the library-best reference score on CDK2, F2, and PGR.

ADRS system-optimization tasks

Five SkyDiscover ADRS tasks, three runs per method on the same hardware. “Mean” averages the final score over the three runs, “Best” is the best run, and “Cost” is the model spend in USD at which the best solution was reached (lower is better). Cloudcast is a cost to minimize; the other four are scores to maximize. Bold marks the best mean or best run on each task.

SystemEPLB ↑PRISM ↑Cloudcast ↓Txn Scheduling ↑LLM_SQL ↑
MeanBestCost $MeanBestCost $MeanBestCost $MeanBestCost $MeanBestCost $
Codex0.14410.14410.4525.6025.650.23677.10618.100.24306935580.190.68770.71090.22
ReAct0.14400.14400.2826.2426.260.54621.11618.100.49393240000.300.70010.70690.64
EvoX0.13180.14339.7026.2626.262.00618.10618.080.15400941849.510.71490.719612.58
ScientistOne0.13190.1440—26.2626.26—624.10618.10—25333962—0.70340.7137—
Arbor0.14330.14356.8626.2526.262.31618.80618.102.52366837594.790.71300.72181.23
CORAL0.14360.143612.0126.2626.260.30619.00619.000.244077420110.890.71150.71631.23
SHELDON (ours)0.14410.14410.8026.2626.261.28618.09618.076.84419842749.840.72040.724514.42

SHELDON has the best or tied-best mean and best run on every task. The clearest margins are on transaction scheduling and LLM_SQL, where the discovered programs use qualitatively new strategies rather than tuned variants of the seed. Several baselines tie on EPLB, PRISM and Cloudcast, which suggests those tasks have little remaining headroom. ScientistOne is evaluated on its released programs, so its cost is not available. Without any seed program, SHELDON still recovers 95 to 100 percent of the seeded improvement on four of the six optimization tasks.

How the winning LLM_SQL program was found. Early gains refine known ideas; the largest gain arrives late, when Explore opens a new region and combines it with the refined grouping strategy.

What the ablations say

VariantDiagnosisArenaORCA-BenchARC-AGI-2QNN topology
SHELDON (full)84.0066.6759.171.00
without comparative analysis78.0056.6750.000.883
without observation space70.0053.3350.000.867
without continued exploration82.0053.3357.700.733

Removing the observation space costs the most consistently. On ORCA-Bench, SHELDON's hypotheses cover 59 of 77 ground-truth causes (76.6%) versus 36 (46.8%) for GoS, and it investigates 58 of its 59 (98.3%) versus 23 of 36 (63.9%). Of the 59 causes found, 40 came from Explore and 19 from refining existing hypotheses.

Citation

@article{joshi2026sheldon,
  title   = {{SHELDON}: Reasoning through Dual Search over Hypothesis and Observation Spaces},
  author  = {Joshi, Harshit and Shethia, Priyank and Lam, Monica S.},
  journal = {arXiv preprint},
  year    = {2026},
  url     = {https://github.com/stanford-oval/sheldon}
}