Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SciDesignBench: Benchmarking and Improving Language Models for Scientific Inverse Design

About

Many of the most important problems in science and engineering are inverse problems: given a desired outcome, find a design that achieves it. Evaluating whether a candidate meets the spec is often routine; a binding energy can be computed, a reactor yield simulated, a pharmacokinetic profile predicted. But searching a combinatorial design space for inputs that satisfy those targets is fundamentally harder. We introduce SciDesignBench, a benchmark of 520 simulator-grounded tasks across 14 scientific domains and five settings spanning single-shot design, short-horizon feedback, long-horizon refinement, and seed-design optimization. On the 10-domain shared-core subset, the best zero-shot model reaches only 29.0% success despite substantially higher parse rates. Simulator feedback helps, but the leaderboard changes with horizon: Sonnet 4.5 is strongest in one-turn de novo design, whereas Opus 4.6 is strongest after 20 turns of simulator-grounded refinement. Providing a starting seed design reshuffles the leaderboard again, demonstrating that constrained modification requires a fundamentally different capability from unconstrained de novo generation. We then introduce RLSF, a simulator-feedback training recipe. An RLSF-tuned 8B model raises single-turn success rates by 8-17 percentage points across three domains. Together, these results position simulator-grounded inverse design as both a benchmark for scientific reasoning and a practical substrate for amortizing expensive test-time compute into model weights.

David van Dijk, Ivan Vrkic• 2026

Related benchmarks

TaskDatasetResultRank
De Novo Design10 de novo shared-core domains--
21
Design Optimization10 optimization shared-core domains--
14
Long-Horizon Evaluation with Simulator Feedback10 manifest-defined shared-core domains (de novo)--
7
Scientific DesignSciDesignBench de novo domains v1--
7
Showing 4 of 4 rows

Other info

Follow for update