Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making

About

Scientific simulators are increasingly being integrated into LLM-driven systems for high-stakes simulation-driven decision-making. However, existing frameworks primarily use LLMs to generate, calibrate, or execute simulators, treating them as black-box interfaces rather than as structured mechanistic systems that can be reasoned about. As a result, current approaches lack the ability to identify, represent, and reason about the assumptions and mechanisms underlying simulator behavior, limiting transparency, auditability, and decision justification. We introduce MechSim, a mechanism-grounded neuro-symbolic reasoning framework for executable scientific simulators. Unlike prior neuro-symbolic approaches that primarily reason over static symbolic structures, MechSim enables LLM agents to reason about the mechanisms, assumptions, and execution behavior of scientific simulators. Our framework represents simulators through a shared structured schema capturing assumptions, variables, mechanism dependencies, and execution traces. On top of this representation, LLM agents operate as constrained reasoning engines that generate structured, evidence-grounded explanations linking simulator outcomes to their underlying mechanisms. We evaluate our approach across multiple high-stakes domains and show that it improves mechanism-level explanation quality, simulator analysis, and downstream decision-making reliability.

Yuhan Yang, Ruipu Li, Alexander Rodr\'iguez• 2026

Related benchmarks

TaskDatasetResultRank
Policy SelectionCOVID-19
Precision@383.3
24
Policy SelectionSupply Chain
Precision@397
24
Policy SelectionMeasles
Precision@378
24
Simulator selectionCOVID-19
Top-1 Regret1.65
24
Simulator selectionSupply Chain
Top-1 Regret0.24
24
Simulator selectionMeasles
Top-1 Regret0.7
24
Mechanism-grounded reasoningSupply Chain
Completeness4.6
12
Mechanism-grounded reasoningMeasles
Completeness4.2
12
Mechanism-grounded reasoningCOVID-19
Completeness3.8
12
Showing 9 of 9 rows

Other info

Follow for update