Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations

About

Modern artificial intelligence excels at prediction but cannot explain. From large language models to AI-for-science systems, today's machines answer what by recombining patterns already present in the human literature, yet they cannot reason out why a phenomenon must arise from underlying principles even though explanation, not prediction, lies at the heart of scientific discovery. Here we ask whether the structure of scientific explanation can be operationalized to guide how a machine generates hypotheses. We introduce DN-Hypo-Pipeline, a hypothesis-generation framework that adopts a layered, explanation-theoretic scaffold: Hempel's deductive-nomological (DN) model supplies the output form and deductive validity of a hypothesis, Salmon's causal-process account supplies an organizing constraint on where to search for the governing laws, and Armstrong's view of laws as relations between universals supplies the bridge from a phenomenon's constituent processes to the laws that may be associated with it. Rather than searching the space of what has been written, the framework searches the space of what principles govern a phenomenon: given an explanandum, it abstracts the universals instantiated in the phenomenon's formation process, retrieves the laws relating those universals, and deductively reconstructs a new, testable explanation. Evaluated in data-science modeling and judged by both LLMs and human experts, hypotheses generated through this principled reasoning significantly outperform those from direct prompting. Crucially, we translated the two highest-scoring hypotheses into novel algorithms one that reduces the Transformer's theoretical complexity with only minimal performance loss, and another that achieves competitive accuracy with substantially fewer parameters.

Lei Lin, Xinlong Pan, Ronghao Wang, Chunbao Zhou, Jue Wang, Yangang Wang, Ivana Rasovska• 2026

Related benchmarks

TaskDatasetResultRank
Language ModelingWikiText-2 (test)
PPL25.68
2416
Language ModelingWikiText2 (val)
Perplexity (PPL)24.47
436
Word SimilarityWordSim-353
Spearman Rho0.686
121
Word SimilarityRW-STANFORD
Spearman Correlation0.253
13
Word AnalogyGoogle analogy task
Semantic Accuracy32.8
7
Sequence PredictionWikiText-2 (average)
Average PPL25.07
5
Scientific Modeling Idea GenerationScientific Modeling Ideas 12 evaluation instances
Wilcoxon Statistic66
3
Showing 7 of 7 rows

Other info

Follow for update