Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences

About

In this report, we present LOGOS (Language Of Generative Objects in Science), a scientific generative language model that unifies heterogeneous tasks across the natural sciences within a single autoregressive framework based on a shared scientific grammar. It encodes diverse scientific objects and their spatial interactions as token sequences over a common vocabulary. By representing spatial contact and constraint patterns as discrete tokens, the model captures complex structural interactions in a purely sequential manner, without relying on explicit coordinates or geometric neural networks. This unified representation enables a wide range of downstream tasks to be formulated consistently as next-token prediction in the same grammar space, creating strong alignment between continued multi-domain pre-training and downstream objectives. Across diverse tasks, LOGOS consistently matches or outperforms domain-specific baselines, providing preliminary evidence for the feasibility of "one model fits all" in the natural sciences. We train LOGOS models at different scales (1B, 3B, and 8B parameters) and find a consistent positive correlation between model size and performance. This suggests that the future of AI for Science (AI4S) may not lie in building an independent technical stack that is separated from large language models (LLMs). Instead, it may depend on deeply aligning scientific foundation models with LLMs through shared architectures, shared training paradigms, and shared inference infrastructure, so that LLMs can truly become a new entry point for AI4S. We release the model weights and associated resources to facilitate further research.

Mingyang Li, Yurou Liu, Jieping Ye, Bing Su, Ji-Rong Wen, Zheng Wang• 2026

Related benchmarks

TaskDatasetResultRank
Retrosynthesis predictionUSPTO-50k (test)
Top-1 Accuracy74.8
48
Ligand Binding Site PredictionCOACH420--
14
Interaction-Aware Ligand Design for Binding PocketsPDBbind Core Set v2016
Vina Score-7.76
12
Protein editingAAV Medium Difficulty
Fitness69
10
Protein editingAAV Hard Difficulty
Fitness70
10
Protein editingGFP (Medium Difficulty)
Fitness93
10
Protein editingGFP Hard Difficulty
Fitness93
10
Antibody CDR DesignSAbDab CDR-H1 (standard train test)
AAR79.82
9
Antibody CDR DesignSAbDab CDR-H2 (train test)
AAR66.83
9
Antibody CDR DesignSAbDab CDR-L1
AAR85.18
9
Showing 10 of 15 rows

Other info

GitHub

Follow for update