Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

What Does a Chemical Language Model Know About Molecules?

About

Chemical language models (cLMs) are widely assumed to learn surface-level syntactic patterns rather than learning meaningful molecular semantics. Here, we apply sparse autoencoders (SAEs) to MolFormer, an encoder-only cLM, to mechanistically examine how molecular representations are built across layers. We discover that early layers rely on position-tracking latents to parse molecular grammar, while later layers encode atom-in-substructure and pharmacologically relevant features. Additionally, we show that non-canonical SMILES produce more disruptive representation shifts than invalid SMILES, driven by position-latent disruption propagating across layers. To support further exploration, we develop InterMol, an interactive visualizer for SAE activations on molecular strings and structures.

Christian Kenneth, Etowah Adams, Liam Bai, Gerard JP van Westen• 2026

Related benchmarks

TaskDatasetResultRank
ADMET Properties PredictionTDC AMES
AUROC0.82
20
Drug discovery classificationClinTox
ROC-AUC91
15
ADMET Properties PredictionTDC HIA Hou
AUROC98
15
ADMET Properties PredictionTDC Pgp Broccatelli
AUROC0.93
15
ADMET Properties PredictionTDC Bioavailability Ma
AUROC0.72
15
Absorption RegressionFreeSolv (train-test)
RMSE1.2
8
Absorption RegressionLipophilicity AstraZeneca (train test)
RMSE0.75
8
Absorption RegressionAqSolDB (train-test)
RMSE1.65
8
ADMET ClassificationPAMPA NCATS
ROC-AUC76
8
ADMET ClassificationBBB
ROC-AUC90
8
Showing 10 of 21 rows

Other info

Follow for update