Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

MolScribe: Robust Molecular Structure Recognition with Image-To-Graph Generation

About

Molecular structure recognition is the task of translating a molecular image into its graph structure. Significant variation in drawing styles and conventions exhibited in chemical literature poses a significant challenge for automating this task. In this paper, we propose MolScribe, a novel image-to-graph generation model that explicitly predicts atoms and bonds, along with their geometric layouts, to construct the molecular structure. Our model flexibly incorporates symbolic chemistry constraints to recognize chirality and expand abbreviated structures. We further develop data augmentation strategies to enhance the model robustness against domain shifts. In experiments on both synthetic and realistic molecular images, MolScribe significantly outperforms previous models, achieving 76-93% accuracy on public benchmarks. Chemists can also easily verify MolScribe's prediction, informed by its confidence estimation and atom-level alignment with the input image. MolScribe is publicly available through Python and web interfaces: https://github.com/thomas0809/MolScribe.

Yujie Qian, Jiang Guo, Zhengkai Tu, Zhening Li, Connor W. Coley, Regina Barzilay• 2022

Related benchmarks

TaskDatasetResultRank
Chemical Structure Recognitionhand-drawn images (test)
Acc (T=1)10.2
20
Molecular structure recognitionJPO (450)
Accuracy76.2
19
Optical Chemical Structure RecognitionUSPTO-10K
Accuracy (Exact Match)96
12
Optical Chemical Structure RecognitionWildMol-10K
Exact Match Accuracy66.4
12
Optical Chemical Structure RecognitionCLEF 992
Exact Match Accuracy88.9
12
Optical Chemical Structure RecognitionUSPTO (5719)
Exact Match Accuracy92.6
12
Optical Chemical Structure RecognitionUOB (5740)
Exact Match Accuracy87.9
12
Molecular structure recognitionUOB Synthetic
Exact Matching Accuracy87.9
11
Chemical Structure RecognitionChemPix out of domain (test)
Exact Match Accuracy22.8
11
Molecular structure recognitionCLEF Synthetic
Exact Match Accuracy88.9
10
Showing 10 of 27 rows

Other info

Follow for update