Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Ai2 Scholar QA: Organized Literature Synthesis with Attribution

About

Retrieval-augmented generation is increasingly effective in answering scientific questions from literature, but many state-of-the-art systems are expensive and closed-source. We introduce Ai2 Scholar QA, a free online scientific question answering application. To facilitate research, we make our entire pipeline public: as a customizable open-source Python package and interactive web app, along with paper indexes accessible through public APIs and downloadable datasets. We describe our system in detail and present experiments analyzing its key design decisions. In an evaluation on a recent scientific QA benchmark, we find that Ai2 Scholar QA outperforms competing systems.

Amanpreet Singh, Joseph Chee Chang, Chloe Anastasiades, Dany Haddad, Aakanksha Naik, Amber Tanaka, Angele Zamarron, Cecile Nguyen, Jena D. Hwang, Jason Dunkleberger, Matt Latzke, Smita Rao, Jaron Lochner, Rob Evans, Rodney Kinney, Daniel S. Weld, Doug Downey, Sergey Feldman• 2025

Related benchmarks

TaskDatasetResultRank
Deep ResearchResearchQA
Score75
42
Long-form researchDRB
Score36.1
39
Deep ResearchHealthBench
Score32
38
Science Question AnsweringResearchQA
Accuracy (ResearchQA)75
37
Deep Research Report GenerationDRB
Overall Score36.1
24
Search-based Question AnsweringSQA v2
Overall Score87.7
21
Aggregate Deep Research PerformanceSQA, ResearchQA, and DRB v2
Average Score66.3
21
Deep ResearchHealthBench ResearchQA DRB Macro Average
Average Score47.7
21
Deep ResearchDeepResearchBench (DRB)
Overall Score36.1
21
Deep ResearchSQA v2
Score87.7
18
Showing 10 of 14 rows

Other info

Follow for update