AI-based Cognitive-linguistic Features for Dementia Assessment in Picture Description
About
Picture descriptions provide valuable insights into several clinical constructs related to cognitive-linguistic abilities. However, operationalizing these constructs into quantitative measures remains challenging, limiting interpretability and clinical utility. We introduced seven constructs tailored to the Cookie Theft picture description task and prompted large language models (LLMs) to evaluate them, generating severity scores and example-based explanations. Among the examined LLMs, Claude 3.5 Sonnet performed the best, producing severity scores that significantly distinguish cognitively impaired individuals from healthy controls. The model achieves a high accuracy of 85% on the ADReSS dataset. Expert evaluation of Claude's scores and explanations yields a 3.99/5 average agreement. The findings demonstrate the potential of LLMs to operationalize clinical constructs and generate interpretable evaluations, offering a promising approach for accessible cognitive screening tools.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Dementia detection | ADReSS Challenge (test) | Precision90 | 8 | |
| Salience of information assessment | DementiaBank manual transcripts | Mean Score (Control Group)2.09 | 8 | |
| Causal and temporal relations assessment | DementiaBank manual transcripts | Control Mean2.02 | 4 | |
| General cognition and perception assessment | DementiaBank manual transcripts | Mean (Control Group)1.93 | 4 | |
| Mental state language assessment | DementiaBank manual transcripts | Control Mean Score2.29 | 4 | |
| Semantic categories assessment | DementiaBank manual transcripts | Mean Control Score1.59 | 4 | |
| Structural language and speech assessment | DementiaBank manual transcripts | Control Group Mean1.84 | 4 |