KARMA: Knowledge graph-based Automated Reasoning Materialization and Alignment
About
Template-based contrastive synthesis is scalable, but its candidates often differ only in a few entity-slots while sequence-level optimization spreads supervision over mostly shared templates. We formalize this as the Resolution Mismatch Problem and propose KARMA, which enumerates schema-constrained paths over domain knowledge graphs and verbalizes them into slot-aligned contrastive candidates. Slot-Parallel Alignment (SPA) then applies a decoupled slot-level objective to route preference supervision to discriminative entity-slots, with slot-aware masked attention serving as an optional packed-evaluation implementation. Across biomedical, computer-science, and chemistry benchmarks, KARMA outperforms base LLM and same-data SFT baselines, and compares favorably with sequence and token-level preference methods.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Medical Reasoning | PubMedQA | Accuracy52.2 | 48 | |
| Domain Reasoning | MedQA Biomedical | Accuracy54.1 | 4 | |
| Domain Reasoning | MMLU Biomedical | Accuracy70.2 | 4 | |
| Domain Reasoning | MMLU-Pro Biomedical | Accuracy (MMLU-Pro Biomedical)53.2 | 4 | |
| Domain Reasoning | MMLU CS (Computer Science) | MMLU CS Domain Reasoning Accuracy70.1 | 4 | |
| Domain Reasoning | MMLU-Pro Computer Science | MMLU-Pro CS Domain Reasoning Accuracy25.6 | 4 | |
| Domain Reasoning | MMLU Chemistry | Accuracy (MMLU Chemistry Domain Reasoning)63 | 4 | |
| Domain Reasoning | MMLU-Pro Chemistry | Accuracy19.5 | 4 | |
| Domain Reasoning | GPQA-Chem | Accuracy20.4 | 4 | |
| Domain Reasoning | GPQA Bio | Accuracy (GPQA Bio)66.7 | 4 |