Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Impacts of Aggregation on Model Diversity and Consumer Utility

About

Consider a marketplace of AI tools, each with slightly different strengths and weaknesses. By picking the right model for the task at hand, a user can do better than simply using the same model for everything. Routers operate under a similar principle, where sophisticated model selection can increase overall performance. However, aggregation is often noisy, reflecting imperfect user choices or routing decisions. This leads to two main questions: first, what does a "healthy marketplace" of models look like for maximizing consumer utility? Secondly, how can we incentivize producers to create such models? We show that winrate, a standard benchmark in LLM evaluation, can incentivize model creators to homogenize for both types of model changes, reducing consumer welfare. We propose a new mechanism, weighted winrate, which rewards models for answers that are higher quality, and show that it provably improves incentives for producers to specialize and increases consumer welfare. We conclude by exploring the impact of our theoretical results in empirical benchmark datasets and discussing implications for benchmark design.

Kate Donahue, Manish Raghavan• 2026

Related benchmarks

TaskDatasetResultRank
Value Allocation OptimizationMT-Bench 101
CM Score10
3
Code GenerationMBPP
Winrate60
1
Code GenerationLCB
Winrate48
1
Emotion DetectionEmory
Winrate42
1
Emotion DetectionMELD
Winrate54
1
General KnowledgeKOR
Winrate36
1
General KnowledgeMP
Winrate48
1
General ReasoningBBH
Winrate53
1
Logic reasoningK&K
Winrate25
1
Mathematical Problem SolvingAIME
Winrate45
1
Showing 10 of 15 rows

Other info

Follow for update