Impacts of Aggregation on Model Diversity and Consumer Utility
About
Consider a marketplace of AI tools, each with slightly different strengths and weaknesses. By picking the right model for the task at hand, a user can do better than simply using the same model for everything. Routers operate under a similar principle, where sophisticated model selection can increase overall performance. However, aggregation is often noisy, reflecting imperfect user choices or routing decisions. This leads to two main questions: first, what does a "healthy marketplace" of models look like for maximizing consumer utility? Secondly, how can we incentivize producers to create such models? We show that winrate, a standard benchmark in LLM evaluation, can incentivize model creators to homogenize for both types of model changes, reducing consumer welfare. We propose a new mechanism, weighted winrate, which rewards models for answers that are higher quality, and show that it provably improves incentives for producers to specialize and increases consumer welfare. We conclude by exploring the impact of our theoretical results in empirical benchmark datasets and discussing implications for benchmark design.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Value Allocation Optimization | MT-Bench 101 | CM Score10 | 3 | |
| Code Generation | MBPP | Winrate60 | 1 | |
| Code Generation | LCB | Winrate48 | 1 | |
| Emotion Detection | Emory | Winrate42 | 1 | |
| Emotion Detection | MELD | Winrate54 | 1 | |
| General Knowledge | KOR | Winrate36 | 1 | |
| General Knowledge | MP | Winrate48 | 1 | |
| General Reasoning | BBH | Winrate53 | 1 | |
| Logic reasoning | K&K | Winrate25 | 1 | |
| Mathematical Problem Solving | AIME | Winrate45 | 1 |