Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CUPID in the Model Zoo: Online Matchmaking for Selecting Your Dream LLM

About

Users increasingly face the challenge of selecting an appropriate LLM for a given task from a rapidly growing pool of LLMs, each with distinct but often opaque latent properties. Compounding this challenge, users may lack the vocabulary or awareness to explicitly articulate the characteristics they value in an LLM's responses or deployment. We propose an interaction-efficient active learning framework in which a dueling bandit algorithm iteratively selects pairs of LLMs, collects user feedback about their responses, and updates its belief about the user's latent preferences. We introduce a novel belief-aware upper confidence bound strategy that balances exploration of the model pool with exploitation of inferred preferences, enabling efficient alignment between user needs and LLM capabilities under user-specified cost and time budgets. Through diverse experiments on LLMs and human studies, we experimentally verify that our model can efficiently match well-aligned LLMs to users at a lower cost.

Son Nguyen, Xinyuan Liu, Ransalu Senanayake• 2026

Related benchmarks

TaskDatasetResultRank
Multi-objective preference selectionOpenAI models Setup A2: K=7, M*=10
Cost ($)0.035
8
Multi-objective preference selectionOpenAI models Setup A3: K=9, M*=2
Cost ($)0.022
8
Multi-objective preference selectionOpenAI models Setup A1: K=5, M*=2
Cost ($)0.056
8
Multi-objective preference selectionOpenAI models Setup A4: K=5, M*=5
Cost ($)0.058
8
Budget complianceHuman-study paradigms H1–H4
Win Rate70
6
Quality ratingHuman-study paradigms H1–H4
Win Rate66.7
6
LLM Selection17-model zoo diverse providers
Cost ($)0.031
3
Showing 7 of 7 rows

Other info

Follow for update