Differentiable Efficient Operator Search
About
Efficient multimodal foundation models often rely on manually designed token-reduction operators, such as pruning, merging, pooling, and adaptive reweighting. Although these operators appear different, we show that they can be interpreted as distinct regimes of a shared operator space. Based on this view, we introduce Efficient Operator Search, a differentiable framework that jointly searches where to reduce tokens, how many tokens to retain, and how reduced token information should be processed. The proposed search space parameterizes layer activation, retention budget, and operator behavior, while the search policy optimizes task performance under one-sided budget and cost constraints. This formulation recovers representative hand-designed baselines as special cases and further discovers hybrid operators beyond isolated manual designs. Experiments on multimodal benchmarks show that the searched operators achieve competitive accuracy-efficiency trade-offs, especially under aggressive visual-token reduction. These results suggest that efficient multimodal inference can be reframed from manual operator design to differentiable operator search.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Object Hallucination Evaluation | POPE | -- | 2056 | |
| Diagram Question Answering | AI2D | -- | 509 | |
| Multimodal Perception and Cognition | MME | Overall Score1.65e+3 | 344 | |
| Multimodal Evaluation | MMStar | -- | 177 | |
| OCR Performance Evaluation | OCRBench | Score25.5 | 98 | |
| Chart Question Answering | ChartQA | Score15.36 | 48 | |
| Science Question Answering | SQA | SQA Score63.83 | 36 | |
| Visual Question Answering | RealWorldQA (RWQA) | Score52.94 | 26 | |
| Multimodal Understanding | SEED | Score59.35 | 25 | |
| Multimodal Reasoning and Understanding | MMB EN | Score58.62 | 10 |