Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

SEEM: Exploiting Black-Box Text Attacks to Manipulate Tool Selection

About

Tool learning has emerged as a powerful auxiliary mechanism that extends the capabilities of large language models (LLMs), enabling them to address complex tasks that demand real-time relevance or high-precision operations. However, beneath this strength lie significant security risks. Prior studies have primarily concentrated on corrupting the outputs of invoked tools, while largely overlooking the vulnerability of the tool selection process itself. To bridge this gap, we introduce a black-box, text-based attack that substantially increases the likelihood of a target tool being selected. We propose SEEM, a two-level coarse-to-fine perturbation method that operates at both the word and character levels. Through comprehensive experiments, we show that merely perturbing the textual information of tools can markedly raise the probability of the target tool being prioritized and ranked higher among candidates. Our findings expose critical weaknesses in the tool selection mechanism and lay the groundwork for developing defenses to secure this essential process.

Liuji Chen, Hao Gao, Jinghao Zhang, Qiang Liu, Shu Wu, Liang Wang• 2025

Related benchmarks

TaskDatasetResultRank
Tool selectionScienceQA
P(Use)99.29
14
Tool RetrievalToolBench I1
Hit@12.686
6
Tool RetrievalToolBench I3 conditional attack setting
Hit@170.51
6
Showing 3 of 3 rows

Other info

Follow for update