Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Ignore Previous Prompt: Attack Techniques For Language Models

About

Transformer-based large language models (LLMs) provide a powerful foundation for natural language tasks in large-scale customer-facing applications. However, studies that explore their vulnerabilities emerging from malicious user interaction are scarce. By proposing PromptInject, a prosaic alignment framework for mask-based iterative adversarial prompt composition, we examine how GPT-3, the most widely deployed language model in production, can be easily misaligned by simple handcrafted inputs. In particular, we investigate two types of attacks -- goal hijacking and prompt leaking -- and demonstrate that even low-aptitude, but sufficiently ill-intentioned agents, can easily exploit GPT-3's stochastic nature, creating long-tail risks. The code for PromptInject is available at https://github.com/agencyenterprise/PromptInject.

F\'abio Perez, Ian Ribeiro• 2022

Related benchmarks

TaskDatasetResultRank
System Prompt ExfiltrationCline Agent Environment
Pseudo-Recall0.00e+0
50
Retrieval-Augmented GenerationMS Marco--
45
RAG AttackHotpotQA--
41
Question AnsweringHotpotQA
Exact Match (EM)76
36
Question Answering2WikiMultihopQA
EM59
36
Multihop Question Answering2WikiMultihopQA
EM56
36
Multihop Question AnsweringHotpotQA
EM70
36
Question AnsweringMuSiQue
Exact Match (EM)22
36
Multihop Question AnsweringMuSiQue
EM25
36
Question Answering2WikiMultihopQA
Guard Rate100
32
Showing 10 of 35 rows

Other info

Follow for update