Prompt Injection
Prompt injection is an attack in which text supplied to a language model application is crafted to override the instructions the developer gave the model. It works because the model receives developer instructions and untrusted content as one stream of tokens and has no reliable way to tell which part carries authority. It differs from a jailbreak, which tries to get a model to produce content its safety training refuses, whereas prompt injection targets the application built on the model and what that application is allowed to do. In agentic systems, the most serious form is indirect injection, where the malicious text arrives inside a web page, email, document, or tool result that the agent reads while working on a task.
The risk grows with the agent's permissions, since an injected instruction can make an agent send data out through a tool call or take an action the user never requested. No known filter removes the risk completely, so defenses are layered: least-privilege tool access, separation of trusted and untrusted content, human approval for sensitive actions, checks on output before it leaves the system, and monitoring of tool calls. OWASP lists prompt injection first in its 2025 Top 10 for LLM applications.
Related terms
Related services: LLM development, Agentic AI consulting.
Ready to put agentic AI to work?
Book a free 45-minute consultation. We'll map one real process worth automating with production-grade AI.