Indirect prompt injection
Malicious instructions embedded in content an agent retrieves or opens rather than typed by the user. Treating remote content as data, not authority, is the core boundary.
Glossary · area 17 of 19
Where an agent, a tool or a page can be turned against its operator, and the controls that limit the damage without proving safety.
39 terms. Each opens its own page with what it can and cannot support, how the scanner measures it, and where it comes up in the guides.
Malicious instructions embedded in content an agent retrieves or opens rather than typed by the user. Treating remote content as data, not authority, is the core boundary.
User-supplied text attempting to override application or system instructions. Refusal behavior helps but is not a complete authorization control.
A prompt or interaction pattern designed to bypass model restrictions. Success against one version may fail against another, so jailbreak rates require dated and versioned tests.
Unauthorized transfer of secrets or sensitive information from a system. A model can facilitate it through tools, outputs, encoded text, or attacker-controlled destinations.
Exposure of private, confidential, regulated, or security-relevant data in model output or logs. Correctly answering a prompt can still be a disclosure failure.
Accidental exposure of credentials, keys, tokens, or internal values. Redaction at display time does not remove secrets already sent to a provider or log.
Passing model-generated content into interpreters, browsers, databases, or shells without appropriate validation. The output is untrusted even when it came from a trusted model endpoint.
Granting an agent more tools, permissions, autonomy, or action scope than needed. Accuracy improvements do not compensate for avoidable blast radius.
Giving each component only the permissions required for its current task. It limits damage but requires continuous maintenance as tools and workflows change.
The point where data or control moves between components with different assurance levels. Every crossing needs explicit validation rather than inherited trust.
Checking supplied values against allowed types, formats, ranges, and business rules. Schema validation is one layer and cannot detect every harmful valid value.
Checking generated or tool-returned data before display, storage, or execution. It should be tailored to the downstream sink, not treated as generic sanitization.
An explicit set of permitted values, operations, hosts, or paths. It reduces ambiguity but becomes dangerous when broad wildcards silently expand it.
A set of known prohibited patterns or targets. It blocks recognized cases and predictably misses novel encodings, variants, and attack paths.
Server-side request forgery, where attacker-controlled input makes a server request unintended internal or external resources. URL parsing, redirects, DNS changes, and metadata endpoints all matter.
Remote code execution, where an attacker causes code to run on another system. An agent with a shell tool can create equivalent impact without exploiting a traditional software bug.
The resources and actions a credential authorizes. A short-lived credential can still be overpowered if its scope is broad.
A named permission requested and granted within an OAuth authorization flow. The label’s actual force depends on the resource server enforcing it correctly.
A shared secret used to identify or authorize API access. It usually identifies an application rather than a person and should not be embedded in public client code.
A credential accepted from whoever possesses it. Transport security, storage, expiration, and audience restrictions are therefore critical.
Mutual TLS, where client and server authenticate with certificates during connection establishment. It verifies certificate possession, not whether an application request is safe.
Rules controlling how much traffic a principal may send during a period. It protects capacity but can penalize shared addresses and bursty legitimate work.
A record of security-relevant identities, actions, targets, outcomes, and times. Logs support investigation only when complete, protected, retained, and attributable.
Personally identifiable information that can identify or reasonably link to a person. Definitions vary by jurisdiction and context, so one static field list is insufficient.
Collecting, sending, and retaining only data necessary for a stated purpose. It reduces exposure but requires knowing which later uses are genuinely necessary.
Restricting data use to specified, legitimate purposes. Broad consent or vague product improvement language weakens the practical boundary.
The time data remains stored before deletion or anonymization. A stated period is not evidence that every cache, backup, and derived record follows it.
The geographic location where data is stored or processed. Residency does not by itself determine jurisdiction, access, ownership, or privacy compliance.
Controls preventing one customer’s data or actions from crossing into another customer’s environment. Shared models, logs, caches, and vector indexes are all possible boundaries.
A structured account of assets, actors, entry points, trust boundaries, and plausible attacks. It guides controls but becomes stale as architecture and capabilities change.
A DNS record telling receiving mail servers what to do with a message that claims your domain and fails authentication: none, quarantine or reject. It evaluates SPF and DKIM results for alignment with the visible From domain; it authenticates nothing itself, and p=none is monitoring, not protection.
The requirement that the domain which passed SPF or DKIM be the same organisational domain as the visible From address. A sender can pass SPF or DKIM on its own domain and still fail DMARC because nothing aligned, which is how a marketing platform or form tool that signs with its own domain lands in spam under quarantine.
Sender Policy Framework, a DNS record listing the servers allowed to send mail for a domain, checked against the envelope sender's address. It breaks on any forwarding hop, because the forwarder is not in the list, which is why SPF alone cannot make a quarantine policy safe.
DomainKeys Identified Mail, a signature over a message's headers and body made with a private key whose public half is published in DNS under a selector. It survives forwarding because it is computed over the content, and its d= domain is what DMARC aligns against; a domain with mail and no DKIM has one point of failure.
A policy, published at a well-known URL and announced in DNS, telling sending servers that mail to this domain must use TLS to the named mail exchangers. It closes the downgrade attack that plain SMTP allows; testing mode reports, enforce mode refuses, and TLS-RPT is where the reports go.
A DNS record naming an address that receives daily reports of TLS failures when other servers deliver mail to your domain. It is how an MTA-STS policy is monitored before it is enforced; the address has to exist, or the reports bounce into nothing.
A DNS record naming which certificate authorities may issue certificates for a domain. An authority that is not listed must refuse, which limits the damage of a compromised validation path; it does not affect certificates already issued and does nothing if no record exists.
Signatures over DNS records, chained from the root through the registry to the zone, so that a resolver can detect a forged answer. It protects the lookup, not the site: a signed zone can still point at a compromised server. It is switched on at the DNS host and completed at the registrar, and a broken chain makes the domain unreachable.
A permanent delivery failure: the mailbox does not exist. Sending systems respond by suppressing the address so they stop hammering it, which means every later message is accepted by the sending API and dropped, with no error, until the suppression is cleared. It is the failure that looks like success from the sending side.