Tech●●●●●Difficulty 5 of 5

Why can a web page or an email hijack an AI agent?

A model reads instructions and data in the same stream of words, so text on a page it was only meant to read can end up giving the orders.

β–Ά Start the story

Because a language model cannot reliably tell instructions from data. Prompt injection is an attack in which innocuous-looking inputs are designed to cause unintended behavior in a model. It works because the model's inputs contain instructions and data together in the same context, so the algorithm cannot distinguish between them. In its original, direct form, a user's input is mistaken for the developer's instructions.

For agents the danger grows. In indirect injection, the instruction sits in external data such as emails and documents, and the AI may mistake it as coming from the user or the developer. Web pages work the same way: adversarial instructions are embedded in a website's content, and if the model retrieves and processes the page, it may execute them as if they were legitimate commands. A mild everyday case: a job-seeker hides white text in a resume so that an AI rating the resume gives a good rating while ignoring the content. In December 2024 The Guardian reported that ChatGPT's search tool was vulnerable to hidden webpage content manipulating its responses.

Two kinds of injection

Direct

  • User input is mistaken for a developer instruction
  • The original form of the attack, a threat to the developer from the user

Indirect

  • Instructions hide in external data such as emails, documents or web pages
  • A threat from the data's author to the user

The defences are partial. In August 2023 the UK National Cyber Security Centre said prompt injection may simply be an inherent issue with this technology, and that as yet there were no surefire mitigations. OWASP's suggestions include least privilege access, human oversight for sensitive operations, and isolating external content. Telling the model in its system prompt to be careful has limited effectiveness, and OWASP says retrieval-augmented generation and fine-tuning do not eliminate the threat.

Quiz me

0/3

  1. 1.Why can a model be hijacked by text on a web page it only meant to read?
  2. 2.What is the difference between direct and indirect prompt injection?
  3. 3.Which defence does OWASP list among its suggestions, according to Wikipedia?

Recap

Models read instructions and data in one stream, so anything they read can try to give orders.

πŸ’‘ A trick to remember it Β· If the robot reads it, the robot might obey it: so give it no more power than you would give a stranger's letter.

Surprising fact Β· A job-seeker's invisible resume text can steer an AI rater.

Sources (3)

No source, no claim. Every fact in this lesson (16 claims) cites at least one of these.

  1. [1]Prompt injection Β· Wikipedia
  2. [2]Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection Β· arXiv (Greshake et al.)
  3. [3]Model Context Protocol Β· Wikipedia
More lessons in πŸ’» Tech (3) See all tech lessons β†’

One more light on your map.

Get one lesson like this every day, about the things you love. Free, in two or five minutes.

Get the share card for this lesson β†—