Skip to content

Agent Security

  • Prompt injection via files, web pages, and tool output
  • The “lethal trifecta”: private data + untrusted content + an exfiltration channel
  • Defenses: permission prompts, sandboxing, least privilege, human review
  • Discussion: how would you threat-model an agent?