Independent security research · updated daily

The enterprise AI security desk

Exploited vulnerabilities, model releases, research and regulation for the people who have to secure AI in production. Mapped to OWASP, MITRE ATLAS and NIST; every item links to its source.

News Vulnerabilities Models

Top stories, this week

All news →

SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs

Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby…

Aletheia: Permission-Minimality Testing for Coding-Agent Rules

Repository instruction files guide coding agents, but also expose them to prompt injection. Malicious rules can request credential access or data transfer while the agent produces a correct patch. We present Aletheia, a framework for permission-minimality testing. Aletheia…

pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation

Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core…

ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents

Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack…

Self-Evolving Defense: Continual Security Policy Learning for LLM Agents

Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and…