News

AI-security reporting from vendor research teams, model providers, government advisories, arXiv and analysts. Tagged by what it means for you: attacks on AI systems, attacks using AI, vulnerabilities found by AI, regulation and releases.

Updated · 24 of 32 sources answered on the last run

Meta, Amazon, And The Real Question About Agentic Commerce

Recent media reports have highlighted Amazon’s decision to block Meta’s Muse AI agent from making purchases on its marketplace. Muse has quickly gained attention for its ability to research products, complete tasks, and shop on behalf of consumers, while Amazon appears focused…

Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI

Fine-tuning teaches a small search agent your tools and environment, giving it the reliability of a frontier model at lower latency and cost. In this post, we fine-tune an LLM-powered search agent with multi-turn reinforcement learning (MTRL) on Amazon SageMaker AI and share the…

Add secure Web Search to Claude Desktop with Amazon Bedrock AgentCore

Claude Desktop on Amazon Bedrock is limited to the model's knowledge cutoff without web search. In this post, we walk through connecting Claude Desktop to Web Search using Amazon Bedrock AgentCore Gateway, with JWT-based inbound authentication through AWS IAM Identity Center and…

A model guide for the GPT-6 family

Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.

The eternal complement

Advanced AI may matter most for the routine work behind breakthrough ideas. Explore why execution could shape the next economy and the pace of progress.

The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching

With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks…

Simplify dashboard drill-down with the Amazon Quick Sight hierarchy filter

Amazon Quick Sight is a fully managed, cloud-native business intelligence (BI) capability for building and publishing interactive dashboards. The new hierarchy filter gives dashboard authors rich, multi-level filtering in a single compact control, reducing clutter and guiding…

Serve live, governed data in AI-built apps with Amazon Quick

With Live Data in Apps in Amazon Quick, AI-built apps query your governed Quick Sight datasets in real time instead of static, build-time snapshots. Each query runs as the person viewing the app, so row-level and column-level security apply per reader. Learn how to build…

Scaling cloud migrations with agentic AI on Amazon Bedrock AgentCore

Learn how AWS Professional Services uses a multi-agent framework built on Amazon Bedrock AgentCore to automate enterprise cloud migrations end to end. Purpose-built AI agents handle discovery, infrastructure as code generation, portfolio governance, and post-migration…

PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents

Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking…

Implementing Multi-Environment Access for Claude Platform on AWS

Learn how to configure secure, multi-environment access to Claude Platform on AWS from a single subscription: cross-account SigV4 for AWS workloads, workspace-scoped API keys for developers, and OIDC federation for external environments, with workspace-level isolation in a…

How To Stop Rogue AI

A model does not have to be smarter than us to be a problem. It needs a goal that it is certain about, the resources to keep pursuing it, and no one watching. Two of those three exist already. We took one doom scenario apart to find out how worried enterprises should be.

How AI Is Changing the Roles Required in the Security Operations Center

As AI takes on more of the enrichment, correlation, and initial assessment inside the SOC, roles, skills, and KPIs still require deliberate redesign. Security leaders need to decide where automation is dependable, where human judgment should remain decisive, and how teams should…

Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors

Learn how to use Amazon S3 Vectors as the persistent memory layer within the NVIDIA NeMo Agent Toolkit (NAT), deployed on Amazon Elastic Kubernetes Service (Amazon EKS). This post shows how NAT's memory subsystem works and how to implement Amazon S3 Vectors as a custom memory…

The AI Doomsday Circus: Don’t “Step Right Up!”

Don’t let AI doomsday headlines dictate your strategy. Stay grounded in what’s real, manage the risks you can control, and focus on the investments most likely to create new value.

SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs

Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby…

Query claims in natural language with Amazon Bedrock Knowledge Bases

This technical how-to builds a conversational claims assistant on Amazon Bedrock Knowledge Bases that answers natural-language questions with citations. It covers ingesting claim documents from Amazon S3, querying with the AgenticRetrieveStream API, multi-turn follow-ups…

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

Helping small businesses put AI to work

OpenAI is partnering with America’s SBDC to expand hands-on AI training and local support for small businesses, alongside a new report on how small teams are using AI.

From SELECT to SYSADMIN with SQL Copilot (CVE-2026-65669)

Two weeks back I presented at BlueHat Asia 2026 about my research on Microsoft’s Copilot in SSMS, the SQL Server Management Studio. This post is a write up about the talk, which covered CVE-2026-65669 , a SQL Server Elevation of Privilege Vulnerability rated critical by…

Aletheia: Permission-Minimality Testing for Coding-Agent Rules

Repository instruction files guide coding agents, but also expose them to prompt injection. Malicious rules can request credential access or data transfer while the agent produces a correct patch. We present Aletheia, a framework for permission-minimality testing. Aletheia…

pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation

Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core…

ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents

Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack…

Self-Evolving Defense: Continual Security Policy Learning for LLM Agents

Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and…

Introducing dots

Dots by OpenAI are proactive assistants that can keep working across complex projects and everyday tasks. Learn how dots help you stay in control while work moves forward.

Introducing GPT-6.1 Sol

Meet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.

DevDay 2026 Recap

Explore more than 20 announcements from OpenAI DevDay 2026, including GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.

CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes…

Controlled Decoding Attacks on Black-Box LLMs

Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled…

Towards safety cases for frontier AI training

Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents

Render Before Reading: Visual Rendering as a Prompt Injection Defense

Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are…

How we will do better for Australia

OpenAI apologises for incidents involving Australian government websites and outlines stronger safeguards and support to strengthen Australia’s cyber defences.

CoSec: Benchmarking Agent Security in Communities

LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks…

Certified Multi-Source Integrity for Structured Agent Actions

LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself…

ORBIT: A Framework for Multi-Agent Safety and Security Evaluations

Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose…

Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money

AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it…

API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary

Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where…

TempQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks

Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V…

Storm-3168: Agentic-driven cloud attacks using compromised service principals

Microsoft details JADEPUFFER-linked Azure reconnaissance, resource deletion, and credential access using compromised service principals, identifying the activity as associated with Storm-3168 and providing guidance for defenders. The post Storm-3168: Agentic-driven cloud attacks…

Prompt Injection Detection for Email Agents Through Attack Chain Modeling

Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary…

AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents

AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or…

Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure

A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce…

ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions…

The Closed Quorum: Inside the first reported autonomous AI C2 implant

CLOSEDQUORUM, a malware binary discovered through Cisco Talos’ CAIRN project, exhibits fully autonomous command and control (C2). It represents a shift in effort displacement for attackers, in which expanding portions of the attack chain can be executed without operator…

Auditing in the age of (good enough) AI

Security firms have published numerous blog posts describing how they pointed their agent harness at a codebase and found dozens of bugs ( we’re one of them ). However, these posts tend to focus on agentic code review, which is just one aspect of how we use AI in our security…

Should you care about an “AI slowdown?”

In this week's Threat Source, David talks about why focusing on your security basics is still your best bet, even in a world with rapid AI advancements.

Self-generated prompt injections in compaction summaries

Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerning model behavior we’ve observed in the last six months". This one here is my favorite: they caught some of their…

Making global data easier to explore

Google and the UN system have launched the UN System Data Commons, a new open platform making global statistics accessible and easy to search.

AI Threat Landscape Digest: July–August 2026

The defining development of the period came not from attackers but from the AI labs themselves, whose models broke out of controlled evaluations and reached real systems. In the wild, the criminal and state use of AI continued to mature along the lines tracked in earlier…

Securing the unpatchable in an age of AI-driven vulnerabilities

Advances in AI technology will continue to identify vulnerabilities that in some circumstances are difficult, or effectively impossible, to patch. Appropriate network segmentation, rigorous visibility, and the deployment of NGFW/IPS combinations can provide a powerful…

Building AI to accelerate science and improve lives

The true measure of AI is who it helps. Here’s how it’s impacting lives today. We're focused on key areas where advanced technology can help make extraordinary progress …

AI for everyone in every language

We’re moving beyond traditional text translation to build models that understand the world’s rich, living languages exactly as they are expressed.

AI for Societal Impact

Explore this collection to see how experts and local leaders are using AI breakthroughs to ensure everyone can share the opportunity of AI.

1Password's AI patching benchmark is misleading

1Password’s FLAWED report , published on August 6, 2026, gives defenders a misleading picture of AI patching. Its headline says models produced clean fixes only 26% of the time. That figure includes experiments that deliberately instructed agents to apply the wrong fix, along…

DevFest is back

DevFest 2026 is back and here’s how you can connect with one of the more than 800 global events to build, secure, and scale in the agentic AI era.

Metasploit Wrap Up: This One Goes to Sixteen!

This One Goes to Sixteen! Another banger from Metasploit with sixteen new modules, including ten exploit modules, with five on the CISA KEV list. Cisco, Papercut, Sonicwall, Jetbrains, and Langflow all have exploit modules, and not to be outdone, we even have a Metasploit…

PuzzleMask: Abusing Plain Prose as a Covert AI Attack Vector

Executive Summary In this research we introduce a prompt-crafting technique for bypassing quick LLM-based policy checks — using plain English (no emojis, base64, invisible formatting, etc.) A policy-violating payload (e.g. ”encrypt files in ~/Documents”, “give me a biohazard…

A “proof” of Fermat’s Last Theorem that fits the margin

Fermat famously claimed to have a “truly marvelous proof” of his Last Theorem , but he never wrote it down, insisting the margin of his page was too narrow to contain it. A few centuries later, Anthropic announced a complete formalization of Fermat’s Last Theorem using 13…

The Shared Clipboard Inside the Sandbox: Cross-Account Data Leakage in ChatGPT

Research by: Alexey Bukhteyev Key Takeaways Introduction Over the past several years, AI assistants have moved far beyond text generation. Modern systems can execute code, install additional dependencies, analyze user files, and access data through connected services. These…

An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation

Using autonomous AI agents, an attacker breached an enterprise network in a matter of hours. Understand how to address and defend against agentic attacks. The post An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation appeared first on Unit 42 .

Breaking Claude Code Opus 5 Auto Mode

Breaking Claude Code Opus 5 Auto Mode Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently made that the default and have made bold claims about its effectiveness. Johann…

Breaking Claude Code Opus 5 Auto Mode

In this post, we explore how a simple website summary request hijacks Claude Code Opus 5 in Auto Mode and achieves code execution with 60-80% attack success rate using a small sample size. This is interesting because a third-party evaluation commissioned by Anthropic showed a…

Recovering Encrypted LLM Reasoning Traces

A few days ago, a paper named “Stealing Reasoning Traces from Proprietary LLM APIs” was published. It describes a simple, yet super elegant way to recover encrypted LLM reasoning traces. Naturally, I had to try it. Background AI labs like OpenAI and Anthropic send reasoning…

Stealing Reasoning Traces from Proprietary LLM APIs

Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name ( stolen-thoughts.com ) for a neat paper : Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced…

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode , to the point that they are making it the default setting for new sessions in most Claude Code plans starting on August 14th. This was one of the…

Incident Report: unsanctioned agent behaviour during cyber testing

Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their…

Headlines and short snippets only; each links to the original. Tags are assigned automatically from the text and can be wrong — the source is authoritative.