News
AI-security reporting from vendor research teams, model providers, government advisories, arXiv and analysts. Tagged by what it means for you: attacks on AI systems, attacks using AI, vulnerabilities found by AI, regulation and releases.
Updated · 24 of 32 sources answered on the last run
Setting Grok Bot loose on procurement We gave Grok Bot access to vendor spend, contracts, and usage data. It found more than $100,000 in direct savings. Sep 4, 2026
Sep 23, 2026 Science Claude discovers a novel enzyme system with CRISPR-like repeats
Sep 18, 2026 Introducing Grok Voice Transcribe 2.0
Sep 18, 2026 Announcements Partnering with Accenture on embedded evaluation
Sep 17, 2026 Announcements Introducing the Life Sciences Verification Program
Sep 1, 2026 Announcements Developing Enterprise Frontier Safeguards with our customers
Responsible Scaling Policy
Reimagining Independence: How Meta’s AI Models Are Helping the University of Pittsburgh Transform Assistive Robotics
Product · Sep 28, 2026 Team Bots: shared AI teammates that learn as they work
Product · Sep 22, 2026 How SpaceXAI is using Grok Bot to scale customer support
Product · Sep 16, 2026 Memory in Grok Build
Product Agentic Search. More accurate and efficient results from your AI systems. The retrieval layer that helps AI systems navigate, read, and verify information inside and across even the most…
PhantomRaven: An LLM-Generated Information Stealer Developed for Bug Bounty Hunting
Oct 2, 2026 Announcements Anthropic invests $100 million to train 10,000 engineers and tackle the enterprise AI talent gap
Oct 1, 2026 Announcements Barclays scales Claude to upgrade operations and improve client experience
Mistral and Mozilla are bringing open, private and multilingual AI to your web browser
Mistral Small 4
Mistral Medium 3.5
Introducing Muse Spark 1.1
Introducing Muse Image and Muse Video
How Meta’s AI Models Are Powering the First Wave of Genesis Mission Projects
Hallo, Deutschland!
Grok Bot now works with X Grok Bot now has a tighter integration with X. Aug 29, 2026
Grok Bot is now included with more plans Grok Bot is now available for SuperGrok, Cursor Pro, and all Cursor Teams plans. Aug 26, 2026
Grok Bot for Enterprise Grok Bot is now available for enterprises. Grok and Cursor Enterprise customers have free usage for the next two weeks, and can invite their whole organization, including…
Grok 4.7 Sep 21, 2026 Introducing Grok 4.7 SpaceXAI's most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models. Read More
Grok 4.6 on Microsoft Foundry Grok 4.6 is now available via Microsoft Foundry. Aug 26, 2026
From Brain Waves to Words: Brain2Qwerty Offers a New Path to Communication Without Surgery
Designing Grok Bot for a world of persistent agents How we designed Grok Bot for agents that persist beyond a single session — from a chat history to a Bot roster, presence, a computer of the Bot’s…
CrowdStrike Delivers the Next Evolution of the Agentic SOC
CrowdStrike Accelerates Real-Time Data Classification with On-Device AI
Company Mistral x HUMAIN August 24, 2026 By Mistral
Company Mistral raises €3B to make sovereign, open-weight AI the technology frontier Mistral today announced that it has raised €3 billion in a Series D funding round at a post-money valuation of…
Company Mistral and Mozilla are bringing open, private and multilingual AI to your web browser Mozilla and Mistral AI are partnering to bring open, private and multilingual AI to Firefox Smart…
Company Hallo, Deutschland! Mistral Opens German Hub in Munich to Advance Industrial AI in Europe’s Largest Economy September 28, 2026 By Mistral
Company Cloudera and Mistral Partner to Bring Specialized, Sovereign Intelligence to Enterprise Data September 10, 2026 By Mistral
Cloudera and Mistral Partner to Bring Specialized, Sovereign Intelligence to Enterprise Data
Biosecurity at the frontier LatchBio evaluated Grok's performance on biosecurity monitoring and adversarial biological tasks. They found that Grok 4.6 detects and refuses dangerous queries more…
Aug 31, 2026 Announcements Improving our alignment and security efforts
Aug 27, 2026 Announcements Previewing the Model Hardware Standard
Aug 27, 2026 Announcements Expanding our support for scientists
Aug 25, 2026 Announcements Funding better evaluations of AI’s impact on wellbeing
A Win for Defenders: CrowdStrike and NVIDIA Extend Security Across the AI Stack
The Agent Said It Was Done. The Database Disagreed.
Meta, Amazon, And The Real Question About Agentic Commerce
Recent media reports have highlighted Amazon’s decision to block Meta’s Muse AI agent from making purchases on its marketplace. Muse has quickly gained attention for its ability to research products, complete tasks, and shop on behalf of consumers, while Amazon appears focused…
There’s A Consumer AI Trust Gap In Financial Services, But The Solution Is Surprisingly Simple
Today, as a society, we agree on little: Politics. Climate. Health. Whether a hot dog is a sandwich. But one topic is bringing us back together: artificial intelligence! According to Forrester’s Consumer Benchmark Survey, 2026, nearly nine in 10 online adults in the US and the…
The latest AI news we announced in September 2026
Here are Google’s latest AI updates from September 2026
Sweep thousands of leases for compliance using Amazon Quick and the Adjudicated Query pattern
The Adjudicated Query pattern pairs the Amazon Quick chat agent with a bounded MCP server over a deterministic rules engine to deliver provably complete, defensible compliance answers. This post walks through the reference architecture and a deployable AWS CDK sample, using…
Open-sourcing AstaBrief, the fast report-generation model in Asta
Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI
Fine-tuning teaches a small search agent your tools and environment, giving it the reliability of a frontier model at lower latency and cost. In this post, we fine-tune an LLM-powered search agent with multi-turn reinforcement learning (MTRL) on Amazon SageMaker AI and share the…
Chatham scales its capital markets expertise with OpenAI
Chatham Financial uses Codex and GPT-5.6 to build technology and redesign workflows, cutting trade validation from 30 minutes to under 4.
AutoSynthData: Generating Training Data for Enterprise Agents
Add secure Web Search to Claude Desktop with Amazon Bedrock AgentCore
Claude Desktop on Amazon Bedrock is limited to the model's knowledge cutoff without web search. In this post, we walk through connecting Claude Desktop to Web Search using Amazon Bedrock AgentCore Gateway, with JWT-based inbound authentication through AWS IAM Identity Center and…
A model guide for the GPT-6 family
Learn how startups can choose GPT-6 models, tune reasoning effort, improve prompts and skills, coordinate tools, and prepare workflows for production.
Uplifting conversion across the acquisition funnel with personalization using contextual bandits on AWS
Generative AI makes it cheap to produce personalized content at scale, but which variation do you show each customer? Amazon Payments used a multi-objective contextual bandit on Amazon SageMaker AI to personalize an acquisition funnel, achieving a high single-digit conversion…
The eternal complement
Advanced AI may matter most for the routine work behind breakthrough ideas. Explore why execution could shape the next economy and the pace of progress.
The Innocent Courier: Covert Exfiltration Through Legitimate LLM Web Fetching
With the increasing capabilities of Large-Language-Models (LLMs) and LLM-based agents, users are increasingly using them to solve everyday problems, such as answering e-mails or providing programming support. Existing work has extensively investigated security and privacy risks…
The Den frees up 10-15 hours a week to grow with ChatGPT Work
As it opens a new location, the social club prepares grant applications in 2 hours instead of 3 days and liquor-license materials in 3 hours instead of 4 days.
Simplify dashboard drill-down with the Amazon Quick Sight hierarchy filter
Amazon Quick Sight is a fully managed, cloud-native business intelligence (BI) capability for building and publishing interactive dashboards. The new hierarchy filter gives dashboard authors rich, multi-level filtering in a single compact control, reducing clutter and guiding…
Serve live, governed data in AI-built apps with Amazon Quick
With Live Data in Apps in Amazon Quick, AI-built apps query your governed Quick Sight datasets in real time instead of static, build-time snapshots. Each query runs as the person viewing the app, so row-level and column-level security apply per reader. Learn how to build…
Scaling cloud migrations with agentic AI on Amazon Bedrock AgentCore
Learn how AWS Professional Services uses a multi-agent framework built on Amazon Bedrock AgentCore to automate enterprise cloud migrations end to end. Purpose-built AI agents handle discovery, infrastructure as code generation, portfolio governance, and post-migration…
PACE: Provenance-Aware Capability Enforcement for Tool-Using LLM Agents
Tool-using large language model (LLM) agents turn generated text into real side effects, so poisoned tool metadata, retrieved pages, memory, and reusable skills can steer the next call. Vetting an artifact before admission does not settle this. A safe variant and a leaking…
Oktane 2026 Recap: Okta Announces A Unified IAM Control Plane, With Agent Governance And Identity Details Remaining Unclear
At last week’s Oktane conference in Las Vegas, Okta unveiled its ambitious vision for transforming the Okta platform into a control plane for AI agents and extending deeper into AI security. The strategy was reflected throughout the event’s keynotes and product announcements…
MOMAT: Mixture of Multiple Atlases for Low-Power Jailbreak Defense of Quantized LLMs
Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address…
Lessons Learned From The Forrester Wave™: Conversational AI Platforms For Customer Service, Q2 2026
Learn key lessons about deploying conversational AI from practitioners who have delivered self-service applications for their organization.
Implementing Multi-Environment Access for Claude Platform on AWS
Learn how to configure secure, multi-environment access to Claude Platform on AWS from a single subscription: cross-account SigV4 for AWS workloads, workspace-scoped API keys for developers, and OIDC federation for external environments, with workspace-level isolation in a…
How uniopen customized Amazon Nova to their retail moderation policies for production deployment
See how uniopen, a retail platform from Taiwan's Uni-President Enterprises Group, adapted Amazon Nova 2 Lite to its content-moderation policies using supervised fine-tuning in Amazon SageMaker AI and prompt optimization. Business-relevant evaluation and release gates kept…
How To Stop Rogue AI
A model does not have to be smarter than us to be a problem. It needs a goal that it is certain about, the resources to keep pursuing it, and no one watching. Two of those three exist already. We took one doom scenario apart to find out how worried enterprises should be.
How Albertsons Companies is reimagining retail from the inside out
Albertsons Cos. is using ChatGPT Enterprise and the OpenAI API to help teams work faster and make grocery shopping easier for millions of customers.
How AI Is Changing the Roles Required in the Security Operations Center
As AI takes on more of the enrichment, correlation, and initial assessment inside the SOC, roles, skills, and KPIs still require deliberate redesign. Security leaders need to decide where automation is dependable, where human judgment should remain decisive, and how teams should…
Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows
Ambient agents respond to events such as an Amazon S3 upload, a schedule, or an alert instead of waiting for a chat prompt. This post walks through building framework-agnostic ambient agents on Amazon Bedrock AgentCore using Amazon SQS, AWS Lambda, and Amazon DynamoDB, with a…
Build agent memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors
Learn how to use Amazon S3 Vectors as the persistent memory layer within the NVIDIA NeMo Agent Toolkit (NAT), deployed on Amazon Elastic Kubernetes Service (Amazon EKS). This post shows how NAT's memory subsystem works and how to implement Amazon S3 Vectors as a custom memory…
Secure what’s next: Your guide to Microsoft Security at Microsoft Ignite 2026
This year at Microsoft Ignite, we spotlight our AI-first, end-to-end security platform designed to protect identities, devices, data, applications, clouds, infrastructure, and the AI agents now working alongside your teams. The post Secure what’s next: Your guide to Microsoft…
The AI Doomsday Circus: Don’t “Step Right Up!”
Don’t let AI doomsday headlines dictate your strategy. Stay grounded in what’s real, manage the risks you can control, and focus on the investments most likely to create new value.
SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs
Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby…
Query claims in natural language with Amazon Bedrock Knowledge Bases
This technical how-to builds a conversational claims assistant on Amazon Bedrock Knowledge Bases that answers natural-language questions with citations. It covers ingesting claim documents from Amazon S3, querying with the AgenticRetrieveStream API, multi-turn follow-ups…
Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning
Introducing SynthID Bio
Proof of concept for watermarking AI-generated proteins while preserving biological function.
Helping small businesses put AI to work
OpenAI is partnering with America’s SBDC to expand hands-on AI training and local support for small businesses, alongside a new report on how small teams are using AI.
Gemini 4 Argon: our next era of frontier intelligence
From SELECT to SYSADMIN with SQL Copilot (CVE-2026-65669)
Two weeks back I presented at BlueHat Asia 2026 about my research on Microsoft’s Copilot in SSMS, the SQL Server Management Studio. This post is a write up about the talk, which covered CVE-2026-65669 , a SQL Server Elevation of Privilege Vulnerability rated critical by…
From A2A Attacks to Envelope-Layer Defense: Red-Teaming Evaluation of LLM Agents and a Three-Layer Isomorphic Attack-Defense Model
Agent interaction protocols such as ACP and A2A have moved LLM-based agents toward multi-agent collaboration, introducing new security threats. A task sent by a remote peer over A2A is treated as a legitimate request, providing a natural channel for indirect prompt injection…
Disrupting a coordinated model-distillation campaign
Learn how OpenAI disrupted a campaign to extract protected model reasoning and is strengthening defenses against adversarial distillation.
CodeMimicry: Exploiting Safety Generalization Lag in Large Language Models via Structured Code Completion
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained…
Aletheia: Permission-Minimality Testing for Coding-Agent Rules
Repository instruction files guide coding agents, but also expose them to prompt injection. Malicious rules can request credential access or data transfer while the agent produces a correct patch. We present Aletheia, a framework for permission-minimality testing. Aletheia…
pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation
Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core…
Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance
When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does…
ToolFence: Fine-Grained Authorization for Secure Tool-Using LLM Agents
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack…
Self-Evolving Defense: Continual Security Policy Learning for LLM Agents
Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and…
NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction
Introducing dots
Dots by OpenAI are proactive assistants that can keep working across complex projects and everyday tasks. Learn how dots help you stay in control while work moves forward.
Introducing GPT-6.1 Sol
Meet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices.
Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse
A Technical Report: Operating a large language model (LLM) as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of…
Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents
Does the Unsafe Gradient Survive a Conversation? On the Fragility of Gradient-Based Jailbreak Detection in Multi-Turn Dialogue
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an…
DevDay 2026 Recap
Explore more than 20 announcements from OpenAI DevDay 2026, including GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders.
CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering
Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes…
Controlled Decoding Attacks on Black-Box LLMs
Manipulating next-token probabilities during generation can bypass the safety alignment of large language models. Existing approaches, however, rely on access to model weights or numerical token probabilities and therefore do not apply to interfaces that return only sampled…
ContractWarden: Kernel-Enforced Damage Boundaries for AI Agents via Human-Authorized Contracts
Large language model agents can execute commands, create subprocesses, and directly access files and networks, allowing prompt injection or planning errors to become operating-system side effects. We present ContractWarden, a Linux reference monitor that enforces a…
Watch the winning trailer from the Future Vision XPRIZE, The Gifted.
Watch the winning trailer from the Future Vision XPRIZE, The Gifted.
Towards safety cases for frontier AI training
Our early guidelines for safety cases in frontier AI training cover technical safeguards, operational practices, and investigating misalignment incidents
Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The…
Render Before Reading: Visual Rendering as a Prompt Injection Defense
Large language models are vulnerable to prompt injection attacks, where third-party adversarial content can hijack the model's behavior. In this paper, we study the role played by the adversarial data's input modality, and identify a systematic asymmetry: multimodal LLMs are…
How we will do better for Australia
OpenAI apologises for incidents involving Australian government websites and outlines stronger safeguards and support to strengthen Australia’s cyber defences.
Holo4: powering generalist computer-use agents
Continuous Assurance of Agentic Security Auditors for Software Delivery Decision Gates
Large language model (LLM)-based repository auditors are increasingly deployed as security controls within continuous integration (CI) pipelines, where their findings admit, block, or delay software changes. As Agentic Software Development Life Cycle (SDLC) Security Controls…
CoSec: Benchmarking Agent Security in Communities
LLM agents operate in persistent collaborative environments involving multiple users, communities, memories, files, and tools. Community boundaries may remain fixed or evolve with changes in membership, roles, composition, and relationships. Agents must complete legitimate tasks…
CoDeL: Co-Evolutionary Defense against Indirect Prompt Injection in LLM-based Agents
Large language model (LLM)-based agents increasingly rely on external tools and content, exposing them to indirect prompt injection (IPI). This threat has motivated a wide range of defenses, among which training-based defenses are often regarded as most reliable. However…
Certified Multi-Source Integrity for Structured Agent Actions
LLM agents increasingly take privileged, often irreversible structured actions, such as paying an invoice. They assemble each action from action-critical fields in documents and tool outputs that an adversary can corrupt, and indirect prompt injection can drive the model itself…
ORBIT: A Framework for Multi-Agent Safety and Security Evaluations
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose…
Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models select from predefined answers and report probabilities that software can use to allow, block, or review inputs, but the…
Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning
Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting…
Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money
AI agents now hold spend authority and settle payments without per-action human confirmation. The resulting loss is often not a security failure: a counterparty with the correct domain, the correct settlement address and a genuinely delivered service can charge more than it…
API Secrets Should Never Become Tokens in the LLM's Vocabulary: A Threat Analysis of API Credential Handling in LLM Agent Systems and an Empirical Evaluation of a Vault-Mediated Execution Boundary
Tool-using large language model (LLM) agents turn credential hygiene from a storage problem into an execution-security problem. A key pasted into a prompt, or embedded in a system prompt or tool configuration, crosses from an authentication boundary into a data pipeline, where…
Silent Failures in Agentic Security Evaluation: A Validated Harness for Tool-Call Mediation Under Indirect Prompt Injection
LLM agents that invoke privileged tools are vulnerable to indirect prompt injection (IPI), in which adversarial instructions embedded in retrieved data hijack the agent's actions. A growing body of work evaluates defenses against IPI, but the validity of that evaluation is…
TempQ-Jail: Query-Constrained Candidate Ranking for Text-to-Video Jailbreak Attacks
Existing text-to-video (T2V) jailbreak methods mainly seek more effective or stealthier attack candidates. In guarded T2V systems, however, video generation and security evaluation are costly, so an attacker often cannot test a large candidate pool. We therefore formulate T2V…
Storm-3168: Agentic-driven cloud attacks using compromised service principals
Microsoft details JADEPUFFER-linked Azure reconnaissance, resource deletion, and credential access using compromised service principals, identifying the activity as associated with Storm-3168 and providing guidance for defenders. The post Storm-3168: Agentic-driven cloud attacks…
Resource-Optimized and Energy-Aware Agentic AI Framework Anchored on Blockchain for Secure Software Supply Chains
This paper proposes a blockchain-backed agentic security framework designed to safeguard the complete software development lifecycle (SDLC) while also securing the agentic AI components responsible for monitoring it. The framework coordinates a set of specialised security…
Prompt Injection Detection for Email Agents Through Attack Chain Modeling
Large language model email assistants are particularly vulnerable to indirect prompt injection because untrusted email content can be retrieved into the model context and influence subsequent tool use. Existing prompt injection detectors mainly formulate this problem as binary…
MetaPermit: Scalable and Auditable Access Control for AI Agents via LLM-Inferred Meta-Attributes
The rise of autonomous AI agents equipped with tools has introduced significant security risks, ranging from unintended tool misuse to adversarial manipulation through Indirect Prompt Injection (IPI) attacks. In practice, deployed agent systems such as OpenAI Codex and Claude…
Crypto-bound identity-verified capability tokens for coordinating distributed AI agents: A proposal
The prospect of fully autonomous transactional agents did not appear on the horizon until the advent of high capability language models. With such models, the operational benefits of adaptive task orchestration and independent (but constrained) decision making are tantalizing…
AgentXploit: Autonomous Repository-to-Runtime Red-Teaming for AI Agents
AI agents combine language models with external data and tools that can modify files, call APIs, or execute code. Security failures can arise when adversarial content changes an agent's tool use or when the surrounding software contains vulnerabilities such as path traversal or…
PrivDrift: Auditing User-Secret Leakage Under Topic Drift in Active LLM Conversations
Large language models increasingly operate as persistent assistants in user-facing, shared-session, and tool-augmented settings. When users disclose sensitive information during an active conversation, that information may remain behaviorally recoverable through later prompts…
Introducing Gemini 3.8 Live with Live Avatar
Instrumental Monitor Evasion Emerges Under Ordinary Task Pressure
A central concern in AI safety is that agents may treat oversight as an obstacle when it conflicts with completing their goals. We study instrumental evasion, the propensity of LLM agents to circumvent runtime monitoring as a means of completing ordinary tasks. We introduce…
ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation
Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions…
Accelerating vision-language models with LFM2.5-VL-DSpark
Google Beam expands with new regions, partners, and customers
We’re expanding Google Beam to five new countries, and partnering with Industrious for an extended network.
Gemini 3.8 text-to-speech says hello
Advancing Private AI Compute with secure, server-side memory
Introducing private, server-side memory to Private AI Compute for personal AI.
Transformers now runs llama.cpp quants
The Closed Quorum: Inside the first reported autonomous AI C2 implant
CLOSEDQUORUM, a malware binary discovered through Cisco Talos’ CAIRN project, exhibits fully autonomous command and control (C2). It represents a shift in effort displacement for attackers, in which expanding portions of the attack chain can be executed without operator…
Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community
Introducing CAIRN: Frontier tracking for AI-integrated malware
Talos is releasing CAIRN, a research toolkit for hunting, classifying, and tracking emerging AI-integrated malware.
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
tokenizers v1: encode, decode and scaling, measured
New experts join Google’s AI & Economy team
We are expanding our AI & Economy team with world-class academic advisors, fellows, and core internal researchers.
Co-creating the future of fashion with Google
Google worked side-by-side with designers Jane Wade and Sergio Hudson to custom-design Google Flow tools to prep for NYFW.
Auditing in the age of (good enough) AI
Security firms have published numerous blog posts describing how they pointed their agent harness at a codebase and found dozens of bugs ( we’re one of them ). However, these posts tend to focus on agentic code review, which is just one aspect of how we use AI in our security…
A Vault with a Heap-View: The Uncomfortable Space Between AgentCore Harness and Identity
Analysis of how default configurations in AWS AgentCore Harness allow prompt injection to exfiltrate credentials, and key steps to secure your agents. The post A Vault with a Heap-View: The Uncomfortable Space Between AgentCore Harness and Identity appeared first on Unit 42 .
Should you care about an “AI slowdown?”
In this week's Threat Source, David talks about why focusing on your security basics is still your best bet, even in a world with rapid AI advancements.
Self-generated prompt injections in compaction summaries
Self-generated prompt injections in compaction summaries In Our framework for reporting model misalignment OpenAI provide "six reports on unexpected or concerning model behavior we’ve observed in the last six months". This one here is my favorite: they caught some of their…
Ransomware incidents in Japan in the first half of 2026: Investigation of The Gentlemen’s infrastructure and evidence of Qilin's AI use
Ransomware incidents in Japan rose 4.7% year over year. The Gentlemen was the most active group, with leak-site listings more than doubling from January to July. Qilin ranked second and appeared to use AI, while SMEs with capital under JPY 1 billion represented 80% of victims.
Making global data easier to explore
Google and the UN system have launched the UN System Data Commons, a new open platform making global statistics accessible and easy to search.
AI Threat Landscape Digest: July–August 2026
The defining development of the period came not from attackers but from the AI labs themselves, whose models broke out of controlled evaluations and reached real systems. In the wild, the criminal and state use of AI continued to mature along the lines tracked in earlier…
Securing the unpatchable in an age of AI-driven vulnerabilities
Advances in AI technology will continue to identify vulnerabilities that in some circumstances are difficult, or effectively impossible, to patch. Appropriate network segmentation, rigorous visibility, and the deployment of NGFW/IPS combinations can provide a powerful…
Agents at Large | Tracing Illicit OpenAI Agent Activity on Hugging Face
Two Hugging Face accounts reveal that OpenAI's agents staged relay code, internal probes and ChatGPT account registration beyond the published timeline.
New insights from Google’s AI & Economy ATLAS
We’ve translated ATLAS’s millions of global data points into an interactive, open-access experience.
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Building AI to accelerate science and improve lives
The true measure of AI is who it helps. Here’s how it’s impacting lives today. We're focused on key areas where advanced technology can help make extraordinary progress …
AI for everyone in every language
We’re moving beyond traditional text translation to build models that understand the world’s rich, living languages exactly as they are expressed.
AI for Societal Impact
Explore this collection to see how experts and local leaders are using AI breakthroughs to ensure everyone can share the opportunity of AI.
1Password's AI patching benchmark is misleading
1Password’s FLAWED report , published on August 6, 2026, gives defenders a misleading picture of AI patching. Its headline says models produced clean fixes only 26% of the time. That figure includes experiments that deliberately instructed agents to apply the wrong fix, along…
Watch astronaut Christina Koch and Google’s James Manyika discuss space, technology, and discovery.
Christina Koch sits down with James Manyika, Google’s Senior Vice President of Research, Labs, Technology & Society.
DevFest is back
DevFest 2026 is back and here’s how you can connect with one of the more than 800 global events to build, secure, and scale in the agentic AI era.
Metasploit Wrap Up: This One Goes to Sixteen!
This One Goes to Sixteen! Another banger from Metasploit with sixteen new modules, including ten exploit modules, with five on the CISA KEV list. Cisco, Papercut, Sonicwall, Jetbrains, and Langflow all have exploit modules, and not to be outdone, we even have a Metasploit…
PuzzleMask: Abusing Plain Prose as a Covert AI Attack Vector
Executive Summary In this research we introduce a prompt-crafting technique for bypassing quick LLM-based policy checks — using plain English (no emojis, base64, invisible formatting, etc.) A policy-violating payload (e.g. ”encrypt files in ~/Documents”, “give me a biohazard…
A “proof” of Fermat’s Last Theorem that fits the margin
Fermat famously claimed to have a “truly marvelous proof” of his Last Theorem , but he never wrote it down, insisting the margin of his page was too narrow to contain it. A few centuries later, Anthropic announced a complete formalization of Fermat’s Last Theorem using 13…
The Shared Clipboard Inside the Sandbox: Cross-Account Data Leakage in ChatGPT
Research by: Alexey Bukhteyev Key Takeaways Introduction Over the past several years, AI assistants have moved far beyond text generation. Modern systems can execute code, install additional dependencies, analyze user files, and access data through connected services. These…
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
AlphaGenome Atlas maps the molecular effects of 9 billion single-letter DNA variants across the human genome.
Introducing WeatherNext 3, our most advanced and accurate global weather AI model
Attackers Expose Ongoing AI Tool Use Targeting Organizations in Latin America
Explore how attackers targeting Latin American entities use AI for data exfiltration and how basic OpSec errors allow defenders to disrupt operations. The post Attackers Expose Ongoing AI Tool Use Targeting Organizations in Latin America appeared first on Unit 42 .
Proactive cyber defense for governments and enterprises
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation
Using autonomous AI agents, an attacker breached an enterprise network in a matter of hours. Understand how to address and defend against agentic attacks. The post An AI-Assisted Cyber Attack: Inside a Unit 42 Investigation appeared first on Unit 42 .
Introducing agentic video understanding with Gemini
Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety
New research reveals that AI safety refusal lives in a thin neural layer, highlighting the critical need for external, multi-layered security. The post Perturbation Probing: A New Diagnostic for the Fragility of LLM Safety appeared first on Unit 42 .
Gemini Omni 1.1 Flash lets you build with more control
Breaking Claude Code Opus 5 Auto Mode
Breaking Claude Code Opus 5 Auto Mode Anthropic are putting a great deal of faith in Claude Code's auto mode for protecting their coding agent users against prompt injection attacks. They recently made that the default and have made bold claims about its effectiveness. Johann…
Breaking Claude Code Opus 5 Auto Mode
In this post, we explore how a simple website summary request hijacks Claude Code Opus 5 in Auto Mode and achieves code execution with 60-80% attack success rate using a small sample size. This is interesting because a third-party evaluation commissioned by Anthropic showed a…
The State of AI-Enabled Malware August 2026: From Brand Abuse to Agentic Execution
Explore Unit 42 research on AI-enabled malware. Learn how existing behavioral detection and endpoint analytics stop AI-authored code before execution. The post The State of AI-Enabled Malware August 2026: From Brand Abuse to Agentic Execution appeared first on Unit 42 .
Recovering Encrypted LLM Reasoning Traces
A few days ago, a paper named “Stealing Reasoning Traces from Proprietary LLM APIs” was published. It describes a simple, yet super elegant way to recover encrypted LLM reasoning traces. Naturally, I had to try it. Background AI labs like OpenAI and Anthropic send reasoning…
The Model Is the Malware | What Four Agentic Intrusions Tell Defenders
OpenAI, Anthropic and Meta disclosed agents reaching external systems. The tools didn't matter, and that changes the playbook for investigating intrusions.
Stealing Reasoning Traces from Proprietary LLM APIs
Stealing Reasoning Traces from Proprietary LLM APIs A vanity domain name ( stolen-thoughts.com ) for a neat paper : Anthropic, OpenAI, and Google return encrypted chain-of-thought blocks to clients that can be replayed across sessions, users, and models. We take a trace produced…
Auto mode is now the default in Claude Code for Pro, Max, and Team plans
Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode , to the point that they are making it the default setting for new sessions in most Claude Code plans starting on August 14th. This was one of the…
Incident Report: unsanctioned agent behaviour during cyber testing
Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other companies while running an evaluation with models with the safety filters turned off. From their…
Headlines and short snippets only; each links to the original. Tags are assigned automatically from the text and can be wrong — the source is authoritative.