Attack on AIResearch
arXiv cs.CR01 Oct 2026LLM01:2025
Quantized large language models are increasingly deployed on edge devices for their low latency and energy efficiency. However, model quantization weakens alignment safeguards, leaving qLLMs (quantized large language models) highly vulnerable to jailbreak attacks. To address…
Attack on AIResearch
arXiv cs.CR30 Sep 2026LLM01:2025
Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby…
Attack using AIProvider
Hugging Face30 Sep 2026
Attack on AIResearch
arXiv cs.CR30 Sep 2026LLM01:2025
Agent interaction protocols such as ACP and A2A have moved LLM-based agents toward multi-agent collaboration, introducing new security threats. A task sent by a remote peer over A2A is treated as a legitimate request, providing a natural channel for indirect prompt injection…
Attack on AIResearch
arXiv cs.CR30 Sep 2026LLM01:2025
Large language models have achieved remarkable capabilities across diverse domains, yet their safety alignment remains vulnerable to jailbreak attacks. In this work, we identify a previously underexplored failure mode - safety generalization lag - where alignment trained…
Attack on AIResearch
arXiv cs.CR30 Sep 2026LLM01:2025
Repository instruction files guide coding agents, but also expose them to prompt injection. Malicious rules can request credential access or data transfer while the agent produces a correct patch. We present Aletheia, a framework for permission-minimality testing. Aletheia…
Attack on AIResearch
arXiv cs.CR29 Sep 2026LLM01:2025
Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core…
Attack on AIResearch
arXiv cs.CR29 Sep 2026LLM01:2025
When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does…
Attack on AIResearch
arXiv cs.CR29 Sep 2026LLM01:2025
Tool-using LLM agents remain vulnerable to indirect prompt injection because trusted instructions and untrusted observations share one context, allowing malicious content to steer consequential input-filtering defenses. Multi-path consensus defenses still leave a high attack…
Attack on AIResearch
arXiv cs.CR29 Sep 2026LLM01:2025
Large language models (LLMs) increasingly power agents that access sensitive information, use external tools, and modify software repositories. Although these capabilities offer substantial benefits, they also create security risks such as jailbreaks, prompt injection, and…
Attack on AIResearch
arXiv cs.CR29 Sep 2026LLM01:2025
A Technical Report: Operating a large language model (LLM) as a service requires more than inference infrastructure: the provider must also defend against adversarial interactions that seek to exploit the service, including jailbreaking for harmful use, sophisticated denial of…
Attack on AIResearch
arXiv cs.CR29 Sep 2026LLM01:2025
Safety-aligned language models are commonly deployed as multi-turn assistants, which lets adversaries spread unsafe intent across several user turns instead of a single prompt. Gradient-based jailbreak detectors such as GradSafe were developed for single prompts: they score an…