Study Finds 17.6% of Nearly 1,000 Real AI Agent Skills Carry Exploitable Prompt-Injection Flaws

Researchers built an AI red team to audit live skills from the marketplace skills.sh, and their own scanner beat seven other LLMs at catching what it found.

EduFabTech · 28 September 2026 · 4 min read · 261 views
A donut chart shows 17.6% of the 954 audited AI agent skills flagged for exploitable prompt-injection flaws, with a cracked skill-file icon above.
EduFabTech · Own work

A study published this month found that 168 out of 954 publicly available "skills" for AI coding agents, 17.6%, contain exploitable prompt-injection vulnerabilities that could let a malicious file hijack an agent's actions. The audit, run with a new open framework called SkillSecurer, is one of the first systematic security assessments of a distribution model that has expanded rapidly over the past year with no standard review process attached to it.

The paper, "SkillSecurer: Detecting and Patching Prompt-Injection Vulnerabilities in AI Agent Skills," was submitted to arXiv on September 12, 2026 by Donato Mecca, Alberto Verna, Youness Bouchari, Nikhil Jha and Marco Mellia of Politecnico di Torino. "Skills" are the reusable folders of instructions, scripts and reference files that let an AI agent load domain-specific expertise on demand rather than relying only on what is in its prompt. Anthropic introduced the concept for Claude in October 2025 and later published it as an open, cross-platform standard, and marketplaces such as skills.sh now let the same packages be installed into Claude Code, Cursor, Copilot, Windsurf and other agentic tools.

That portability is also the risk the study targets. A skill is mostly prose: a markdown file telling the agent what to do, sometimes bundled with executable scripts. Conventional code scanners are built to flag suspicious code, not to judge whether a paragraph of natural-language instructions quietly tells an agent to exfiltrate credentials or run an unrelated command once invoked. Because agents are built to trust the skills they load, a single poisoned file can influence every task the skill is used for.

A horizontal bar chart ranks the eight LLM backends by injection detection rate, with Claude Sonnet 5's 100% score highlighted in red at the top.
A horizontal bar chart ranks the eight LLM backends by injection detection rate, with Claude Sonnet 5's 100% score highlighted in red at the top.EduFabTech · Own work

How the audit worked

SkillSecurer runs as a "fully agentic" red-team/blue-team pipeline. A red agent generates malicious variants of a skill across nine defined prompt-injection threat categories, and a blue agent then tries to detect, localise and patch the injected content, with the authors manually cross-validating the results at each stage. The researchers then compared eight different large language models used as the blue agent's detection backend, measuring what they call the injection detection rate (IDR) against the red agent's synthetic attacks.

Backend modelInjection detection rate
Claude Sonnet 5100.0%
Gemini 3.6 Flash99.4%
GPT-5.6 Sol99.4%
Kimi K398.2%
DeepSeek Pro96.3%
Gemma 31B92.7%
GPT-oss 120B86.1%
DeepSeek Flash83.6%

With its best-performing backend, the authors report SkillSecurer was the only scanner in their comparison to reach a 100% detection rate on the synthetic test set. They then pointed the same pipeline at real skills pulled from skills.sh: of 954 skills for which the audit produced complete data, 168 were flagged for latent security issues, and the flagged skills accounted for 289 distinct security-relevant findings in total, not all in every skill.

Part of a wider pattern

The paper lands alongside a broader push to formalize security review for this category of software. The OWASP Agentic Skills Top 10 (AST10), published by the OWASP Foundation in 2026, is described by its authors as the first comprehensive framework aimed specifically at the risks in agent skills rather than at the model-to-tool protocols that carry them. The project distinguishes between how an agent talks to its tools and what the loaded skill instructs those tools to actually do, and maps its ten risk categories onto the Cloud Security Alliance's layered threat model for agentic systems.

A three-stage diagram shows SkillSecurer's audit pipeline: a red agent injects attacks, a blue agent detects and patches them, then humans cross-validate the results.
A three-stage diagram shows SkillSecurer's audit pipeline: a red agent injects attacks, a blue agent detects and patches them, then humans cross-validate the results.EduFabTech · Own work

Neither the SkillSecurer authors nor OWASP claim their figures describe the entire skills ecosystem: the arXiv study covers one marketplace, skills.sh, at a single snapshot in time, and the detection-rate comparison was run against the red agent's own synthetic attacks rather than against attacks designed by an independent third party. Even so, a roughly one-in-six vulnerability rate among skills already published and available for installation is a concrete number for an ecosystem that, until this year, had no dedicated audit tooling or top-ten risk list at all.

What it means for teams using agent skills

For engineering teams adopting agent skills, the practical takeaway is that a skill should be treated the same way as any other third-party dependency: reviewed before installation, pinned to a known version, and re-checked when it updates, rather than trusted by default because it came from a marketplace. The SkillSecurer authors have released their detection-and-patching pipeline openly, giving marketplace operators and individual teams a way to run comparable checks on the skills they already have installed, though the paper stops short of naming which specific skills or publishers on skills.sh were flagged.


References
  1. Donato Mecca, Alberto Verna, Youness Bouchari, Nikhil Jha, Marco Mellia. SkillSecurer: Detecting and Patching Prompt-Injection Vulnerabilities in AI Agent Skills. arXiv preprint, 2026. link
  2. OWASP Foundation. OWASP Agentic Skills Top 10 (AST10). OWASP Foundation, 2026. link
  3. Anthropic. Introducing Agent Skills. Anthropic, 2025. link