# Are agent skills safe? What a SKILL.md can reach, and how to check one

Canonical: https://clawdocx.com/blog/are-agent-skills-safe
Author: Mira Castellan
Published: 2026-10-07
Updated: 2026-10-07

> A skill is instructions plus code that runs with your agent's own authority. What a SKILL.md can reach on each platform, and how to check one.

A skill is safe in the way a downloaded program is safe: it does nothing on its own, and quite a lot once your agent runs it. Anthropic's own documentation is blunt about the trust model, saying to "use Skills only from trusted sources: those you created yourself or obtained from Anthropic", because a malicious skill "can direct Claude to invoke tools or execute code in ways that don't match the Skill's stated purpose". So the useful question is not whether skills are safe in general. It is what this skill can reach where you are about to run it, and the two things that set that are the runtime you install it into and the access your agent already has there.

## What does a skill actually add to an agent?

A skill is a folder with a `SKILL.md` at its root: YAML frontmatter carrying a `name` and a `description`, then Markdown instructions, plus any supporting files the author bundles. Anthropic's overview describes three levels that load at different times. The frontmatter is "always loaded", and Claude "loads this metadata at startup and includes it in the system prompt". The body loads when the description matches your request. Bundled resources and scripts load only when something references them.

That last level is where the security story starts. Anthropic describes bundled "executable scripts (fill_form.py, validate.py) that Claude runs using bash, providing deterministic operations without loading their code into context", and spells out the consequence: "When Claude runs `validate_form.py`, the script's code never loads into the context window. Only its output ... consumes tokens."

Read that as a reviewer rather than as a performance note. The agent executes bundled code that it never reads, so nothing in the loop has looked at the script before it runs. Whatever review happens has to be yours, before you install.

## Where does the real risk come from?

Not from a clever exploit. Palo Alto Networks' Unit 42, in research published on 23 June 2026, puts the mechanism plainly: "While conventional malware often faces limitations from language runtimes or containers, malicious skills use semantic instruction hijacking to bypass technical constraints." The paper continues that by misusing the model's natural language interpretation, "malicious skills can exploit the agent's operational context, including file systems, shells and credential managers, without requiring a conventional exploit".

Its sharpest sentence is about authority rather than code: "The lack of isolation between skill logic and agent authority means that installation results in complete control over the agent's identity. This allows a malicious skill to perform unauthorized actions through the agent's own authenticated sessions."

Anthropic's security section names the same hazards from the other side, as things to look for: tool misuse, where a skill invokes file operations, bash commands or code execution in harmful ways; data exposure, where a skill with access to sensitive data leaks it; and external dependencies, where "Skills that fetch data from external URLs pose particular risk, as fetched content may contain malicious instructions". It adds the point that a skill you vetted once is not vetted forever: "Even trustworthy Skills can be compromised if their external dependencies change over time."

## What can a skill reach on each platform?

The same `SKILL.md` carries different risk depending on where it runs, and the differences are documented. Anthropic's runtime constraints and the MCP extension's host rules give four concrete cases.

| Where the skill runs | Network access it gets | Package installs | Who can manage it centrally |
|---|---|---|---|
| Claude API, in the code execution container | None: "no network access" | "No runtime package installation" | Workspace-wide: all workspace members can access uploaded skills |
| claude.ai, with code execution enabled | "full, partial, or no network access", depending on user and admin settings | Not documented as available | Nobody: claude.ai "does not support centralized admin management or org-wide distribution of custom Skills" |
| Claude Code, on your own machine | "the same network access as any other program on the user's computer" | Local installs only; global installs discouraged | Nobody by default; skills are files in `~/.claude/skills/` or `.claude/skills/` |
| Served to a host over MCP | Set by the host, which must treat skill content as untrusted input | Set by the host | The host, which must require per-skill approval for code execution and tool grants |

The first and third rows are the two ends of the range. A skill on the API runs in a container with the network switched off, so an infostealer in it has nowhere to send anything. The same skill in Claude Code runs with your user's reach over your filesystem and your network. If you want to try something you have not fully read, the sandboxed runtime is where to try it.

Enterprise scanning narrows the gap without closing it. Anthropic documents that Claude Enterprise organizations "can also turn on Skill content scanning for custom Skills uploaded in claude.ai and Claude Cowork", then names the hole in the same breath: "Scanning doesn't cover Skills uploaded through the Skills API or the Claude Console."

## Have malicious skills actually been found?

Yes, and the best documented cases are on ClawHub, the marketplace for OpenClaw skills. Unit 42 describes it as a supply chain: "OpenClaw is an AI agent that executes third-party skills from ClawHub, its dedicated marketplace. Skills are markdown-driven packages with broad local system access, making ClawHub a critical link in the agentic software supply chain."

In its own analysis, which it dates from February to May 2026, Unit 42 reported: "We identified five unblocked skills. We reported all five to ClawHub for takedown. OpenClaw banned the accounts mentioned and deleted all of the skills." It sorts them into two macOS infostealers reaching command and control infrastructure, one skill padded to evade scanners, and two agentic techniques it calls runtime agentic affiliate injection and agentic front-running.

The padding case is worth the detail, because it shows how little sophistication was needed. Unit 42 writes that in one skill "the malicious payload appears at the start, followed by 22 MB of padding characters", and that "this padding inflates the file size beyond the limits that many content-analysis pipelines enforce before declining to process a file". The result, in its summary: "One skill has an inflated file size to exceed scanner thresholds, bypassing both ClawScan and VirusTotal detection."

Unit 42 also relays earlier numbers from other researchers, and they are worth keeping attached to the people who produced them. It reports that in early February 2026 Bitdefender Labs found that "approximately 17% of OpenClaw skills they analyzed in the first few weeks of the platform's release carried malicious payloads", that "Koi Security's ClawHavoc disclosure documented 341 malicious skills", and that Trend Micro separately confirmed skills distributing Atomic macOS stealer malware. Those are three different studies of one marketplace in its first months, not a measurement of skills everywhere.

ClawHub has not stood still. Unit 42 notes that the February findings "prompted ClawHub to integrate VirusTotal" and its own scanner, and that "OpenClaw is now also collaborating with NVIDIA to provide documentation of what each skill does, and to run NVIDIA's analysis tool on all skills". ClawHub's own checks, statuses and CLI are covered step by step in [how to audit a SKILL.md](/docs/skill-security-audit), which is the page to follow if OpenClaw is the agent you are using.

## Does a marketplace audit mean a skill is safe?

No, and the registries are the ones saying so. ClawHub's security audits page states that "audits are strong safety signals, but they are not a guarantee that a release is risk-free", and that "a Pass is reassuring, but it does not replace your own judgment". The skills.sh documentation answers the same question about its own listings: "We do our best to maintain a safe ecosystem, but we cannot guarantee the quality or security of every skill listed on skills.sh. We encourage you to review skills before installing and use your own judgment."

Unit 42's findings put a number on that caution. Of its five skills, it writes that "each case passed existing detection tools at the time of our analysis", and in one campaign ClawHub's automated auditing "returned a verdict of Pass" for one skill and no verdict for another, with "neither skill triggered detection, despite containing a verbatim paste-site prerequisite lure".

So a clean audit status tells you that nothing was caught, which is a weaker statement than nothing is there. It is a reason to keep reading, not a reason to stop.

## What does the MCP skills extension require of a host?

Skills served over the Model Context Protocol are the one case here whose security rules are normative requirements on the host rather than guidance to the reader, set out in the `io.modelcontextprotocol/skills` extension whose SEP-2640 the documentation marks as Final. They bind the host rather than vetting the skill, and four of them are worth knowing as a user.

- Skill content is untrusted by default. Hosts "MUST" "treat skill content as untrusted input", and "host-side code execution and permission grants such as `allowed-tools` require explicit per-skill user approval".
- Names are not identities. Hosts must "prevent skills with the same name from silently replacing one another", and must preserve both the originating server and the skill URI, because "names are labels and are not guaranteed to be unique".
- Files are pinned. A host must verify "each file's raw byte size and SHA-256 digest" against the manifest it approved, and "a changed, added, or removed file revokes that approval", requiring fresh approval before loading or executing.
- One approval is not blanket consent. Hosts must "obtain fresh user consent before activating a nested skill", and cross-server reads "require explicit per-call approval naming both servers".

The specification is careful not to oversell any of it. Its own note says that "digests establish consistency with the server's manifest, not trust in its content". Integrity checking tells you the file is the one you approved. It says nothing about whether approving it was wise.

## How do you check a skill before you install it?

The full procedure, including the registry commands and what to verify after install, is in [how to audit a SKILL.md](/docs/skill-security-audit); the containment side, including where credentials belong, is in [SKILL.md security best practices](/docs/skill-security-guide). What follows is the short version that applies whichever agent you use, and ClawDocx presents no third-party skill as vetted or endorsed.

1. **Read the description against the body.** The description is the part that is always in the system prompt and is what the agent matches your request against. A narrow description over a broad body is the mismatch worth finding.
2. **Open every bundled file, not just `SKILL.md`.** Anthropic's instruction is to "review all files bundled in the Skill: SKILL.md, scripts, images, and other resources", looking for "unusual patterns such as unexpected network calls, file access patterns, or operations that don't match the Skill's stated purpose". The scripts are the files the agent will run without reading.
3. **Find every outbound fetch.** A URL in the body is content that arrives later and can carry instructions, which is the dependency risk Anthropic names. A fetch from a host you do not recognise is a reason to stop.
4. **Check what it asks to be allowed.** Any declared tool permissions, allowed commands or environment variables are the skill's own statement of what it needs. Credentials it uses but never declares mean the label does not match the contents.
5. **Pick the narrowest runtime that still does the job.** The table above is the decision: a sandboxed container with no network beats your own laptop for anything you have not fully read.
6. **Re-check on update.** Approval of one version is not approval of the next, which is exactly why the MCP extension revokes an approval when any file changes.

## What should you do after it is installed?

Assume the skill has whatever reach the agent has, because that is Unit 42's finding about agent identity, and watch accordingly. Keep the agent's own tool policy and sandbox settings as the real boundary rather than relying on the skill to behave; those controls, not the skill's text, are what limit a command once it runs. If a skill does something you did not ask for, remove it and rotate every credential it could have read, on the assumption that it could read all of them.

A last point about scope. Most of the published evidence of malicious skills comes from one marketplace in its first year, and that is a fact about where researchers have looked as much as about where the risk is. The format is deliberately portable: the client showcase for the [Agent Skills standard](/blog/what-is-skill-md) listed 46 products when we counted it on 7 October 2026, Claude Code, ChatGPT and Codex, Cursor, GitHub Copilot, VS Code, Gemini CLI, Goose, OpenHands and OpenClaw among them. Nothing about a `SKILL.md` makes it safer on a different agent. What changes between agents is the runtime, which is the one variable you decide.

## Frequently asked questions

**Are agent skills safe to install?**

A skill is as safe as the person who wrote it and the place you run it. Anthropic's documentation says to use skills only from trusted sources, meaning ones you wrote yourself or got from Anthropic, because a malicious skill can direct the agent to invoke tools or execute code in ways that do not match the skill's stated purpose. Treat installing one the way you would treat installing software.

**Can a SKILL.md run code?**

Yes, if the skill bundles scripts and the agent has a way to run them. Anthropic's overview describes executable scripts that Claude runs using bash, and notes that a script's code never enters the context window, only its output. The model therefore runs code it has not read, which is why reading the bundled files yourself is the check that matters.

**Does a Pass from a marketplace audit mean a skill is safe?**

No, and both registries say so themselves. ClawHub's security audits page says audits are strong safety signals, but they are not a guarantee that a release is risk-free, and that a Pass does not replace your own judgment. The skills.sh docs say the site cannot guarantee the quality or security of every skill listed on it.

**Where is the safest place to run a skill you do not fully trust?**

The most restricted runtime you can put it in. Anthropic documents that skills on the Claude API run in a sandboxed container with no network access and no runtime package installation, while skills in Claude Code have the same network access as any other program on your computer. The same SKILL.md carries very different risk in those two places.

**Have malicious agent skills actually been found in the wild?**

Yes, on OpenClaw's ClawHub marketplace. Palo Alto Networks' Unit 42 reported on 23 June 2026 that it identified five unblocked malicious skills in an analysis running from February to May 2026, covering infostealers, scanner evasion and two agentic monetisation techniques, and that each case passed existing detection tools at the time of the analysis.

**Does the MCP skills extension make skills safer?**

It sets a baseline for hosts rather than vetting skills. The extension requires a host to treat skill content as untrusted input, to require explicit per-skill user approval before host-side code execution or permission grants, and to verify each file's size and SHA-256 digest against a manifest. The specification notes that digests establish consistency with the server's manifest, not trust in its content.

## Sources

- [Anthropic: Agent Skills overview](https://platform.claude.com/docs/en/agents-and-tools/agent-skills/overview)
- [Model Context Protocol: Skills extension overview](https://modelcontextprotocol.io/extensions/skills/overview)
- [Unit 42: OpenClaw's Skill Marketplace and the Emerging AI Supply Chain Threat](https://unit42.paloaltonetworks.com/openclaw-ai-supply-chain-risk/)
- [ClawHub: Security Audits](https://docs.openclaw.ai/clawhub/security-audits)
- [skills.sh: Documentation](https://skills.sh/docs)
- [Agent Skills specification](https://agentskills.io/specification)
- [Agent Skills: Client Showcase](https://agentskills.io/clients.md)