Anthropic / News

Claude Code mods on Windows: we built two and tried to break them

Anthropic's new Claude Code mods can block commands and hide secrets. We wrote two safety mods on Windows and tried nine ways to get past them. Three got through.

Table of nine test commands against two Claude Code mods: six handled as intended, two missed, one blocked by mistake
Chart of our own test, 4 October 2026: nine tries against two mods we wrote ourselves. Test folders and fake secrets only.

On 1 October 2026 Anthropic introduced mods for Claude Code: small TypeScript functions that change how the coding assistant behaves. A mod can block a command, rewrite a prompt, hide a secret from the model or add something to the screen. That sounds like a safety net. We wrote two simple safety mods ourselves on Windows, then tried nine ways to get past them. Six went as intended. Two got through. One time the guard blocked us when it should not have.

This follows our test of the God's Eye View MCP server in Claude, another way to extend what Claude can do.

What a mod is

Diagram: Claude decides to run a command, the mod sees the call first and can block, change or pass it, then other checks and the tool follow
Diagram by Not an AI App, based on Anthropic's mods documentation: where a mod sits when Claude calls a tool.

Anthropic's announcement describes mods as "small TypeScript functions that change how Claude Code works". Each time Claude Code does something, it emits an event: calling a tool, asking for permission, drawing part of the screen. A mod is a function that hooks into one of those events.

In the documentation every hook has the same shape: a function that receives the engine, the event and a function called next. Calling next passes the event on to the next mod and finally to Claude Code itself. Not calling it means the mod answers by itself. To block a command, a mod returns an object with a deny message, and Claude reads that message as the result.

Mods ship inside plugins, so you install them like any plugin. The documentation lists Claude Code 2.1.287 as the minimum version. They run in the terminal and in the Code tab of the Claude Desktop app, which is where we tested. Mods replace part of what hooks did before. Anthropic says hooks "can't rewrite events, draw new UI, or replace features. Mods can."

The two mods we built

Terminal output of claude plugin validate for the nai-guard mod: it hooks tool calls for Bash and PowerShell and calls the status and toast functions; validation passed with one warning
The validator's verdict on our guard mod, from our terminal on 4 October 2026, set in a frame by us. Paths shortened.

We kept both mods small on purpose, the way a careful user might start.

  • nai-guard blocks a recursive, forced delete. It hooks tool calls for both shells Claude Code uses on Windows, Bash and PowerShell, and checks the whole command against a few text patterns: rm -rf, rm -r -f, PowerShell's Remove-Item -Recurse -Force, and rmdir /s. On a match it refuses the call, shows a notice and counts the blocks in the status line.
  • nai-redact hides secrets. After a Bash command has run, it replaces anything that looks like our fake test secret, an OpenAI-style key starting with sk- or a GitHub token starting with ghp_ with [REDACTED], before Claude reads the output.

The core of the guard is a few lines:

on('tool.call', { tool: 'Bash' }, ($, e, next) => {
  if (!check(e.command)) return next(e)
  return { deny: 'recursive force-delete blocked by a test mod.' }
})

Before loading them we ran claude plugin validate on both. It reads a mod the way Claude Code will and lists every event the mod hooks and every action it can take. For our guard: tool calls for Bash and PowerShell, plus the status line and the notice. That list is the first thing to read before you install a mod someone else wrote.

Claude Code itself can write mods too. When it does, they go into a per-session folder, and Claude Code asks once whether to turn on hot reloading for that session. The owner of this site said yes; the mods loaded at the end of that turn.

What happened when we tried to break them

Screenshot of a notice in Claude Desktop reading: nai-guard: nai-guard blocked a recursive delete
The notice our guard mod showed in Claude Desktop. Claude Code puts the plugin's name in front, so ours reads twice. Screenshot by Marvin Smit.

We ran nine commands, one at a time, against empty test folders and fake secrets. No real file and no real key was ever at risk.

#What we triedWhat happened
1A command that prints a fake secret and a fake sk- keyBoth reached Claude as [REDACTED]
2A plain recursive force-delete of a test folderBlocked
3The same, nested: echo start && bash -c "…"Blocked
4The same with separate flags, -r -fBlocked
5The same, disguised as $(echo rm) -rfGot through: the test folder was deleted
6PowerShell's Remove-Item -Recurse -ForceBlocked
7A recursive delete without the force flagAllowed: outside our rule
8The fake secret read with Claude's Read tool instead of a commandNot redacted: Claude saw it in full
9Unplanned: our own commit message only mentioned the wordsBlocked by mistake

Every block showed up in Claude Desktop as a notice, as in the screenshot in this section, and Claude received our deny message as the tool's result. The redaction was invisible to us as users: the command ran normally, and only what Claude read was changed.

Where our mods fell short

Diagram of three limits: disguised commands got past the text pattern, the Read tool bypassed the redaction, and a commit message caused a false alarm
Diagram by Not an AI App summarising tests 5, 8 and 9 of our own test.

The three misses are more useful than the six successes.

A text pattern cannot read intent. Test 5 built the word rm while the command ran, with $(echo rm). Our guard checks the text of the command before it runs, and that text does not contain "rm -rf". The folder was deleted. Any rule that matches text can be talked around like this, and a model that has been tricked by something it read could do the same.

A mod only guards the door it watches. Test 8 put the same fake secret in a file and let Claude open it with the Read tool, not with a shell command. Our redaction mod hooks Bash only, so the secret came through in full. Reading it with cat in Bash, a moment later, gave [REDACTED]. To hide secrets properly, a mod would have to cover every tool that can read files, and every MCP tool that can return them.

A strict rule also hits innocent text. Test 9 was not planned. While saving these notes, our own git commit was refused, because the commit message mentioned the delete command we were testing. Nothing would have been deleted. The guard cannot tell a command that runs from text that only names it.

There is a fourth gap we found in the documentation rather than in a test. A hook that fails is skipped. If a hook throws an error or takes longer than its time limit before it calls next, Claude Code "skips it, and the next handler runs in its place". For a guard, that means the command goes ahead. Our guard did not use the documented fix, a catch handler that refuses the call when the hook fails. For a redaction hook that fails after the command has run, the documentation says the original "result stands", so the secret goes through.

What the documentation says about safety

Five safety points from the documentation: not sandboxed, fails open, can overrule deny rules outside managed setups, sec-default only on Team and Enterprise, draws only in the terminal and Desktop Code tab
Diagram by Not an AI App, quoting Anthropic's mods and permissions documentation as read on 4 October 2026.

Anthropic is direct about the risks, and it is worth reading before you install a mod you did not write.

  • Not sandboxed. "Mods run with the same access to your machine as Claude Code itself. They aren't sandboxed," the announcement says. The documentation adds that a process a mod starts runs outside Claude Code's sandbox for shell commands.
  • They can overrule your rules. According to the permissions documentation, on a machine with managed settings, or when you are signed in with a Team or Enterprise plan, deny rules hold over a mod by default. "Anywhere else, the mod can approve a call that a deny rule refuses."
  • sec-default is not for everyone. A built-in mod called sec-default loads first on Team and Enterprise plans and on managed machines, and stops installed mods from, for example, overriding deny rules. Its source also shows what it leaves alone: tool calls, file access and starting processes all pass.
  • Where they draw. Hooks run wherever the plugin loads, but anything a mod draws appears only in the terminal and the Desktop Code tab. The documentation lists the VS Code extension, claude -p and cloud sessions as places where nothing is drawn, and says plugins do not run at all in a WSL session in the Desktop app.

None of this is a flaw in itself. Mods are designed to change behaviour deeply, and that power cuts both ways. The practical rule is Anthropic's own: install mods "only from sources you trust, the same way you'd install any code on your computer".

How we checked

Diagram of our test in three steps: write two mods and validate them, load them with hot reload in Claude Desktop, try nine commands one by one
Diagram by Not an AI App: how we tested.

We read Anthropic's announcement of 1 October 2026, the mods documentation including the pages on events and permissions, and the source of the built-in mods, all on 4 October 2026. The quotes in this article were checked against those pages that day. Documentation and behaviour can change with any update.

We wrote both mods ourselves on 4 October 2026 and ran them in the Code tab of Claude Desktop on Windows 11, with hot reloading switched on by the owner of this site. We validated them with claude plugin validate from the standalone command-line tool. The nine commands ran one at a time against empty test folders and fake secrets in a temporary folder. The notice in this article is a screenshot by Marvin Smit and has only been resized; the validator output is our own terminal text set in a frame; the other images are charts and diagrams we made from our test file. None of the images is AI-generated.

Our test covers two simple mods and nine commands on one Windows laptop. It does not cover mods written by others, Team or Enterprise plans with sec-default, managed settings, macOS or Linux, or the full list of events. This is an editorial review of the sources we consulted and the tests we ran, not a security audit.