Open Source / News

Matt Pocock's skills tested in Claude Code: tdd and diagnosing-bugs

We installed Matt Pocock's skills plugin and tested tdd and diagnosing-bugs on our secret coding jobs. Haiku 4.5 got a few more checks right; Sonnet 5.5 stayed perfect but took twice as long.

Table of our test: Haiku 4.5 155 of 186 secret checks without a skill and 160 of 186 with one; Sonnet 5.5 186 of 186 and 186 of 186, with time and cost per job
Our own test, 10 and 11 October 2026: Claude Code with a clean setup, three jobs, three tries each. Chart by Not an AI App.

Matt Pocock's "Skills for Real Engineers" is one of the most starred projects on GitHub, with about 285,000 stars, and it is in Anthropic's official plugin marketplace for Claude Code. A skill is a set of instructions that tells the AI how to do one kind of work, like fixing bugs or writing tests. We installed the plugin and tested two of its skills on our secret coding jobs. They clearly changed how Claude worked. For Haiku 4.5, the bug-fixing skill gave a few more correct checks. For Sonnet 5.5 the score stayed perfect, but the test-first skill doubled the time and made each job about 76% more expensive.

What is in the plugin?

List of the 25 skills in mattpocock-skills 1.2.3: 18 engineering skills such as tdd, diagnosing-bugs and code-review, and 7 productivity skills such as grill-me and handoff
From the plugin's own plugin.json, 11 October 2026. Image by Not an AI App.

The plugin, version 1.2.3, has 25 skills. Its own description calls them skills "for real engineering: grilling, spec/ticket flows, TDD, code review, domain modelling and more".

  • 18 engineering skills, such as tdd (write tests first), diagnosing-bugs (a step-by-step way to find a bug), code-review and domain-modeling.
  • 7 productivity skills, such as grill-me, which questions you about a plan, and handoff, which writes notes for the next session.

The skills are plain text files. Claude reads one when you call it, or when it decides it fits the job. Matt Pocock is a TypeScript teacher; his GitHub profile calls him a "TypeScript wizard" who built the course Total TypeScript. The repository is free under the MIT licence.

How to install it

Commands: claude plugin install mattpocock-skills@claude-plugins-official, claude plugin update, claude plugin list, and /mattpocock-skills:tdd inside Claude Code
Image by Not an AI App.

In a terminal, one command is enough:

claude plugin install mattpocock-skills@claude-plugins-official

Then, inside Claude Code, you call a skill with its name, for example /mattpocock-skills:tdd followed by what you want. The README also has install steps for Codex, GitHub Copilot, Gemini and other tools. It says "A plugin updates itself"; to update by hand, use claude plugin update.

We looked at one skill from this collection before, Grill Me, in our test of eight GitHub repos.

Our test: two skills, two models

Diagram of our test: three coding jobs, Haiku 4.5 and Sonnet 5.5, with and without the skill, the same secret checks, test lines, time and cost
Diagram by Not an AI App.

We used the three jobs from our private coding test. The tests in our public test kit are known by now, so these stay secret:

  • D. Fix a money reader with several bugs. Here we used diagnosing-bugs. 28 secret checks.
  • E. Split a restaurant bill from a spec. Here we used tdd. 17 checks.
  • F. Read a CSV file from a spec, also with tdd. 17 checks.

Each job ran three times with Claude Haiku 4.5, a smaller model that makes mistakes, and three times with Claude Sonnet 5.5, a strong one. We compared with the same jobs without a skill, run the day before in the same way, from our CLAUDE.md test.

Claude Code ran with a clean setup, with only this plugin installed, so our own settings could not affect the result.

Did the skills help?

Table of secret checks per job: with diagnosing-bugs, Haiku 4.5 passed 66 of 84 on the bug fix instead of 60 of 84; Sonnet 5.5 passed everything with and without
Our own test, three tries per job added up.

A little, for the weaker model:

  • Haiku 4.5: 155 of 186 secret checks without a skill, 160 of 186 with one. Most of the gain was on the bug fix: 60 of 84 without, 66 of 84 with diagnosing-bugs.
  • Sonnet 5.5: 186 of 186 with and without. It was already perfect, so there was nothing to gain.

The skills were really used: in 8 of 9 Haiku jobs and 6 of 9 Sonnet jobs, Claude used the skill's own words, such as "phase" for the bug skill and "red" and "green" for the test-first skill. Haiku's results also varied a lot from try to try, so a difference of a few checks is not a big one.

What it cost

Bar chart: with the tdd and diagnosing-bugs skills, Sonnet 5.5 took 67 seconds and $0.24 per job instead of 33 seconds and $0.14
Token value at Anthropic's list prices. Our own test, chart by Not an AI App.

The skills made Claude work more carefully, step by step, and that takes time:

  • Sonnet 5.5: 33 seconds per job without a skill, 67 seconds with one. It took about 22 steps instead of 10, and the token value per job went from $0.14 to $0.24.
  • Haiku 4.5: about the same time, 74 seconds and 69 seconds, and about the same cost.

The tdd skill asks Claude to write one test, then just enough code, and repeat. That is a good habit for real projects, where tests protect you later. For small jobs like ours, it mainly added work.

One more thing we noticed: tdd says to agree on what to test "with the user" first. When Claude works on its own, as in our test, there is nobody to ask, so it just went ahead.

Should you install it?

Overview: use diagnosing-bugs when stuck on a bug, tdd when you want tests you can keep, and no skill for quick small jobs with a strong model
Chart by Not an AI App, based on our test of 11 October 2026.

It is free, quick to install and easy to remove, so trying it costs little. Our advice, based on this small test:

  • Use `diagnosing-bugs` when you are stuck on a bug. It gave a weaker model a better plan.
  • Use `tdd` when you want tests you can keep, not to finish quick jobs faster.
  • Expect more steps and a higher bill with a strong model.

We also tested a single CLAUDE.md with general coding rules, and it changed nothing measurable. Skills that you call for one specific job did more. This was a small test with three short jobs in fresh folders; on a big project the skills may matter more.

Questions people ask

List of questions people ask Google about Matt Pocock's skills, with the 4 questions we answer marked
"People also ask" questions from Google (US), 11 October 2026, via DataForSEO. Chart by Not an AI App.

What is Matt Pocock known for?

Teaching TypeScript. His GitHub profile calls him a "TypeScript wizard" who built the course Total TypeScript. He also shares a popular, free set of skills for AI coding agents, with about 285,000 stars on GitHub in October 2026.

What are Matt Pocock's best skills for Codex?

We tested his skills in Claude Code, not in Codex. The same skills install in Codex too, the README says. Of the two we tried, diagnosing-bugs helped a smaller model fix more of a bug, and tdd made the work slower but more careful.

What are Matt Pocock's code review skills?

The plugin has a code-review skill. It checks the changes since a chosen point in two ways: whether the code follows the project's coding standards, and whether it does what the spec asked. We did not test this skill.

How do I update Matt Pocock's skills?

If you installed them as a Claude Code plugin, the README says "A plugin updates itself". You can also update by hand with claude plugin update. If you copied the files into your project another way, you update them yourself.

How we checked

Diagram: with the skill, an empty Claude Code settings folder with only mattpocock-skills; without, another empty folder with no plugins; our own settings were loaded in neither
Diagram by Not an AI App.

On 11 October 2026 we installed mattpocock-skills 1.2.3 from Anthropic's official marketplace with claude plugin install, into a separate, empty Claude Code settings folder, and read the plugin's own files and the README. We read Matt Pocock's GitHub profile the same day.

We ran 18 jobs with the skills on our Windows 11 PC with Claude Code 2.1.280, effort "medium", logged in with a token from the site owner's Claude plan: 3 jobs, 3 tries, 2 models. Each job started with the skill's slash command. The 18 jobs without a skill come from our CLAUDE.md test of 10 October 2026, with the same setup. Claude could read, edit and run node; on Windows it sometimes also tried PowerShell, which was not allowed, in all setups alike. The secret checks ran only after each job. Costs are worked out from Claude Code's token counts and Anthropic's list prices.

The charts were made by us. None of the images was made by AI.