Matt Pocock's skills tested in Claude Code: tdd and diagnosing-bugs
We installed Matt Pocock's skills plugin and tested tdd and diagnosing-bugs on our secret coding jobs. Haiku 4.5 got a few more checks right; Sonnet 5.5 stayed perfect but took twice as long.
Matt Pocock's "Skills for Real Engineers" is one of the most starred projects on GitHub, with about 285,000 stars, and it is in Anthropic's official plugin marketplace for Claude Code. A skill is a set of instructions that tells the AI how to do one kind of work, like fixing bugs or writing tests. We installed the plugin and tested two of its skills on our secret coding jobs. They clearly changed how Claude worked. For Haiku 4.5, the bug-fixing skill gave a few more correct checks. For Sonnet 5.5 the score stayed perfect, but the test-first skill doubled the time and made each job about 76% more expensive.
What is in the plugin?
The plugin, version 1.2.3, has 25 skills. Its own description calls them skills "for real engineering: grilling, spec/ticket flows, TDD, code review, domain modelling and more".
- 18 engineering skills, such as
tdd(write tests first),diagnosing-bugs(a step-by-step way to find a bug),code-reviewanddomain-modeling. - 7 productivity skills, such as
grill-me, which questions you about a plan, andhandoff, which writes notes for the next session.
The skills are plain text files. Claude reads one when you call it, or when it decides it fits the job. Matt Pocock is a TypeScript teacher; his GitHub profile calls him a "TypeScript wizard" who built the course Total TypeScript. The repository is free under the MIT licence.
How to install it
In a terminal, one command is enough:
claude plugin install mattpocock-skills@claude-plugins-official
Then, inside Claude Code, you call a skill with its name, for example /mattpocock-skills:tdd followed by what you want. The README also has install steps for Codex, GitHub Copilot, Gemini and other tools. It says "A plugin updates itself"; to update by hand, use claude plugin update.
We looked at one skill from this collection before, Grill Me, in our test of eight GitHub repos.
Our test: two skills, two models
We used the three jobs from our private coding test. The tests in our public test kit are known by now, so these stay secret:
- D. Fix a money reader with several bugs. Here we used
diagnosing-bugs. 28 secret checks. - E. Split a restaurant bill from a spec. Here we used
tdd. 17 checks. - F. Read a CSV file from a spec, also with
tdd. 17 checks.
Each job ran three times with Claude Haiku 4.5, a smaller model that makes mistakes, and three times with Claude Sonnet 5.5, a strong one. We compared with the same jobs without a skill, run the day before in the same way, from our CLAUDE.md test.
Claude Code ran with a clean setup, with only this plugin installed, so our own settings could not affect the result.
Did the skills help?
A little, for the weaker model:
- Haiku 4.5: 155 of 186 secret checks without a skill, 160 of 186 with one. Most of the gain was on the bug fix: 60 of 84 without, 66 of 84 with
diagnosing-bugs. - Sonnet 5.5: 186 of 186 with and without. It was already perfect, so there was nothing to gain.
The skills were really used: in 8 of 9 Haiku jobs and 6 of 9 Sonnet jobs, Claude used the skill's own words, such as "phase" for the bug skill and "red" and "green" for the test-first skill. Haiku's results also varied a lot from try to try, so a difference of a few checks is not a big one.
What it cost
The skills made Claude work more carefully, step by step, and that takes time:
- Sonnet 5.5: 33 seconds per job without a skill, 67 seconds with one. It took about 22 steps instead of 10, and the token value per job went from $0.14 to $0.24.
- Haiku 4.5: about the same time, 74 seconds and 69 seconds, and about the same cost.
The tdd skill asks Claude to write one test, then just enough code, and repeat. That is a good habit for real projects, where tests protect you later. For small jobs like ours, it mainly added work.
One more thing we noticed: tdd says to agree on what to test "with the user" first. When Claude works on its own, as in our test, there is nobody to ask, so it just went ahead.
Should you install it?
It is free, quick to install and easy to remove, so trying it costs little. Our advice, based on this small test:
- Use `diagnosing-bugs` when you are stuck on a bug. It gave a weaker model a better plan.
- Use `tdd` when you want tests you can keep, not to finish quick jobs faster.
- Expect more steps and a higher bill with a strong model.
We also tested a single CLAUDE.md with general coding rules, and it changed nothing measurable. Skills that you call for one specific job did more. This was a small test with three short jobs in fresh folders; on a big project the skills may matter more.
Questions people ask
What is Matt Pocock known for?
Teaching TypeScript. His GitHub profile calls him a "TypeScript wizard" who built the course Total TypeScript. He also shares a popular, free set of skills for AI coding agents, with about 285,000 stars on GitHub in October 2026.
What are Matt Pocock's best skills for Codex?
We tested his skills in Claude Code, not in Codex. The same skills install in Codex too, the README says. Of the two we tried, diagnosing-bugs helped a smaller model fix more of a bug, and tdd made the work slower but more careful.
What are Matt Pocock's code review skills?
The plugin has a code-review skill. It checks the changes since a chosen point in two ways: whether the code follows the project's coding standards, and whether it does what the spec asked. We did not test this skill.
How do I update Matt Pocock's skills?
If you installed them as a Claude Code plugin, the README says "A plugin updates itself". You can also update by hand with claude plugin update. If you copied the files into your project another way, you update them yourself.
How we checked
On 11 October 2026 we installed mattpocock-skills 1.2.3 from Anthropic's official marketplace with claude plugin install, into a separate, empty Claude Code settings folder, and read the plugin's own files and the README. We read Matt Pocock's GitHub profile the same day.
We ran 18 jobs with the skills on our Windows 11 PC with Claude Code 2.1.280, effort "medium", logged in with a token from the site owner's Claude plan: 3 jobs, 3 tries, 2 models. Each job started with the skill's slash command. The 18 jobs without a skill come from our CLAUDE.md test of 10 October 2026, with the same setup. Claude could read, edit and run node; on Windows it sometimes also tried PowerShell, which was not allowed, in all setups alike. The secret checks ran only after each job. Costs are worked out from Claude Code's token counts and Anthropic's list prices.
The charts were made by us. None of the images was made by AI.

