Open Source / News

Karpathy's CLAUDE.md tested: does it make Claude Code better?

The Karpathy-inspired CLAUDE.md has about 218,000 stars. We tested Claude Code with and without it on three coding jobs: Haiku 4.5 155 of 186 vs 152 of 186, Sonnet 5.5 186 of 186 vs 186 of 186.

Table of our test: Haiku 4.5 155 of 186 secret checks without the CLAUDE.md and 152 of 186 with it; Sonnet 5.5 186 of 186 without and 186 of 186 with it, plus lines of code, time and cost
Our own test on 10 October 2026: Claude Code with a clean setup, three jobs, three tries each. Chart by Not an AI App.

One of the most starred projects on GitHub is a single small file: a CLAUDE.md "derived from Andrej Karpathy's observations" on AI coding mistakes. About 218,000 people starred it. Claude Code reads a CLAUDE.md at the start of every session, so the idea is simple: put four good rules in it, and the AI makes fewer mistakes. We tested that. In our test the file did not make Claude Code better. Haiku 4.5 got 3 fewer checks passed, Sonnet 5.5 got the same score, and the code did not get clearly simpler.

What is Karpathy's CLAUDE.md?

The four rules of the CLAUDE.md file, word for word: Think Before Coding, Simplicity First, Surgical Changes, Goal-Driven Execution, each with its one-line summary
Typed out from the CLAUDE.md in multica-ai/andrej-karpathy-skills on GitHub, 10 October 2026. Image by Not an AI App.

A CLAUDE.md is a text file with instructions for Claude Code, Anthropic's AI coding helper. Claude Code reads it at the start of every session, as Anthropic's docs explain.

This one has four rules, each with a short list underneath:

1. Think Before Coding: "Don't assume. Don't hide confusion. Surface tradeoffs." 2. Simplicity First: "Minimum code that solves the problem. Nothing speculative." 3. Surgical Changes: "Touch only what you must. Clean up only your own mess." 4. Goal-Driven Execution: "Define success criteria. Loop until verified."

The whole file is short: 65 lines. The repository also has versions for Cursor and a plugin for Claude Code.

Who made it?

Facts about the repository: made by multica-ai, not Andrej Karpathy; based on Karpathy's post of 26 January 2026; created 27 January 2026; about 218,000 stars; no licence listed
From GitHub, 10 October 2026. Image by Not an AI App.

Not Andrej Karpathy. The repository belongs to multica-ai, and its README says the file is "derived from" Karpathy's observations. Karpathy is a well-known AI researcher who helped start OpenAI.

The repository was created on 27 January 2026, one day after Karpathy's post. On 10 October 2026 it had about 218,000 stars and 22,000 forks on GitHub. Stars are like bookmarks: they show interest, not that it works.

One more detail: GitHub lists no licence for it. Without a licence, the code is not open source in the legal sense, even though anyone can read it.

What Karpathy actually said

Screenshot of the repo's README, The Problems: quotes from Andrej's post about models making wrong assumptions, overcomplicating code and changing code as side effects
The README of multica-ai/andrej-karpathy-skills on GitHub. Screenshot from 10 October 2026, cropped.

Karpathy wrote about his experience with Claude Code in a long post on X on 26 January 2026. He said the models "make wrong assumptions on your behalf and just run along with them without checking", "really like to overcomplicate code", and sometimes change code they "don't sufficiently understand" as side effects.

But there is a sentence the README does not quote. Karpathy wrote that all this "happens despite a few simple attempts to fix it via instructions in CLAUDE.md". So he had already tried what this repository offers. That made us curious whether it works for us.

Our test: with and without the file

Diagram of our test: three coding jobs, four setups (Haiku 4.5 and Sonnet 5.5, each with and without the CLAUDE.md), the same secret checks, lines of code, files touched and time
Diagram by Not an AI App.

We used the three jobs from our private coding test, the same ones as in our Cursor test. The tests in our public test kit are known by now, so these stay secret:

  • D. Fix a money reader: code that turns "€ 1.234,56" into cents, with several bugs. 28 secret checks.
  • E. Split a restaurant bill: shared dishes and a tip, with cents that must add up exactly. 17 checks.
  • F. Read a CSV file: a spreadsheet saved as text, where commas, quotes and line breaks can hide inside a field. 17 checks.

We ran Claude Code with two models: Haiku 4.5, a smaller model that makes some mistakes, and Sonnet 5.5, a strong one. Each model did each job three times without the file and three times with the file in the project folder. Besides the secret checks, we counted the lines of code and checked which files Claude changed.

Did it get better?

Bar chart of the lines of code per job, for each job and each setup, with and without the CLAUDE.md
Our own test on 10 October 2026. Chart by Not an AI App.

No, not in our test:

  • Haiku 4.5: 155 of 186 secret checks without the file, 152 of 186 with it.
  • Sonnet 5.5: 186 of 186 without, 186 of 186 with.

The jobs also took about as long: 74 seconds and 70 seconds for Haiku, 33 seconds and 35 seconds for Sonnet. The token value per job, at Anthropic's prices, was $0.12 and $0.13 for Haiku, $0.14 and $0.14 for Sonnet.

The mistakes Haiku made were the same kind with and without the file: amounts like "1234,5" and "1.234.567,89" in the money job, and a few CSV rules. A file with general advice did not fix those.

Was the code simpler?

Overview of jobs where Claude added or changed a file outside the solution and test folder, per setup
Our own test on 10 October 2026. Chart by Not an AI App.

"Simplicity First" asks for the "minimum code that solves the problem". So we counted the non-blank lines in each solution.

  • Haiku 4.5: with the file, the code was longer: 97 lines instead of 85, on average.
  • Sonnet 5.5: with the file, the code was shorter: 59 lines instead of 63.

The tests Claude wrote itself also stayed about the same size. So in our jobs, the rule about simple code made no clear difference: a few lines less for Sonnet, more for Haiku.

Did it stay inside the task?

Diagram: we ran Claude Code with an empty config folder and a login token, so our own personal CLAUDE.md, plugins and skills were not loaded
Diagram by Not an AI App.

"Surgical Changes" asks Claude to touch only what it must. We compared every file before and after each job.

  • Haiku 4.5: 1 of 9 jobs without the file and 2 of 9 with it added a file outside the normal places. In each case it was a test file next to the code instead of in the test folder.
  • Sonnet 5.5: 0 of 9 without and 0 of 9 with it.

So the file did not make the changes neater either. Note that our jobs were small and in a fresh folder; in a big project there is more to touch, and the rule may matter more there.

Should you use it?

It does not hurt to try: the file is short and costs almost nothing in extra text. But do not expect it to fix an AI's mistakes. In our test it changed nothing measurable, and Karpathy himself said such instructions did not fix it for him.

Anthropic's own docs explain why: Claude Code treats a CLAUDE.md "as context, not enforced configuration", and "the more specific and concise your instructions, the more consistently Claude follows them". So rules about your own project, like how to run the tests or which folder to use, are likely to help more than general advice. To create a starting file for your project, type /init in Claude Code.

We looked at other popular Claude Code add-ons in our GitHub repos test and at its permission modes in our skip permissions test.

Questions people ask

List of questions people ask Google about CLAUDE.md files, with the 5 questions we answer marked
"People also ask" questions from Google (US), 10 October 2026, via DataForSEO. Chart by Not an AI App.

What are Andrej Karpathy's coding rules?

The popular file with "Karpathy" in its name has four rules: think before coding, keep it simple, change only what you must, and work towards a goal you can check. It was written by multica-ai, based on a post by Karpathy, not by Karpathy himself. In our test it did not make Claude Code score better.

How do I make an md file for CLAUDE?

The easiest way is to type /init in Claude Code. Claude then looks at your project and writes a starting CLAUDE.md with things like build and test commands. You can also create a text file called CLAUDE.md yourself and write your instructions in it.

Where to save a Claude MD file?

For a project, save it as CLAUDE.md in the project folder, or in .claude/CLAUDE.md. For instructions that apply to all your projects, use ~/.claude/CLAUDE.md in your home folder. Claude Code reads these at the start of a session, Anthropic's docs say.

Is it best practice to commit CLAUDE MD?

For the project file, yes: Anthropic's docs say a project CLAUDE.md is shared with your team through version control, so it should hold project rules like test commands. Personal preferences belong in a CLAUDE.local.md that you leave out of git.

What are the best practices for creating Claude.md files?

Anthropic's docs advise specific and short instructions, and less than 200 lines per file, because Claude treats the file as context, not as hard rules. In our test, a file with general advice changed nothing measurable, so rules about your own project are likely more useful.

How we checked

On 10 October 2026 we read the repository on GitHub (CLAUDE.md as of commit 2c60614), Karpathy's post on X and Anthropic's CLAUDE.md docs. The date of the post comes from its link on X. The screenshot is our own capture, cropped only.

We ran 36 coding jobs the same day on our Windows 11 PC with Claude Code 2.1.280, effort "medium": 3 jobs, 3 tries, 2 models, with and without the file. Claude could read, edit and run node, nothing else. We started Claude Code with an empty settings folder and a login token, so that our own personal CLAUDE.md, plugins and skills were not loaded. Otherwise "without the file" would not really be without instructions. The secret checks ran only after each job. Costs are worked out from Claude Code's token counts and Anthropic's list prices.

This was a small test with three short jobs in fresh folders. In big projects, with more room for side effects, the file may behave differently. The charts were made by us. None of the images was made by AI.