LiteLLM tested: one API for every AI model, does it hold up?
We ran LiteLLM's proxy with three makers' models on our secret coding test. No measurable delay, 18 of 18 jobs perfect, and its cost tracking matched OpenRouter's bill 18 of 18 times.
LiteLLM is one of the most used open-source tools for working with AI models: about 60,900 stars on GitHub, and 22,000 searches a month in the US. It promises one way to talk to more than 100 models. We installed it in a sandbox, a separate test computer we deleted afterwards, and checked three things: does it slow things down, does it work with models from different makers, and does its cost tracking add up? It added no delay we could measure, all 18 coding jobs passed every secret check, and its cost was exactly right 18 out of 18 times.
What is LiteLLM?
An API is the connection a program uses to talk to an AI model. Every maker has its own, with its own small differences. LiteLLM puts one layer in front of all of them. Its docs say it "gives you a single, unified interface to call 100+ LLMs" in the format OpenAI uses.
You can use it in two ways:
- As a library: a piece of Python code inside your own program.
- As a proxy, or gateway: a small server that your apps talk to. It passes requests on to the right model and keeps track of keys, budgets and costs.
We tested the proxy, because that is what most teams use it for.
Who makes it?
LiteLLM is made by a company called BerriAI. The repository started in July 2023 and had about 60,900 stars on GitHub on 10 October 2026.
Most of it is open source under the MIT licence, which means you may use and change it for free. One folder, for enterprise features, has its own licence. We used version 1.104.2, the newest on PyPI, Python's package store, that day.
Is LiteLLM safe? The March 2026 hack
In March 2026, LiteLLM was hit by a supply chain attack: someone published hacked versions of the real package. LiteLLM's own security update says: "The compromised PyPI packages were litellm==1.82.7 and litellm==1.82.8." They were online on 24 March 2026 for about 40 minutes before PyPI blocked them.
The hidden code tried to steal passwords and keys from the computer it ran on. LiteLLM says the attack came in through a security scanner used in its own build process. People who used the official LiteLLM Docker image were not affected, LiteLLM says.
What this means for you: never install 1.82.7 or 1.82.8, and pin the exact version you use. If you did install one of them, LiteLLM advises treating every password and key on that computer as stolen and changing them. We installed a much later version, in a throwaway sandbox.
Our test
We ran LiteLLM's proxy in a fresh Linux sandbox and connected it to three models from three makers, all through OpenRouter, a service that sells access to many models:
- Claude Haiku 5.5 (Anthropic)
- GPT-6 Luna (OpenAI)
- Gemini 3.8 Flash (Google)
Then we did three tests:
1. Speed: 20 small requests straight to OpenRouter and 20 through LiteLLM, taking turns. 2. Real work: the three jobs from our private coding test, two tries each, for all three models, all through LiteLLM. The jobs are the same as in our Haiku 5.5 test; see our methods page for how we test. 3. Fallback: a model name that points to a model that does not exist, with Haiku 5.5 as backup.
Does LiteLLM slow things down?
No, not that we could measure. For a tiny request, the middle time was:
- Straight to OpenRouter: 1010 milliseconds.
- Through LiteLLM: 984 milliseconds.
That small difference is less than the normal ups and downs of the internet. LiteLLM also reports its own work time in a header of each answer: 25 milliseconds for one of our requests. So the extra step costs very little time.
Note that the proxy ran on the same machine as our test. If your proxy runs somewhere else, the trip to it adds time.
One API, three makers
This is what LiteLLM is for. Our test program sent every job in the same OpenAI format to the same address. Only the model name changed. All 18 jobs passed every secret check:
- Claude Haiku 5.5: about 16 seconds and 0.20 cents per job.
- GPT-6 Luna: about 33 seconds and 0.13 cents per job.
- Gemini 3.8 Flash: about 181 seconds and 10.9 cents per job.
Gemini took longest and cost the most, because it "thinks" a lot before it answers. LiteLLM itself made no difference to the answers.
Does the cost tracking add up?
LiteLLM keeps track of what every request costs, so you can set budgets per team or per key. We checked its numbers against OpenRouter's own record of each request.
They matched exactly in 18 of 18 jobs, to the last digit. For example, one Haiku job cost $0.0013339 according to LiteLLM, and OpenRouter billed $0.0013339. So for these three models, you can trust LiteLLM's cost numbers.
That may be different for a brand-new model LiteLLM does not know the price of yet. We did not test that.
Fallbacks: when a model fails
A fallback is a backup model. If the first one fails, LiteLLM tries the next one, so your app keeps working.
We tested it with a model name pointing to a model that does not exist. LiteLLM switched to Haiku 5.5 by itself, and the answer came back in 1.0 seconds. In the response, a header said x-litellm-attempted-fallbacks: 1, so you can see that the backup was used.
Is LiteLLM free?
The open-source version is free. LiteLLM's website lists it at "$0" and "Free forever", with more than 140 provider integrations, virtual keys, budgets, load balancing and guardrails.
There is also an Enterprise version for big teams, with things like single sign-on, audit logs and support. Its price is on request. You still pay the AI makers or a service like OpenRouter for the models you use. In our test, that was 0.20 cents per job with Haiku 5.5.
LiteLLM or OpenRouter? They do different things. OpenRouter is a service you pay that gives one key for many models. LiteLLM is software you run yourself, in front of any provider, including OpenRouter, as in our test. Many teams use them together. Tools like OpenCode can also connect straight to OpenRouter.
Questions people ask
What is LiteLLM used for?
To talk to many AI models through one API. Programs send every request in the same format, and LiteLLM passes it on to the right model, such as Claude, GPT or Gemini. As a proxy, it also keeps track of keys, budgets and costs, and can switch to a backup model when one fails.
Is LiteLLM free?
The open-source version is free: LiteLLM's website lists it at $0, "Free forever". There is also a paid Enterprise version for big teams, with the price on request. You still pay for the AI models you use; in our test that was about 0.2 cents per coding job with Claude Haiku 5.5.
Is LiteLLM any good?
In our small test, yes. It added no delay we could measure, all 18 coding jobs through it passed every secret check, its cost numbers matched OpenRouter's bill exactly, and the fallback to a backup model worked. We did not test it under heavy traffic.
Is LiteLLM expensive?
The software itself is free in its open-source form. What costs money is the AI models. In our test, one coding job cost about 0.1 to 11 cents, depending on the model. LiteLLM did not add anything to that bill.
How we checked
On 10 October 2026 we read LiteLLM's GitHub page, docs, security update and website. The screenshots are our own captures, cropped only.
The same day we ran LiteLLM 1.104.2 as a proxy in a Vercel Sandbox with Linux and Python 3.13, with all three models through our own OpenRouter key: 40 small requests, 18 coding jobs and one fallback test. The secret checks never went into the sandbox: we copied the answers back and checked them on our own PC. Costs are what OpenRouter recorded for each request. The sandbox was deleted afterwards.
This was a small test of the basic features. We did not test LiteLLM with heavy traffic, its enterprise features, or its other ways of running, such as Docker. The charts were made by us. None of the images was made by AI.

