Open Source / News

LiteLLM tested: one API for every AI model, does it hold up?

We ran LiteLLM's proxy with three makers' models on our secret coding test. No measurable delay, 18 of 18 jobs perfect, and its cost tracking matched OpenRouter's bill 18 of 18 times.

Summary of our LiteLLM test: 1010 milliseconds direct vs 984 milliseconds through LiteLLM; 18 of 18 coding jobs passed every check; LiteLLM's cost matched OpenRouter's bill in 18 of 18 jobs; the fallback worked
Our own test in a Linux sandbox, 10 October 2026. Chart by Not an AI App.

LiteLLM is one of the most used open-source tools for working with AI models: about 60,900 stars on GitHub, and 22,000 searches a month in the US. It promises one way to talk to more than 100 models. We installed it in a sandbox, a separate test computer we deleted afterwards, and checked three things: does it slow things down, does it work with models from different makers, and does its cost tracking add up? It added no delay we could measure, all 18 coding jobs passed every secret check, and its cost was exactly right 18 out of 18 times.

What is LiteLLM?

Screenshot of LiteLLM's docs: LiteLLM is an open-source library that gives you a single, unified interface to call 100+ LLMs (OpenAI, Anthropic, Vertex AI, Bedrock, and more) using the OpenAI format
LiteLLM's documentation. Screenshot from 10 October 2026, cropped.

An API is the connection a program uses to talk to an AI model. Every maker has its own, with its own small differences. LiteLLM puts one layer in front of all of them. Its docs say it "gives you a single, unified interface to call 100+ LLMs" in the format OpenAI uses.

You can use it in two ways:

  • As a library: a piece of Python code inside your own program.
  • As a proxy, or gateway: a small server that your apps talk to. It passes requests on to the right model and keeps track of keys, budgets and costs.

We tested the proxy, because that is what most teams use it for.

Who makes it?

Facts about LiteLLM: made by BerriAI, MIT licence except an enterprise folder, about 60,900 stars on GitHub, started July 2023, we used version 1.104.2
From GitHub, the LICENSE file and PyPI, 10 October 2026. Image by Not an AI App.

LiteLLM is made by a company called BerriAI. The repository started in July 2023 and had about 60,900 stars on GitHub on 10 October 2026.

Most of it is open source under the MIT licence, which means you may use and change it for free. One folder, for enterprise features, has its own licence. We used version 1.104.2, the newest on PyPI, Python's package store, that day.

Is LiteLLM safe? The March 2026 hack

Screenshot of LiteLLM's security update: the compromised PyPI packages were litellm 1.82.7 and 1.82.8, live on March 24, 2026 from 10:39 UTC for about 40 minutes before being quarantined by PyPI
LiteLLM's own security update of March 2026. Screenshot from 10 October 2026, cropped.

In March 2026, LiteLLM was hit by a supply chain attack: someone published hacked versions of the real package. LiteLLM's own security update says: "The compromised PyPI packages were litellm==1.82.7 and litellm==1.82.8." They were online on 24 March 2026 for about 40 minutes before PyPI blocked them.

The hidden code tried to steal passwords and keys from the computer it ran on. LiteLLM says the attack came in through a security scanner used in its own build process. People who used the official LiteLLM Docker image were not affected, LiteLLM says.

What this means for you: never install 1.82.7 or 1.82.8, and pin the exact version you use. If you did install one of them, LiteLLM advises treating every password and key on that computer as stolen and changing them. We installed a much later version, in a throwaway sandbox.

Our test

Diagram of our test: 20 small requests direct and 20 through LiteLLM; three coding jobs, two tries, three makers' models through one API; a broken model with a backup
Diagram by Not an AI App.

We ran LiteLLM's proxy in a fresh Linux sandbox and connected it to three models from three makers, all through OpenRouter, a service that sells access to many models:

  • Claude Haiku 5.5 (Anthropic)
  • GPT-6 Luna (OpenAI)
  • Gemini 3.8 Flash (Google)

Then we did three tests:

1. Speed: 20 small requests straight to OpenRouter and 20 through LiteLLM, taking turns. 2. Real work: the three jobs from our private coding test, two tries each, for all three models, all through LiteLLM. The jobs are the same as in our Haiku 5.5 test; see our methods page for how we test. 3. Fallback: a model name that points to a model that does not exist, with Haiku 5.5 as backup.

Does LiteLLM slow things down?

Bar chart: the middle time for one small request was 1010 milliseconds direct to OpenRouter and 984 milliseconds through the LiteLLM proxy
Our own test, 10 October 2026, 20 requests each. Chart by Not an AI App.

No, not that we could measure. For a tiny request, the middle time was:

  • Straight to OpenRouter: 1010 milliseconds.
  • Through LiteLLM: 984 milliseconds.

That small difference is less than the normal ups and downs of the internet. LiteLLM also reports its own work time in a header of each answer: 25 milliseconds for one of our requests. So the extra step costs very little time.

Note that the proxy ran on the same machine as our test. If your proxy runs somewhere else, the trip to it adds time.

One API, three makers

Table: through LiteLLM, Claude Haiku 5.5, GPT-6 Luna and Gemini 3.8 Flash all passed every secret check; 0.20 cents, 0.13 cents and 10.9 cents per job
Our own test, 10 October 2026, two tries per job. Chart by Not an AI App.

This is what LiteLLM is for. Our test program sent every job in the same OpenAI format to the same address. Only the model name changed. All 18 jobs passed every secret check:

  • Claude Haiku 5.5: about 16 seconds and 0.20 cents per job.
  • GPT-6 Luna: about 33 seconds and 0.13 cents per job.
  • Gemini 3.8 Flash: about 181 seconds and 10.9 cents per job.

Gemini took longest and cost the most, because it "thinks" a lot before it answers. LiteLLM itself made no difference to the answers.

Does the cost tracking add up?

Table of cost per job as LiteLLM reported it and as OpenRouter billed it, identical in 18 of 18 jobs
Our own test, 10 October 2026. Chart by Not an AI App.

LiteLLM keeps track of what every request costs, so you can set budgets per team or per key. We checked its numbers against OpenRouter's own record of each request.

They matched exactly in 18 of 18 jobs, to the last digit. For example, one Haiku job cost $0.0013339 according to LiteLLM, and OpenRouter billed $0.0013339. So for these three models, you can trust LiteLLM's cost numbers.

That may be different for a brand-new model LiteLLM does not know the price of yet. We did not test that.

Fallbacks: when a model fails

The fallback in our test: we asked for a broken model; Claude Haiku 5.5 answered OK in about one second, and LiteLLM reported one attempted fallback
Response and headers from our test, 10 October 2026. Image by Not an AI App.

A fallback is a backup model. If the first one fails, LiteLLM tries the next one, so your app keeps working.

We tested it with a model name pointing to a model that does not exist. LiteLLM switched to Haiku 5.5 by itself, and the answer came back in 1.0 seconds. In the response, a header said x-litellm-attempted-fallbacks: 1, so you can see that the backup was used.

Is LiteLLM free?

Screenshot of LiteLLM's website: open source $0, free forever, with 140+ LLM provider integrations, virtual keys, budgets, load balancing and guardrails; Enterprise with support, SSO and audit logs, price on request
LiteLLM's website. Screenshot from 10 October 2026, cropped.

The open-source version is free. LiteLLM's website lists it at "$0" and "Free forever", with more than 140 provider integrations, virtual keys, budgets, load balancing and guardrails.

There is also an Enterprise version for big teams, with things like single sign-on, audit logs and support. Its price is on request. You still pay the AI makers or a service like OpenRouter for the models you use. In our test, that was 0.20 cents per job with Haiku 5.5.

LiteLLM or OpenRouter? They do different things. OpenRouter is a service you pay that gives one key for many models. LiteLLM is software you run yourself, in front of any provider, including OpenRouter, as in our test. Many teams use them together. Tools like OpenCode can also connect straight to OpenRouter.

Questions people ask

List of questions people ask Google about LiteLLM, with the 4 questions we answer marked
"People also ask" questions from Google (US), 10 October 2026, via DataForSEO. Chart by Not an AI App.

What is LiteLLM used for?

To talk to many AI models through one API. Programs send every request in the same format, and LiteLLM passes it on to the right model, such as Claude, GPT or Gemini. As a proxy, it also keeps track of keys, budgets and costs, and can switch to a backup model when one fails.

Is LiteLLM free?

The open-source version is free: LiteLLM's website lists it at $0, "Free forever". There is also a paid Enterprise version for big teams, with the price on request. You still pay for the AI models you use; in our test that was about 0.2 cents per coding job with Claude Haiku 5.5.

Is LiteLLM any good?

In our small test, yes. It added no delay we could measure, all 18 coding jobs through it passed every secret check, its cost numbers matched OpenRouter's bill exactly, and the fallback to a backup model worked. We did not test it under heavy traffic.

Is LiteLLM expensive?

The software itself is free in its open-source form. What costs money is the AI models. In our test, one coding job cost about 0.1 to 11 cents, depending on the model. LiteLLM did not add anything to that bill.

How we checked

On 10 October 2026 we read LiteLLM's GitHub page, docs, security update and website. The screenshots are our own captures, cropped only.

The same day we ran LiteLLM 1.104.2 as a proxy in a Vercel Sandbox with Linux and Python 3.13, with all three models through our own OpenRouter key: 40 small requests, 18 coding jobs and one fallback test. The secret checks never went into the sandbox: we copied the answers back and checked them on our own PC. Costs are what OpenRouter recorded for each request. The sandbox was deleted afterwards.

This was a small test of the basic features. We did not test LiteLLM with heavy traffic, its enterprise features, or its other ways of running, such as Docker. The charts were made by us. None of the images was made by AI.