One gateway for every model: setting up LiteLLM for your team
Route OpenAI-compatible calls to local and hosted models through one endpoint, with keys and logs in one place.
On this page
Every AI feature starts with one model and one SDK. Six months later there are three providers, two local models, keys pasted into four services and no idea which feature spends what. An AI gateway puts one endpoint in front of all of them. This post sets up LiteLLM, an open-source gateway that speaks the OpenAI API, in front of a local Ollama model.
What a gateway gives you
Your applications talk to one URL with one client library. The gateway decides which model actually answers.
- Swap models without code changes. Applications ask for a model name such as
chat-default; you decide in config which real model that is. - One place for keys. Provider keys live in the gateway, not in every service.
- Logs and spend in one place. Every call passes through the same door, so it can be recorded once.
Install and configure LiteLLM
LiteLLM runs as a small Python server. Install the proxy extra and make sure Ollama is running with a model pulled, as in the RAG tutorial.
pip install "litellm[proxy]"
ollama pull llama3.2The config file maps the names your apps use to real models:
model_list:
- model_name: chat-default
litellm_params:
model: ollama/llama3.2
api_base: http://localhost:11434
general_settings:
master_key: sk-change-meThe master_key protects the proxy. Use a long random value, keep it out of git, and load it from an environment variable in anything beyond a laptop experiment.
Start the proxy and call it
litellm --config config.yamlBy default the proxy listens on port 4000. Any OpenAI client can now use it by changing two settings, the base URL and the key:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:4000", api_key="sk-change-me")
reply = client.chat.completions.create(
model="chat-default",
messages=[{"role": "user", "content": "Say hello in one short sentence."}],
)
print(reply.choices[0].message.content)To move chat-default to a hosted model later, add that provider's entry to model_list and restart the proxy. The application code stays the same.
Keys, budgets and logs
The master key is enough for one developer. For a team, LiteLLM can issue virtual keys per person or per service, with their own budgets and rate limits. That feature stores its data in Postgres, so it needs a DATABASE_URL in the proxy's environment; the LiteLLM documentation covers the setup.
Name models by job, not by vendor: chat-default, chat-cheap, embed-default. When a better model appears, you change one line of config instead of searching every repository for a model id.
Where to go next
With one gateway in place, comparing models becomes a config change. Try pointing chat-default at different local runners from the Ollama, LM Studio and llama.cpp comparison and keep whichever answers your real questions best.