AI Tools

One gateway for every model: setting up LiteLLM for your team

Route OpenAI-compatible calls to local and hosted models through one endpoint, with keys and logs in one place.

On this page

Every AI feature starts with one model and one SDK. Six months later there are three providers, two local models, keys pasted into four services and no idea which feature spends what. An AI gateway puts one endpoint in front of all of them. This post sets up LiteLLM, an open-source gateway that speaks the OpenAI API, in front of a local Ollama model.

What a gateway gives you

Your applications talk to one URL with one client library. The gateway decides which model actually answers.

  • Swap models without code changes. Applications ask for a model name such as chat-default; you decide in config which real model that is.
  • One place for keys. Provider keys live in the gateway, not in every service.
  • Logs and spend in one place. Every call passes through the same door, so it can be recorded once.
your appLiteLLM proxyollama/llama3.2hosted model
Applications call the proxy with the OpenAI API; the proxy routes each model name to a real backend.

Install and configure LiteLLM

LiteLLM runs as a small Python server. Install the proxy extra and make sure Ollama is running with a model pulled, as in the RAG tutorial.

Terminal
pip install "litellm[proxy]"
ollama pull llama3.2

The config file maps the names your apps use to real models:

config.yamlYAML
model_list:
  - model_name: chat-default
    litellm_params:
      model: ollama/llama3.2
      api_base: http://localhost:11434

general_settings:
  master_key: sk-change-me
Warning

The master_key protects the proxy. Use a long random value, keep it out of git, and load it from an environment variable in anything beyond a laptop experiment.

Start the proxy and call it

Terminal
litellm --config config.yaml

By default the proxy listens on port 4000. Any OpenAI client can now use it by changing two settings, the base URL and the key:

client.pyPython
from openai import OpenAI

client = OpenAI(base_url="http://localhost:4000", api_key="sk-change-me")

reply = client.chat.completions.create(
    model="chat-default",
    messages=[{"role": "user", "content": "Say hello in one short sentence."}],
)
print(reply.choices[0].message.content)

To move chat-default to a hosted model later, add that provider's entry to model_list and restart the proxy. The application code stays the same.

Keys, budgets and logs

The master key is enough for one developer. For a team, LiteLLM can issue virtual keys per person or per service, with their own budgets and rate limits. That feature stores its data in Postgres, so it needs a DATABASE_URL in the proxy's environment; the LiteLLM documentation covers the setup.

Tip

Name models by job, not by vendor: chat-default, chat-cheap, embed-default. When a better model appears, you change one line of config instead of searching every repository for a model id.

Where to go next

With one gateway in place, comparing models becomes a config change. Try pointing chat-default at different local runners from the Ollama, LM Studio and llama.cpp comparison and keep whichever answers your real questions best.

Rahul Agarwal

Instructor, SkillByAI

Web developer for 15 years. Teaches developers and founders to build real products with open-source AI tools.

Get the next tutorial by email

One email when a new post lands, with the code ready to run. No spam.

By subscribing you agree to receive emails from SkillByAI. Unsubscribe anytime.