Ollama vs LM Studio vs llama.cpp: which local runner should you use?
All three run open models on your own machine. Here is where each one fits, where it gets in your way, and how to choose in five minutes.
On this page
You want to run an open model on your own machine. Three names come up every time: Ollama, LM Studio and llama.cpp. They overlap a lot (all three can run models in the GGUF format and serve an OpenAI-compatible API), so the choice is less about what they can do and more about how you like to work.
The short answer
- Pick Ollama if you are a developer who wants a model running behind an API with one command.
- Pick LM Studio if you want a desktop app to browse, download and chat with models before you write any code.
- Pick llama.cpp if you need full control over how the model runs, or you are building the runner into something else.
Side by side
| Ollama | LM Studio | llama.cpp | |
|---|---|---|---|
| What it is | Command-line tool and background server | Desktop app with a built-in server | C/C++ library with command-line tools and a server |
| Licence | Open source (MIT) | Free to use, not open source | Open source (MIT) |
| Getting a model | ollama pull <name> from its library | Search and download inside the app | Download GGUF files yourself |
| API | Native API plus OpenAI-compatible endpoints | OpenAI-compatible local server | llama-server, OpenAI-compatible |
| Default port | 11434 | 1234 | 8080 |
| Best at | Scripting, Docker, team setups | Exploring models with a GUI | Tuning, embedding, unusual hardware |
Check each project's documentation before you rely on a detail in this table; all three move quickly.
Ollama
Ollama hides almost every decision behind sensible defaults: ollama pull llama3.2 downloads a model in a sensible quantisation, and the server is already running. It is the runner the RAG tutorial uses, because it keeps the focus on your code.
Where it gets in your way: when you want a specific quantisation or runtime flag that its defaults do not expose, you end up writing a Modelfile or reaching for llama.cpp directly.
LM Studio
LM Studio is the friendliest way to look at models: search, compare sizes, see whether a model fits in your memory, chat with it, then switch on the local server. On Apple Silicon Macs it can also run models in Apple's MLX format.
Where it gets in your way: it is a desktop application, so it does not fit servers, containers or CI. The app is also not open source, which matters if your team only runs open-source tools.
llama.cpp
llama.cpp is the engine underneath much of the local-model world. You choose the build options, the exact model file, the context size, how many layers go to the GPU and how the server behaves.
Where it gets in your way: you manage everything yourself, including downloading the right GGUF file and keeping builds up to date.
Put a gateway in front of whichever runner you choose. Your apps then call one URL, and switching runners later is a config change. The LiteLLM tutorial shows how.
How to choose in five minutes
- Will this run on a server or in Docker? Choose Ollama or llama.cpp.
- Do you mostly want to try models before writing code? Start with LM Studio.
- Do you need a setting that Ollama does not expose? Use llama.cpp for that model.
There is no wrong first choice. The model you run matters far more than the runner, and with an OpenAI-compatible API on all three, moving between them costs minutes, not days.