Comparisons

Ollama vs LM Studio vs llama.cpp: which local runner should you use?

All three run open models on your own machine. Here is where each one fits, where it gets in your way, and how to choose in five minutes.

On this page

You want to run an open model on your own machine. Three names come up every time: Ollama, LM Studio and llama.cpp. They overlap a lot (all three can run models in the GGUF format and serve an OpenAI-compatible API), so the choice is less about what they can do and more about how you like to work.

The short answer

  • Pick Ollama if you are a developer who wants a model running behind an API with one command.
  • Pick LM Studio if you want a desktop app to browse, download and chat with models before you write any code.
  • Pick llama.cpp if you need full control over how the model runs, or you are building the runner into something else.

Side by side

OllamaLM Studiollama.cpp
What it isCommand-line tool and background serverDesktop app with a built-in serverC/C++ library with command-line tools and a server
LicenceOpen source (MIT)Free to use, not open sourceOpen source (MIT)
Getting a modelollama pull <name> from its librarySearch and download inside the appDownload GGUF files yourself
APINative API plus OpenAI-compatible endpointsOpenAI-compatible local serverllama-server, OpenAI-compatible
Default port1143412348080
Best atScripting, Docker, team setupsExploring models with a GUITuning, embedding, unusual hardware

Check each project's documentation before you rely on a detail in this table; all three move quickly.

Ollama

Ollama hides almost every decision behind sensible defaults: ollama pull llama3.2 downloads a model in a sensible quantisation, and the server is already running. It is the runner the RAG tutorial uses, because it keeps the focus on your code.

Where it gets in your way: when you want a specific quantisation or runtime flag that its defaults do not expose, you end up writing a Modelfile or reaching for llama.cpp directly.

LM Studio

LM Studio is the friendliest way to look at models: search, compare sizes, see whether a model fits in your memory, chat with it, then switch on the local server. On Apple Silicon Macs it can also run models in Apple's MLX format.

Where it gets in your way: it is a desktop application, so it does not fit servers, containers or CI. The app is also not open source, which matters if your team only runs open-source tools.

llama.cpp

llama.cpp is the engine underneath much of the local-model world. You choose the build options, the exact model file, the context size, how many layers go to the GPU and how the server behaves.

Where it gets in your way: you manage everything yourself, including downloading the right GGUF file and keeping builds up to date.

Tip

Put a gateway in front of whichever runner you choose. Your apps then call one URL, and switching runners later is a config change. The LiteLLM tutorial shows how.

How to choose in five minutes

  1. Will this run on a server or in Docker? Choose Ollama or llama.cpp.
  2. Do you mostly want to try models before writing code? Start with LM Studio.
  3. Do you need a setting that Ollama does not expose? Use llama.cpp for that model.

There is no wrong first choice. The model you run matters far more than the runner, and with an OpenAI-compatible API on all three, moving between them costs minutes, not days.

Rahul Agarwal

Instructor, SkillByAI

Web developer for 15 years. Teaches developers and founders to build real products with open-source AI tools.

Get the next tutorial by email

One email when a new post lands, with the code ready to run. No spam.

By subscribing you agree to receive emails from SkillByAI. Unsubscribe anytime.