Tutorials

Build a RAG chatbot on your own laptop with Ollama and pgvector

Index your own docs, store embeddings in Postgres, and answer questions with a local model. No API keys, no cloud bill, about an hour of work.

Retrieval-augmented generation: the model answers from your documents, not from memory.
On this page

Most "chat with your docs" tutorials start with an API key and end with a monthly bill. This one builds a RAG chatbot that runs entirely on your machine. By the end you will have two small Python scripts that answer questions about any folder of Markdown files, using only open-source tools: Ollama for the models and Postgres with pgvector for search.

Why run RAG locally?

A language model only knows what it was trained on. Retrieval-augmented generation (RAG) fixes that by finding the most relevant pieces of your documents and handing them to the model together with the question. Running it locally means your documents never leave your laptop, and you can experiment as much as you like without watching a usage meter.

If you can explain the answer from three paragraphs of your own docs, a small local model is usually enough.

Note

You need about 8 GB of free RAM for the chat model used here. On a smaller machine, swap in a lighter model; the code does not change.

What you'll build

docs/*.mdchunkspgvectortop 4 matchesllama3.2answer
Figure 1. Documents are split, embedded and stored once; each question retrieves the closest chunks before the model answers.
  • An ingest script that splits Markdown files into chunks and stores an embedding for each chunk.
  • A Postgres table with the vector type, so similarity search is one SQL query.
  • An ask script that retrieves context and calls a local model.

An embedding is a list of numbers that represents the meaning of a piece of text. Texts with similar meaning get similar numbers, which is what lets a database find "the paragraphs closest to this question" without any keyword matching.

Set up Ollama and Postgres

  1. Install Ollama and check that ollama --version prints a version.
  2. Pull a chat model and an embedding model.
  3. Start Postgres with the pgvector extension in Docker, then install the Python packages.
Terminal
ollama pull llama3.2
pulling manifest
success
ollama pull nomic-embed-text
success
docker run -d --name pgvector -e POSTGRES_PASSWORD=dev -p 5432:5432 pgvector/pgvector:pg16
pip install ollama "psycopg[binary]" pgvector numpy
Tip

The copy button on a terminal block copies only the commands, never the output lines, so you can paste straight into your terminal.

If port 5432 is already taken by another Postgres on your machine, map a different host port, for example -p 5433:5432, and change the connection string below to match.

Store embeddings in pgvector

Create the table

The nomic-embed-text model returns 768 numbers per chunk, so the column is vector(768). Connect with any SQL client (or docker exec -it pgvector psql -U postgres) and run:

schema.sqlSQL
CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE docs (
  id        bigserial PRIMARY KEY,
  source    text NOT NULL,
  body      text NOT NULL,
  embedding vector(768)
);

The source column keeps the file name next to each chunk, so later you can show where an answer came from.

Chunk and embed your docs

Split each file into overlapping chunks of about 800 characters, embed each chunk, and insert it. The overlap means a sentence cut at a chunk boundary still appears whole in one of the two neighbouring chunks.

ingest.pyPython
from pathlib import Path

import numpy as np
import ollama
import psycopg
from pgvector.psycopg import register_vector

conn = psycopg.connect("postgresql://postgres:dev@localhost:5432/postgres", autocommit=True)
register_vector(conn)


def embed(text: str) -> list[float]:
    return ollama.embed(model="nomic-embed-text", input=text)["embeddings"][0]


def chunk(text: str, size: int = 800, overlap: int = 120):
    for start in range(0, len(text), size - overlap):
        yield text[start:start + size]


# one row per chunk
for path in Path("docs").glob("*.md"):
    for piece in chunk(path.read_text()):
        conn.execute(
            "INSERT INTO docs (source, body, embedding) VALUES (%s, %s, %s)",
            (path.name, piece, np.array(embed(piece))),
        )

Put a few Markdown files in a docs folder next to the script and run python ingest.py. Re-running it inserts the chunks again, so empty the table first with TRUNCATE docs; when you re-ingest.

Warning

If you switch embedding models later, the vector size changes. Re-create the table and re-ingest everything; one column cannot mix sizes, and vectors from different models are not comparable anyway.

Which embedding model should you pick? For most docs, start with the default and only change it if answers miss obvious matches.

ModelDimensionsGood for
nomic-embed-text768A solid default for long technical docs
mxbai-embed-large1024Higher retrieval quality, slower to ingest
all-minilm384Tiny and fast, good for quick prototypes

Ask questions over your docs

Embed the question the same way, fetch the four closest chunks with the <=> cosine distance operator, and pass them to the model as context.

ask.pyPython
import sys

import numpy as np
import ollama
import psycopg
from pgvector.psycopg import register_vector

conn = psycopg.connect("postgresql://postgres:dev@localhost:5432/postgres")
register_vector(conn)

question = sys.argv[1] if len(sys.argv) > 1 else "How do I rotate the API keys?"
query = np.array(ollama.embed(model="nomic-embed-text", input=question)["embeddings"][0])

rows = conn.execute(
    "SELECT source, body FROM docs ORDER BY embedding <=> %s LIMIT 4",
    (query,),
).fetchall()
context = "\n\n".join(f"[{source}]\n{body}" for source, body in rows)

reply = ollama.chat(model="llama3.2", messages=[
    {"role": "system", "content": f"Answer only from this context. If the answer is not there, say so.\n\n{context}"},
    {"role": "user", "content": question},
])
print(reply["message"]["content"])
print("\nSources:", ", ".join(sorted({source for source, _ in rows})))

Run it with a question in quotes, for example python ask.py "How do I reset my password?". The answer is printed first, followed by the files the context came from.

"Answer only from this context" does more for accuracy than any model upgrade.

The one-line prompt that keeps answers grounded

When answers are wrong

Most bad answers come from retrieval, not from the model. Before you try a bigger model, print rows and read what was actually retrieved:

  • The right chunk is missing. Try smaller chunks, a larger LIMIT, or a stronger embedding model from the table above.
  • The right chunk is there but cut in half. Increase the overlap.
  • The model ignores the context. Keep the system prompt short and explicit, and put the context in the system message as shown.
Covered in the course

Turning this script into a chatbot with a web UI is the focus of Hour 11 of Learning Development with AI, and serving it to a team comes in Hour 12.

Where to go next

You now have the core of every RAG system: ingest, retrieve, generate. From here, the biggest wins come from better chunking (split on headings instead of character counts), showing sources next to each answer, and putting one gateway in front of all your models so you can swap them freely. The LiteLLM gateway tutorial covers that last step, and the comparison of local runners helps if Ollama is not the right fit for your machine.

Rahul Agarwal

Instructor, SkillByAI

Web developer for 15 years. Teaches developers and founders to build real products with open-source AI tools.

Get the next tutorial by email

One email when a new post lands, with the code ready to run. No spam.

By subscribing you agree to receive emails from SkillByAI. Unsubscribe anytime.