Lesson 22 / 27

Serving Tuned Models: Merged Weights and Adapters

Choose between one merged model per task and shared base with swappable adapters.

Merge for simplicity, swap for scale

A LoRA adapter can be merged into the base weights, giving a normal model with no extra latency but one full copy per task. Alternatively keep one shared base in GPU memory and load small adapters per task or customer on demand; several serving stacks support this, which is why LoRA suits many specialised variants. Plan for GPU memory, cold-start time when an adapter is loaded, batching across adapters, and a fallback model. With a hosted service you may only see a model id, with the provider handling this; check its price, rate limits and retirement policy.

Serving choices

Trade-offs at a glance.

merged weights        one standalone model per task     simple, no adapter overhead   more GPU memory per task
shared base + adapters one base + small adapters         many variants cheaply         adapter load / routing logic
hosted fine-tune       provider runs it, you call an id  least ops work              price, limits, retirement set by provider

Keep a fallback

Keep the untuned prompted model as a fallback route so a bad tuned release can be switched off in minutes.

Quick check: Why is LoRA convenient for many specialised variants?

  • It deletes the base model
  • It makes every model free
  • It removes the need for GPUs
  • Small adapters can be swapped over one shared base model
Answer

Small adapters can be swapped over one shared base model — Small per-task deltas over a shared base save memory.