Lesson 22 / 27
Serving Tuned Models: Merged Weights and Adapters
Choose between one merged model per task and shared base with swappable adapters.
Merge for simplicity, swap for scale
A LoRA adapter can be merged into the base weights, giving a normal model with no extra latency but one full copy per task. Alternatively keep one shared base in GPU memory and load small adapters per task or customer on demand; several serving stacks support this, which is why LoRA suits many specialised variants. Plan for GPU memory, cold-start time when an adapter is loaded, batching across adapters, and a fallback model. With a hosted service you may only see a model id, with the provider handling this; check its price, rate limits and retirement policy.
Serving choices
Trade-offs at a glance.
merged weights one standalone model per task simple, no adapter overhead more GPU memory per task
shared base + adapters one base + small adapters many variants cheaply adapter load / routing logic
hosted fine-tune provider runs it, you call an id least ops work price, limits, retirement set by providerKeep a fallback
Keep the untuned prompted model as a fallback route so a bad tuned release can be switched off in minutes.
Quick check: Why is LoRA convenient for many specialised variants?
- It deletes the base model
- It makes every model free
- It removes the need for GPUs
- Small adapters can be swapped over one shared base model
Answer
Small adapters can be swapped over one shared base model — Small per-task deltas over a shared base save memory.