# Serving Tuned Models: Merged Weights and Adapters — Fine-tuning vs Prompting

Source: https://www.skillbyai.com/en/fine-tuning/o-serving

> Choose between one merged model per task and shared base with swappable adapters.

## Merge for simplicity, swap for scale

A LoRA adapter can be **merged** into the base weights, giving a normal model with no extra latency but one full copy per task. Alternatively keep one **shared base** in GPU memory and load **small adapters** per task or customer on demand; several serving stacks support this, which is why LoRA suits many specialised variants. Plan for GPU memory, cold-start time when an adapter is loaded, batching across adapters, and a fallback model. With a hosted service you may only see a model id, with the provider handling this; check its price, rate limits and retirement policy.

## Serving choices

Trade-offs at a glance.

```text
merged weights        one standalone model per task     simple, no adapter overhead   more GPU memory per task
shared base + adapters one base + small adapters         many variants cheaply         adapter load / routing logic
hosted fine-tune       provider runs it, you call an id  least ops work              price, limits, retirement set by provider
```

## Keep a fallback

Keep the untuned prompted model as a fallback route so a bad tuned release can be switched off in minutes.

**Quiz:** Why is LoRA convenient for many specialised variants?

- [ ] It deletes the base model
- [ ] It makes every model free
- [ ] It removes the need for GPUs
- [x] Small adapters can be swapped over one shared base model

*Answer:* Small adapters can be swapped over one shared base model. Small per-task deltas over a shared base save memory.
