# The Quality-Cost Frontier — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/t-pareto

> Do not pay for quality you do not need.

## Dominated versus frontier options

Prompt variants differ in quality and cost (prompt length, number of examples, model size, number of calls). An option is **dominated** if another is at least as good and no more expensive. The remaining options form the **Pareto frontier**: each is the best quality available at its cost. Choose a point on the frontier that meets your quality bar at acceptable cost, and include latency if it matters. Recompute when prices or models change.

## Choose wisely, protect what works

Pick prompts on the quality-cost frontier, block regressions, and avoid overfitting to the score.

![Three ideas: Pareto frontier, regression gates, Goodhart's law.](assets/figures/prompt-quality-scoring/section-7-map.svg) — Figure 7.1 — Frontier, gates and Goodhart.

## Finding the frontier among six variants, run

I ran this with Python 3 (scipy 1.18.1 where imported) on example or seeded simulated data, not results from a real product. Of six example variants, five are on the frontier; few-shot with a medium model is dominated by few-shot with a small model, which is slightly better and half the price. Whether to pay $7.80 per thousand for 0.86 instead of $1.30 for 0.80 depends on the quality bar.

```python
variants = {  # name: (quality score, cost per 1k requests in $)
    "short prompt, small model": (0.71, 0.40), "long prompt, small model": (0.78, 0.90),
    "short prompt, large model": (0.83, 4.10), "long prompt, large model": (0.86, 7.80),
    "few-shot, small model": (0.80, 1.30), "few-shot, medium model": (0.79, 2.60)}
frontier = [n for n, (q, c) in variants.items()
            if not any(q2 >= q and c2 <= c and (q2, c2) != (q, c) for q2, c2 in variants.values())]
for n, (q, c) in sorted(variants.items(), key=lambda x: x[1][1]):
    print(f"{n:<27} quality {q:.2f}  ${c:.2f}/1k  {'frontier' if n in frontier else 'dominated'}")
```

Output:

```
short prompt, small model   quality 0.71  $0.40/1k  frontier
long prompt, small model    quality 0.78  $0.90/1k  frontier
few-shot, small model       quality 0.80  $1.30/1k  frontier
few-shot, medium model      quality 0.79  $2.60/1k  dominated
short prompt, large model   quality 0.83  $4.10/1k  frontier
long prompt, large model    quality 0.86  $7.80/1k  frontier
```

## Set the quality bar first

Decide the minimum acceptable score before looking at costs, then pick the cheapest frontier option above it.

**Quiz:** When is a prompt variant "dominated"?

- [ ] It has the highest quality
- [x] Another variant is at least as good and no more expensive
- [ ] It is the cheapest
- [ ] It was written first

*Answer:* Another variant is at least as good and no more expensive. Never choose a dominated option.
