Prompt Quality Scoring

Measure prompts instead of guessing: prompt linting, format checks, weighted rubrics, calibrated LLM judges, bias checks, pairwise ratings, significance tests, cost-quality frontiers and regression gates, with every calculation run.

Start course →

What you'll learn

  • Explain what to score in an LLM feature: the prompt text, its outputs, and the whole system.
  • Lint prompts for missing structure, vague wording and conflicting instructions.
  • Score outputs with deterministic checks, weighted rubrics and must-pass criteria.
  • Calibrate LLM judges against humans and detect length and position bias.
  • Compare prompt versions with pairwise ratings, McNemar tests and repeated sampling.
  • Choose prompts on the quality-cost frontier and protect them with regression gates.

Syllabus

Why Score Prompts

  1. The Problem With Eyeballing Prompts
  2. What to Score: Prompt, Output, System
  3. A Scoring Pipeline

Static Prompt Checks

  1. A Prompt Structure Checklist
  2. Vague Wording and Conflicting Instructions
  3. Template Hygiene: Variables, Versions and Length

Deterministic Output Checks

  1. Format and Constraint Checks
  2. Reference-Based Metrics
  3. Assertion-Based Test Cases

Rubrics and Human Grading

  1. Designing a Rubric
  2. Weighted Scores and Must-Pass Criteria
  3. Running Human Grading Well

LLM Judges

  1. Designing a Judge Prompt
  2. Calibrating a Judge Against Humans
  3. Length Bias
  4. Position Bias in Pairwise Judging

Comparing Prompt Versions

  1. Pairwise Comparisons and Ratings
  2. Paired Significance Tests
  3. Run-to-Run Variance and Repeated Sampling

Trade-Offs, Regression Gates and Goodhart

  1. The Quality-Cost Frontier
  2. Regression Gates in CI
  3. Goodhart's Law: Overfitting to the Score

Putting Scoring Into Practice

  1. Building Test Sets and a Scoring Dashboard
  2. Case Study: Scoring a Ticket-Summary Prompt
  3. A Prompt Scoring Checklist