Prompt Quality Scoring
Measure prompts instead of guessing: prompt linting, format checks, weighted rubrics, calibrated LLM judges, bias checks, pairwise ratings, significance tests, cost-quality frontiers and regression gates, with every calculation run.
What you'll learn
- Explain what to score in an LLM feature: the prompt text, its outputs, and the whole system.
- Lint prompts for missing structure, vague wording and conflicting instructions.
- Score outputs with deterministic checks, weighted rubrics and must-pass criteria.
- Calibrate LLM judges against humans and detect length and position bias.
- Compare prompt versions with pairwise ratings, McNemar tests and repeated sampling.
- Choose prompts on the quality-cost frontier and protect them with regression gates.
Syllabus
Why Score Prompts
Static Prompt Checks
- A Prompt Structure Checklist
- Vague Wording and Conflicting Instructions
- Template Hygiene: Variables, Versions and Length
Deterministic Output Checks
Rubrics and Human Grading
LLM Judges
- Designing a Judge Prompt
- Calibrating a Judge Against Humans
- Length Bias
- Position Bias in Pairwise Judging