# Running Human Grading Well — Prompt Quality Scoring

Source: https://www.skillbyai.com/en/prompt-quality-scoring/r-human

> Consistent graders, blind comparisons, measured agreement.

## Guidelines, blinding, agreement

Human grades are the reference for everything else, so make them reliable: written guidelines with examples, **blind** grading (graders do not know which prompt version produced an output), randomised order, a calibration session, and **double-grading** a sample to measure agreement between graders. Track grader time and fatigue; short sessions with breaks give better results. Use domain experts for specialised content and protect any personal data in the outputs.

## Exam marking

Exam boards give markers a mark scheme, anonymise scripts, and have a second marker check a sample; prompt grading needs the same habits.

## Double-grade at least 10%

Measure agreement between graders on a sample; low agreement means the rubric, not the prompt, needs work.

**Quiz:** Why grade outputs blind to the prompt version?

- [ ] To make grading faster
- [x] To avoid graders favouring the version they expect to be better
- [ ] Because versions have no names
- [ ] To hide the rubric

*Answer:* To avoid graders favouring the version they expect to be better. Blinding removes expectation bias.
