Lesson 12 / 25
Running Human Grading Well
Consistent graders, blind comparisons, measured agreement.
Guidelines, blinding, agreement
Human grades are the reference for everything else, so make them reliable: written guidelines with examples, blind grading (graders do not know which prompt version produced an output), randomised order, a calibration session, and double-grading a sample to measure agreement between graders. Track grader time and fatigue; short sessions with breaks give better results. Use domain experts for specialised content and protect any personal data in the outputs.
Exam marking
Exam boards give markers a mark scheme, anonymise scripts, and have a second marker check a sample; prompt grading needs the same habits.
Double-grade at least 10%
Measure agreement between graders on a sample; low agreement means the rubric, not the prompt, needs work.
Quick check: Why grade outputs blind to the prompt version?
- To make grading faster
- To avoid graders favouring the version they expect to be better
- Because versions have no names
- To hide the rubric
Answer
To avoid graders favouring the version they expect to be better — Blinding removes expectation bias.