Lesson 12 / 25

Running Human Grading Well

Consistent graders, blind comparisons, measured agreement.

Guidelines, blinding, agreement

Human grades are the reference for everything else, so make them reliable: written guidelines with examples, blind grading (graders do not know which prompt version produced an output), randomised order, a calibration session, and double-grading a sample to measure agreement between graders. Track grader time and fatigue; short sessions with breaks give better results. Use domain experts for specialised content and protect any personal data in the outputs.

Exam marking

Exam boards give markers a mark scheme, anonymise scripts, and have a second marker check a sample; prompt grading needs the same habits.

Double-grade at least 10%

Measure agreement between graders on a sample; low agreement means the rubric, not the prompt, needs work.

Quick check: Why grade outputs blind to the prompt version?

  • To make grading faster
  • To avoid graders favouring the version they expect to be better
  • Because versions have no names
  • To hide the rubric
Answer

To avoid graders favouring the version they expect to be better — Blinding removes expectation bias.