How much your raters actually agree—agreement, kappa and ICC—plus how to fix disagreement at the rubric level.
Inter-rater reliability quantifies agreement between independent scorers of the same work. Simple agreement % is the entry level; chance-corrected measures (kappa, ICC) are the honest ones, because two raters can agree by drift alone.
% agreement; Cohen's κ = (po − pe)/(1 − pe); ICC for continuous scores.
Two managers score the same 20 interview notes: 85% identical ratings, κ = .62—good agreement, but the borderline band still needs a calibration session.
Mandatory anywhere humans score subjective work: interview panels, 360 programs, exam grading, rubric assessment.
Low reliability is usually a rubric bug, not a rater bug—fix anchors and examples before blaming the people.
Not by itself—high agreement can be pure chance when categories are skewed. Report a chance-corrected statistic alongside.
ICC for 1-5 numeric ratings; weighted kappa when the categories are ordered but few.
Turn the concept into a live assessment:Inter-Rater Reliability
Browse templates →