Cohen's kappa corrects raw agreement for chance. The formula, the .61-.80 sweet spot, and the prevalence paradox.
Cohen's kappa is raw agreement minus the agreement two raters would reach by chance, divided by the maximum possible surplus. It answers: do my raters agree more than a coin flip would?
κ = (po − pe) / (1 − pe), po = observed agreement, pe = expected chance agreement.
A screening quiz flags 90% of the same candidates as the expert panel, but panel labels are 85% one-sided—κ = .41 despite high raw agreement.
Two raters, two categories (or ordered few), small-to-medium samples: pass/fail grading, risk flags, label annotation.
The paradox: with heavily skewed labels, kappa collapses while agreement soars. Report both, or switch to PABAK/Gwet's AC when skewed.
.61-.80 is substantial, .81-1.0 almost perfect—below .60 with skewed data, check the base rate before blaming the raters.
Percent agreement flatters; kappa is the conservative read. Always report n and the label distribution too.
Turn the concept into a live assessment:Cohen's Kappa
Browse templates →