📖 Glossary · Reliability

Cohen's Kappa

Cohen's kappa corrects raw agreement for chance. The formula, the .61-.80 sweet spot, and the prevalence paradox.

Home›Glossary›Cohen's Kappa

Cohen's kappa is raw agreement minus the agreement two raters would reach by chance, divided by the maximum possible surplus. It answers: do my raters agree more than a coin flip would?

🧮 Formula

κ = (po − pe) / (1 − pe), po = observed agreement, pe = expected chance agreement.

✏️ Worked example

A screening quiz flags 90% of the same candidates as the expert panel, but panel labels are 85% one-sided—κ = .41 despite high raw agreement.

🎯 When to use it

Two raters, two categories (or ordered few), small-to-medium samples: pass/fail grading, risk flags, label annotation.

⚠️ Watch out

The paradox: with heavily skewed labels, kappa collapses while agreement soars. Report both, or switch to PABAK/Gwet's AC when skewed.

Questions practitioners ask

What kappa value is trustworthy?

.61-.80 is substantial, .81-1.0 almost perfect—below .60 with skewed data, check the base rate before blaming the raters.

Kappa vs percent agreement?

Percent agreement flatters; kappa is the conservative read. Always report n and the label distribution too.

Turn the concept into a live assessment:Cohen's Kappa

Browse templates →

Glossary

Likert Scale Cronbach's Alpha Test-Retest Reliability Inter-Rater Reliability ∞