Two item metrics that pull against each other—and how to keep both in the healthy band when designing a test.
Difficulty (p) is the share of respondents who answer correctly; discrimination (D) is how well the item splits high from low performers. The sweet spot: moderate difficulty with high discrimination—items at the extremes of p can't discriminate much whatever their quality.
p = correct/total; D = p(top 27%) − p(bottom 27%). Healthy: p .30-.80, D ≥ .30.
A certification aims for p≈.60 on core items; take-home items at p=.95 move to practice quizzes instead of the exam.
Designing any scored test—balance the item mix across difficulty bands before launch, then verify with real data after.
Don't chase difficulty for its own sake: a hard item that doesn't discriminate is just confusing.
No—a good test maximizes information, which lives at moderate difficulty. Piling on hard items mostly adds noise.
Partial credit raises usable difficulty range; keep rubric bands documented for reliability.
Turn the concept into a live assessment:Difficulty & Discrimination
Browse templates →