📖 Glossary · Item quality

Difficulty & Discrimination

Two item metrics that pull against each other—and how to keep both in the healthy band when designing a test.

Home›Glossary›Difficulty & Discrimination

Difficulty (p) is the share of respondents who answer correctly; discrimination (D) is how well the item splits high from low performers. The sweet spot: moderate difficulty with high discrimination—items at the extremes of p can't discriminate much whatever their quality.

🧮 Formula

p = correct/total; D = p(top 27%) − p(bottom 27%). Healthy: p .30-.80, D ≥ .30.

✏️ Worked example

A certification aims for p≈.60 on core items; take-home items at p=.95 move to practice quizzes instead of the exam.

🎯 When to use it

Designing any scored test—balance the item mix across difficulty bands before launch, then verify with real data after.

⚠️ Watch out

Don't chase difficulty for its own sake: a hard item that doesn't discriminate is just confusing.

Questions practitioners ask

Is a hard test a good test?

No—a good test maximizes information, which lives at moderate difficulty. Piling on hard items mostly adds noise.

What about all-or-nothing scoring?

Partial credit raises usable difficulty range; keep rubric bands documented for reliability.

Turn the concept into a live assessment:Difficulty & Discrimination

Browse templates →

Glossary

Likert Scale Cronbach's Alpha Test-Retest Reliability Inter-Rater Reliability ∞