📖 Glossary · Item quality

Item Analysis

Using difficulty and discrimination to find weak items in any quiz or survey—and the thresholds that flag them.

Home›Glossary›Item Analysis

Item analysis inspects each question after scoring: how hard was it (difficulty) and how well did it separate strong from weak performers (discrimination). Together they expose broken, ambiguous or mis-keyed items that aggregate scores hide.

🧮 Formula

p = proportion correct (difficulty); D = %top-group correct − %bottom-group correct; r(item-total) as correlation variant.

✏️ Worked example

One exam item has p=.94 and D=.02—everyone passes it and it finds nothing. Flagged, rewritten, re-run.

🎯 When to use it

After any scored assessment with 20+ respondents per item; before you reuse a question bank.

⚠️ Watch out

On attitude scales, low discrimination may mean the item is fine but the construct is narrow—don't auto-delete, review wording first.

Questions practitioners ask

What discrimination is acceptable?

D ≥ .20 usable, ≥ .30 good; below .10 the item is dead weight or actively harmful.

Small sample—can I still do it?

Top/bottom 27% grouping works from ~30 responses; item-total correlation needs more. Treat early runs as directional.

Turn the concept into a live assessment:Item Analysis

Browse templates →

Glossary

Likert Scale Cronbach's Alpha Test-Retest Reliability Inter-Rater Reliability ∞