Using difficulty and discrimination to find weak items in any quiz or survey—and the thresholds that flag them.
Item analysis inspects each question after scoring: how hard was it (difficulty) and how well did it separate strong from weak performers (discrimination). Together they expose broken, ambiguous or mis-keyed items that aggregate scores hide.
p = proportion correct (difficulty); D = %top-group correct − %bottom-group correct; r(item-total) as correlation variant.
One exam item has p=.94 and D=.02—everyone passes it and it finds nothing. Flagged, rewritten, re-run.
After any scored assessment with 20+ respondents per item; before you reuse a question bank.
On attitude scales, low discrimination may mean the item is fine but the construct is narrow—don't auto-delete, review wording first.
D ≥ .20 usable, ≥ .30 good; below .10 the item is dead weight or actively harmful.
Top/bottom 27% grouping works from ~30 responses; item-total correlation needs more. Treat early runs as directional.
Turn the concept into a live assessment:Item Analysis
Browse templates →