📖 Glossary · Reliability

Test-Retest Reliability

Administer the same assessment twice and correlate the scores: how stable your measurement really is, and what interval to use.

Home›Glossary›Test-Retest Reliability

Test-retest reliability measures score stability: the same people take the same instrument twice, and the correlation between the two rounds is the answer. Correlations above .80 mean the instrument is not just capturing the mood of the day.

🧮 Formula

r = Pearson (or ICC) correlation between time-1 and time-2 scores of the same respondents.

✏️ Worked example

A team health check re-run after 4 weeks correlates .86 with the first round—changes beyond ±10 points are real signal, not noise.

🎯 When to use it

Use for anything positioned as a trait or state you will track over time: pre/post programs, quarterly checks, coaching outcomes.

⚠️ Watch out

Too short an interval and memory inflates stability; too long and real change masquerades as unreliability. Match the interval to the decision cadence.

Questions practitioners ask

What correlation counts as good?

.80+ is the common bar for group-level tracking; individual coaching decisions want .90+ or a different design.

Does practice effect matter?

Yes—alternate forms or a 2+ week gap reduce it. Note it in your report either way.

Turn the concept into a live assessment:Test-Retest Reliability

Browse templates →

Glossary

Likert Scale Cronbach's Alpha Inter-Rater Reliability Cohen's Kappa ∞