Administer the same assessment twice and correlate the scores: how stable your measurement really is, and what interval to use.
Test-retest reliability measures score stability: the same people take the same instrument twice, and the correlation between the two rounds is the answer. Correlations above .80 mean the instrument is not just capturing the mood of the day.
r = Pearson (or ICC) correlation between time-1 and time-2 scores of the same respondents.
A team health check re-run after 4 weeks correlates .86 with the first round—changes beyond ±10 points are real signal, not noise.
Use for anything positioned as a trait or state you will track over time: pre/post programs, quarterly checks, coaching outcomes.
Too short an interval and memory inflates stability; too long and real change masquerades as unreliability. Match the interval to the decision cadence.
.80+ is the common bar for group-level tracking; individual coaching decisions want .90+ or a different design.
Yes—alternate forms or a 2+ week gap reduce it. Note it in your report either way.
Turn the concept into a live assessment:Test-Retest Reliability
Browse templates →