Assessment & Survey Glossary

Plain definitions, worked examples and the mistakes that matter—each term is a starting point, with the deep dive one click away.

Measurement

Likert ScaleA Likert scale measures agreement or frequency with a statement using 4-7 ordered options. How to build one that yields reliable data.

Reliability

Cronbach's AlphaCronbach's alpha estimates how consistently a set of items measures the same construct. Targets, formula and when it misleads. Test-Retest ReliabilityAdminister the same assessment twice and correlate the scores: how stable your measurement really is, and what interval to use. Inter-Rater ReliabilityHow much your raters actually agree—agreement, kappa and ICC—plus how to fix disagreement at the rubric level. Cohen's KappaCohen's kappa corrects raw agreement for chance. The formula, the .61-.80 sweet spot, and the prevalence paradox.

Validity

Content ValidityContent validity is coverage, not statistics—how to prove your assessment samples the whole domain with expert review and CVR. Construct ValidityThe umbrella question of measurement—convergent, discriminant and factorial evidence, and what a validation story looks like. Criterion ValidityConcurrent and predictive validity: correlating your assessment with real outcomes, and the correlation numbers worth bragging about.

Foundations

PsychometricsWhat psychometrics is—measurement theory for tests and surveys—and the four questions it answers for any assessment.

Item quality

Item AnalysisUsing difficulty and discrimination to find weak items in any quiz or survey—and the thresholds that flag them. Difficulty & DiscriminationTwo item metrics that pull against each other—and how to keep both in the healthy band when designing a test.

Advanced

Factor AnalysisHow factor analysis reveals the hidden dimensions inside a questionnaire—and the sample sizes and pitfalls behind honest results.

Statistics

Normal DistributionWhy the bell curve matters for grading, benchmarking and norming—and when real assessment data refuses to obey it. z-ScoreThe z-score converts any raw score into SDs from the mean—comparable across scales, tests and dimensions. T-ScoreThe T-score rescales z into a mean-50, SD-10 ruler—no negatives, easy bands, and the standard in health and competency reporting. Percentile RankPercentile rank converts scores into position—"better than 80% of the group"—with the small-sample traps to avoid. Standard DeviationSD quantifies how far scores scatter from the mean—the backbone of curves, bands, reliability and benchmark width.

Scoring

Weighted ScoreWeighted scoring multiplies each item's result by its importance before summing—how to set weights defensibly and keep them explainable. Reverse ScoringReverse-worded items need their scale flipped before averaging—why we use them, and the attention-check trap they hide.

Design

RubricA rubric is a scoring map—criteria, performance levels and concrete anchors. How to write one two raters can't disagree with. Competency ModelA competency model turns "good leadership" into observable behaviors in levels—the spine of 360s, hiring scorecards and development plans.

Methods

360-Degree FeedbackMulti-rater feedback from manager, peers, reports, self and clients—design rules that keep it developmental, not political. Training Needs AnalysisTNA separates skill gaps from process gaps before you build anything—three levels, one question, and the data that ends the guessing. Kirkpatrick ModelReaction, Learning, Behavior, Results—the four-level ladder for proving training worked, and what each level actually requires. Formative AssessmentFormative assessment checks learning while there's still time to change it—quizzes, muddiest-point, mid-course pulses and the feedback loop. Summative AssessmentSummative assessment certifies what was learned at the end—exam design, scoring thresholds, certificates and the defensibility rules. Benchmark ReportHow to build an industry benchmark report from assessment data—percentiles, cohorts, the distribution chart that gets cited.

Metrics

Net Promoter Score (NPS)NPS subtracts detractors from promoters on a 0-10 recommend question—the math, the follow-up question that matters, and its limits. CSAT (Customer Satisfaction Score)CSAT measures satisfaction with one specific interaction on a 5-point scale—the question design, timing window and reporting that make it work. Customer Effort Score (CES)CES measures the effort a customer spent to get something done—why effort beats satisfaction at predicting loyalty, and where to ask it.