Here's a scenario that plays out in every L&D department: you run a two-day leadership training. The feedback forms come back glowing—"great facilitator," "useful content," "would recommend." You compile the results into a dashboard, send it to leadership, and check the "training effectiveness" box.
Six months later, someone asks: "Did that training actually change anything?" You pull up the dashboard. Average satisfaction: 4.6 out of 5. And that's all you have.
This is the gap that Donald Kirkpatrick identified back in 1959 and that most organizations still haven't closed. We measure what's easy to measure (did people like it?) and call it "training effectiveness evaluation." But liking isn't learning, and learning isn't performance.
The Four Levels Aren't a Checklist
Kirkpatrick's model has four levels. Most people know them by name. Fewer understand the relationship between them.
The key insight that gets missed: these four levels are a causal chain, not a menu. Reaction is supposed to lead to learning. Learning is supposed to lead to behavior change. Behavior change is supposed to produce results. If the chain breaks at any link, the final result won't happen—no matter how much you spent on the training.
Where the Chain Breaks
Here's the uncomfortable truth I've seen across dozens of organizations: the chain breaks at Level 3 more often than anywhere else.
People attend the training (L1: they liked it). They pass the knowledge test (L2: they learned it). Then they go back to their desks, and... nothing changes. Same workflows, same habits, same output. L3 never happens, which means L4 is impossible.
Why does L3 break so often? Three reasons:
- The training taught knowledge, not behavior. Knowing "active listening means paraphrasing before responding" is different from actually doing it in a tense meeting. If the training didn't include practice with feedback, L2 doesn't naturally translate to L3.
- The environment doesn't support the new behavior. A manager learns coaching skills in training, goes back to a team that expects directives, and reverts to old habits within a week. Environment eats training for breakfast.
- Nobody checked. Without measurement, there's no accountability and no feedback loop. If you don't measure L3, you're essentially hoping behavior change happens by magic.
Practical takeaway: Before you design training, ask: "What specific behavior should change after this program?" If you can't name the behavior in observable terms, you can't measure L3—and without L3, L4 is a wish, not a result.
How Deep Should You Go?
A common mistake is treating all four levels as mandatory for every training program. They're not. The depth of evaluation should match the stakes of the training.
Low-stakes training (e.g., a one-hour software orientation): L1 is sufficient. You need to know if the content was clear and the format worked. Spending time on L2-L4 for a tool tutorial is overkill.
Medium-stakes training (e.g., a two-day communication skills workshop): L1 + L2. Check satisfaction and run a pre-post knowledge test. This gives you enough data to iterate on content without going overboard.
High-stakes training (e.g., a six-month leadership development program costing $50K+): All four levels. This is where L3 and L4 earn their keep. If you're investing serious money, you need to prove the investment produced observable change and business outcomes.
"The biggest waste in training evaluation isn't measuring too little—it's measuring the wrong level for the stakes. A $500 workshop doesn't need a 12-week behavior tracking study. But a $50,000 leadership program that only collects smile sheets? That's malpractice."
Practical Implementation
Here's how to actually set up the four levels in practice:
Level 1: Design the right questions
Stop using generic satisfaction scales ("How would you rate this training? 1-5"). They produce data that looks good in dashboards but means nothing. Instead, ask three targeted questions:
- "What is one thing from this training you plan to apply in the next two weeks?" (If they can't answer, the training didn't connect to their work.)
- "What part of this training felt least relevant to your actual job?" (This surfaces content gaps that average ratings hide.)
- "What would you tell a colleague who's considering attending this training?" (This is a better predictor of actual satisfaction than a 1-5 rating.)
Level 2: Use pre-post design
Test before training. Test after. The difference is your learning evidence. A standalone post-test tells you what they know—it doesn't tell you what the training caused. (See our dedicated pre-post assessment design article for the full methodology.)
Level 3: Plan the follow-up before training starts
The single most important L3 design decision: schedule the behavior assessment before the training happens. If you wait until after the training to figure out how you'll measure behavior change, you've already lost the baseline. Send a manager or peer assessment 4-6 weeks post-training, anchored to the specific behaviors the training targeted.
Level 4: Connect to existing business metrics
You don't need to create new metrics for L4. Use what the business already tracks—if the training targets sales skills, look at conversion rates; if it targets safety, look at incident reports; if it targets leadership, look at team engagement scores. The key is defining the connection before the training, not mining for correlations afterward.
The Phillips Fifth Level: ROI
Jack Phillips extended the model with a fifth level: ROI. This converts L4 business results into monetary value and compares them to training costs. It's powerful when done well—and misleading when done poorly. (Our training ROI calculation article covers this in depth.)
The temptation with Level 5 is to claim everything as a training benefit. Sales went up 15%? Training did it! But maybe the market recovered. Maybe a competitor exited. Maybe the sales team just got a new CRM tool. Without a comparison group, attributing business results to training is guessing with confidence intervals.
Start Where You Are
If your organization currently only does Level 1 (smile sheets), don't try to jump to Level 4 overnight. The infrastructure, stakeholder buy-in, and data literacy needed for L3-L4 take time to build.
A realistic progression: Get Level 1 right first (ask better questions, not just ratings). Then add Level 2 (pre-post tests). Then pilot Level 3 on one high-stakes program. Then expand. Rome wasn't built in a day, and neither was any organization's training measurement system.
🛠️ Build Your Evaluation System in FormLM
The hardest part of Kirkpatrick evaluation isn't the theory—it's the execution. FormLM covers the full chain:
- Scale fields: L1 satisfaction surveys and L2 knowledge tests in one platform
- Insight pre-post comparison: Same questionnaire sent twice, automatic delta calculation for L2 and L3
- Multi-role collection: L3 behavior assessments from managers and peers
- AI reports: Auto-generated evaluation reports ready for leadership
- Statistics overview: Average scores, standard deviations, and trend tracking across programs
✅ Key Takeaways
- The four levels are a causal chain, not a menu—if L3 breaks, L4 is impossible
- The chain breaks most often at Level 3: people learn but don't change behavior
- Match evaluation depth to stakes: L1 for low-stakes, all four for high-stakes programs
- Schedule L3 follow-up before training starts—you need the baseline
- L4 should connect to existing business metrics, not new ones invented post-hoc
- Progress incrementally: fix L1, add L2, pilot L3, then expand
