You run a sales training program. Post-training sales increase 12%. You report: "Training had a significant effect."

Leadership asks: "Marketing ran a promotion during the same period. How do you know the 12% is from training and not the promotion?"

You don't have an answer.

This is the thorniest problem in training evaluation: attribution. Scores went up, business metrics improved—but was it the training, or other factors? Comparison group design is the method that solves this problem.

The Logic of Comparison Group Design

The logic is simple: take two groups. One attends training (treatment), one doesn't (control). Measure both groups' metrics before and after training. The treatment group's change minus the control group's change equals the net training effect.

It works because it controls for "time passing" and "other interventions." If the control group's sales also increased 8% during the same period, and the treatment group increased 12%, the training's net contribution is 4 percentage points—not 12.

Timeline
Treatment Attends training
Control No training
Pre-training
Pre-test A₁
Pre-test B₁
↓ Interval
Training delivered
Business as usual
Post-training
Post-test A₂
Post-test B₂
Effect
Net Training Effect = (A₂−A₁) − (B₂−B₁)

This formula is called "Difference-in-Differences" (DiD). The treatment group's change (A₂−A₁) includes training effect + time effect + other intervention effects. The control group's change (B₂−B₁) includes time effect + other intervention effects but not training. Subtract one from the other, and the training effect is isolated.

How to Select a Control Group

The formula isn't the hard part—the control group's quality is. If the treatment and control groups were already different before training, the comparison is meaningless.

Three principles for selecting a control group:

1. Baseline match. The two groups should have roughly equal key metrics before training. You can't use a top-performing team as the treatment group and a newly-formed team as the control—their pre-training sales ability is already different, so post-training differences can't be attributed to training.

2. Similar environment. Both groups should face comparable market conditions, management practices, and resource levels. If the treatment group is in Region A and the control group is in Region B, and the two regions have very different market dynamics, the comparison is contaminated.

3. No cross-contamination. Treatment group participants shouldn't "infect" the control group by sharing training materials. If treatment participants pass their training notes to control group colleagues, the control group also learns, and the comparison breaks down.

Practical tip: The cleanest control group is a parallel team in a different region or department. For example, train the East Coast sales team and use the West Coast team as the control—similar business structure but no personnel overlap, low contamination risk. If you can't find a perfect control group, "imperfect but usable" still beats "no control group at all." Statistical methods like propensity score matching can partially correct for baseline differences.

When to Use Comparison Group Design

Comparison group design isn't for every training. It has costs—finding a control group, managing dual data collection, running more complex analysis. It's worth the effort when:

Not suitable for: onboarding programs (no viable control group), small-scale training (sample too small for statistical significance), pure knowledge training (L2 pre-post is sufficient—no need for L4-level comparison groups).

Practical Challenges and Solutions

Challenge 1: Business units won't let you hold back training.

"Why do they get training and we don't?" is the most common objection. Solution: tell the business unit that the control group gets priority enrollment in the next training cohort. It's not "you don't get training"—it's "you serve as the baseline, and you're first in line next round."

Challenge 2: The control group self-educates.

Treatment participants share training materials, control group members study on their own—both cause contamination. Solution: set clear expectations with the treatment group that materials aren't to be shared externally, and keep the evaluation window short (4-8 weeks) to minimize the influence of natural learning.

Challenge 3: Sample size is too small.

Fewer than 15 people per group means insufficient statistical power—even real differences may not reach significance. Solution: if you truly can't find enough control participants, relax matching criteria and use statistical correction. Or pool data from multiple training cohorts.

Comparison group design has a bonus benefit: if the control group's pre-training metrics matched the treatment group's, but the gap widened post-training—that divergence chart is the most powerful "proof of training effect" you can show leadership. A line chart with two groups' pre-post comparison is more persuasive than any satisfaction score. When leadership sees "the group that got training clearly outperformed the group that didn't," they don't need you to explain Kirkpatrick levels. The chart speaks for itself.

When You Can't Build a Perfect Comparison Group

In reality, a perfect randomized controlled trial is rarely achievable. But a quasi-experiment is always better than no comparison at all. Two fallback methods:

Historical comparison: Use data from a similar team during the same period last year as the baseline. The前提 is that this year's business environment is roughly similar to last year's. Advantage: no need to designate a control group. Disadvantage: can't control for environmental differences.

Self-comparison: Compare a person's training-related metric change against a non-training-related metric change in the same person. For example, after sales training, a rep's sales revenue increased 12%—but their client visit count didn't change. This suggests the revenue increase wasn't from working harder, but from better technique (i.e., training effect).

These methods aren't as rigorous as a true comparison group, but to leadership they're far more credible than "we have post-training data with no baseline."

🛠️ Support Comparison Group Evaluation in FormLM

Comparison group data collection can be implemented with FormLM:

  • Use scale fields to send pre and post assessments to both treatment and control groups, with data automatically categorized by group
  • Use data export to pull Excel data for difference-in-differences analysis in statistical software
  • Use Insight pre-post comparison to view change trends for treatment and control groups separately
  • Use AI reports to generate assessment reports for each group and compare the differences
Build your comparison group evaluation →

✅ Key Takeaways

  • Core logic: treatment change − control change = net training effect (Difference-in-Differences)
  • Three control group selection principles: baseline match, similar environment, no cross-contamination
  • Not for every program—use for high-investment, KPI-aligned, or pilot programs needing validation
  • Business unit resistance can be managed with "priority enrollment next round"
  • When perfect comparison isn't possible, historical and self-comparison are better than nothing
  • A two-group pre-post line chart is the most persuasive training evidence you can show leadership
← Previous: Reporting Training Value Next Track: Workshop Design →