Comparing groups

Research Skills — Week 7

Recap

  • Why not just do lots of t-tests?
  • How many tests?
  • ANOVA
  • Run an ANOVA
  • Effect sizes
  • HolmesCo’s report card
  • Wrap-up

Where we are

Last week: t-tests for two groups. The base rate fallacy.

This week: what happens when you have more than two groups — and more than one test.

Questions?

Submit questions:

PollEv.com/geol

text geol to 07480 781235

Why not just do lots of t-tests?

  • Recap

🎓 Concept block 1

  • How many tests?
  • ANOVA
  • Run an ANOVA
  • Effect sizes
  • HolmesCo’s report card
  • Wrap-up

The problem

Three groups (e.g., wind farms in three regions).

You could run three pairwise t-tests: A vs B, A vs C, B vs C.

Each test has a 5% false positive rate.

Three tests: P(at least one false positive) = 1 − 0.95³ ≈ 14%

Ten tests: 40%

Twenty tests: 64%

The escalation

Number of tests P(≥1 false positive)
1 5%
3 14%
10 40%
20 64%
100 99.4%

Run enough tests and you’re guaranteed to “find” something.

HolmesCo’s mineral survey

“Geological Solutions Since 2019”

Press Release — Multi-Element Anomaly Discovered!

“We tested soil samples from 20 sites for enrichment in 20 elements. We found statistically significant enrichment (p < 0.05) for three elements at one site!”

If there’s truly nothing there, how many significant results would you expect by chance?

The answer

20 × 20 = 400 tests at α = 0.05

Expected false positives: 20

HolmesCo found 3.

That’s fewer than expected by chance. The “discovery” is noise.

How many tests?

  • Recap
  • Why not just do lots of t-tests?

💬 Exercise 1

  • ANOVA
  • Run an ANOVA
  • Effect sizes
  • HolmesCo’s report card
  • Wrap-up

Your project

How many comparisons are you planning to make?

If you run all of them at α = 0.05, what’s your family-wise false positive rate?

Formula: 1 − (1 − α)k where k = number of tests

ANOVA

  • Recap
  • Why not just do lots of t-tests?
  • How many tests?

🎓💻 Concept block 2

  • Run an ANOVA
  • Effect sizes
  • HolmesCo’s report card
  • Wrap-up

One test for multiple groups

ANOVA (Analysis of Variance) asks:

“Does any group differ from the others?”

Without specifying which.

The logic

  1. Total variation = between groups + within groups
  2. If between-group variation is large relative to within-group → at least one group is different
  3. The F-statistic captures this ratio
  4. The p-value tells you how surprising this F is under the null

Same data, four ways

Before any test: how you show a comparison changes what you see.

Four regions, same 100 sites, same order (low mean to high) — shown four different ways.

Same data, four ways: the bar

A bar chart of mean capacity factor for four UK regions, ordered low to high: Southern England, Wales, Scotland, Northern England. Four bars, nothing else — no error bars, no points.

Same data, four ways: the box plot

A box plot of the same four regions in the same order. The boxes for Scotland and Northern England sit higher and overlap each other; Southern England and Wales sit lower and overlap each other. The two pairs overlap only a little across the middle of the plot.

Same data, four ways: the raw points

The 25 individual sites per region, jittered, with the group mean and its 95% confidence interval marked. The Scotland and Northern England intervals sit clearly above the Wales and Southern England intervals, with visible overlap in the raw points between neighbouring regions.

Takeaway: a bar of means can hide a bad idea; the points behind it usually don’t.

Same data, four ways: the raincloud

For each region, in the same low-to-high order: the 25 sites as jittered points on the left, the mean with its 95% confidence interval in the middle, and a half-violin of the distribution on the right. Northern England's half-violin has two bumps, a detail the bar, box and interval views all hide.

w07-groups-raincloud.R · after Heiss (2026)

“Relative to within”

Two panels of capacity factors for three regions, with a horizontal bar at each group mean. The means sit at identical heights in both panels. On the left the points cluster tightly around them and F is 83; on the right they scatter widely and overlap between regions, and F is 6.

Live demo

If significant → which pairs differ?

Corrections for multiple comparisons

Method How When
Tukey’s HSD All pairwise, designed for ANOVA After a significant ANOVA
Bonferroni Divide α by number of tests Simple, conservative, any context

Both reduce the false positive rate by being stricter about what counts as significant.

Run an ANOVA

  • Recap
  • Why not just do lots of t-tests?
  • How many tests?
  • ANOVA

✏️💻 Exercise 2

  • Effect sizes
  • HolmesCo’s report card
  • Wrap-up

In WebR

Provided dataset: wind farm output in four regions.

  1. Visualize with violins
  2. Run aov()
  3. Is it significant?
  4. Run TukeyHSD() — which pairs differ?
  5. Write a one-sentence conclusion

Attendance

Effect sizes

  • Recap
  • Why not just do lots of t-tests?
  • How many tests?
  • ANOVA
  • Run an ANOVA

🎓 Concept block 3

  • HolmesCo’s report card
  • Wrap-up

“Significant” ≠ “important”

A p-value tells you whether an effect is detectable.

It does not tell you whether it’s big enough to matter.

The HolmesCo temperature problem

“Geological Solutions Since 2019”

Two boreholes, 1,000 readings each.

Mean difference: 0.02°C. p-value: 0.001.

“Highly significant temperature anomaly detected!”

Is 0.02°C meaningful for any practical purpose?

Cohen’s d

A standardized measure of how far apart two groups are:

d Interpretation
0.2 Small
0.5 Medium
0.8 Large
effectsize::cohens_d(y ~ group, data = df)

What those numbers look like

Three panels, each with two bell curves and their overlap shaded. At d = 0.2 the curves are almost on top of each other; at d = 0.5 they are visibly offset but still mostly shared; at d = 0.8 they are clearly apart and still share about 69% of their range.

For ANOVA: η² (eta-squared)

Proportion of variance explained by the grouping variable.

η² Interpretation
0.01 Small
0.06 Medium
0.14 Large

Always report both the p-value and an effect size.

A policy-maker needs to know “how big?” — not just “is it real?”

HolmesCo’s report card

  • Recap
  • Why not just do lots of t-tests?
  • How many tests?
  • ANOVA
  • Run an ANOVA
  • Effect sizes

✏️💬 Integrative exercise

  • Wrap-up

Three test results

“Geological Solutions Since 2019”

Technical Summary

  1. Test A: p = 0.002, Cohen’s d = 1.2
  2. Test B: p = 0.04, Cohen’s d = 0.08
  3. Test C: p = 0.08, 95% CI = [−0.3, 6.1]

For each: should we act on this? Why or why not?

Wrap-up

  • Recap
  • Why not just do lots of t-tests?
  • How many tests?
  • ANOVA
  • Run an ANOVA
  • Effect sizes
  • HolmesCo’s report card

Key points

  1. Multiple tests inflate false positives — correct for them
  2. ANOVA compares multiple groups in one test; TukeyHSD finds which differ
  3. Significance ≠ importance — always report effect sizes
  4. “Is that a big number?” now has a formal answer

Any questions we missed?

Submit questions:

PollEv.com/geol

text geol to 07480 781235

Next time

Application session: “Deepening your analysis”

Apply ANOVA or refine your t-test with effect sizes and assumption checks.