Hypotheses and checking assumptions before we test
2026-10-01
filter(), select(), mutate(), arrange() as the core tidyverse verbsgroup_by() + summarize()stat_summary()Note
✅ Key idea for today
Last time we ran a t-test. Today we slow down and do it properly: state the hypotheses, check the assumptions, and only then run the test.
Tools added today:
stat_qq() + stat_qq_line() — QQ plotsshapiro.test() — normality testcar::leveneTest() — variance testReading:
Today’s naming:
_plot_modelNote
car is the Companion to Applied Regression package. It contains leveneTest(), which checks whether two groups have similar spread.
Install once: install.packages("car")
We will cover: why comparing means isn’t enough, the five steps we’ll use for every test this semester, what a two-sample t-test does, and why we default to Welch’s.
🛑 After this chunk: Activity Parts 1–2 (load data, recap the descriptive stats).
# Remind ourselves what the data look like ---------------
recap_plot <- leaf_df %>%
ggplot(aes(x = shade, y = mass_g, color = shade)) +
geom_jitter(width = 0.15, alpha = 0.5, size = 2) +
stat_summary(fun = mean, geom = "point", size = 4) +
stat_summary(
fun.data = mean_se,
geom = "errorbar",
width = 0.15,
linewidth = 0.9
) +
labs(x = "Shade", y = "Leaf Mass (g)") +
theme_minimal() +
theme(legend.position = "none")
recap_plot
The two means look close — and with this much scatter, they could easily be the same:
A p-value answers:
“If there were really no difference, how often would we see a gap at least this big just by chance?”
Large p → a gap this size is nothing unusual.
Every hypothesis test this semester — t-test, ANOVA, regression — follows the same steps:
| Step | Action | When |
|---|---|---|
| 1 | State H₀ and Hₐ | Today |
| 2 | Check normality | Today |
| 3 | Check equal variance | Today |
| 4 | Run the test | Today (a first look) |
| 5 | Interpret and report | Next lecture |
Important
Check assumptions BEFORE running the test.
If the data badly break the assumptions, the p-value can’t be trusted.
With ~25 per group our main worry is badly non-normal data. The t-test handles mild departures fine.
A two-sample t-test compares the means of two independent groups.
It asks: is the gap between the two means big compared with the scatter inside each group?
\[t = \frac{\text{difference between the means}}{\text{noise (scatter) in the data}}\]
R does the arithmetic — your job is to set the test up correctly and read the answer.
| Piece | Our leaves |
|---|---|
| Group 1 | Shady leaves |
| Group 2 | Sunny leaves |
| Measured | mass_g |
| H₀ | \(\mu_{shady} = \mu_{sunny}\) |
| Hₐ | \(\mu_{shady} \neq \mu_{sunny}\) |
There are two versions of the two-sample t-test:
When the spreads really are equal, both give almost the same answer. When they are not, the standard test gives too many false positives.
Tip
Best practice: use Welch’s as your default for every two-group comparison.
In R: var.equal = FALSE
🛑 Do Activity Parts 1–2 now
Load the data and recompute the group means and SDs. Predict each output, type it, run it.
We will cover: writing H₀ / Hₐ, checking normality three ways (histogram, QQ plot, Shapiro-Wilk), and checking equal variance (Levene’s).
🛑 After this chunk: Activity Parts 3–7 (hypotheses, histogram, box plot + QQ plot, Shapiro-Wilk, Levene’s).
Biological prediction: shady-side leaves will be larger and heavier — they need more surface area to capture the limited light under the canopy.
| Formal statement | |
|---|---|
| H₀ — null | \(\mu_{shady} = \mu_{sunny}\) — no difference in mean leaf mass |
| Hₐ — alternate | \(\mu_{shady} \neq \mu_{sunny}\) — mean leaf mass differs by side |
Two-tailed — detects a difference in either direction. α = 0.05 — reject H₀ if p < 0.05.
Why two-tailed when we predicted shady > sunny?
Why write hypotheses first?
ncol = 1 stacks the panels. Both share one x-axis, so you can see at a glance:
What to look for:
With ~25 leaves the bars are lumpy — that’s normal for real data. It’s hard to judge “bell-shaped” by eye, which is why we add a QQ plot.

A quantile is a value that a certain percent of the data falls below.
The box plot already shows three of them:
A QQ plot uses every leaf as a quantile, not just three — and asks whether they are spaced the way a normal (bell-curve) distribution would space them.

What R does for you:
Straight line = normal.
The dashed lines connect the box plot’s Q1, median, and Q3 to the same leaves on the QQ plot. The middle half of the dots is the box.
Here the heaviest few shady leaves bend up above the line — a long right tail.

Made-up data, 30 values each. Top: histogram. Bottom: QQ plot of the same values.
A little wiggle at the very ends is fine — look for a clear bend.
# QQ plot per group — dots should follow the line -------
qq_plot <- leaf_df %>%
ggplot(aes(sample = mass_g, color = shade)) +
stat_qq() +
stat_qq_line(color = "black") +
facet_wrap(~shade) +
labs(x = "Theoretical Quantiles",
y = "Leaf Mass (g)") +
theme_minimal() +
theme(legend.position = "none")
qq_plot
New pieces:
aes(sample = mass_g) — a QQ plot takes sample =, not x = / y =stat_qq() — draws one dot per leafstat_qq_line() — draws the “perfectly normal” lineWhat do you see?
Note
🔮 Predict first: Shapiro-Wilk’s H₀ is “the data are normal.” From the QQ plots, predict for each side: will p be above or below 0.05?
# A tibble: 2 × 2
shade shapiro_p
<chr> <dbl>
1 shady 0.0427
2 sunny 0.288
Same group_by() + summarize() as for means — it runs the test once per side. $p.value pulls out just the p-value.
Shapiro-Wilk H₀: the data are normal.
Important
The plot comes first; the test backs it up.
Shady is just under 0.05 — matching the bend we saw in the QQ plot. A mild tail like this is not a reason to abandon the t-test; it is robust to small departures.
Levene's Test for Homogeneity of Variance (center = median)
Df F value Pr(>F)
group 1 2.9667 0.09105 .
51
---
Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
Look back at the stacked histograms: which group was wider?
Levene’s H₀: the two groups have equal variance (spread).
Either way we use Welch’s (var.equal = FALSE). We run Levene’s to understand our data, not to pick the test.
🛑 Go to Activity 9 — Parts 3–7
State your hypotheses, then check normality (histograms, box plot + QQ plot, Shapiro-Wilk) and variance (Levene’s). Predict every p-value before you run it.
We will cover: running Welch’s t-test now that every assumption is checked. Reading the output line by line is next lecture.
🛑 After this chunk: Activity Part 8 (run the t-test).
Note
🔮 Predict first: the means were close and the groups overlap heavily. Above or below 0.05?
Welch Two Sample t-test
data: mass_g by shade
t = 0.35579, df = 42.527, p-value = 0.7238
alternative hypothesis: true difference in means between group shady and group sunny is not equal to 0
95 percent confidence interval:
-0.08392774 0.11987117
sample estimates:
mean in group shady mean in group sunny
0.5230357 0.5050640
Reading the formula mass_g ~ shade:
“mass explained by shade” — the number on the left, the groups on the right.
For today, find three numbers:
Next lecture: every line of this output, writing a results sentence, and the paired t-test.
🛑 Go to Activity 9 — Part 8
Run the t-test and record t, df, and p. Then start Part 9 — sugar maple (Part B of your skeleton): that is what you hand in.
Next lecture you run two tests on the sugar maple leaves. The checks you do today decide which version is allowed:
| Test | Must be roughly normal | Equal spread needed? |
|---|---|---|
| Two-sample (Welch’s) | each group | No — Welch’s handles it |
| Paired | the differences within each pair | No — only one set of numbers |
If normality clearly fails → use the rank-based partner instead (next lecture).
Important
Check first, test second.
A p-value from a test whose assumptions are broken can’t be trusted — and you can’t un-see a result. Deciding the test before you run it keeps you honest.
Your homework (Part B) is exactly this: check the sugar maple data and decide whether a t-test is OK.
var.equal = FALSE is the safe defaultncol = 1), QQ plot, Shapiro-WilkHand in: Part B of your skeleton — check the assumptions for the sugar maple leaves, written by you.
Up next — Lecture 10, T-Tests II:
t.test() output line by line