Lecture 09 — T-Tests I: Setting Up the Test

Hypotheses and checking assumptions before we test

tidyverse
descriptive-stats

Why comparing means isn’t enough, why Welch’s is the safe default, checking normality (histogram, QQ plot, Shapiro-Wilk) and variance (Levene’s), then a first run of the t-test.

Author

Bill Perry

Published

October 1, 2026

Where we left off (Lectures 05–07)

  • Wrangling — filter(), select(), mutate(), arrange() as the core tidyverse verbs
  • Descriptive stats — mean, SD, SE with group_by() + summarize()
  • Box plot + jittered points; mean ± SE with stat_summary()
  • Quarto report — you already ran a Welch’s t-test once, at the end
  • Prediction: shady-side leaves would be heavier than sunny-side leaves
Note

✅ Key idea for today

Last time we ran a t-test. Today we slow down and do it properly: state the hypotheses, check the assumptions, and only then run the test.

Goals for Today

  • Understand why comparing means is not enough — we need a test
  • Meet the two-sample t-test and why we default to Welch’s
  • State null and alternate hypotheses
  • Check assumptions before running any test
    • normality: histogram, QQ plot, Shapiro-Wilk
    • variance: Levene’s test
  • Run the t-test (we decode it next lecture)

Tools added today:

  • stat_qq() + stat_qq_line() — QQ plots
  • shapiro.test() — normality test
  • car::leveneTest() — variance test

Reading:

  • 📖 Whitlock & Schluter, Ch. 12 — Comparing two means
  • 📖 Whitlock & Schluter, Ch. 13 — Handling violations of assumptions

Today’s naming:

  • plots → _plot
  • test results → _model

Load Libraries and Data

# Load all packages at the top ----------------------------
library(readxl)    # reading Excel files
library(tidyverse) # data manipulation + ggplot2
library(janitor)   # clean_names()
library(car)       # Levene's test for equal variance
# Load the leaf data from the data folder -----------------
leaf_df <-
  read_excel("data/2026_09_03_data_sci_leaf_area.xlsx") %>%
  clean_names()
Note

car is the Companion to Applied Regression package. It contains leveneTest(), which checks whether two groups have similar spread.

Install once: install.packages("car")

🧩 Chunk 1 of 3 · Why We Test

We will cover: why comparing means isn’t enough, the five steps we’ll use for every test this semester, what a two-sample t-test does, and why we default to Welch’s.

🛑 After this chunk: Activity Parts 1–2 (load data, recap the descriptive stats).

Why Comparing Means Is Not Enough

# Remind ourselves what the data look like ---------------
recap_plot <- leaf_df %>%
  ggplot(aes(x = shade, y = mass_g, color = shade)) +
  geom_jitter(width = 0.15, alpha = 0.5, size = 2) +
  stat_summary(fun = mean, geom = "point", size = 4) +
  stat_summary(
    fun.data = mean_se,
    geom = "errorbar",
    width = 0.15,
    linewidth = 0.9
  ) +
  labs(x = "Shade", y = "Leaf Mass (g)") +
  theme_minimal() +
  theme(legend.position = "none")
recap_plot

Jittered points of leaf mass in grams for shady and sunny sides, with a black point and error bar showing the mean plus or minus one standard error for each side; the two means sit close together with overlapping error bars.

The two means look close — and with this much scatter, they could easily be the same:

  • About 25 leaves per group
  • Lots of natural variability within each group
  • The groups overlap heavily

A p-value answers:

“If there were really no difference, how often would we see a gap at least this big just by chance?”

Large p → a gap this size is nothing unusual.

Five Steps to a Hypothesis Test

Every hypothesis test this semester — t-test, ANOVA, regression — follows the same steps:

Step Action When
1 State H₀ and Hₐ Today
2 Check normality Today
3 Check equal variance Today
4 Run the test Today (a first look)
5 Interpret and report Next lecture
Important

Check assumptions BEFORE running the test.

If the data badly break the assumptions, the p-value can’t be trusted.

With ~25 per group our main worry is badly non-normal data. The t-test handles mild departures fine.

What Is a Two-Sample t-Test?

A two-sample t-test compares the means of two independent groups.

It asks: is the gap between the two means big compared with the scatter inside each group?

\[t = \frac{\text{difference between the means}}{\text{noise (scatter) in the data}}\]

  • Big gap, little scatter → big t → small p-value
  • Small gap, lots of scatter → t near 0 → large p-value

R does the arithmetic — your job is to set the test up correctly and read the answer.

Piece Our leaves
Group 1 Shady leaves
Group 2 Sunny leaves
Measured mass_g
H₀ \(\mu_{shady} = \mu_{sunny}\)
Hₐ \(\mu_{shady} \neq \mu_{sunny}\)

Why Welch’s? — Don’t Assume Equal Variance

There are two versions of the two-sample t-test:

  • Standard (Student’s) t-test — assumes both groups have the same spread
  • Welch’s t-test — lets each group have its own spread

When the spreads really are equal, both give almost the same answer. When they are not, the standard test gives too many false positives.

Tip

Best practice: use Welch’s as your default for every two-group comparison.

In R: var.equal = FALSE

→ ACTIVITY 9 Parts 1–2 now

🛑 Do Activity Parts 1–2 now

Load the data and recompute the group means and SDs. Predict each output, type it, run it.

🧩 Chunk 2 of 3 · Steps 1–3: Hypotheses & Assumptions

We will cover: writing H₀ / Hₐ, checking normality three ways (histogram, QQ plot, Shapiro-Wilk), and checking equal variance (Levene’s).

🛑 After this chunk: Activity Parts 3–7 (hypotheses, histogram, box plot + QQ plot, Shapiro-Wilk, Levene’s).

Step 1 — Hypotheses

Biological prediction: shady-side leaves will be larger and heavier — they need more surface area to capture the limited light under the canopy.

Formal statement
H₀ — null \(\mu_{shady} = \mu_{sunny}\) — no difference in mean leaf mass
Hₐ — alternate \(\mu_{shady} \neq \mu_{sunny}\) — mean leaf mass differs by side

Two-tailed — detects a difference in either direction. α = 0.05 — reject H₀ if p < 0.05.

Why two-tailed when we predicted shady > sunny?

  • A one-tailed test only looks in one direction
  • Two-tailed is more conservative and the standard in most journals

Why write hypotheses first?

  • Keeps the analysis honest
  • Prevents “HARKing” — Hypothesizing After the Results are Known

Step 2 — Normality: Histograms

# Histograms stacked in one column — same x-axis --------
hist_norm_plot <- leaf_df %>%
  ggplot(aes(x = mass_g, fill = shade)) +
  geom_histogram(binwidth = 0.05) +
  facet_wrap(~shade, ncol = 1) +
  labs(x = "Leaf Mass (g)", y = "Count") +
  theme_minimal() +
  theme(legend.position = "none")
hist_norm_plot

Two histograms of leaf mass in grams stacked one above the other, shady on top and sunny below, sharing the same x-axis so the ranges can be compared directly.

ncol = 1 stacks the panels. Both share one x-axis, so you can see at a glance:

  • where each group is centered
  • how wide each group spreads
  • which group has the long tail

What to look for:

  • Roughly bell-shaped — one hump, not badly lopsided
  • No big outliers far from the rest

With ~25 leaves the bars are lumpy — that’s normal for real data. It’s hard to judge “bell-shaped” by eye, which is why we add a QQ plot.

First: What Is a Quantile? (you already know one)

Box plot of shady-side leaf mass with jittered points, with labels pointing to the lower quartile (25 percent of leaves below), median (50 percent below), and upper quartile (75 percent below).

A quantile is a value that a certain percent of the data falls below.

The box plot already shows three of them:

  • Q1 (25th percentile) — bottom of the box
  • Median (50th) — the line in the middle
  • Q3 (75th) — top of the box

A QQ plot uses every leaf as a quantile, not just three — and asks whether they are spaced the way a normal (bell-curve) distribution would space them.

How a QQ Plot Works

QQ plot of shady-side leaf mass with a box plot of the same data drawn on its left edge. Dashed horizontal lines run from the box plot's quartiles across the QQ plot, meeting the points near theoretical values of minus 0.67, 0 and 0.67. Most points follow the straight reference line, but the largest few leaves bend upward above it.

What R does for you:

  1. Sorts the leaves from lightest to heaviest (the y-axis)
  2. Works out where each leaf would sit if the data were perfectly normal (the x-axis, in SD units: −2 … 0 … +2)
  3. Plots actual vs. expected — one dot per leaf

Straight line = normal.

The dashed lines connect the box plot’s Q1, median, and Q3 to the same leaves on the QQ plot. The middle half of the dots is the box.

Here the heaviest few shady leaves bend up above the line — a long right tail.

What Good and Bad QQ Plots Look Like

Three columns of made-up data. Top row histograms, bottom row QQ plots. Left: normal data, bell-shaped histogram, points on the line. Middle: right-skewed data, histogram with a long right tail, points curving upward above the line at the right end. Right: data with outliers, points on the line except two far above it at the right end.

Made-up data, 30 values each. Top: histogram. Bottom: QQ plot of the same values.

  • Normal: dots hug the line
  • Skewed: dots curve away at one end (a banana shape)
  • Outliers: a few dots jump off the line at an end

A little wiggle at the very ends is fine — look for a clear bend.

Step 2 — Normality: Our QQ Plots

# QQ plot per group — dots should follow the line -------
qq_plot <- leaf_df %>%
  ggplot(aes(sample = mass_g, color = shade)) +
  stat_qq() +
  stat_qq_line(color = "black") +
  facet_wrap(~shade) +
  labs(x = "Theoretical Quantiles",
       y = "Leaf Mass (g)") +
  theme_minimal() +
  theme(legend.position = "none")
qq_plot

QQ plots of leaf mass for shady and sunny leaves side by side. Sunny points follow the reference line closely. Shady points mostly follow the line but the largest few curve upward above it.

New pieces:

  • aes(sample = mass_g) — a QQ plot takes sample =, not x = / y =
  • stat_qq() — draws one dot per leaf
  • stat_qq_line() — draws the “perfectly normal” line

What do you see?

  • Sunny: dots follow the line well
  • Shady: mostly on the line, but the heaviest leaves bend up — the same right tail you saw in the histogram

Step 2 — Normality: Shapiro-Wilk Test

Note

🔮 Predict first: Shapiro-Wilk’s H₀ is “the data are normal.” From the QQ plots, predict for each side: will p be above or below 0.05?

# Shapiro-Wilk p-value for each side -------------------
leaf_df %>%
  group_by(shade) %>%
  summarize(shapiro_p = shapiro.test(mass_g)$p.value)
# A tibble: 2 × 2
  shade shapiro_p
  <chr>     <dbl>
1 shady    0.0427
2 sunny    0.288 

Same group_by() + summarize() as for means — it runs the test once per side. $p.value pulls out just the p-value.

Shapiro-Wilk H₀: the data are normal.

  • p > 0.05 → no evidence against normality → go ahead
  • p < 0.05 → evidence of non-normality → look hard at the plots
Important

The plot comes first; the test backs it up.

Shady is just under 0.05 — matching the bend we saw in the QQ plot. A mild tail like this is not a reason to abandon the t-test; it is robust to small departures.

Step 3 — Check Variance: Levene’s Test

# Levene's test — are the spreads similar? --------------
leveneTest(mass_g ~ shade, data = leaf_df)
Levene's Test for Homogeneity of Variance (center = median)
      Df F value  Pr(>F)  
group  1  2.9667 0.09105 .
      51                  
---
Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Look back at the stacked histograms: which group was wider?

Levene’s H₀: the two groups have equal variance (spread).

  • p > 0.05 → spreads are not significantly different
  • p < 0.05 → spreads are different

Either way we use Welch’s (var.equal = FALSE). We run Levene’s to understand our data, not to pick the test.

→ ACTIVITY 9 Parts 3–7 now

🛑 Go to Activity 9 — Parts 3–7

State your hypotheses, then check normality (histograms, box plot + QQ plot, Shapiro-Wilk) and variance (Levene’s). Predict every p-value before you run it.

🧩 Chunk 3 of 3 · Step 4: Run the Test

We will cover: running Welch’s t-test now that every assumption is checked. Reading the output line by line is next lecture.

🛑 After this chunk: Activity Part 8 (run the t-test).

Step 4 — Run Welch’s t-Test

Note

🔮 Predict first: the means were close and the groups overlap heavily. Above or below 0.05?

# Welch's two-sample t-test -----------------------------
leaf_model <- t.test(mass_g ~ shade, data = leaf_df,
                     var.equal = FALSE)
leaf_model

    Welch Two Sample t-test

data:  mass_g by shade
t = 0.35579, df = 42.527, p-value = 0.7238
alternative hypothesis: true difference in means between group shady and group sunny is not equal to 0
95 percent confidence interval:
 -0.08392774  0.11987117
sample estimates:
mean in group shady mean in group sunny 
          0.5230357           0.5050640 

Reading the formula mass_g ~ shade:

“mass explained by shade” — the number on the left, the groups on the right.

For today, find three numbers:

  • t — the signal-to-noise ratio
  • df — degrees of freedom
  • p-value — the decision number

Next lecture: every line of this output, writing a results sentence, and the paired t-test.

→ ACTIVITY 9 Part 8 now

🛑 Go to Activity 9 — Part 8

Run the t-test and record t, df, and p. Then start Part 9 — sugar maple (Part B of your skeleton): that is what you hand in.

Why the Checks Matter — Next Lecture’s Choice

Next lecture you run two tests on the sugar maple leaves. The checks you do today decide which version is allowed:

Test Must be roughly normal Equal spread needed?
Two-sample (Welch’s) each group No — Welch’s handles it
Paired the differences within each pair No — only one set of numbers

If normality clearly fails → use the rank-based partner instead (next lecture).

Important

Check first, test second.

A p-value from a test whose assumptions are broken can’t be trusted — and you can’t un-see a result. Deciding the test before you run it keeps you honest.

Your homework (Part B) is exactly this: check the sugar maple data and decide whether a t-test is OK.

What We Learned Today

  • Why we need a hypothesis test — means alone are not enough
  • The five steps we’ll reuse all semester
  • Welch’s t-test — var.equal = FALSE is the safe default
  • Checking assumptions before testing:
    • normality — histograms (ncol = 1), QQ plot, Shapiro-Wilk
    • variance — Levene’s
  • The QQ plot: dots on the line = normal

Hand in: Part B of your skeleton — check the assumptions for the sugar maple leaves, written by you.

Up next — Lecture 10, T-Tests II:

  • Decode the t.test() output line by line
  • Make a decision and write a results sentence
  • Two-sample vs. paired t-tests