Activity 10 — T-Tests II: Two-Sample vs. Paired
Check the assumptions, run both tests, plot and report the result
Hands-on companion to T-Tests II: read and report a t-test, then on the sugar maple leaves average pseudoreplicates, check assumptions, run two-sample and paired t-tests, plot the paired result, and try the rank-based backups. Students hand in the same workflow on leaf mass.
In-class Activity 10: Two-Sample vs. Paired
Today’s Objectives
- Read a
t.test()output, make a decision, and write a results sentence - Run a two-sample and a paired t-test on our own leaf data
- Spot pseudoreplication — and fix it with one
summarize()at the top of the script - Rerun the same code on the fixed data and see how far the answer moves
- Check the assumptions first, then test the sugar maples the same way
- Plot and report a paired result
- Know the rank-based backups for when normality fails
- Skeleton script:
10_t_tests_2_skeleton.R→ save it asscripts/10_t_tests_2.R. - Leaf data:
2026_09_03_data_sci_leaf_area.xlsx→data/(you should already have it). - Sugar maple data:
2026_09_02_sugar_maple_leaf_area.xlsx→data/(you have it from Activity 09). - Make sure your project has a
figures/folder (Files pane → New Folder) — Part 11 saves a plot there.
Same setup as Activity 09:
- Part A — in class. The code is already in your skeleton. Run it chunk by chunk as you work through Parts 1–12 below, and answer the QA boxes. Parts 1–6 use our own leaf data; Parts 7–12 use sugar maple leaf area.
- Part B — hand in. The same maple workflow on leaf mass — you write the code by copying the matching Part A chunk and changing what needs changing. Answer the QB boxes. See Part 13.
🧩 Chunk 1 — The t-test code (after lecture Chunk 1)
Parts 1–6: run the two-sample and paired tests on every leaf, find the pseudoreplication, fix it with one
summarize()at the top, and rerun the same code.
Part 1 · Load libraries and data
Already at the top of your skeleton: readxl, tidyverse, janitor, car, then leaf_df and maple_df. Run those lines first.
Part 2 · Two-sample t-test on every leaf
▶ Run this:
# Welch's two-sample t-test on the leaf data -----------
leaf_model <- t.test(mass_g ~ shade, data = leaf_df,
var.equal = FALSE)
leaf_model
# Pull out t, df, and p for the results sentence -------
tibble(
t = leaf_model$statistic,
df = leaf_model$parameter,
p = leaf_model$p.value
)| Output line | What it tells you |
|---|---|
t = |
gap between the means ÷ noise |
df = |
degrees of freedom (Welch’s gives decimals — normal) |
p-value = |
how often you’d see a gap this big if H₀ were true |
95 percent confidence interval |
plausible range for the true difference |
mean in group ... |
the two group means |
✏️ QA2: Reject or fail to reject H₀? Does the CI include zero? Write the results sentence using the template [Finding] (Welch’s t-test: t(df) = X.XX, p = X.XX).
Part 3 · Where did these leaves come from?
Before the paired test, look at the design of your own data.
▶ Run this:
# How many leaves did each team turn in? ---------------
leaf_df %>%
count(teams, shade) %>%
pivot_wider(names_from = shade, values_from = n)Each team went to one tree and picked leaves off the sunny side and the shady side of that tree.
✏️ QA3: How many teams are there? Did every team turn in the same number of leaves? Which teams did not?
Part 4 · Paired t-test on every leaf
The two sides are matched within a team, so let’s run the paired test. A paired test needs the two groups side by side in one row, so we reshape first.
▶ Run this:
# Number the leaves within each team and side ----------
leaf_wide_df <- leaf_df %>%
group_by(teams, shade) %>%
mutate(leaf_n = row_number()) %>%
ungroup() %>%
select(teams, leaf_n, shade, mass_g) %>%
pivot_wider(names_from = shade, values_from = mass_g) %>%
mutate(diff = shady - sunny)
leaf_wide_df
# Paired t-test on all the leaves ----------------------
leaf_paired_model <- t.test(leaf_wide_df$shady,
leaf_wide_df$sunny,
paired = TRUE)
leaf_paired_modelrow_number()numbers the leaves 1, 2, 3… inside each team and sidepivot_wider()puts the shady and sunny leaf with the same number in the same row
⚠️ Watch out! Look at the
shadyandsunnycolumns for the teams with uneven counts — some rows haveNA. Those leaves had no partner, andt.test()dropped them without telling you.
✏️ QA4: Record t, df, and p. How many pairs does df say you have? Now look at your table — why does shady leaf #1 get paired with sunny leaf #1? Is there any real reason those two leaves belong together?
Part 5 · Pseudoreplication — and the one-line fix
The 6 shady leaves one team turned in all came off one tree. They share that tree’s soil, water, light, and age — they are not independent. Treating all 53 leaves as independent is pseudoreplication: it pretends you have far more information than you collected.
You did not sample 53 trees. You sampled 5.
The fix: go back to the top of your script — to where you loaded the data — and add the averaging step to the end of that chain. Keep the same name (leaf_df) and the same column name (mass_g).
▶ Edit the chunk at the top of your script so it looks like this, then run it:
# Load the leaf data and average each team's side ------
leaf_df <-
read_excel("data/2026_09_03_data_sci_leaf_area.xlsx") %>%
clean_names() %>%
group_by(teams, shade) %>%
summarize(mass_g = mean(mass_g),
.groups = "drop")
leaf_dfBecause you reused the names leaf_df and mass_g, every line of code below keeps working — this is how real analysis goes. You don’t write a new script; you fix how the data are built and rerun.
✏️ QA5: How many rows does leaf_df have now? What is one row — a leaf, a side, a team, or a tree?
Part 6 · Rerun the same code and compare
▶ Rerun your Part 2 and Part 4 chunks, unchanged. Scroll up and run them again — nothing to retype.
# Keep a copy of the pseudoreplicated results ----------
pseudo_2s <- leaf_model
pseudo_paired <- leaf_paired_model⚠️ Run those two lines before you rerun Parts 2 and 4, or the old results are gone.
Then put all four side by side:
# All four results from the leaf data ------------------
tibble(
analysis = c("Two-sample, every leaf",
"Paired, every leaf",
"Two-sample, team means",
"Paired, team means"),
t = c(pseudo_2s$statistic, pseudo_paired$statistic,
leaf_model$statistic, leaf_paired_model$statistic),
df = c(pseudo_2s$parameter, pseudo_paired$parameter,
leaf_model$parameter, leaf_paired_model$parameter),
p = c(pseudo_2s$p.value, pseudo_paired$p.value,
leaf_model$p.value, leaf_paired_model$p.value)
)✏️ QA6: What happened to df? Did you lose data, or stop double-counting it? Now look at the sign of t: shady leaves looked heavier before averaging and lighter after. Given your answer to QA3 about uneven leaf counts, why would averaging change the direction of the difference?
🧩 Chunk 2 — The maple test (after lecture Chunk 2)
Parts 7–12, all in one go: average the leaves on each side of each tree, check the assumptions, run the two-sample and the paired t-test, plot and report the paired result, then the rank-based backups. Then start Part B.
Part 7 · Average per tree and side
Same problem, same fix. Each tree gave 2 leaves per side, and those 2 leaves are not independent (same tree, same soil, same light) — exactly like the 6 leaves one team pulled off one tree. So we start the same way: average up to the real replicate, the tree.
One difference worth noting: this design is balanced — exactly 2 leaves on every tree-side, nothing uneven, nothing missing. So averaging won’t flip the difference the way it did for our leaf data. Every tree already had an equal say.
▶ Run this:
# Average the 2 leaves per tree and side ---------------
maple_tree_df <- maple_df %>%
group_by(tree, side) %>%
summarize(area = mean(total_area_cm2),
.groups = "drop")
maple_tree_df✏️ QA7: Why average? How many values per side do you have now? How many leaves went in?
Part 8 · Two-sample — check, then test
💡 Check → decide → test. A two-sample t-test needs each side to be roughly normal.
▶ Run this:
# Normality of each side (tree means) ------------------
maple_tree_df %>%
group_by(side) %>%
summarize(shapiro_p = shapiro.test(area)$p.value)
# Equal spread? ----------------------------------------
leveneTest(area ~ side, data = maple_tree_df)🔮 Predict first: assumptions OK? Then: with 8 trees per side, will the two-sample test be significant?
# Welch's two-sample t-test on the tree means ----------
maple_2s_model <- t.test(area ~ side,
data = maple_tree_df,
var.equal = FALSE)
maple_2s_model✏️ QA8: Assumptions OK? t, df, p? Reject or fail to reject?
Part 9 · The slope plot
▶ Run this:
# One line per tree: sunny -> shady --------------------
maple_tree_df %>%
ggplot(aes(x = side, y = area, group = tree)) +
geom_line() +
geom_point() +
scale_x_discrete(limits = c("sunny", "shady"))group = tree joins the two points from the same tree with a line.
✏️ QA9: How many lines go up? Why can this pattern be so clear when the two-sample test was borderline?
Part 10 · Paired — check, then test
A paired test needs one row per tree and checks that the differences are roughly normal.
▶ Run this:
# One row per tree + the shady - sunny difference ------
maple_wide_df <- maple_tree_df %>%
pivot_wider(names_from = side, values_from = area) %>%
mutate(diff = shady - sunny)
maple_wide_df
# QQ plot of the differences ---------------------------
maple_wide_df %>%
ggplot(aes(sample = diff)) +
stat_qq() +
stat_qq_line()
# Shapiro-Wilk on the differences ----------------------
shapiro.test(maple_wide_df$diff)pivot_wider()turns thesidecolumn into ashadyand asunnycolumn — a sneak peek; you’ll learn it in the Pivoting lecture.- No Levene’s test for a paired test — there is only one set of numbers, the differences.
🔮 Predict first: will the paired p be bigger or smaller than the two-sample p?
# Paired t-test ----------------------------------------
maple_paired_model <- t.test(maple_wide_df$shady,
maple_wide_df$sunny,
paired = TRUE)
maple_paired_model⚠️ Watch out!
t.test(area ~ side, paired = TRUE)gives an error — R doesn’t allowpairedwith a formula. That’s why we made the wide table.
✏️ QA10: Differences normal enough? Paired t, df, p, and mean difference? Why is the paired test so much stronger than the two-sample test?
Part 11 · Plot and report the paired result
▶ Run this:
# Label with the paired test result ---------------------
paired_label <- paste0(
"Paired t-test: t = ",
round(maple_paired_model$statistic, 2),
", df = ", maple_paired_model$parameter, ", ",
scales::pvalue(maple_paired_model$p.value,
add_p = TRUE)
)
# Each tree in grey + mean ± SE in color ---------------
maple_results_plot <- maple_tree_df %>%
ggplot(aes(x = side, y = area)) +
geom_line(aes(group = tree), color = "grey70") +
geom_point(color = "grey70") +
stat_summary(fun.data = mean_se, geom = "errorbar",
width = 0.1, color = "darkgreen") +
stat_summary(fun = mean, geom = "point",
size = 4, color = "darkgreen") +
scale_x_discrete(limits = c("sunny", "shady")) +
labs(x = "Side of tree",
y = "Mean leaf area (cm²)",
subtitle = paired_label) +
theme_minimal()
maple_results_plot
# Save it -----------------------------------------------
ggsave("figures/maple_area_paired_plot.png",
plot = maple_results_plot,
width = 5, height = 4, dpi = 300)✏️ QA11: Write the results sentence: [Finding] (paired t-test: t(df) = X.XX, p = X.XXX; mean difference = X.X cm²). If p is tiny, write “p < 0.001”.
Part 12 · Rank-based backups
Use these when the QQ plot and Shapiro-Wilk say the data are clearly not normal.
▶ Run this:
# Mann-Whitney U: two-sample partner -------------------
wilcox.test(area ~ side, data = maple_tree_df)
# Wilcoxon signed-rank: paired partner -----------------
wilcox.test(maple_wide_df$shady, maple_wide_df$sunny,
paired = TRUE)| Roughly normal | Clearly not normal | |
|---|---|---|
| Two independent groups | Welch’s two-sample t-test | Mann-Whitney U |
| Paired | Paired t-test | Wilcoxon signed-rank |
✏️ QA12: p for each? Which one “misses” the effect the t-test caught? Would you report these here — why or why not?
Part 13 · Your turn: leaf mass (hand in)
Start in class, finish before next class. This is Part B of your skeleton and it is what gets graded.
Every step in Part B matches a Part A step. For each # TODO:, copy the matching Part A chunk, paste it under the TODO, and change what needs changing:
- the variable is
mass_g, nottotal_area_cm2 - call the averaged column
mass(notarea) — then every laterareabecomesmass - use new names so you don’t overwrite Part A:
mass_tree_df,mass_wide_df,mass_2s_model,mass_paired_model,mass_results_plot - fix the plot’s y-axis label and the file name in
ggsave()
⚠️ Watch out! If you get
object 'area' not found, you renamed the column tomassbut missed anareafurther down.
Leaf mass tells a different story from leaf area. Pay attention to the assumption checks and to what each test says — QB10 asks which test you trust and why.
When you’re done, run the whole script top to bottom (Ctrl/Cmd + Shift + Enter) — it must run with no errors — and hand in scripts/10_t_tests_2.R.
End of Activity 10. Next: Lecture 11 — linear regression.