Grouped summaries, facets, and the pipe into ggplot
2026-10-01
data/, scripts/, figures/install.packages() once, library() every sessionread_excel(), clean_names() into leaf_df%>% reads as “then”geom_point()Note
✅ Key idea from Lecture 02
You can already load data and make one honest point on a page. Today that one geom becomes a whole toolkit: custom colors, side-by-side panels, and a proper mean ± SE summary.
leaf_df %>% ggplot(...)scale_color_manual()facet_wrap() and facet_grid()stat_summary()Tip
🖐 Try it yourself
By the end you will make a publication-style figure: raw leaf data, grouped and colored on purpose, with the mean and its uncertainty layered on top.
This lecture runs in three short chunks. After each chunk you switch to the activity and type the code yourself into your R script.
For every code block, do three things:
Note
✅ Why bother?
facet_wrap() and facet_grid() look interchangeable until you’ve typed both and watched one wrap into rows while the other insists on a strict grid.scale_color_manual(values = c(...)) yourself is what makes you notice the color names have to match your group names exactly.Rows: 53
Columns: 8
$ twig_id <chr> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA…
$ leaf_id <chr> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA…
$ teams <chr> "12345", "12345", "12345", "12345", "12345…
$ shade <chr> "sunny", "sunny", "sunny", "sunny", "sunny…
$ mass_g <dbl> 0.4000, 0.4800, 0.3400, 0.6500, 0.2700, 0.…
$ petiole_mm <dbl> 79.0000, 63.0000, 68.0000, 35.0000, 34.000…
$ thickness_mm <dbl> 0.15, 0.14, 0.14, 0.15, 0.11, 0.16, 0.14, …
$ paper_mass_g <dbl> 0.2100, 0.2100, 0.2100, 0.2400, 0.1500, 0.…
We will cover: starting a plot with the pipe, and choosing colors with scale_color_manual().
🛑 After this chunk you will do Activity Parts 1–2.
# Pipe the data frame straight into ggplot() -----------
leaf_df %>%
ggplot(aes(x = shade, y = mass_g, fill = shade)) +
geom_boxplot(alpha = 0.6, outlier.shape = NA) +
geom_jitter(width = 0.15, alpha = 0.6) +
labs(
title = "Leaf Mass by Shade",
x = "Shade",
y = "Leaf Mass (g)",
fill = "Shade"
) +
theme_minimal()
leaf_df %>% ggplot(...) reads as “take leaf_df, then plot it” — the same pipe you’ve already usedggplot() is added with +, not %>% — ggplot layers use plus, pipelines use the pipefill = shade maps a column to color — ggplot chooses the colors for youWarning
⚠️ Watch out!
%>% connects data-wrangling steps. + connects plot layers. Mixing them up is the single most common ggplot error.
scale_color_manual()Note
🔮 Predict first: scale_fill_manual() needs one color per group. We have two groups (sunny, shady). How many colors do we need to supply?
# scale_fill_manual() assigns exact colors to exact groups
leaf_df %>%
ggplot(aes(x = shade, y = mass_g, fill = shade)) +
geom_boxplot(alpha = 0.6, outlier.shape = NA) +
geom_jitter(width = 0.15, alpha = 0.6) +
scale_fill_manual(values = c(
"sunny" = "goldenrod",
"shady" = "forestgreen"
)) +
labs(
title = "Leaf Mass by Shade",
x = "Shade",
y = "Leaf Mass (g)",
fill = "Shade"
) +
theme_minimal() +
theme(legend.position = "none")
scale_fill_manual() — for fill = aesthetics (boxes, bars, columns)scale_color_manual() — for color = aesthetics (points, lines)c() must exactly match the values in your data ("sunny", "shady") — case and spelling both countWarning
⚠️ Watch out!
If a name in values = c(...) doesn’t match a value in your data exactly, that group’s color falls back to grey and ggplot prints a warning.
🛑 Do Activity Parts 1–2 now
Pipe into ggplot(), then assign sunny/shady their own colors with scale_fill_manual(). Predict, type, run.
We will cover: facet_wrap() and facet_grid() — splitting one plot into a panel per group.
🛑 After this chunk you will do Activity Part 3.
facet_wrap() — One Panel per Group# facet_wrap() splits into one panel per level ---------
leaf_df %>%
ggplot(aes(x = mass_g, fill = shade)) +
geom_histogram(binwidth = 0.05, color = "white") +
facet_wrap(~shade) +
scale_fill_manual(values = c("sunny" = "goldenrod",
"shady" = "forestgreen")) +
labs(x = "Leaf Mass (g)", y = "Count",
title = "Mass Distribution by Shade") +
theme_minimal() +
theme(legend.position = "none")
~shade reads as “by shade” — the ~ is requiredfacet_wrap(~shade, scales = "free_y") lets each panel’s y-axis rescale independentlyfacet_grid() — a Strict Rows-and-Columns Grid# facet_grid(rows ~ cols) — use "." for "no grouping" --
leaf_df %>%
ggplot(aes(x = mass_g, fill = shade)) +
geom_histogram(binwidth = 0.05, color = "white") +
facet_grid(shade ~ .) +
scale_fill_manual(values = c("sunny" = "goldenrod",
"shady" = "forestgreen")) +
labs(x = "Leaf Mass (g)", y = "Count") +
theme_minimal() +
theme(legend.position = "none")
facet_grid(shade ~ .) stacks panels in rows; facet_grid(. ~ shade) arranges them in columns. means “nothing on this side of the grid”facet_grid() really shines with two grouping variables (rows ~ cols) — we only have one (shade) today, but you’ll use the two-variable form the moment your data has a second grouping columnNote
💡 facet_wrap vs facet_grid
Use facet_wrap() for one grouping variable — it wraps panels into a grid automatically. Use facet_grid() when you have two variables and want rows and columns to each mean something specific.
🛑 Do Activity Part 3 now
Build both a facet_wrap() and a facet_grid() version of the histogram. Predict, type, run.
stat_summary()We will cover: layering a group mean and its standard error directly on top of raw data — no summary table needed.
🛑 After this chunk you will do Activity Parts 4–5.
stat_summary() — the Mean as a Layer# stat_summary() computes the mean straight from the data
leaf_df %>%
ggplot(aes(x = shade, y = mass_g, color = shade)) +
geom_jitter(width = 0.15, alpha = 0.4, size = 2) +
stat_summary(fun = mean, geom = "point", size = 4,
color = "black") +
scale_color_manual(values = c("sunny" = "goldenrod",
"shady" = "forestgreen")) +
labs(
title = "Leaf Mass with Group Means",
x = "Shade", y = "Leaf Mass (g)"
) +
theme_minimal() +
theme(legend.position = "none")
stat_summary() computes a statistic from the raw data and draws it — no pre-built summary table neededfun = mean — any function that returns one number (mean, median, max, …)geom = "point" — draw that number as a pointgeom_jitter() before stat_summary() so the mean sits on top of the raw points, not buried under themstat_summary() — Mean ± SENote
🔮 Predict first: fun.data = mean_se returns three numbers per group instead of one (the mean, and the bar’s top and bottom). What do you think those three are called?
# fun.data = returns THREE values: y, ymin, ymax -------
leaf_mean_se_plot <- leaf_df %>%
ggplot(aes(x = shade, y = mass_g, color = shade)) +
geom_jitter(width = 0.15, alpha = 0.35, size = 2) +
stat_summary(fun = mean, geom = "point", size = 4) +
stat_summary(
fun.data = mean_se,
geom = "errorbar",
width = 0.15,
linewidth = 0.9
) +
scale_color_manual(values = c("sunny" = "goldenrod",
"shady" = "forestgreen")) +
labs(
title = "Mean ± SE Leaf Mass by Shade",
x = "Shade", y = "Leaf Mass (g)"
) +
theme_minimal() +
theme(legend.position = "none")
leaf_mean_se_plot
Two stat_summary() layers, stacked:
| call | what it draws |
|---|---|
fun = mean |
a point at the mean |
fun.data = mean_se |
error bars for ± 1 SE |
alpha = 0.35)📖 Whitlock & Schluter, Ch. 3 — Describing Data
🛑 Go to Activity 3 — Parts 4–5
Close the slides. Build your own mean ± SE plot, then save it to figures/.
leaf_df %>% ggplot(...)scale_color_manual() / scale_fill_manual() — colors you choose, not colors ggplot guessesfacet_wrap(~var) — one panel per group, wraps automaticallyfacet_grid(rows ~ cols) — strict grid, best with two grouping variablesstat_summary() — mean and mean ± SE, drawn directly from raw data