New data, safe zooming, distributions, themes, and patchwork
2026-10-01
leaf_df %>% ggplot(...)scale_color_manual() / scale_fill_manual() — colors you choosefacet_wrap() and facet_grid() — one panel per groupstat_summary() — mean and mean ± SE, drawn from raw dataNote
✅ Key idea from Lecture 03
You already know how to build a grouped, colored, faceted figure — but only on leaf_df. Today’s real test: can you do it on data you’ve never seen before?
coord_cartesian()patchworkTip
🖐 Try it yourself
By the end you will recognize the single most common ggplot mistake — zooming with ylim() / scale_y_continuous(limits = ) — before you ever make it yourself.
References:
Both PDFs are in the course readings/ folder.
This lecture runs in five short chunks. After each chunk you switch to the activity and type the code yourself into your R script.
Note
✅ Why bother?
ylim() and coord_cartesian() produce plots that look identical for a simple scatter — the difference only shows up once you compute something (a boxplot, a smooth). Typing both yourself and comparing outputs is the only way this sticks.
We will cover: loading palmerpenguins, and rebuilding the two plot types you already know on data you’ve never seen.
🛑 After this chunk you will do Activity Parts 1–2.
Rows: 342
Columns: 8
$ species <fct> Adelie, Adelie, Adelie, Adelie, Adeli…
$ island <fct> Torgersen, Torgersen, Torgersen, Torg…
$ bill_length_mm <dbl> 39.1, 39.5, 40.3, 36.7, 39.3, 38.9, 3…
$ bill_depth_mm <dbl> 18.7, 17.4, 18.0, 19.3, 20.6, 17.8, 1…
$ flipper_length_mm <int> 181, 186, 195, 193, 190, 181, 195, 19…
$ body_mass_g <int> 3750, 3800, 3250, 3450, 3650, 3625, 4…
$ sex <fct> male, female, female, female, male, f…
$ year <int> 2007, 2007, 2007, 2007, 2007, 2007, 2…
penguins comes inside the package — no read_excel() needed, just library(palmerpenguins)drop_na() removes the handful of rows missing a measurement — same idea as the NAs in leaf_dfNote
🔮 Predict first: flipper length and body mass should move together — longer flippers, heavier birds. Will three species separate into three clouds, or overlap into one?
# The exact same recipe you used on leaf_df ------------
penguins_df %>%
ggplot(aes(x = flipper_length_mm, y = body_mass_g,
color = species)) +
geom_point(alpha = 0.7, size = 2) +
labs(
title = "Body Mass vs. Flipper Length",
x = "Flipper Length (mm)",
y = "Body Mass (g)",
color = "Species"
) +
theme_minimal()
aes() + geom_point() + labs() + theme_minimal() you already typed for leaves# boxplot + jitter, exactly like the leaf_df pattern ---
penguins_df %>%
ggplot(aes(x = species, y = body_mass_g, fill = species)) +
geom_boxplot(alpha = 0.6, outlier.shape = NA) +
geom_jitter(width = 0.15, alpha = 0.4) +
labs(x = "Species", y = "Body Mass (g)",
title = "Body Mass by Species") +
theme_minimal() +
theme(legend.position = "none")
Tip
💡 Where would you find a dataset like this on your own?
library(palmerpenguins) or library(FSA) (fish data, Derek Ogle) — real data bundled inside an R package, zero downloading.csv, load it with read_csv() like any other fileBoth a simple XY numeric plot and a simple categorical plot (like today’s two) work on almost any tidy dataset you find.
🛑 Do Activity Parts 1–2 now
Load palmerpenguins, then rebuild the XY scatter and the boxplot+jitter — same recipe, new data. Predict, type, run.
We will cover: coord_cartesian() to zoom — and the ylim() / scale_y_continuous(limits = ) trap to avoid.
🛑 After this chunk you will do Activity Part 3.
coord_cartesian() — Zoom the View, Keep the Data# coord_cartesian() zooms the VIEW, keeps all the data -
penguins_df %>%
ggplot(aes(x = species, y = body_mass_g, fill = species)) +
geom_boxplot(alpha = 0.6, outlier.shape = NA) +
coord_cartesian(ylim = c(3000, 5000)) +
labs(x = "Species", y = "Body Mass (g)",
title = "Zoomed with coord_cartesian()") +
theme_minimal() +
theme(legend.position = "none")
ylim() / scale_y_continuous(limits = ) Delete Rows FirstNote
🔮 Predict first: Gentoo penguins run heavier than the other two species — many sit above 5000 g. If we drop every penguin outside 3000–5000 g before drawing the boxes, what happens to the Gentoo box specifically?
# ylim() DELETES rows outside the range first ----------
penguins_df %>%
ggplot(aes(x = species, y = body_mass_g, fill = species)) +
geom_boxplot(alpha = 0.6, outlier.shape = NA) +
ylim(3000, 5000) +
labs(x = "Species", y = "Body Mass (g)",
title = "ylim() — boxes recomputed on fewer penguins") +
theme_minimal() +
theme(legend.position = "none")
Warning
⚠️ Watch out!
ylim() — and its long form scale_y_continuous(limits = c(3000, 5000)) — silently delete every row outside the range before anything is computed. Compare the Gentoo box above to the coord_cartesian() version: same window, different box.
Rule: zoom with coord_cartesian(). Only use ylim() / scale_y_continuous(limits = ) when you actually want those rows gone from the analysis.
🛑 Do Activity Part 3 now
Zoom with coord_cartesian(), then run the ylim() version and count exactly how many penguins it drops. Predict, type, run.
We will cover: geom_histogram() and geom_density() — the two standard ways to show how one numeric variable is distributed.
🛑 After this chunk you will do Activity Part 4.
geom_histogram() — Bars of Countsbinwidth sets the width of each bar in the units of x — here, 200 g per binbinwidth values before settling on onegeom_density() — a Smooth Curve Instead of BarsTip
Coming soon: in the Summary Statistics lecture we’ll come back to this exact curve and shade in the 95% range underneath it. Today, just get comfortable building both plot types.
🛑 Do Activity Part 4 now
Build a histogram and a density plot of the same variable. Predict, type, run.
We will cover: sourcing a custom theme file and applying it — no need to build one from scratch today.
🛑 After this chunk you will do Activity Part 5.
source() a Theme File Instead of Repeating theme()# source() loads functions from another .R file -------
source("themes/r_themes_for_3_sizes.R")
penguins_df %>%
ggplot(aes(x = species, y = body_mass_g, fill = species)) +
geom_boxplot(alpha = 0.6, outlier.shape = NA) +
labs(x = "Species", y = "Body Mass (g)",
title = "Body Mass by Species") +
theme_regular() +
theme(legend.position = "none")
source("themes/r_themes_for_3_sizes.R") runs that file once and loads theme_small(), theme_regular(), theme_large() into your sessiontheme_regular()# theme() and scale_fill_manual() still work as usual --
penguins_df %>%
ggplot(aes(x = species, y = body_mass_g, fill = species)) +
geom_boxplot(alpha = 0.6, outlier.shape = NA) +
scale_fill_manual(values = c(
"Adelie" = "darkorange", "Chinstrap" = "purple",
"Gentoo" = "cyan4"
)) +
labs(x = "Species", y = "Body Mass (g)",
title = "Body Mass by Species") +
theme_regular() +
theme(
legend.position = "none",
plot.title = element_text(size = rel(1.4), face = "bold")
)
theme() call — + theme(...) after it overrides individual pieces, same as any built-in themescale_fill_manual() still controls colors; the theme controls everything else (fonts, lines, margins)Note
✅ Key idea
One theme file, sourced at the top of every script, keeps every figure in a project looking consistent — change the file once, every plot updates.
🛑 Do Activity Part 5 now
Download the theme file, source it, and apply theme_regular() to one of your plots. Then tweak a font size and a color. Predict, type, run.
patchworkWe will cover: patchwork — putting two ggplot objects side by side as one figure.
🛑 After this chunk you will do Activity Part 6.
patchwork — Two Plots, One Figurelibrary(patchwork) # combine ggplot objects with + and /
xy_plot <- penguins_df %>%
ggplot(aes(x = flipper_length_mm, y = body_mass_g,
color = species)) +
geom_point(alpha = 0.7) +
labs(x = "Flipper Length (mm)", y = "Body Mass (g)",
color = "Species") +
theme_regular()
box_plot <- penguins_df %>%
ggplot(aes(x = species, y = body_mass_g, fill = species)) +
geom_boxplot(alpha = 0.6, outlier.shape = NA) +
labs(x = "Species", y = "Body Mass (g)") +
theme_regular() +
theme(legend.position = "none")
xy_plot + box_plot +
plot_annotation(
title = "Penguin Body Mass, Two Ways",
tag_levels = "A"
)
+ places two patchwork-ready plots side by side; / stacks themplot_annotation() adds one title and A/B/C tags across the whole combined figure, not any single panel🛑 Go to Activity 4 — Part 6
Close the slides. Combine two of your plots with patchwork and plot_annotation(), then save the combined figure.
palmerpenguins, FSA, or a .csv from Kagglecoord_cartesian() zooms safely; ylim() / scale_y_continuous(limits = ) silently deletes rows firstgeom_histogram() and geom_density() — two views of one distributionsource() a theme file once, then call theme_regular() (or _small() / _large()) everywherepatchwork (+ / /) combines saved plots; plot_annotation() titles and tags the whole figureReferences:
Up next — Lecture 05, Wrangling:
filter(), select(), mutate(), arrange()