Activity 04 - GGPlot II

New data, safe zooming, distributions, themes, and patchwork

ggplot
tidyverse

Hands-on companion to the GGPlot II lecture. Rebuild your plotting recipe on a new dataset, compare coord_cartesian() to the ylim() trap, build a histogram and density plot, apply a custom theme file, and combine panels with patchwork.

Author

Bill Perry

Published

October 1, 2026

New Data, Safe Zooming, Distributions, and Style

Recap from Activity 03

  • Piped leaf_df straight into ggplot()
  • Assigned exact colors with scale_fill_manual() / scale_color_manual()
  • Split plots with facet_wrap() and facet_grid()
  • Built a mean ± SE plot with stat_summary()

Today’s Objectives

  1. Rebuild the XY-scatter and boxplot+jitter recipe on a dataset you’ve never seen — palmerpenguins
  2. Zoom a plot safely with coord_cartesian() — and recognise the ylim() / scale_y_continuous(limits = ) trap
  3. Build a geom_histogram() and a geom_density() of the same variable
  4. Source a reusable theme file and tweak a font and a color
  5. Combine two plots into one figure with patchwork and plot_annotation()

How this activity works

You are building one R script this whole class, and you turn it in. Create it now: scripts/04_ggplot_2.R.

  • Every line you run goes in that script — not in the Console, not typed into this page. Type it, don’t paste it.
  • Start every code chunk in your script with a short # comment that says what it does. The comments are part of the grade.
  • The top of your script, in this order: a title comment, then your library() calls, then the line that loads the data into penguins_df.
  • Code marked ▶ Run this is typed into your script exactly as shown. Code marked ✏️ Your turn is a change you make in that same script and run.

🔮 Predict before you run

Before you run any ▶ Run this block, cover the output and predict what R will draw. Predicting is what turns “I saw it on a slide” into “I can write it.”


Part 1 · Load libraries and a new dataset

Tip📂 No download needed today

palmerpenguins ships its data inside the package — once it’s installed (it’s in install_packages.R from Week 1), library(palmerpenguins) is all you need. No file to put in data/.

▶ Type this at the very top of scripts/04_ggplot_2.R:

# ---- Activity 04: GGPlot II ------------------------------
# your name, today's date

# ---- Libraries -------------------------------------------
library(tidyverse)      # dplyr + ggplot2
library(palmerpenguins) # penguins data, Palmer Station LTER
# ---- Load data --------------------------------------------
penguins_df <- penguins %>%
  drop_na(species, bill_length_mm, bill_depth_mm,
          flipper_length_mm, body_mass_g)

glimpse(penguins_df)     # look at the data right after loading

Part 2 · The same two plot types, on new data

🔮 Predict first: flipper length and body mass should move together. Will the three species separate into three clouds, or overlap into one?

▶ Run this — the exact scatter recipe you used on leaf_df:

# scatter of body mass vs flipper length, colored by species
penguins_df %>%
  ggplot(aes(x = flipper_length_mm, y = body_mass_g,
             color = species)) +
  geom_point(alpha = 0.7, size = 2) +
  labs(
    title = "Body Mass vs. Flipper Length",
    x = "Flipper Length (mm)", y = "Body Mass (g)",
    color = "Species"
  ) +
  theme_minimal()

▶ Run this — the exact boxplot+jitter recipe you used on leaf_df:

# boxplot + jitter of body mass by species
penguins_df %>%
  ggplot(aes(x = species, y = body_mass_g, fill = species)) +
  geom_boxplot(alpha = 0.6, outlier.shape = NA) +
  geom_jitter(width = 0.15, alpha = 0.4) +
  labs(x = "Species", y = "Body Mass (g)",
       title = "Body Mass by Species") +
  theme_minimal() +
  theme(legend.position = "none")

✏️ Your turn — in your script: Swap body_mass_g for bill_length_mm in the boxplot above (x stays species). Run it.

Which species has the longest bills?
Tip

💡 Finding a dataset like this on your own

You’ll pick one of these for the Extension at the end of this activity.

✅ You can already do a lot: the same two plot recipes work on any tidy dataset. Parts 3–6 add new skills — zooming, distributions, themes, and combining panels.


Part 3 · Zoom in safely with coord_cartesian()

Gentoo penguins run heavier than the other two species. coord_cartesian() zooms the view and leaves the data alone — every penguin still counts toward each box’s median and quartiles.

▶ Run this:

# zoom the view; all 342 penguins still contribute to each box
penguins_df %>%
  ggplot(aes(x = species, y = body_mass_g, fill = species)) +
  geom_boxplot(alpha = 0.6, outlier.shape = NA) +
  coord_cartesian(ylim = c(3000, 5000)) +
  labs(title = "Zoomed with coord_cartesian()",
       x = "Species", y = "Body Mass (g)") +
  theme_minimal() +
  theme(legend.position = "none")

✏️ Your turn — in your script: Change the window to coord_cartesian(ylim = c(3500, 4500)) and run it. The view gets tighter — do the box shapes themselves change?

Did the boxes change shape?
Warning

⚠️ The classic mistake: ylim() / scale_y_continuous(limits = ) delete rows

ylim() looks like it just zooms, but it silently drops every row outside the range before anything is calculated. A box built after ylim() is fit only on the penguins that survived — so the median and quartiles can shift, with no warning that data went missing. scale_y_continuous(limits = c(3000, 5000)) is the exact same trap written out in full.

Rule: zoom with coord_cartesian(). Only reach for ylim() / scale_y_continuous(limits = ) when you actually want those rows removed from the analysis.

▶ Run this once to see the trap:

# same plot, but ylim() drops the outside
# rows FIRST — watch the Gentoo box
penguins_df %>%
  ggplot(aes(x = species, y = body_mass_g, fill = species)) +
  geom_boxplot(alpha = 0.6, outlier.shape = NA) +
  ylim(3000, 5000) +
  labs(title = "ylim() — boxes recomputed on fewer penguins",
       x = "Species", y = "Body Mass (g)") +
  theme_minimal() +
  theme(legend.position = "none")

▶ Run this to see exactly how many rows ylim() silently drops:

# count penguins inside vs outside the 3000-5000 g window
nrow(penguins_df)

penguins_df %>%
  filter(body_mass_g >= 3000, body_mass_g <= 5000) %>%
  nrow()

penguins_df %>%
  filter(body_mass_g >= 3000, body_mass_g <= 5000) %>%
  count(species)
How many penguins did ylim(3000, 5000)
drop, and which species lost the most?

Part 4 · Two views of one distribution

▶ Run this:

# histogram — bars of counts per bin
penguins_df %>%
  ggplot(aes(x = body_mass_g)) +
  geom_histogram(binwidth = 200, fill = "steelblue",
                 color = "white") +
  labs(x = "Body Mass (g)", y = "Count",
       title = "Distribution of Body Mass") +
  theme_minimal()
# density — a smoothed curve of the same shape
penguins_df %>%
  ggplot(aes(x = body_mass_g)) +
  geom_density(fill = "steelblue", alpha = 0.5) +
  labs(x = "Body Mass (g)", y = "Density",
       title = "Density of Body Mass") +
  theme_minimal()

✏️ Your turn — in your script: Rebuild both plots for flipper_length_mm instead of body_mass_g (pick a sensible new binwidth). Run them.

Does flipper length look more or less
"bumpy" (multiple humps) than body mass?
Tip

Coming soon: in the Summary Statistics lecture we’ll come back to this exact density curve and shade in the 95% range underneath it. Today, just get comfortable building both plot types.


Part 5 · A reusable theme file

Tip📂 Get the theme file

r_themes_for_3_sizes.R → put it in a new themes/ folder in your project.

▶ Run this:

# source() loads theme_small(), theme_regular(), theme_large()
source("themes/r_themes_for_3_sizes.R")

body_mass_plot <- penguins_df %>%
  ggplot(aes(x = species, y = body_mass_g, fill = species)) +
  geom_boxplot(alpha = 0.6, outlier.shape = NA) +
  labs(x = "Species", y = "Body Mass (g)",
       title = "Body Mass by Species") +
  theme_regular() +
  theme(legend.position = "none")

body_mass_plot

Tweak a font and a color

body_mass_plot_styled <- penguins_df %>%
  ggplot(aes(x = species, y = body_mass_g, fill = species)) +
  geom_boxplot(alpha = 0.6, outlier.shape = NA) +
  scale_fill_manual(values = c(
    "Adelie" = "darkorange", "Chinstrap" = "purple",
    "Gentoo" = "cyan4"
  )) +
  labs(x = "Species", y = "Body Mass (g)",
       title = "Body Mass by Species") +
  theme_regular() +
  theme(
    legend.position = "none",
    plot.title = element_text(size = rel(1.4), face = "bold")
  )

body_mass_plot_styled

✏️ Your turn — in your script: Swap in theme_small() instead of theme_regular() and run it. What changes besides the text size?

What else changed?

Part 6 · Combine panels with patchwork

▶ Run this:

library(patchwork)   # combine ggplot objects with + and /

xy_plot <- penguins_df %>%
  ggplot(aes(x = flipper_length_mm, y = body_mass_g,
             color = species)) +
  geom_point(alpha = 0.7) +
  labs(x = "Flipper Length (mm)", y = "Body Mass (g)",
       color = "Species") +
  theme_regular()

combined_plot <- xy_plot + body_mass_plot_styled +
  plot_annotation(
    title = "Penguin Body Mass, Two Ways",
    tag_levels = "A"
  )

combined_plot
# save the combined figure
ggsave("figures/penguin_combined.png",
       plot = combined_plot, width = 9, height = 4.5,
       units = "in", dpi = 300)

✏️ Your turn — in your script: Change xy_plot + body_mass_plot_styled to xy_plot / body_mass_plot_styled (stack instead of side-by-side) and run it.

Which layout — side by side or stacked
— reads better for these two panels?

Part 7 · Review and checkpoint

At this point you should be able to:

✏️ Your turn — before you move on: Run your entire script top to bottom with Ctrl/Cmd + Shift + Enter (Source). Does it complete without errors?

Ran cleanly?  Y / N
If not, what error appeared:

What your project folder should contain

leaf_project/
├── figures/
│   └── penguin_combined.png    <- Part 6
├── themes/
│   └── r_themes_for_3_sizes.R  <- Part 5
└── scripts/
    └── 04_ggplot_2.R           <- your script (turn this in)

Extension — out of class (~20 min)

Add this to the bottom of scripts/04_ggplot_2.R and turn it in with the rest. Put written answers in # comments right under the code they go with.

E1 · Find your own dataset (5 pts)

Pick one option:

  • A Kaggle dataset (UFO sightings, Bigfoot sightings, or any other you find) — download the .csv into your project’s data/ folder and load it with read_csv()
  • library(FSA) — pick one of its built-in fish datasets (e.g. ?FSA::Mirex) the same way you loaded penguins

Build one simple plot from it — either an XY scatter (geom_point()) if you have two numeric columns, or a categorical plot (geom_boxplot() + geom_jitter(), or a bar chart) if you have a grouping column. Label your axes.

E2 · Explain your choice (5 pts)

In two or three sentences: why did you pick an XY plot or a categorical plot for this dataset — what did the columns look like that made that the right choice? Then note one thing about this new dataset that surprised you or that you’d want to check before trusting it (missing values? a weird outlier? too few rows in a group?).


Getting unstuck

  1. Read the error message out loud. R usually names the line and the problem.
  2. Box shape looks wrong after zooming? Check whether you used ylim() / scale_y_continuous(limits = ) (deletes data) instead of coord_cartesian() (keeps it).
  3. source() can’t find the theme file? Check that themes/r_themes_for_3_sizes.R is spelled exactly right and sits in a themes/ folder at your project root.
  4. patchwork plots look squished? Save at a wider width in ggsave() — combined figures usually need more horizontal room than a single panel.
  5. Cheat sheets — https://posit.co/resources/cheatsheets/
  6. Bring the exact error (copy-paste it) to class, Canvas, or office hours.

End of the GGplot II activity. Next: wrangling with filter, select, mutate, and arrange.