Week 2: Your First ggplot

From base R to the grammar of graphics

Introduction

In the content session, you learned about summarizing data (mean, median, spread), honest data visualization, and the grammar of graphics — the idea that a plot is built from layers: data, aesthetics, and geometries.

These exercises let you practise the ggplot2 syntax before the application session, where you’ll build briefing-quality figures.

Goals:

  • Build a histogram with geom_histogram()
  • Build a plot with colour mapping and a legend
  • Get comfortable with the ggplot(data, aes(...)) + geom_*() pattern

The code boxes start empty, except for comments to guide you through the code that’s expected.

Work one comment at a time. If you get stuck:

  1. Hint 1 restates what you are trying to do.
  2. Hint 2 sketches the shape the code might take.
  3. Hint 3 gives you the code with the key pieces blanked out.
  4. Solution gives you the lot.

Take a real attempt before opening a hint: you will learn by doing, not by reading. The demonstrators are here to help.

Three things worth knowing:

  • The boxes share one R session, exactly like a script. Anything you create in Exercise 1 is still there in Exercise 2.
  • Variable names are up to you. The last thing your block prints is what gets checked, so end each block with the answer.
  • Exercises have extensions underneath them. If you finish early, give them a stab. These are designed to equip you to tackle real coding problems you might face, using the documentation. We can help – but check the manuals first!
TipStuck?

Ask on the class board, in Code Q&A. Someone else is likely stuck on the same thing, so asking in public helps them too. Demonstrators check it most weekdays, but they give classmates a chance to answer first.

The board is private to this class, so sign in to GitHub first. If the link shows “404 – page not found”, you’re not signed in, or you haven’t accepted the Classroom 50 invitation yet.

In GEOL1151 you built plots imperatively — one command per element:

plt.plot(year, bio, label="Bioenergy")
plt.xlabel("Year")
plt.legend()
plt.show()

ggplot2 works declaratively: you describe what the plot is, and ggplot2 figures out how to draw it.

Concept matplotlib ggplot2
Start a plot plt.figure() ggplot(data, aes(...))
Map columns to axes plt.plot(x, y) aes(x = year, y = value)
Choose a geometry plt.scatter(...) + geom_point()
Add layers Separate calls + operator
Labels plt.xlabel(...) + labs(x = ...)
Legend Manual label= args Automatic from colour/fill mapping

The biggest shift: aes() does the mapping. You tell ggplot2 which columns mean what (x position, y position, colour), then add a geometry that knows how to draw those mappings. No quoting column names inside aes() — just bare names like aes(x = year, y = bioenergy_twh).

Exercise 1: The shape of the data

Last week you plotted bioenergy against time. Now throw time away and look only at the values: how are the 26 yearly figures distributed?

Build a histogram of bioenergy_twh, labelled well enough that someone else could read it.

NoteHint 1

Every ggplot has the same three parts: the data, the mapping from columns to visual properties, and at least one geometry that draws something.

A histogram is unusual in needing only an x mapping. The y axis is a count that ggplot works out from the data — which is why you never name a y column for one.

NoteHint 2

The shape of it:

ggplot(data, aes(x = column)) +
  geom_histogram(bins = 10) +
  labs(x = "...", y = "...", title = "...")

Note the + at the end of each line, not the start of the next one. R reads line by line, so a line that could be complete is treated as complete — a leading + gives you a plot with no layers and no error.

NoteHint 3
ggplot(______, aes(x = ______)) +
  geom_histogram(bins = 10) +
  labs(x = "Bioenergy generation (TWh)",
       y = "______",
       title = "Distribution of UK bioenergy generation (2000-2025)")
TipSolution
ggplot(elec, aes(x = bioenergy_twh)) +
  geom_histogram(bins = 10) +
  labs(x = "Bioenergy generation (TWh)",
       y = "Count",
       title = "Distribution of UK bioenergy generation (2000-2025)")

The distribution is lumpy and right-skewed: a cluster of low values from the early 2000s, when bioenergy was small, and a second group above 30 TWh from the years after 2015.

That shape is a warning. These 26 numbers are not 26 measurements of the same thing — they are a quantity that grew. Summarizing them with a mean describes a year that never happened. The histogram is doing something more useful than showing a distribution: it is telling you that a distribution is the wrong model.

Extension: the bins are an argument

Optional, and all in the documentation: ?geom_histogram, ?geom_density, ?geom_violin, ?geom_rug.

  1. Redraw with 3 bins, then 30. One version says “two clusters”, one says “steady spread”, one says almost nothing. Same data, three stories — and nothing in the output marks any of them as wrong.
  2. bins and binwidth are two ways to say the same thing. Which would you use if you wanted charts of two different datasets to be comparable, and why?
  3. Show the same distribution without geom_histogram(). At least three other geoms will do it. Which of them still shows the gap in the middle?
  4. Add the individual observations along the bottom axis, so a reader can see there are only 26 of them.

Quick check

For these 26 values the mean and the median are some way apart. Print the mean minus the median.

Then look back at your histogram and work out which side the difference comes from.

TipSolution
mean(elec$bioenergy_twh) - median(elec$bioenergy_twh)

4.88 TWh: the mean is 21.3, the median 16.4.

When the mean exceeds the median, the high values sit further from the centre than the low ones. Here that is not an outlier — it is the growth trend, showing up as skew because we threw the time axis away.

Exercise 2: Two fuels, one legend

Now put bioenergy and coal on the same axes, with a legend built by ggplot rather than by hand.

The trick is that elec is wide: one column per fuel. ggplot expects one column to map to one aesthetic, so with data shaped like this you need one layer per fuel — and you create the legend by mapping colour to a constant string inside aes().

It is clumsy. Notice how clumsy, because the application session shows you the fix.

NoteHint 1

Anything inside aes() is a mapping and gets a legend; anything outside it is a setting and does not.

So geom_point(colour = "red") makes red points and no legend, while geom_point(aes(colour = "red")) makes a legend entry labelled “red” and lets ggplot pick the colour. That looks like a bug the first time you see it. Here we are using it deliberately.

NoteHint 2

The shape of it:

ggplot(elec, aes(x = year)) +
  geom_point(aes(y = first_column, colour = "First label")) +
  geom_point(aes(y = second_column, colour = "Second label")) +
  labs(x = "...", y = "...", colour = "...")

Each layer inherits x from the top-level aes() and supplies its own y. The labs(colour = ) argument sets the legend’s title, not its entries — those come from the strings you mapped.

NoteHint 3
ggplot(elec, aes(x = ______)) +
  geom_point(aes(y = bioenergy_twh, colour = "Bioenergy")) +
  geom_point(aes(y = ______, colour = "______")) +
  labs(x = "Year",
       y = "Generation (TWh)",
       colour = "Fuel",
       title = "UK electricity: bioenergy vs coal")
TipSolution
ggplot(elec, aes(x = year)) +
  geom_point(aes(y = bioenergy_twh, colour = "Bioenergy")) +
  geom_point(aes(y = coal_twh, colour = "Coal")) +
  labs(x = "Year",
       y = "Generation (TWh)",
       colour = "Fuel",
       title = "UK electricity: bioenergy vs coal")

Bioenergy rises while coal collapses — the Week 1 picture, with a legend ggplot wrote for you.

Now count the work. Two fuels took two nearly identical layers. Six fuels would take six, and adding a seventh would mean editing the code rather than the data. That is the signal that the data are the wrong shape for the question, and reshaping them is the first thing you will do in the application session.

Extension: make the legend earn its place

Optional. ?scale_colour_manual, ?theme, ?guides, ?geom_line, ?scale_y_continuous.

  1. Connect the dots: add a line layer for each fuel so the trends read at a glance. You will need the colour mapping on the lines too, or the legend will disagree with the chart.
  2. ggplot picked the colours. Choose them yourself, so that coal is dark and biomass is green — then ask whether that choice is informative or persuasive.
  3. Move the legend inside the plot area, where the empty top-right corner is. Then decide whether that is an improvement or just cleverness.
  4. Bioenergy in 2025 is 41 TWh and coal is zero. On this axis the coal line is flat against the bottom for years. Redraw with a log y axis (ggplot will warn that it has dropped the zero) and say what it reveals — and what it hides from a reader who does not notice the axis.

Next steps

You’ll use these ggplot2 skills throughout the application session, where you’ll build the figures for your policy briefing.