Week 2: Making the Biomass Case

Building briefing-quality figures with ggplot2

Introduction

You’ve explored the biomass data and learned the ggplot2 syntax. Now you’ll produce briefing-quality figures — the kind you might put in a policy document to make the case about biomass electricity.

Today’s goals:

  • Reshape wide data into the shape ggplot actually wants
  • Compare CO₂ emissions across fuels under different assumptions
  • See where the UK’s wood pellets come from
  • Track bioenergy’s share of total electricity over time
  • Think about how figure choices shape the story

The code boxes start empty, except for comments to guide you through the code that’s expected.

Work one comment at a time. If you get stuck:

  1. Hint 1 restates what you are trying to do.
  2. Hint 2 sketches the shape the code might take.
  3. Hint 3 gives you the code with the key pieces blanked out.
  4. Solution gives you the lot.

Take a real attempt before opening a hint: you will learn by doing, not by reading. The demonstrators are here to help.

Three things worth knowing:

  • The boxes share one R session, exactly like a script. Anything you create in Exercise 1 is still there in Exercise 2.
  • Variable names are up to you. The last thing your block prints is what gets checked, so end each block with the answer.
  • Exercises have extensions underneath them. If you finish early, give them a stab. These are designed to equip you to tackle real coding problems you might face, using the documentation. We can help – but check the manuals first!
TipStuck?

Ask on the class board, in Code Q&A. Someone else is likely stuck on the same thing, so asking in public helps them too. Demonstrators check it most weekdays, but they give classmates a chance to answer first.

The board is private to this class, so sign in to GitHub first. If the link shows “404 – page not found”, you’re not signed in, or you haven’t accepted the Classroom 50 invitation yet.

Exercise 1: One layer, not six

In the content session you drew two fuels with two nearly identical geom_point() layers, and the clumsiness was the point. elec is wide: one column per fuel. ggplot wants long data: one row per observation, with a column saying which fuel it is.

Reshape elec so that coal, gas, wind and bioenergy sit in a single twh column, tagged by a fuel column. Then draw all four with one geom_line().

Finish the block with your reshaped data frame so it can be checked.

NoteHint 1

“Long” means one row per measurement. Here a measurement is one fuel in one year, so 26 years and 4 fuels gives 104 rows and three columns: the year, which fuel, and how many TWh.

pivot_longer() does this. You tell it which columns to stack, what to call the column holding their names, and what to call the column holding their values.

NoteHint 2

The shape of it:

long <- pivot_longer(data,
                     cols = c(col_a, col_b, ...),
                     names_to = "fuel",
                     values_to = "twh")

ggplot(long, aes(x = year, y = twh, colour = fuel)) +
  geom_line() +
  labs(...)

Two traps. cols must name only the fuels — sweep in total_twh and you will draw the total as though it were a fifth fuel. And the names you get are the column names, coal_twh and friends, not tidy labels; the extension deals with that.

NoteHint 3
elec_long <- pivot_longer(
  elec,
  cols = c(coal_twh, gas_twh, wind_twh, ______),
  names_to = "______",
  values_to = "twh"
)

ggplot(elec_long, aes(x = year, y = ______, colour = ______)) +
  geom_line() +
  labs(x = "Year",
       y = "Generation (TWh)",
       colour = "Fuel",
       title = "UK electricity generation by fuel type",
       caption = "Source: DUKES 2026")

elec_long
TipSolution
elec_long <- pivot_longer(
  elec,
  cols = c(coal_twh, gas_twh, wind_twh, bioenergy_twh),
  names_to = "fuel",
  values_to = "twh"
)

ggplot(elec_long, aes(x = year, y = twh, colour = fuel)) +
  geom_line() +
  labs(x = "Year",
       y = "Generation (TWh)",
       colour = "Fuel",
       title = "UK electricity generation by fuel type",
       caption = "Source: DUKES 2026")

elec_long

The big stories: coal’s collapse from about 120 TWh to zero, wind’s rise to about 86 TWh, closing on gas, which stays dominant throughout, and bioenergy’s growth real but modest next to wind.

The code is also the lesson. Adding nuclear and solar now means adding two names to cols — nothing else changes, and the legend updates itself. Wide data is convenient for a human reading a table; long data is what a plotting library can reason about. Most of the friction people blame on ggplot2 is really data in the wrong shape.

Extension: the labels are still wrong

Optional. ?scale_colour_discrete, ?labs, ?sub, ?factor, ?scale_x_continuous, ?annotate.

  1. The legend says bioenergy_twh. Fix it so it says Bioenergy, and do it in at least two different ways — one that edits the data, one that only changes how the plot is labelled. Which would you rather maintain?
  2. The legend is ordered alphabetically, so it does not match the order the lines appear in. Reorder it to match, and note that this is a question about factor levels, not about ggplot.
  3. Coal and wind cross somewhere around 2016. Mark that on the chart.
  4. Add nuclear and solar. If that took you more than ten seconds, look again at where your fuel names are written down.

Exercise 2: The chart that makes the case

Different fuels emit very different amounts of CO₂ per unit of electricity — and for biomass, the answer depends entirely on what you agree to count. factors holds emission factors under several accounting rules.

Build the bar chart a government press office would publish: the official factor for each of the five fuels. Biomass, coal and gas have a scenario called official; wind and solar have one called lifecycle, because they have no combustion to account for.

End the block with the data frame you plotted.

NoteHint 1

You need rows where the scenario is one of two acceptable values. Testing == "official" alone throws away wind and solar; you want a test that is TRUE for either name.

Look at factors first and count how many rows each fuel has. Biomass has six; that asymmetry is the thing the filter has to handle.

NoteHint 2

The shape of it:

  • x %in% c("a", "b") is TRUE where x is either — the natural way to say “one of these”.
  • df[condition, ] keeps the matching rows and every column.
  • geom_col() draws bars at the height of your y column. geom_bar() is the one that counts rows instead, which is why using it here gives every fuel a bar of height 1.
NoteHint 3
official <- factors[factors$scenario %in% c("______", "______"), ]

ggplot(official, aes(x = fuel, y = ______)) +
  geom_col() +
  labs(x = "Fuel",
       y = "CO2 (kg per MWh)",
       title = "Official CO2 emission factors by fuel",
       caption = "Sources: DESNZ, IPCC, UNFCCC")

official
TipSolution
official <- factors[factors$scenario %in% c("official", "lifecycle"), ]

ggplot(official, aes(x = fuel, y = co2_kg_per_mwh)) +
  geom_col() +
  labs(x = "Fuel",
       y = "CO2 (kg per MWh)",
       title = "Official CO2 emission factors by fuel",
       caption = "Sources: DESNZ, IPCC, UNFCCC")

official

Under official accounting biomass emits zero — the same as wind and solar, against coal’s 910.

Every number in that chart is real, sourced and correctly transcribed. It is still misleading, because the zero is not a measurement. It is a convention: UNFCCC rules count biomass CO₂ in the country where the tree was cut, not where it was burned, so the UK’s figure is zero by construction. The chimney at those plants releases about 1000 kg/MWh, slightly more than coal.

Week 3 is about what happens to this chart when you count that.

Extension: the same data, arguing the other way

Optional.

  1. Redraw with biomass’s with_supply_chain factor in place of its official zero, leaving everything else identical. One number changes; the conclusion inverts. Put the two charts side by side and decide which caption each needs to be honest.
  2. A bar of height zero is invisible. A reader may read “no bar” as “no data” rather than “zero”. Label the bars with their values so the zero is unmistakable — ?geom_text.
  3. Colour the bars by whether the number includes combustion. You will need a new column to map to fill, and deciding which side biomass falls on is the interesting part.

Exercise 3: Where do the pellets come from?

The UK burns millions of tonnes of wood pellets a year, and grows almost none of them. imports records how much came from where.

Chart the 2024 imports by country of origin, ordered by size — a bar chart in alphabetical order makes a reader do work the chart should have done for them.

Finish by printing the USA’s share of 2024 imports, as a percentage.

NoteHint 1

Two separate jobs. The chart needs the bars sorted; the answer needs one country’s tonnage as a fraction of the total.

ggplot orders a character x axis alphabetically. To order it by size instead, the column has to become a factor whose levels are in the order you want — reorder() exists for exactly this.

NoteHint 2

The shape of it:

  • df[df$year == 2024, ] for the subset.
  • aes(x = reorder(origin, -import_kt), y = import_kt) sorts the bars largest first. The minus sign reverses the order; try it both ways.
  • For the share: df$import_kt[df$origin == "USA"] / sum(df$import_kt), times 100.
NoteHint 3
imports_2024 <- imports[imports$______ == 2024, ]

ggplot(imports_2024, aes(x = reorder(origin, -______), y = ______)) +
  geom_col() +
  labs(x = "Country of origin",
       y = "Imports (thousand tonnes)",
       title = "UK wood pellet imports by origin, 2024",
       caption = "Source: Forest Research / HMRC")

100 * imports_2024$import_kt[imports_2024$origin == "______"] /
  sum(imports_2024$______)
TipSolution
imports_2024 <- imports[imports$year == 2024, ]

ggplot(imports_2024, aes(x = reorder(origin, -import_kt), y = import_kt)) +
  geom_col() +
  labs(x = "Country of origin",
       y = "Imports (thousand tonnes)",
       title = "UK wood pellet imports by origin, 2024",
       caption = "Source: Forest Research / HMRC")

100 * imports_2024$import_kt[imports_2024$origin == "USA"] /
  sum(imports_2024$import_kt)

About 74% — 6,895 of 9,317 kt — from the United States alone. Latvia and the rest of the EU are distant runners-up.

Two things follow. The transatlantic voyage is unavoidable rather than incidental, which is what Week 3 puts a number on. And a policy the UK counts as zero-carbon depends on how forests are managed in a jurisdiction the UK has no authority over. “Other” at 186 kt, meanwhile, is a category with an 8000 km transport distance and no named country: worth asking who is in it before you cite the chart.

Extension: the map you cannot draw

Optional.

  1. Redraw the chart as shares of the total rather than absolute tonnage. Does it change which country you would talk about first?
  2. Draw the same origins for 2015 alongside 2024 and see how the mix has shifted. ?facet_wrap if you want them side by side.
  3. Weight each origin by how far its pellets travel — tonnage times transport_km — and redraw. Which origin moves furthest up the ranking, and what does that suggest about which number a shipping emissions estimate should be built on?
  4. Two categories, EU_other and Other, are not countries. Say in one sentence what you would have to assume to put them on a map, and whether you would be willing to defend that assumption.

Exercise 4: Bioenergy’s growing share

In Week 1 you worked out bioenergy’s share of UK electricity for 2025 alone. Now do it for every year and plot the series.

Finish by printing the year in which bioenergy’s share rose the most.

NoteHint 1

A column is just a vector, so you can build a new one from two others in one line and assign it into the data frame with $.

For the last step you want the largest year-on-year change, so you need the changes first, then the position of the largest, then the year at that position. Three steps, and the middle one is where the off-by-one lives.

NoteHint 2

The shape of it:

  • df$new <- 100 * df$a / df$b creates the column.
  • geom_line() + geom_point() draws the series and marks the years.
  • diff(x) gives the changes; which.max() gives the position of the biggest one.
  • diff() returns one value fewer than it was given. So element i of the differences is the change into year i + 1 — which means elec$year[-1] lines up with it, and elec$year does not.
NoteHint 3
elec$bio_pct <- 100 * elec$______ / elec$______

ggplot(elec, aes(x = year, y = ______)) +
  geom_line() +
  geom_point() +
  labs(x = "Year",
       y = "Bioenergy share (%)",
       title = "Bioenergy as a percentage of UK electricity",
       caption = "Source: DUKES 2026")

elec$year[-1][which.max(______(elec$bio_pct))]
TipSolution
elec$bio_pct <- 100 * elec$bioenergy_twh / elec$total_twh

ggplot(elec, aes(x = year, y = bio_pct)) +
  geom_line() +
  geom_point() +
  labs(x = "Year",
       y = "Bioenergy share (%)",
       title = "Bioenergy as a percentage of UK electricity",
       caption = "Source: DUKES 2026")

elec$year[-1][which.max(diff(elec$bio_pct))]

Bioenergy went from about 1% of UK electricity in 2000 to 14% in 2025, with the steepest climb over 2012–2020 as Drax converted its units; the share has been flat at about 14% since 2024.

The single largest jump, though, is 2024 — and it is only partly about biomass. UK total generation that year was 286 TWh, the lowest in the whole record, down from 377 in 2000. A rising share can mean you are generating more, or that everyone else is generating less, and this chart cannot tell the two apart. If you use a share in your briefing, you owe the reader the denominator.

Extension: decompose the rise

Optional.

  1. Plot bioenergy’s TWh and its percentage share on two charts, one above the other, and find a year where the two lines disagree about the direction of travel.
  2. Recompute the share holding total generation fixed at its 2000 value of 377 TWh. The gap between that curve and the real one is the part of the story that belongs to the denominator. How big is it in 2025?
  3. Some people would put both series on one chart with two y axes. Build it, then read ?sec_axis on why ggplot2 makes that deliberately awkward, and decide whether you agree.

Quick check

In 2025, how many times more electricity did wind generate than bioenergy? Print the number.

TipSolution
elec$wind_twh[elec$year == 2025] / elec$bioenergy_twh[elec$year == 2025]

About 2.1.

Exercise 5: Make a misleading figure

Everything so far has been about drawing the data honestly. This one is about learning how dishonesty is done, because you cannot review a figure for tricks you have never used.

Using only real data from these files, make a chart that leaves a reader with a false impression of biomass — in either direction. Then say, in one sentence, exactly what the trick is.

You will need this twice: once in Week 4, when one of you is briefed to mislead the rest, and again every time you read someone else’s chart.

NoteHint 1

The reliable tricks do not touch the numbers at all. They change what the numbers are compared against, or which of them the reader sees:

  • start the y axis somewhere other than zero;
  • choose where the x axis begins and ends;
  • pick the comparator that flatters, and omit the one that does not;
  • plot a share when the absolute number is unflattering, or the other way round;
  • put a confident title on an ambiguous chart.

Every one of these is also, sometimes, the right thing to do. That is what makes them hard to police.

NoteHint 2

For the truncated axis, coord_cartesian(ylim = c(30, 42)) zooms the view without dropping data — which is the honest way to do a dishonest thing, because ylim() silently deletes points outside the range and changes the fitted trend.

Compare what those two produce. One of them is a presentation choice; the other quietly alters the analysis.

TipSolution

One of many:

ggplot(elec[elec$year >= 2020, ], aes(x = year, y = bioenergy_twh)) +
  geom_line(linewidth = 1.5) +
  coord_cartesian(ylim = c(30, 42)) +
  labs(x = "Year",
       y = "Generation (TWh)",
       title = "UK bioenergy generation is surging") +
  theme_minimal()

Two choices do the work: six years instead of twenty-six, and a y axis from 30 to 42 instead of from zero. A 7 TWh wobble fills the frame and reads as a steep climb. Nothing has been fabricated — the same series over the full record is flat since 2018.

The defence is not “always start at zero”, which is wrong often enough to be useless advice. The defence is to ask, of any chart including your own: what would this look like with a different window and a different baseline? If you do not know, you have not read the chart yet.

Extension: brief the opposition

Optional, and good preparation for Week 4.

  1. Make the opposite misleading chart — same data, impression reversed. Both should be defensible line by line.
  2. Swap with someone and try to name their trick in one sentence without being told.
  3. Now make the most honest chart you can of the same question, and write its caption. Notice how much of the honesty lives in the caption rather than the geometry.

Save your work

Copy your best figure code into week2.R in your GitHub repo. Commit and push via GitHub Desktop. Write a commit message that describes what your figures show — not just “week 2 plots”.