Week 3: Changing the Assumptions

Scenario analysis — what happens when the numbers shift?

Introduction

In the content session, you explored sources of uncertainty and audited the assumptions behind biomass carbon accounting. The “official” answer — that biomass emits zero CO₂ — is a policy choice, not a physical measurement. What happens when you change the assumptions?

Today’s goals:

  • Estimate the CO₂ cost of shipping wood pellets across the Atlantic
  • Test how sensitive the biomass emissions number is to what you include
  • Explore carbon payback periods — when (if ever) does biomass break even?
  • Compare cumulative emissions under different scenarios
  • Start your briefing: save a figure and set up briefing.md

The code boxes start empty, except for comments to guide you through the code that’s expected.

Work one comment at a time. If you get stuck:

  1. Hint 1 restates what you are trying to do.
  2. Hint 2 sketches the shape the code might take.
  3. Hint 3 gives you the code with the key pieces blanked out.
  4. Solution gives you the lot.

Take a real attempt before opening a hint: you will learn by doing, not by reading. The demonstrators are here to help.

Three things worth knowing:

  • The boxes share one R session, exactly like a script. Anything you create in Exercise 1 is still there in Exercise 2.
  • Variable names are up to you. The last thing your block prints is what gets checked, so end each block with the answer.
  • Exercises have extensions underneath them. If you finish early, give them a stab. These are designed to equip you to tackle real coding problems you might face, using the documentation. We can help – but check the manuals first!
TipStuck?

Ask on the class board, in Code Q&A. Someone else is likely stuck on the same thing, so asking in public helps them too. Demonstrators check it most weekdays, but they give classmates a chance to answer first.

The board is private to this class, so sign in to GitHub first. If the link shows “404 – page not found”, you’re not signed in, or you haven’t accepted the Classroom 50 invitation yet.

Exercise 1: The CO₂ cost of transport

The UK’s wood pellets travel thousands of kilometres by ship — mostly from the southeastern United States. Burning them is counted as zero. Is the shipping zero too?

Work out the total CO₂ released shipping the UK’s 2024 pellet imports, in millions of tonnes (Mt).

Everything you need is in imports (columns year, origin, import_kt, transport_km) and in the constant shipping_co2_per_tkm, which is 0.005 kg CO₂ per tonne-km.

NoteHint 1

Every tonne of pellets that travels a kilometre releases 0.005 kg of CO₂. So for each origin you need tonnes × kilometres × 0.005, and then you need to add those five numbers together and express the answer in megatonnes.

The hard part is not the multiplication — it is keeping track of which unit you are in at each step.

NoteHint 2

The shape of it:

  • Subsetting rows of a data frame by a condition is the df[condition, ] pattern you used in Week 1.
  • You can create a new column by assigning to a name that does not exist yet: df$new_column <- .... Doing it this way — rather than making a loose vector — keeps each number attached to the origin it belongs to, which matters as soon as you want to look at the table.
  • sum() collapses the column to one number.
  • Then divide to get from kg to Mt.
NoteHint 3
imports_2024 <- imports[imports$______ == 2024, ]

imports_2024$transport_co2_kg <-
  imports_2024$import_kt * ______ *      # kilotonnes to tonnes
  imports_2024$transport_km *
  ______

sum(imports_2024$______) / ______        # kg to Mt
TipSolution
imports_2024 <- imports[imports$year == 2024, ]

imports_2024$transport_co2_kg <-
  imports_2024$import_kt * 1000 *
  imports_2024$transport_km *
  shipping_co2_per_tkm

sum(imports_2024$transport_co2_kg) / 1e9

About 0.25 Mt CO₂ — a quarter of a million tonnes. Is that a big number? Hold that question: shipping turns out to be the smallest of the things the official figure leaves out. The chimney at those same plants releases about 43 Mt (Exercise 2), so transport is under 1% of the story. It is real, it is not zero, and it is a useful warm-up — but if you were hunting for the big missing number, you have not found it yet.

Quick check

Your answer means nothing on its own. Put it on a human scale: how many kilograms of CO₂ does shipping release per tonne of pellets delivered?

Use your 2024 subset again — total shipping CO₂ divided by total tonnes imported.

NoteHint 1

You already have the total CO₂ in kg from Exercise 1 — it is what you had just before dividing by 1e9. Divide that by the total number of tonnes the UK imported in 2024.

TipSolution
sum(imports_2024$transport_co2_kg) / sum(imports_2024$import_kt * 1000)

About 27 kg CO₂ per tonne. The pellet itself releases about 1700 kg when burned. Shipping is a rounding error next to combustion — which is exactly why arguing about shipping is a good way to avoid arguing about combustion.

Extension: would New Zealand be different?

Optional, and more open-ended. This one needs two extra packages, so give it a minute to set up.

The UK’s pellets come mostly from the American southeast. Suppose the same mix of pellets were burned in a plant in New Zealand instead. Shipping distances change completely. Does the conclusion?

Rather than using the transport_km column we handed you, work the distances out yourself: look up a coordinate for each origin, look up one for each destination, and compute great-circle distances.

NoteHint 1

maps::world.cities is a data frame with name, country.etc, lat, long and capital. Subsetting to capital == 1 gives roughly one row per country, and the country.etc values for our origins are spelled "USA", "Canada" and "Latvia" — convenient, and also worth being suspicious of.

geosphere::distHaversine() takes two two-column matrices of (longitude, latitude) — note the order, which is the reverse of how we usually say it — and returns metres.

The moment you compute a distance yourself, you have to decide what you are measuring — and every choice smuggles in an assumption:

  • A capital city is not a port. Wood pellets leave the USA from Savannah and Baton Rouge, not Washington DC. How far wrong does that put you, and in which direction?
  • A great-circle line is not a shipping route. No vessel sails through the isthmus of Panama or across Antarctica. Compare your computed USA–UK distance with the 6200 km in transport_km: which is larger, and why?
  • EU_other and Other have no coordinates at all. What did you do with them? Dropping them is a decision, not a technicality — whose emissions did you just make disappear?
  • A longer voyage may be a more efficient one. Emissions per tonne-km vary with vessel size and route. Holding shipping_co2_per_tkm fixed is itself an assumption.

Write down, in one sentence each, the three assumptions you think most affect your answer. Those sentences are exactly the kind of caveat your Week 4 briefing needs — and exactly the kind a traitor would leave out.

Exercise 2: The number that isn’t counted

Transport is only part of the supply chain. Harvesting, chipping, drying and pelletizing all take energy too — and then there is the chimney itself.

factors holds emission estimates under several accounting rules (columns fuel, scenario, co2_kg_per_mwh). Work out the UK’s 2025 biomass electricity emissions in Mt under the with_supply_chain scenario — the number that official accounting records as zero.

NoteHint 1

Emissions = energy generated × emissions per unit of energy. The factor is per MWh, so the generation has to be in MWh too; the factor is in kg, so the answer starts life in kg.

The fiddly bit is pulling one number out of factors: there is a with_supply_chain row for each of three fuels, so a single condition will not identify it.

NoteHint 2

The shape of it:

  • To pick out one value: subset the co2_kg_per_mwh column by a condition that is TRUE for exactly one row. Two conditions joined with & will do it.
  • elec has one row per year, so elec$bioenergy_twh[elec$year == 2025] is the generation.
  • Then multiply, and divide once at the end to get Mt.

Print each piece as you go.

NoteHint 3
sc_factor <- factors$co2_kg_per_mwh[
  factors$fuel == "______" ______
  factors$scenario == "______"
]

bio_mwh <- elec$bioenergy_twh[elec$year == ______] * ______

bio_mwh * sc_factor / ______
TipSolution
sc_factor <- factors$co2_kg_per_mwh[
  factors$fuel == "biomass" &
  factors$scenario == "with_supply_chain"
]

bio_mwh <- elec$bioenergy_twh[elec$year == 2025] * 1e6

bio_mwh * sc_factor / 1e9

About 43.5 Mt CO₂ per year under the supply chain scenario — counted as zero under official rules. For context, the UK’s total territorial emissions are about 340 Mt. So this accounting gap is roughly 13% of UK emissions. That’s a policy-relevant difference.

Now compare that emission factor to the other fuels in factors. Biomass at the chimney (1000 kg/MWh) is higher than coal (910), not lower. Wood carries less energy per unit of carbon than coal does, so generating a megawatt-hour from pellets releases modestly more CO₂ at the stack. Everything that makes biomass look better than coal comes from the regrowth argument — never from the chimney.

Quick check

Print the biomass rows of factors and look at the spread. Which assumption has the largest effect on the answer: including the supply chain, or choosing a different carbon payback period?

Set answer to 1 if supply chain matters more, or 2 if payback period matters more.

Extension: the same number, six ways

Optional.

  1. You pulled one value out of factors with [ and two conditions. Get the same number without using [ at all. There are at least three other routes — subset(), with(), match() — and each behaves differently when the conditions match no rows. Which of them warns you, and which hands you a silent NA? That difference is worth more to you than the syntax.
  2. Produce 2025’s emissions under all six biomass scenarios at once, as a named vector. (sapply(), or just vectorized arithmetic on the whole column.)
  3. elec goes back to 2000. Under with_supply_chain, which was the first year that UK biomass emitted more CO₂ than UK coal did that same year? You will need coal’s generation as well as biomass’s.

Exercise 3: The range of defensible answers

“Carbon payback” is the idea that a burned tree’s CO₂ comes back out of the atmosphere as a replacement tree grows. The payback period is how long that takes — and factors carries biomass scenarios for 20, 40 and 100 years alongside the official zero and the full supply chain.

Build a chart that shows how far apart those six answers are — with coal’s official factor (910 kg/MWh) drawn on as a reference line, because a number is only big compared to something.

NoteHint 1

ggplot orders a character axis alphabetically, which here is meaningless — official would sit between combustion_only and payback_100yr. Converting the column to a factor and setting its levels in the order you want fixes that, because ggplot respects factor level order.

For the reference line, you want a horizontal line at a fixed y value. ggplot2 has a geom for exactly this.

NoteHint 2

The shape of it:

  • order(x) returns the positions that would sort x, so column[order(other_column)] gives you the level order you want.
  • geom_col() draws bars whose heights are the y values.
  • geom_hline(yintercept = ...) draws the reference line.
  • labs() sets titles and axis labels, and theme(axis.text.x = element_text(angle = 45, hjust = 1)) rotates the tick labels.
NoteHint 3
bio_factors <- factors[factors$fuel == "______", ]

bio_factors$scenario <- factor(
  bio_factors$scenario,
  levels = bio_factors$scenario[______(bio_factors$co2_kg_per_mwh)]
)

ggplot(bio_factors, aes(x = scenario, y = ______)) +
  geom_col() +
  geom_hline(yintercept = ______, linetype = "dashed") +
  labs(x = "Accounting scenario",
       y = expression(CO[2] ~ "(kg per MWh)"),
       title = expression("Biomass" ~ CO[2] * ": the answer depends on the question")) +
  theme(axis.text.x = element_text(angle = 45, hjust = 1))
TipSolution
bio_factors <- factors[factors$fuel == "biomass", ]

bio_factors$scenario <- factor(
  bio_factors$scenario,
  levels = bio_factors$scenario[order(bio_factors$co2_kg_per_mwh)]
)

ggplot(bio_factors, aes(x = scenario, y = co2_kg_per_mwh)) +
  geom_col() +
  geom_hline(yintercept = 910, linetype = "dashed") +
  labs(x = "Accounting scenario",
       y = expression(CO[2] ~ "(kg per MWh)"),
       title = expression("Biomass" ~ CO[2] * ": the answer depends on the question")) +
  theme(axis.text.x = element_text(angle = 45, hjust = 1))

The range is enormous — from 0 to 1060 kg CO₂/MWh. The “official” answer is at one extreme. With a 100-year payback, biomass is nearly carbon-neutral (45 kg/MWh). With a 20-year payback, it’s 365 kg/MWh — about the same as gas. If you count the full supply chain and no regrowth at all, it’s 1060, which is worse than coal (910). Which number you choose depends on what question you’re answering and what timescale you care about — and notice that the choice spans the entire distance from “cleanest thing on the grid” to “dirtier than coal.”

The dashed line is what makes the chart argue rather than merely display: without it, 1060 is just the tallest bar.

Extension: make the chart say it louder

Optional. Each of these is a documentation exercise — the answer is in the help pages, not in these notes. Try ?geom_hline, ?annotate, ?coord_flip, ?scale_x_discrete.

  1. The scenario names (payback_100yr) are variable names, not English. Relabel the axis ticks without editing the data frame.
  2. Make the coal line red and dotted, and label it “coal” on the plot itself — without using labs() or ggtitle().
  3. Colour each bar by whether it sits above or below the coal line.
  4. Long category labels are usually easier to read horizontally. Turn the chart on its side, then decide which version you would put in a briefing.

Extension: best case vs worst case

This exercise is optional — for students who finish early.

Biomass has a range of defensible answers. So, less dramatically, do coal and gas: their supply chains are not free either. Build a chart that puts all three ranges side by side, so a reader can see at a glance which fuel we are least sure about.

The numbers are all in factors — biomass from 45 (100-year payback) to 1060 (full supply chain); coal from 910 to 990; gas from 360 to 410.

NoteHint 1

One approach is a data frame with a fuel column and low/high columns, drawn with geom_errorbar() or geom_linerange(). Another is to filter factors itself and let ggplot summarize with stat_summary(). A third is geom_pointrange().

Try at least two, and decide which communicates best. That judgement — not the code — is the exercise.

TipSolution
range_df <- data.frame(
  fuel = c("Biomass", "Coal", "Gas"),
  low = c(45, 910, 360),
  high = c(1060, 990, 410)
)

range_df$mid <- (range_df$low + range_df$high) / 2

ggplot(range_df, aes(x = fuel, y = mid, ymin = low, ymax = high)) +
  geom_pointrange() +
  labs(x = "Fuel",
       y = expression(CO[2] ~ "(kg per MWh)"),
       title = "Emission factor ranges by fuel",
       caption = "Points show midpoint; bars show range of estimates")

The range for biomass (45–1060 kg/MWh) is vastly wider than for coal (910–990) or gas (360–410). For coal and gas, the uncertainty is about supply chain details — a few percent either way. For biomass, the uncertainty spans the fundamental question of whether it is nearly carbon-free or worse than coal. That’s not a small disagreement — it’s a factor of 24, and the biomass bar overlaps the coal bar entirely.

Notice what that means for the briefing you are about to write: an author who picks the bottom of the biomass range and the top of the coal range can prove almost anything they like, using nothing but real numbers from this file.

This is the core insight of the mini-project: the answer depends on the assumptions, and the assumptions are a choice.

Then push on the presentation. More documentation exercises:

  1. A midpoint is a fiction — nobody has claimed the middle of the biomass range is the right answer. Redraw the chart so no midpoint marker appears at all.
  2. Add each fuel’s official factor as a separate mark, so a reader can see where the headline number sits inside its range. Is it near the middle? For which fuel is it furthest from the middle, and what does that tell you about who chose it?
  3. Sort the fuels by range width rather than alphabetically.

Exercise 4: Cumulative emissions — biomass vs coal

An annual figure hides the thing payback periods are about: time. If a biomass plant runs for fifty years at 2025’s output (41.0 TWh a year), how much CO₂ has it put into the atmosphere by the end — compared with a coal plant doing the same job?

4a. The fifty-year totals

Work out the cumulative emissions after 50 years, in Mt, for three scenarios: coal (910 kg/MWh), biomass with supply chain (1060), and biomass on a 20-year payback (365). Print all three together.

NoteHint 1

Each year emits the same amount, so fifty years emits fifty times as much. You did the one-year version in Exercise 2 — this is that, with three emission factors instead of one.

NoteHint 2

R is vectorized: if you put the three emission factors in a single vector with c(), one arithmetic expression gives all three answers at once. Naming the elements — c(coal = 910, ...) — makes the printed output readable, which is worth doing whenever a result has more than one number in it.

TipSolution
annual_mwh <- 41 * 1e6

c("Coal (910)" = 910,
  "Biomass supply chain (1060)" = 1060,
  "Biomass 20yr payback (365)" = 365) *
  annual_mwh * 50 / 1e9

Coal emits about 1,866 Mt over fifty years. Biomass on a 20-year payback emits about 748 Mt — comfortably better. But biomass with the full supply chain and no regrowth credit emits about 2,173 Mt, which is more than coal.

4b. The picture

Three numbers make the point; a chart makes it memorable. Plot cumulative emissions year by year, one line per scenario, over 50 years.

You will need the data in long form — one row per year per scenario, with columns for the year, the cumulative total, and the scenario name.

NoteHint 1

For one scenario, the running total is cumsum() of a vector holding the same annual figure fifty times — rep() builds that vector.

Then you need to stack the three scenarios into a single data frame. rep() is useful again: the year column repeats 1:50 three times, and the scenario column repeats each label fifty times.

NoteHint 2

The shape of it:

  • cumsum(rep(annual_mt, 50)) for each scenario.
  • c(a, b, c) glues the three running totals end to end.
  • rep(1:50, 3) and rep(labels, each = 50) line up with them.
  • In aes(), mapping colour = scenario draws one line per scenario and builds the legend for you.
NoteHint 3
annual_mwh <- 41 * 1e6
coal_cumul   <- cumsum(rep(annual_mwh * 910 / 1e9, 50))
bio_sc_cumul <- cumsum(rep(annual_mwh * ______ / 1e9, 50))
bio_20_cumul <- cumsum(rep(annual_mwh * ______ / 1e9, 50))

scenarios <- data.frame(
  year = rep(1:50, 3),
  cumul_mt = c(coal_cumul, bio_sc_cumul, bio_20_cumul),
  scenario = rep(c("Coal (910)", "Biomass supply chain (1060)",
                   "Biomass 20yr payback (365)"), each = ______)
)

ggplot(scenarios, aes(x = ______, y = ______, colour = ______)) +
  geom_line() +
  labs(x = "Years", y = expression("Cumulative" ~ CO[2] ~ "(Mt)"), colour = "Scenario")
TipSolution
annual_mwh <- 41 * 1e6

coal_cumul   <- cumsum(rep(annual_mwh * 910 / 1e9, 50))
bio_sc_cumul <- cumsum(rep(annual_mwh * 1060 / 1e9, 50))
bio_20_cumul <- cumsum(rep(annual_mwh * 365 / 1e9, 50))

scenarios <- data.frame(
  year = rep(1:50, 3),
  cumul_mt = c(coal_cumul, bio_sc_cumul, bio_20_cumul),
  scenario = rep(c("Coal (910)", "Biomass supply chain (1060)",
                   "Biomass 20yr payback (365)"),
                 each = 50)
)

ggplot(scenarios, aes(x = year, y = cumul_mt, colour = scenario,
                      linetype = scenario)) +
  geom_line() +
  labs(x = "Years",
       y = expression("Cumulative" ~ CO[2] ~ "(Mt)"),
       colour = "Scenario",
       linetype = "Scenario",
       title = "Cumulative emissions: coal vs biomass scenarios")

That is the whole argument in one chart. Biomass is not better than coal because the chimney is cleaner — it isn’t; it’s slightly dirtier. Biomass is better than coal only if the forest actually regrows, and only on a timescale that matters. Everything turns on a regrowth assumption that this plot takes entirely on trust, and that no measurement in this dataset verifies.

Note: this simplified model assumes constant emission rates — real payback changes over time as the forest regrows, so treat the straight lines as a sketch rather than a projection.

Extension: interrogate the sketch

Optional.

  1. Mapping to both colour and linetype is deliberate: around 1 in 12 men has some degree of colour-vision deficiency, and a chart that encodes its only distinction in hue fails them. Drop the linetype mapping, look at the result, then put it back.
  2. The straight lines are the lie in this chart. Real regrowth is slow at first and faster later. Replace the constant annual figure for one biomass scenario with one that starts at the combustion factor and decays towards the payback factor, and see whether the crossing point moves. seq() will build the sequence.
  3. Mark the year the biomass supply-chain line overtakes coal — without using geom_point(). (?geom_vline, ?annotate, ?geom_segment.)
  4. Fifty years is a choice too. Redo the chart for 10 years and for 200. Does the story change? Which would you put in a briefing, and could you defend that choice out loud?

Start your briefing

Next week you’ll write a two-page policy briefing against the clock. You’ll write it in Markdown: plain text with a few symbols for formatting. GitHub displays a Markdown file as a formatted page, with its figures, and that is how your reviewers will read it. Setting the file up today means next week is just writing.

Step 1: Save a figure

Pick the figure from today that best makes your case. Right-click it and choose Save image as…. Save it inside your repo folder, in a folder called figures, with a name that says what it shows: for example figures/payback-scenarios.png.

In GitHub Desktop, commit and push the figure.

Step 2: Create briefing.md

On github.com, open your repo and choose Add file → Create new file. Name it briefing.md and paste in this skeleton:

# Does UK biomass power help reach net zero, and should we keep paying for it?

## Summary

Your answer, first, in 3–4 sentences.

## Evidence

![What the figure shows, for a reader who can't see it](figures/payback-scenarios.png)

**Figure 1.** Caption: what should the reader notice?

## Caveats

- What does your answer depend on?

## Conclusions

Your recommendation.

Change the image path to match your figure, then click Preview. You should see headings, a bullet, bold text and your figure. If the figure is missing, check the spelling of the path (capitals matter) and that you pushed the figure in Step 1.

Click Commit changes. Then, in GitHub Desktop, click Fetch origin and Pull origin, so the copy on your computer has briefing.md too.

Markdown in one table

You type You get
# Title The title
## Section A section heading
**bold** bold
*italic* italic
- item A bulleted list
1. item A numbered list
![alt text](figures/name.png) The figure
A blank line A new paragraph

The briefing structure guide says what goes in each section.

Save your work

Copy your scenario analysis code into week3.R in your GitHub repo. Commit and push via GitHub Desktop. Write a commit message that explains what you discovered about the assumptions — not just “week 3 exercises”.