Week 1: Meet the Data
Exploring UK electricity generation in R
Before you start, accept the biomass assignment to create your personal repository. Then clone it via GitHub Desktop — your instructor will demonstrate this later in the session.
Introduction
In the content session, we explored what makes a hypothesis testable and started asking whether UK biomass power really helps the UK reach net zero, starting with whether it is carbon-neutral. Now it’s time to dig in to the data.
Today’s goals:
- Load a real dataset into R and inspect its structure
- Make your first plots of UK electricity generation
- Calculate summary statistics and put them in context
- Save your code to a
.Rfile and commit it to GitHub
The code boxes start empty, except for comments to guide you through the code that’s expected.
Work one comment at a time. If you get stuck:
- Hint 1 restates what you are trying to do.
- Hint 2 sketches the shape the code might take.
- Hint 3 gives you the code with the key pieces blanked out.
- Solution gives you the lot.
Take a real attempt before opening a hint: you will learn by doing, not by reading. The demonstrators are here to help.
Three things worth knowing:
- The boxes share one R session, exactly like a script. Anything you create in Exercise 1 is still there in Exercise 2.
- Variable names are up to you. The last thing your block prints is what gets checked, so end each block with the answer.
- Exercises have extensions underneath them. If you finish early, give them a stab. These are designed to equip you to tackle real coding problems you might face, using the documentation. We can help – but check the manuals first!
Ask on the class board, in Code Q&A. Someone else is likely stuck on the same thing, so asking in public helps them too. Demonstrators check it most weekdays, but they give classmates a chance to answer first.
The board is private to this class, so sign in to GitHub first. If the link shows “404 – page not found”, you’re not signed in, or you haven’t accepted the Classroom 50 invitation yet.
You already know how to load data, make plots, and calculate summary statistics — you did all of this in GEOL1151 with NumPy and matplotlib. Today you’ll do the same things in R. The key differences:
| What | Python (GEOL1151) | R (today) |
|---|---|---|
| Load data | np.loadtxt("file.dat") |
read.csv("file.csv") |
| Access a column | data[:, 0] |
df$column_name |
| Mean | np.mean(data[:, 0]) |
mean(df$column_name) |
| Quick plot | plt.plot(x, y) |
plot(x, y) |
| Histogram | plt.hist(x, bins=20) |
hist(x, breaks = 20) |
| Assignment | x = 42 |
x <- 42 |
| Indexing | Starts at 0 | Starts at 1 |
The biggest shift: instead of NumPy arrays, you’ll work with data frames: tables where each column has a name. You access columns with $ (e.g. elec$year) rather than by position.
For a full side-by-side reference, see the Python to R cheat sheet.
Loading the data
The dataset covers UK electricity generation by fuel type from 2000 to 2025. Once you’ve run this block, elec will be loaded into the webR session, and available for you to use in subsequent exercises.
Exercise 1: What have we got?
Never analyse a dataset you have not looked at. Before anything else, find out how big elec is and what is in it.
Finish your block by printing the number of years the data covers.
Three separate questions, each using a distinct function. What do the data look like; what are the columns called; and how many rows are there? Each row is one year, so the row count is the number of years.
The shape of it:
head(df)prints the first six rows.names(df)gives the column names as a character vector.nrow(df)gives the number of rows.
Each of these takes the data frame itself — the whole of elec, not one column of it.
head(______)
names(______)
nrow(______)head(elec)
names(elec)
nrow(elec)26 rows — one per year from 2000 to 2025. The columns are year, generation in TWh for each fuel type (coal_twh, gas_twh, nuclear_twh, wind_twh, solar_twh, bioenergy_twh), and total_twh.
One thing worth noticing straight away: the six fuel columns for 2025 add up to 276.6 TWh, but total_twh says 293.6. The difference is hydro, imports and everything else we did not give a column to. If you ever compute a percentage, that gap decides which denominator is honest.
Extension: three more ways to look
Optional. The answers are in the help pages — try ?str, ?summary, ?dim.
head()shows you the values.str()shows you the structure — run it and work out whatnumandintare telling you about each column.summary(elec)produces a block of numbers per column. Which column has a minimum of zero, and is that a measurement or a gap in the record?- Get the number of years without using
nrow(). There are at least three other routes. Which of them would still work ifelecwere a plain vector rather than a data frame? elec2 <- eleccreates a copy of yourelecdataset. Pretend thatelec2contains an error: row 4 actually describes data from 1999. Fix the error, then determine: does your previous code now report the number of years, the time span, or something else?
Exercise 2: Biomass over time
A table of 26 numbers is hard to read; a picture of 26 numbers is not. Draw bioenergy generation against year, with labelled axes and a title, so someone who has never seen the data could read it.
plot() takes the x values first and the y values second — so year comes first. Both are columns of elec, and you pull a column out with $.
Axis labels are not decoration. “40” means nothing; “40 TWh” means something.
The shape of it:
plot(x_values, y_values,
xlab = "...", ylab = "...", main = "...")xlab and ylab label the axes; main sets the title. All three take a character string in double quotes.
plot(elec$______, elec$______,
xlab = "Year",
ylab = "______",
main = "UK bioenergy electricity generation")plot(elec$year, elec$bioenergy_twh,
xlab = "Year",
ylab = "Bioenergy generation (TWh)",
main = "UK bioenergy electricity generation")Bioenergy grew from about 4 TWh in 2000 to around 40 TWh by the early 2020s — roughly a tenfold increase. Most of that growth happened after 2012: Drax power station converted its first unit from coal to biomass in 2013.
Look at the shape rather than the endpoints, though: the rise is not steady. It is nearly flat to 2010, steep from 2012 to 2015, and flat again after 2018. A single “tenfold increase” headline hides all three phases.
Extension: the same plot, better
Optional, and entirely about the documentation. ?plot is the wrong place to start — try ?par for the styling arguments, and ?mtext and ?title and ?axis for the last one.
- Make it a blue dotted line rather than black circles.
- Then draw the points and the line, so individual years are still visible.
- The y-axis starts near 4, which makes the rise look steeper than it is. Force it to start at zero. Which version would you put in a briefing, and which would a campaigner choose?
- Label both axes without using
plot(xlab = )orplot(ylab = ). There is more than one way; find two.
Exercise 3: Biomass vs coal
Biomass grew while another fuel collapsed. Put both on the same axes — a number only means something next to a comparator.
Draw bioenergy and coal against year, as two distinguishable lines, with a legend saying which is which.
plot() starts a new picture; anything you draw afterwards has to be added to it. So the first series sets the axes, and the axes never change afterwards — if coal reaches 120 and you drew the axes to fit biomass’s 40, coal will be drawn off the top of the plot and silently vanish.
That is why ylim matters here. Check the largest value in both columns before you choose it.
The shape of it:
plot(x, y, type = "l")draws a line rather than points.ylim = c(0, 150)fixes the vertical range.lines(x, y)adds a second series to the plot you already have.colsets colour andltysets line type, on either function.legend("topright", legend = c(...), col = c(...), lty = c(...))builds the key. The three vectors have to be in the same order, and nothing checks that for you.
plot(elec$year, elec$bioenergy_twh,
type = "l", col = "blue", lty = 1,
xlab = "Year", ylab = "Generation (TWh)",
main = "UK electricity: biomass vs coal",
ylim = c(0, ______))
______(elec$year, elec$______, col = "red", lty = 2)
legend("topright", legend = c("Bioenergy", "Coal"),
col = c("blue", "red"), lty = c(1, 2))plot(elec$year, elec$bioenergy_twh,
type = "l", col = "blue", lty = 1,
xlab = "Year", ylab = "Generation (TWh)",
main = "UK electricity: biomass vs coal",
ylim = c(0, 150))
lines(elec$year, elec$coal_twh, col = "red", lty = 2)
legend("topright", legend = c("Bioenergy", "Coal"),
col = c("blue", "red"), lty = c(1, 2))Coal fell from about 120 TWh in 2000 to zero in 2025. Bioenergy rose over the same period, and the two lines cross in 2016–17.
It is very tempting to read that crossing as biomass replaced coal. Look at the scale before you do: coal lost about 120 TWh and biomass gained about 37. Whatever filled the other 83 TWh, it was not bioenergy. Add gas and wind to the same chart and see who the real replacement was.
Extension: the renewables race
Optional.
- Which renewable generates most UK electricity — wind, solar or bioenergy? Put all three on one chart and find out. Did the answer surprise you? Bioenergy takes up far more of the carbon-accounting argument than its share of generation would suggest; ask yourself why that might be.
- Add coal to that chart without using
lines()orpoints().?matplot,?segments,?parand itsnewargument are all routes in. At least one of them will mislead you about the axes — work out which, and how you would spot it. - Generation levels tell you where we are; changes tell you what is happening. Plot how much new generation each source added each year —
diff()gives you the year-on-year changes, and note that it returns one value fewer than it was given, so your x-axis needs thought. Which source has the single biggest one-year jump?
Exercise 4: How much, and how variable?
Pictures show shape; numbers let you argue. Summarize the last five years of bioenergy generation (2021–2025) with its mean and its standard deviation.
Print the two together, as a single named vector, so the output says what each number is.
Three separate jobs: get the right five numbers out of the column, summarize them two ways, then print both summaries together.
Getting the right five is the part that goes wrong. The data end in 2025, so the last five values are the ones you want — and there is a function whose whole job is taking values from the end of a vector.
The shape of it:
tail(x, 5)gives the last five elements of a vector. Give it a column, not the whole data frame, or you will get five rows back.mean()andsd()each take that vector and return one number.c(name = value, name = value)builds a named vector, which prints with a header row telling you which number is which.
last5 <- ______(elec$bioenergy_twh, 5)
bio_mean <- ______(last5)
bio_sd <- ______(last5)
c(mean = ______, sd = ______)last5 <- tail(elec$bioenergy_twh, 5)
bio_mean <- mean(last5)
bio_sd <- sd(last5)
c(mean = bio_mean, sd = bio_sd)About 38.2 TWh on average, varying by about 3 TWh either side.
Those two numbers do different jobs. The mean says how much; the SD says how reliably. A source delivering 38 TWh every year and one delivering 20 TWh in some years and 56 in others have the same mean and very different consequences for the grid — and you cannot tell them apart without the second number.
But is 38 TWh a big number? On its own it means nothing at all. You need a comparator, and that is the next exercise.
Extension: the mean is a choice
Optional. ?median, ?quantile, ?range, ?IQR, ?var.
- Report the median and the range alongside the mean. For these five years they tell much the same story — construct a set of five numbers where they would not.
- Describe the standard deviation without using
sd(). sd()divides by n − 1, not n. Compute both versions for these five values and see how far apart they are. Would the difference ever change a conclusion you would defend out loud?mad()gives a more robust measure of spread. Construct a dataset wheremadandsdtell different stories.
Exercise 5: Is biomass a big deal?
41 TWh sounds like a lot. But to give it some context, we need to compare it to something. Work out bioenergy’s share of all UK electricity in 2025, as a percentage.
A share is one number divided by another; a percentage is that multiplied by 100. The only real work is picking the right two numbers out of a single year.
To keep one row of a data frame, put a condition in the square brackets before the comma: df[condition, ]. The comma matters — it means “these rows, all columns”.
The shape of it:
row_2025 <- elec[elec$year == 2025, ]
100 * row_2025$something / row_2025$something_elseNote == (a test) rather than = (an assignment). Confusing the two is among the most common errors in R, and the message it produces does not obviously point at it.
row_2025 <- elec[elec$______ == 2025, ]
100 * row_2025$______ / row_2025$______row_2025 <- elec[elec$year == 2025, ]
100 * row_2025$bioenergy_twh / row_2025$total_twhAbout 14% of UK electricity came from bioenergy in 2025.
Notice how much work the denominator did. Divide by total_twh and you get 14.0%; divide by the six fuel columns added together and you get 14.8%. Both are computed correctly from the same file. Neither is a lie. When you read a percentage in a press release, the interesting question is almost never the numerator.
Quick check
How many times more bioenergy did the UK generate in 2025 than in 2000? Print the number.
elec$bioenergy_twh[elec$year == 2025] / elec$bioenergy_twh[elec$year == 2000]About 9.5 — call it a tenfold rise.
Extension: the honest denominator
Optional.
- Plot bioenergy’s percentage share for every year, not just 2025. When did the share grow fastest — and is that the same year generation grew fastest? It need not be: the denominator was moving too.
- UK total generation fell from 377 TWh in 2000 to 294 in 2025. Work out how much of bioenergy’s rise in share comes from generating more biomass, and how much from everyone else generating less.
- Write one sentence, fit for a briefing, stating bioenergy’s 2025 share honestly. Then write another that is also true but leaves the reader with the opposite impression. Keep both — you will want them in Week 4.
Save your work
Copy the code you’re most proud of into a file called week1.R in your GitHub repo. Commit and push via GitHub Desktop. Write a commit message that describes what you found — not just “week 1”.