Week 8: Fitting and Breaking Models

What happens when you extrapolate beyond your data?

Introduction

The content session showed you a model failing on data it had never seen. This session does it with a question people actually argue about: how much wind and solar will the UK have in 2040?

You will fit two reasonable models to the same sixteen years of real data. Both will fit well. Their answers will differ by a factor of nearly six, and nothing in either model will tell you which to believe.

Goals:

  • Fit linear and exponential models to UK renewables data
  • Extrapolate both, and check the answers against physical reality
  • Say precisely what each model cannot represent
  • Carry the habit into your own project

The code boxes start empty, except for comments to guide you through the code that’s expected.

Work one comment at a time. If you get stuck:

  1. Hint 1 restates what you are trying to do.
  2. Hint 2 sketches the shape the code might take.
  3. Hint 3 gives you the code with the key pieces blanked out.
  4. Solution gives you the lot.

Take a real attempt before opening a hint: you will learn by doing, not by reading. The demonstrators are here to help.

Three things worth knowing:

  • The boxes share one R session, exactly like a script. Anything you create in Exercise 1 is still there in Exercise 2.
  • Variable names are up to you. The last thing your block prints is what gets checked, so end each block with the answer.
  • Exercises have extensions underneath them. If you finish early, give them a stab. These are designed to equip you to tackle real coding problems you might face, using the documentation. We can help – but check the manuals first!
TipStuck?

Ask on the class board, in Code Q&A. Someone else is likely stuck on the same thing, so asking in public helps them too. Demonstrators check it most weekdays, but they give classmates a chance to answer first.

The board is private to this class, so sign in to GitHub first. If the link shows “404 – page not found”, you’re not signed in, or you haven’t accepted the Classroom 50 invitation yet.

Exercise 1: Fifteen years of growth

UK wind and solar generation has grown sharply since 2010. Build the series and look at it before modelling anything.

Finish by printing 2025’s combined wind and solar generation.

NoteHint 1

Adding two columns together gives a vector; assigning it into the data frame with $ makes it a column.

Filtering to 2010 onward is the df[condition, ] pattern, with >= rather than ==.

Name the filtered data frame recent — the next three exercises refer to it, and the boxes share one session.

NoteHint 2

The shape of it:

df$new <- df$a + df$b

recent <- df[df$year >= 2010, ]

ggplot(recent, aes(x = year, y = new)) +
  geom_point() +
  labs(...)

recent$new[recent$year == 2025]
NoteHint 3
elec$renewable_twh <- elec$______ + elec$______

recent <- elec[elec$year >= ______, ]

ggplot(recent, aes(x = ______, y = ______)) +
  geom_point() +
  labs(x = "Year",
       y = "Renewable generation (TWh)",
       title = "UK wind and solar generation, 2010-2025")

recent$renewable_twh[recent$year == ______]
TipSolution
elec$renewable_twh <- elec$wind_twh + elec$solar_twh

recent <- elec[elec$year >= 2010, ]

ggplot(recent, aes(x = year, y = renewable_twh)) +
  geom_point() +
  labs(x = "Year",
       y = "Renewable generation (TWh)",
       title = "UK wind and solar generation, 2010-2025")

recent$renewable_twh[recent$year == 2025]

From 10.3 TWh in 2010 to 106.5 TWh in 2025 — about tenfold, and over a third of all UK electricity.

Look hard at the shape, because the next two exercises both start here. From 2015 the rise is close to a straight line: roughly the same number of TWh added each year. Before that it is steeper in proportional terms — the jump from 10 to 20 TWh is a doubling; the jump from 80 to 90 is not.

Straight line or compound growth? Both readings are defensible from this picture. They disagree about 2040 by a factor of nearly six.

Extension: is this even the right series?

Optional.

  1. Plot the year-on-year change rather than the level. A straight line in the level means a flat line here. Is it flat?
  2. Add bioenergy and redraw. “Renewables” is a definition, not a measurement, and that one choice moves the 2025 figure by 38%. Note also that hydro has no column at all, though it is unquestionably renewable — it is inside the gap between the fuel columns and total_twh. Which definition would you have to state in a caption?
  3. Start the series in 2000 instead of 2010. The early years are nearly flat at almost zero, which drags a straight line down — refit and see how much the slope changes. Choosing where a series starts is choosing part of the answer.
  4. Plot the same data on a log y axis. Compound growth is a straight line there. Does yours look straight — and over which stretch?

Exercise 2: The straight line into 2040

Fit a linear model to the 2010–2025 series and ask it about 2040 — fifteen years past the last observation.

Print the 2040 prediction.

NoteHint 1

This is Week 8’s Exercise 2 with different columns: fit, summarize, then predict.

predict() needs a data frame whose column is named exactly as the predictor in the model. data.frame(year = 2040) — a data frame of one row, not a bare number.

NoteHint 2

The shape of it:

fit <- lm(outcome ~ year, data = recent)
summary(fit)

predict(fit, newdata = data.frame(year = 2040))

Read the slope as a rate: TWh of new generation per year. Ask whether you believe that rate can be sustained for another fifteen years — the model has no opinion on the question.

NoteHint 3
lin_model <- lm(______ ~ ______, data = recent)
summary(lin_model)

predict(lin_model, newdata = data.frame(year = ______))
TipSolution
lin_model <- lm(renewable_twh ~ year, data = recent)
summary(lin_model)

predict(lin_model, newdata = data.frame(year = 2040))

R² = 0.978, slope 6.64 TWh per year, and a 2040 prediction of 210 TWh — about 72% of all the electricity the UK generated in 2025.

That is the dangerous kind of wrong: plausible. Government targets sit in this region, so the number passes a sniff test and gets quoted.

What the model contains is one straight line. What it does not contain is every reason the line might bend: the best sites get used first, grid connection queues, planning permission, supply chains, subsidy regimes, and the plain fact that electricity demand itself is not fixed. None of those appear in recent, so none can appear in the forecast.

An R² of 0.978 measures how well a line describes sixteen points you already have. It is silent about the fifteenth year after them, and this is the single most common way a competent analysis becomes a bad prediction.

Extension: how uncertain is that line?

Optional. ?predict.lm and its interval argument, ?confint.

  1. Get a 95% prediction interval for 2040. It will be reassuringly narrow — roughly ±17 TWh. Write one sentence explaining why that narrowness is misleading, given everything the model does not know.
  2. Fit to 2010–2017 only, and predict 2025. You can check that one against reality. How far out is it, and is the true value inside the prediction interval?
  3. Do the same from 2010–2020. Does a longer training run help?
  4. Task 2 is the only honest test of an extrapolation available to you: hide data you have, predict it, and look. Nothing else you can do with recent tests the thing that matters.

Exercise 3: The other defensible model

The early years grew proportionally faster, which is what compound growth looks like. An exponential model captures that: fit a straight line to the logarithm of generation, then undo the log.

Fit it, predict 2040, and back-transform.

NoteHint 1

You can put log() directly in the formula; no new column is needed.

Everything the model then produces — coefficients, predictions, intervals — is on the log scale. exp() brings a prediction back, the same back-transformation you used for the confidence interval in Week 6.

The slope on the log scale is a proportional growth rate: 0.140 means roughly 14% more each year, which compounds.

NoteHint 2

The shape of it:

fit <- lm(log(outcome) ~ year, data = recent)
summary(fit)

exp(predict(fit, newdata = data.frame(year = 2040)))

Sanity-check the intermediate value. A log-scale prediction of about 7 is right; if that number looks like a plausible quantity of electricity, you have forgotten the exp().

NoteHint 3
exp_model <- lm(______(renewable_twh) ~ year, data = recent)
summary(exp_model)

______(predict(exp_model, newdata = data.frame(year = ______)))
TipSolution
exp_model <- lm(log(renewable_twh) ~ year, data = recent)
summary(exp_model)

exp(predict(exp_model, newdata = data.frame(year = 2040)))

About 1,190 TWh — about four times the UK’s entire 2025 electricity generation, from wind and solar alone. It is not merely optimistic; there is nowhere to put it and nobody to use it.

The model’s R² is 0.88. Both models fit well. Both are defensible descriptions of 2010–2025. They differ about 2040 by a factor of nearly six, and no diagnostic you have been taught can separate them, because the disagreement is not about the data — it is about whether growth is additive or multiplicative, and the data contain sixteen points that are consistent with both.

Work out when the exponential model crosses total UK generation. The answer is 2031, six years past the last observation and clearly false. That is the check to run: not “does the model fit” but “does its answer collide with something I know independently?” Here the comparator is a physical ceiling. In your project it might be a budget, a mass balance, or a percentage that cannot exceed 100.

Extension: the model that knows about limits

Optional.

  1. Find the year each model crosses 293.6 TWh. Exponential: 2031. Linear: 2053. Neither model knows the number 293.6 exists.
  2. Real growth into a finite resource is usually logistic — an S curve that flattens as it saturates. Sketch one by hand over your scatter plot and mark where you think the ceiling is. Then ask what evidence in recent could possibly tell you.
  3. The honest answer is that sixteen points of rising data cannot distinguish the early part of an S curve from an exponential. Say in one sentence what would distinguish them, and when the UK might have that evidence.
  4. Fit the exponential to 2010–2015 alone and predict 2025. Compare with the truth. This is the cleanest demonstration in the course of why a well-fitting model is not a reliable one.

Exercise 4: Draw the disagreement

Two numbers are an argument. One chart settles what the argument is about.

Plot the data with both fitted curves extended to 2045, and a horizontal reference line at 293.6 TWh — total UK electricity generation in 2025 — so a reader can see where each model leaves the possible.

NoteHint 1

Build one data frame of years and add a prediction column per model, exactly as you did for the polynomial in the content session. Then each curve is a geom_line() layer drawing from that grid, while the points come from recent.

Because the two layers use different data frames, each geom_ call needs its own data = argument.

A colour mapped to a constant string inside aes() builds the legend — the Week 2 trick.

NoteHint 2

The shape of it:

grid <- data.frame(year = 2010:2045)
grid$linear <- predict(lin_model, newdata = grid)
grid$exponential <- exp(predict(exp_model, newdata = grid))

ggplot() +
  geom_point(data = recent, aes(year, renewable_twh)) +
  geom_line(data = grid, aes(year, linear, colour = "Linear")) +
  geom_line(data = grid, aes(year, exponential, colour = "Exponential")) +
  geom_hline(yintercept = 293.6, linetype = "dashed") +
  annotate("text", x = 2011, y = 310,
           label = "Total UK electricity, 2025", hjust = 0, size = 3) +
  labs(x = "Year", y = "Generation (TWh)", colour = "Model")

You will have to decide what to do about the exponential curve leaving the plot. Clipping the y axis makes the other three things readable; leaving it unclipped makes the absurdity unmissable. Both are legitimate, and saying which you chose is the point.

TipSolution
grid <- data.frame(year = 2010:2045)
grid$linear <- predict(lin_model, newdata = grid)
grid$exponential <- exp(predict(exp_model, newdata = grid))

ggplot() +
  geom_point(data = recent, aes(year, renewable_twh), size = 2) +
  geom_line(data = grid, aes(year, linear, colour = "Linear"),
            linewidth = 1) +
  geom_line(data = grid, aes(year, exponential, colour = "Exponential"),
            linewidth = 1) +
  geom_hline(yintercept = 293.6, linetype = "dashed", alpha = 0.5) +
  annotate("text", x = 2011, y = 310,
           label = "Total UK electricity, 2025", hjust = 0, size = 3) +
  coord_cartesian(ylim = c(0, 600)) +
  labs(x = "Year", y = "Generation (TWh)",
       colour = "Model",
       title = "Two models, two futures",
       subtitle = "Both fit the data. Neither predicts the future.") +
  theme(legend.position = "bottom")

exp(predict(exp_model, newdata = data.frame(year = 2040))) /
  predict(lin_model, newdata = data.frame(year = 2040))

5.7 times. The chart makes it visceral: the two curves are almost on top of each other across the observed years and then diverge violently the moment the data stop. The exponential crosses total UK generation in 2031 and keeps going.

The gap between the curves is not uncertainty in the usual sense — no confidence interval contains it, because it comes from the choice of model rather than from noise in the data. Statisticians call this structural uncertainty, and it is almost always larger than the kind that gets reported. Your error bars describe how wrong you might be given your model. They say nothing about the model being the wrong shape.

So when you put a projection in your report, the sentence that earns trust is not “±17 TWh”. It is: here is what I assumed about the shape of the future, and here is what the answer becomes if I assume otherwise.

Extension: make the chart argue honestly

Optional.

  1. Draw it with the y axis unclipped. Now the exponential dominates and the data are an invisible smear at the bottom. Which version tells the truth better, and which would you caption differently?
  2. Shade the region beyond 2025 to mark where the curves stop being fitted and start being guesses. ?annotate with geom = "rect". Almost no published forecast does this, and it takes one line.
  3. Add a third curve: growth that continues at 2025’s rate but decelerates by 5% a year. Where does it land in 2040? You now have three defensible futures and no way to choose between them from the data — which is a finding, and belongs in the report as one.

Exercise 5: What is your model not capturing?

Every analysis has a blind spot: something it structurally cannot represent, no matter how much data you feed it. The linear model cannot represent a ceiling. The exponential cannot represent saturation. HolmesCo’s groundwater model could not represent a season.

Turn that on your own project. There is no code and no mark — this is the habit the whole week exists to build.

NoteHint 1

Blind spots are usually of one of these kinds. Which is yours?

  • A shape you assumed. Straight line, constant rate, no interaction, normal errors.
  • Something absent from the data. A confounder nobody recorded; a variable that exists but has no column.
  • A boundary the model does not know about. A physical ceiling, a number that cannot go below zero, a percentage that cannot exceed
  • A sample that decided itself. The sites that were accessible, the people who replied, the years the instrument worked.
  • Time. A relationship that held over your window and need not hold outside it.

The third question is the one that matters. A named blind spot with a cheap check is a good limitations paragraph. A named blind spot with no check is an apology.

TipSolution

A worked answer, from this page rather than a project:

  1. My extrapolation cannot represent a ceiling on deployment. It assumes 6.6 TWh a year can be added indefinitely, because nothing in sixteen years of rising data told it otherwise.
  2. If deployment saturates — best sites used, grid queue binding — 2040 is nearer 150 TWh than 210, and the recommendation to plan for 72% renewable supply is unsafe.
  3. Compare the rate of additions year by year rather than the level. If the annual increment has begun to fall, saturation has already started, and that is visible in data I already have.

Notice the shape. The blind spot is specific rather than “there may be other factors”. The consequence is quantified. The check is something you could run this afternoon.

That paragraph is what a reviewer reads first and what a policy-maker trusts you for. It is also the paragraph that most student reports either omit or fill with generic hedging — which reads as a disclaimer rather than as evidence you understood your own analysis.

Extension: audit someone else’s blind spot

Optional, and best done in pairs.

  1. Swap projects with another group. Read their question and method, and write down the blind spot you would name. Compare with theirs. The gap between the two lists is the useful part.
  2. Find a forecast in the news this week — energy, population, climate, anything with a number and a year. Name its model shape and its blind spot. Most published forecasts are one of the two models on this page in a suit.
  3. Return to the Week 3 biomass work with this week’s eyes. The cumulative emissions chart projects fifty years of constant generation and constant emission factors. Name three things it cannot represent, and say which would change the conclusion most.

Save your work

Copy the code you’re most proud of into your week8.R file. Commit and push via GitHub Desktop. Write a commit message that describes what you learned — not just “week 8”.

Keep your answer to Exercise 5. It is the first draft of your report’s limitations section, and it is much easier to write now than the night before the deadline.