Week 10: Common Pitfalls
HolmesCo’s greatest hits — and how to keep them out of your report
Introduction
You are about to write your final report, and then to review someone else’s. This page is the checklist for both.
It has two halves. The first is an audit: a paragraph of results that reads perfectly well and contains a number that is wrong. The second is a gallery of six ways a competent analysis turns into a bad report, with the refrains that catch each one.
Ten weeks of this module have been about one habit: checking the thing that looks fine.
The code boxes start empty, except for comments to guide you through the code that’s expected.
Work one comment at a time. If you get stuck:
- Hint 1 restates what you are trying to do.
- Hint 2 sketches the shape the code might take.
- Hint 3 gives you the code with the key pieces blanked out.
- Solution gives you the lot.
Take a real attempt before opening a hint: you will learn by doing, not by reading. The demonstrators are here to help.
Three things worth knowing:
- The boxes share one R session, exactly like a script. Anything you create in Exercise 1 is still there in Exercise 2.
- Variable names are up to you. The last thing your block prints is what gets checked, so end each block with the answer.
- Exercises have extensions underneath them. If you finish early, give them a stab. These are designed to equip you to tackle real coding problems you might face, using the documentation. We can help – but check the manuals first!
Ask on the class board, in Code Q&A. Someone else is likely stuck on the same thing, so asking in public helps them too. Demonstrators check it most weekdays, but they give classmates a chance to answer first.
The board is private to this class, so sign in to GitHub first. If the link shows “404 – page not found”, you’re not signed in, or you haven’t accepted the Classroom 50 invitation yet.
Exercise 1: Audit the paragraph
A classmate sends you this paragraph from their biomass briefing. It is well written, it cites its sources, and every number in it is the right kind of number.
“UK bioenergy generated 41.0 TWh of electricity in 2025, about 14% of total supply. Counted under the full supply-chain scenario rather than the official zero, that generation represents 43.5 Mt CO₂ — more than the 36.7 Mt the same electricity would have produced from coal. Shipping the 2024 pellets across the Atlantic adds a further 2.5 Mt, bringing the total to over 45 Mt.”
Three of those four figures are right. One is not, and it is out by enough to matter.
Find it and print the correct value.
Four claims, four short calculations. Do them in the order the paragraph makes them and stop when one disagrees.
Every one of the four is something you computed in Weeks 1 to 3. If a recomputation feels like more work than it should, you are probably about to find the error in it.
The four, in order:
- 14%: bioenergy over total generation for 2025, times 100.
- 43.5 Mt: 41.0 TWh, converted to MWh, times the biomass
with_supply_chainfactor, converted from kg to Mt. - 37.3 Mt: the same, with coal’s
officialfactor. - 2.5 Mt: the 2024 rows of
imports, tonnes times kilometres times 0.005, converted from kg to Mt.
The last one is Week 3 Exercise 1 exactly. If your answer disagrees with the paragraph, check your own units before you accuse anyone — then check theirs.
# Claim 1: 14% of supply
100 * elec$bioenergy_twh[elec$year == 2025] / elec$total_twh[elec$year == 2025]
# Claim 2: 43.5 Mt under the supply-chain scenario
sc <- factors$co2_kg_per_mwh[factors$fuel == "biomass" &
factors$scenario == "with_supply_chain"]
elec$bioenergy_twh[elec$year == 2025] * 1e6 * sc / 1e9
# Claim 3: 37.3 Mt as coal
coal <- factors$co2_kg_per_mwh[factors$fuel == "coal" &
factors$scenario == "official"]
elec$bioenergy_twh[elec$year == 2025] * 1e6 * coal / 1e9
# Claim 4: shipping
i24 <- imports[imports$year == 2024, ]
sum(i24$import_kt * 1000 * i24$transport_km * shipping_co2_per_tkm) / 1e9The first three check out: 14.0%, 43.5 Mt, 37.3 Mt.
The fourth is 0.25 Mt, not 2.5. A factor of ten, almost certainly a missed conversion between kilotonnes and tonnes or between kg and Mt.
Three things about this error are worth more than the arithmetic.
It points the author’s way. 2.5 Mt makes transport look like a serious contributor and gives the last sentence something to say. 0.25 Mt makes it a rounding error next to combustion. Errors that weaken your argument get caught; errors that strengthen it get published.
Nothing about the paragraph looks wrong. The units are stated, the sources are real, the arithmetic in “bringing the total to over 45 Mt” is correct given the faulty input. Prose fluency is not evidence, and this paragraph is more convincing than a sloppier one that happens to be right.
Only recomputation caught it. Reading harder would not have. When you review a classmate’s briefing this week, recompute at least one number rather than reading all of them — that is the review that finds something.
Extension: audit the rest of your own report
Optional, and the most useful thing you can do with the time.
- Take three numbers from your own draft and recompute each from the raw data, deliberately not looking at how you did it the first time. Independent recomputation is the only check that catches a repeated mistake.
- For each number, ask which direction an error would push your conclusion. Anywhere the answer is “in the direction I am arguing”, check twice.
- Write the factor-of-ten test into your own workflow: for every quantity, state the unit at every step in a comment. Almost every error in this course has been a unit, and units are the one class of mistake you can design out entirely.
Six ways a good analysis becomes a bad report
Each excerpt below is something a competent student could write on the way to a perfectly reasonable conclusion. For each, decide what is wrong before opening the reveal, and name the refrain that catches it.
Discuss as a desk, two minutes per pitfall. Keep notes in the box at the end — you will want them when you review someone’s briefing.
Pitfall 1: The conclusion the data cannot reach
“Our data shows that wind farm output varies by season, with lower generation in summer months. Therefore, wind energy is unreliable and should not receive government subsidies.”
Cover the word “Therefore” and read the two halves separately. Is the first half in dispute? Is the second half about the same subject?
The observation is uncontroversial and well known. The conclusion is about subsidy policy, and nothing between the two supports the journey.
“Varies by season” is not “unreliable” — that word smuggles in a standard nobody has stated. And a subsidy recommendation needs costs, alternatives, storage, and demand patterns, none of which are in the data. The analysis could be flawless and the paragraph would still be indefensible.
Watch for “therefore” and “which proves” in your own draft. They are where an argument most often steps off the evidence, and they are easy to grep for.
Pitfall 2: The figure that explains nothing
“As shown below, there is a clear trend in the data.”
[A scatter plot with no title, no axis labels, no units, no caption and no number. It is not mentioned anywhere else.]
Imagine the figure on its own, on a slide, with the text gone. What could a reader work out from it?
Missing: axis labels with units, a caption saying what the figure shows, a number so it can be cited, a reference in the text that is not “below” — page layout moves — and a statement of what the trend actually is.
The test is that a figure should survive being separated from the text. Yours will be: a reader flicks to the figures first, an examiner reads the caption before the paragraph.
“Clear trend” also does no work. Clear to whom, and how large? That is Pitfall 3 arriving early.
Pitfall 3: The p-value doing a job it cannot do
“The difference in solar panel output between the two regions was highly significant (p = 0.001), confirming that location affects performance.”
Suppose the difference were 0.1 W. Would the sentence still be true? Would it still be worth writing?
How much? The p-value says the difference is unlikely to be zero. It says nothing about size — and with a large enough sample, a difference of 0.1 W would be “highly significant” too. Week 6 built exactly that demonstration.
The sentence needs the effect size, a confidence interval, and a comparator that makes the number mean something.
“Confirming” is a second fault hiding behind the first. A significant test is evidence, not confirmation, and Week 6’s significant result evaporated when its assumptions were checked.
Refrain: is that a big number?
Pitfall 4: The word “prove”
“The results prove that biomass electricity is carbon-neutral under current accounting rules.”
What would have to be true of the evidence for “prove” to be the right word? And what does a reader do next, having read it?
Evidence supports or undermines; it does not prove. Even a result entirely consistent with carbon neutrality cannot rule out the explanations nobody tested.
Better: “the analysis is consistent with…”, or “on these assumptions, the evidence suggests…”.
The reason this matters is not etiquette. “Proven” closes the question, and the whole of this sentence’s content lives in the clause “under current accounting rules” — which is a choice, not a measurement, and is precisely the thing Week 3 showed can be moved from 0 to 1060 kg/MWh without touching a single observation.
Refrain: what assumptions are we making?
Pitfall 5: Methods nobody could follow
“We analysed the data using statistical tests in R and found significant results.”
Hand that sentence, and your data, to the person next to you. Could they produce your numbers?
Almost everything: which data, over what period, which variables; which test and why that one; what assumptions were checked and what happened; what was excluded, transformed or cleaned, and on what rule.
The standard is reproducibility: enough detail that someone with your data and your description gets your numbers. Week 9’s broken script was the same failure in code rather than prose.
One specific thing students omit almost universally: the checks that changed nothing. “Shapiro-Wilk did not reject normality for either group” is evidence you looked, and its absence is indistinguishable from not having looked.
Pitfall 6: The paragraph with no caveats
“Wind generation in the North East is 40% higher than the South East (p < 0.01, Cohen’s d = 1.2). We recommend prioritizing all future wind farm development in the North East.”
The statistics here are reported properly — better than most. So read the second sentence on its own and ask what it assumes that the first sentence did not establish.
The statistics are well reported, which is what makes this the most dangerous example on the page.
What is missing is everything the analysis could not see: confounders (turbine type, terrain, hub height, age); whether the sampled farms represent the regions; grid connection costs and planning constraints, which are not in the dataset and often decide siting; and the leap from “higher average output” to “all development”, which the word “all” cannot survive.
Week 7 is the direct precedent. Region explained 31% of the variation in capacity factor, so 69% was within regions — site selection inside a region mattered more than the choice of region, and an average difference does not license a blanket rule.
Refrains: what is the model not capturing? and compared to what?
Extension: find the pitfalls in the wild
Optional.
- Go back to the Week 4 exemplars — the faithful briefing and the traitor’s. Find one instance of each pitfall. The traitor is easier; the interesting finds are in the faithful one.
- Find a real press release from an energy company or a campaign group and mark it up. Most contain Pitfall 1 and Pitfall 3 in the first paragraph.
- Swap drafts with someone now, before the review session, and mark only for these six. Give them the list rather than a verdict — a review that names the category is one the author can act on.
Exercise 2: Match the refrain to the pitfall
Five questions have recurred since Week 1. They are the whole module compressed into things you can ask out loud:
| Refrain | |
|---|---|
| 1 | “Is that a big number?” |
| 2 | “Compared to what?” |
| 3 | “What assumptions are we making?” |
| 4 | “How plausible was this before we tested?” |
| 5 | “What is the model not capturing?” |
Print a vector of three numbers giving the refrain that most directly catches Pitfalls 3, 4 and 6, in that order.
c(1, 3, 5)- Pitfall 3 → “Is that a big number?” A p-value with no effect size. The refrain demands a magnitude, an interval and a comparator.
- Pitfall 4 → “What assumptions are we making?” Everything the sentence claims rests on “under current accounting rules”, and that rule is a choice worth 1060 kg CO₂ per MWh.
- Pitfall 6 → “What is the model not capturing?” The statistics are fine. The recommendation covers terrain, grid costs and planning, none of which the analysis contained.
The refrains overlap on purpose. “Compared to what?” also catches Pitfall 3, and Pitfall 6 needs it too. Five questions asked reflexively will catch more than any checklist you try to memorize — which is the point of them being short enough to say out loud.
Extension: which refrain do you skip?
Optional, and honest self-assessment is the exercise.
Then find, in your own draft, one place where your lowest-scoring refrain would change what you wrote. Change it. That single edit is worth more than another pass of proofreading.
Before you write
Three things, in this order:
- Recompute three numbers in your draft from the raw data, without looking at how you did them the first time.
- Read your own report for the six pitfalls above. Two of them will be in there.
- Restart R and run your script from the top. Fix what breaks.
Then commit everything to your repo via GitHub Desktop, and bring the draft to the application session for peer review.