AI, reproducibility, and integrity

Research Skills — Week 9

Recap

  • What AI does well — and badly
  • AI audit
  • Reproducibility
  • Reproducibility check
  • Scientific integrity
  • Open discussion
  • Wrap-up

Where we are

Your projects are nearly done. You’ve designed investigations, run tests, fitted models, and identified limitations.

This week: two forces that shape modern research —

AI and the pressure to produce results.

Questions?

Submit questions:

PollEv.com/geol

text geol to 07480 781235

What AI does well — and badly

  • Recap

🎓💬 Concept block 1

  • AI audit
  • Reproducibility
  • Reproducibility check
  • Scientific integrity
  • Open discussion
  • Wrap-up

AI is good at…

  • Generating first-draft code
  • Summarizing text
  • Suggesting structure
  • Spotting syntax errors
  • Explaining error messages

AI is bad at…

  • Knowing whether its answer is correct
  • Understanding your specific data
  • Making judgement calls about methodology
  • Citing sources reliably
  • Distinguishing meaningful from spurious results

Live demo (if possible)

🖥️ Give an AI the HolmesCo gold assay scenario

“We tested 10,000 soil samples with a 95%-accurate assay. 50 tested positive. What’s the probability of a true gold deposit?”

Watch what it does with the base rate.

The discovery is smaller than the error rate

A rising line showing positive tests expected from 10,000 samples as the proportion genuinely over gold increases. It starts at 500 positives with no gold at all. A dashed line far below marks the 50 positives HolmesCo reported.

The key distinction

AI is fluent but not thoughtful.

It can write a convincing paragraph about any result — including a wrong one.

AI audit

  • Recap
  • What AI does well — and badly

💬 Exercise 1

  • Reproducibility
  • Reproducibility check
  • Scientific integrity
  • Open discussion
  • Wrap-up

Your experience

Think about a time in this module when you used AI (or were tempted to). What did it help with? Where did you have to override it?

Common patterns:

  • Useful: debugging code, plot templates, explaining errors
  • Unreliable: interpreting results, choosing methods

Reproducibility

  • Recap
  • What AI does well — and badly
  • AI audit

🎓 Concept block 2

  • Reproducibility check
  • Scientific integrity
  • Open discussion
  • Wrap-up

The replication crisis

Many published findings don’t hold up when others try to reproduce them.

Bar chart of replication success in three fields: psychology 36%, experimental economics 61% (11 of 18), cancer biology 46% (51 of 112). Each bar is annotated with the criterion that project used, and the three criteria differ.

Why?

  • p-hacking — trying many analyses, reporting the one that “works”
  • Underpowered studies — too few observations to detect real effects
  • Selective reporting — filing away null results
  • Data not shared — no one can check your work

Your commit history IS your reproducibility record

If someone cloned your repo and ran your code, would they get the same results?

One number, or the shape it came from

A bootstrap interval — like the one you’ll build this afternoon in exercise w9_1, from uk_electricity.csv — is not a fact about the world. It is a range read off a distribution of 1,000 resampled means.

Two stacked panels sharing an x-axis of mean wind generation in TWh. The top panel shows a single point with a horizontal error bar, labelled 65.6 TWh, 55.2 to 75.1. The bottom panel shows a histogram of 1,000 bootstrap replicate means, roughly bell-shaped, with the same 55.2-to-75.1 range shaded.

Rerun it without the seed

This afternoon’s bootstrap draws random resamples. Run the same code twice with no set.seed() and the 95% interval moves both times:

Run 95% interval (TWh)
No seed, run 1 55.8 to 74.8
No seed, run 2 55.7 to 76.1
set.seed(2847), any run 55.2 to 75.1, always

Which of these three rows would you want printed in your report?

Practical reproducibility habits

  1. Script everything — no “I typed this in the console”
  2. Document your decisions — why this test, why you excluded those data points
  3. Record software versions — sessionInfo() in R
  4. Seed anything random — set.seed() before sample() or replicate(), once, not inside the loop
  5. Share data and code — within ethical limits

Reproducibility check

  • Recap
  • What AI does well — and badly
  • AI audit
  • Reproducibility

💬✏️ Exercise 2

  • Scientific integrity
  • Open discussion
  • Wrap-up

Can you re-run your analysis?

Open your project repo. Start a fresh R session.

Run your analysis script from scratch.

Does it work? What breaks?

Common problems:

  • Hardcoded file paths
  • Missing libraries
  • Steps done in the console but not saved

Fix what you can. 10 minutes.

Attendance

Scientific integrity

  • Recap
  • What AI does well — and badly
  • AI audit
  • Reproducibility
  • Reproducibility check

🎓💬 Concept block 3

  • Open discussion
  • Wrap-up

The spectrum

Three levels

Level What Example
Innocent errors Rounding, wrong column, misread output Your peer reviewer catches these
Questionable practices p-hacking, HARKing, selective reporting Often unintentional — this is what base rate and multiple comparisons sessions were about
Fabrication Making up or altering data Rare, career-ending, easier to detect than you’d think

Where does AI fit?

If AI writes your analysis and you don’t check it → you are responsible for the errors.

If AI generates data or figures that don’t reflect reality → that’s fabrication, even if unintentional.

The HolmesCo connection

HolmesCo doesn’t fabricate data.

But their corner-cutting on design, analysis, and interpretation produces conclusions that are functionally no better than fabrication.

Integrity isn’t just “don’t lie.”

It’s “do the work properly.”

Open discussion

  • Recap
  • What AI does well — and badly
  • AI audit
  • Reproducibility
  • Reproducibility check
  • Scientific integrity

💬 Q&A

  • Wrap-up

Your questions

This is a looser session. What’s on your mind?

  • Questions about your reports?
  • AI use policies?
  • What counts as plagiarism when AI is involved?

Wrap-up

  • Recap
  • What AI does well — and badly
  • AI audit
  • Reproducibility
  • Reproducibility check
  • Scientific integrity
  • Open discussion

Key points

  1. AI is useful for mechanics; unreliable for judgement
  2. Reproducibility = scripted, documented, version-controlled
  3. Integrity = doing the work properly, not just not lying

Any questions we missed?

Submit questions:

PollEv.com/geol

text geol to 07480 781235

Next time

Application session: “AI as critic”

You’ll feed your analysis to an AI and evaluate its feedback.

The question: how much should you trust it?