Advertisement

Subject workflows · Updated 14 September 2026

How to Study Epidemiology With AI (Safely)

The arithmetic takes a minute and checks itself. The sentence you write about it carries the grade — and that sentence is the one thing a chatbot is trained to get wrong.

Advertisement

Epidemiology hands you two jobs that look like one. Getting a risk ratio out of a two-by-two table is a minute of division you can check by hand; deciding what that number entitles you to say — caused, raised the risk of, was associated with — is where the marks are, and it is settled by the study design rather than by the number. A chatbot is a decent calculator and an unreliable epidemiologist, because it learned its causal language from a research literature that does not hold that line either. So the split for this course is unusually clean: do the arithmetic yourself because you can verify it, and never let a model write the interpretation sentence.

Red and blue wooden peg figurines of different sizes standing on a wooden desk in front of a bright window, with three small plain wooden figures in a shallow tray, a potted aloe, a brass magnifying glass, a blue flowering plant and a blue mug

Everything downstream is decided by the comparison group

The free CDC self-study course Principles of Epidemiology in Public Health Practice puts the whole subject in two sentences. "Analytic epidemiology is concerned with the search for causes and effects, or the why and the how," it says, and then: "The key feature of analytic epidemiology is a comparison group."

That second sentence is the one to memorise, because the shape of the comparison determines which measure you are allowed to compute. The CDC course draws the divisions plainly: in an experimental study "the investigator determines through a controlled process the exposure", while in an observational study the epidemiologist "simply observes the exposure and disease status of each study participant." The observational designs then differ in what participants were selected on. A cohort study "records whether each study participant is exposed or not, and then tracks the participants to see if they develop the disease"; a case-control study starts "by enrolling a group of people with disease" and then a group without; a cross-sectional study measures exposures and outcomes "simultaneously."

Read that as a set of permissions. A cohort study follows people forward from exposure, so new cases can be counted in each group and a risk is available to you. A case-control study recruited on the outcome, so the proportion who are ill is an artefact of how many cases the investigators chose to enrol — no risk can be computed from it at all, and what you get instead is an odds ratio. The CDC course names the one circumstance in which substituting one for the other is comfortable: "when the health outcome is uncommon, the odds ratio provides a reasonable approximation of the risk ratio." When the outcome is common it approximates nothing, and a sentence calling it "1.8 times the risk" is simply wrong.

The exam version. Most marks lost here are not lost in the arithmetic. They are lost by computing a legitimate number and then describing it as a measure the design could never have produced.

Why a chatbot is worst at exactly this step

This has actually been measured. In "Can Large Language Models Infer Causation from Correlation?", Jin and colleagues built a benchmark for precisely the operation your course grades: give a model correlational statements and ask it to determine the causal relationship between the variables. They curated "a large-scale dataset of more than 200K samples", evaluated seventeen existing models, and found that "these models achieve almost close to random performance on the task." Fine-tuning helped only in-distribution — the models "fail in out-of-distribution settings generated by perturbing these queries", so a rewording was enough to break them.

Newer and more clinical work points the same way. A study published in NPJ Digital Medicine in April 2026 tested causal reasoning across 99 clinically grounded laboratory scenarios mapped to Pearl's ladder of causation — association, intervention and counterfactual. Both models tested "performed best on intervention questions and worst on counterfactuals", and the authors state the general problem directly: "Current AI models, including LLMs, are not well equipped to handle confounding, selection bias, and other challenges inherent in real-world data."

Counterfactual reasoning is not an exotic corner of this course. It is the question "would these people have got ill anyway?" — which is to say it is confounding, selection and reverse causation, the three things every epidemiology exam asks you about. The weakest rung is the whole syllabus.

The blur is in the training data, which is why it sounds so good

Here is what makes this different from the usual warning about invented facts: the model is not making it up. It is faithfully reproducing a habit that runs through the literature it was trained on.

In 2022, Haber and colleagues published "Causal and Associational Language in Observational Health Research: A Systematic Evaluation" in the American Journal of Epidemiology. They screened 1,170 articles from 18 high-profile journals — the NEJM, The Lancet and JAMA among them — published between 2010 and 2019, and had reviewers rate how much causality each abstract's language implied. Just over half were rated moderate or strong. Yet the dominant linking word was "associate", in 45.7% of abstracts, against 0.8% for "cause".

Two findings matter for how you use a chatbot. The first is the disconnect: for 44.5% of articles the action recommendation implied more causality than the study's own linking sentence, and only 40.3% were commensurate. Papers hedged in one paragraph and told you what to do in the next. The second is that the hedge does not land anyway — over half of reviewers rated "association" as carrying at least some causal implication.

Now put that next to how a language model works. It produces the likely continuation, and in this corpus the likely continuation after a reported association is a mildly causal recommendation written in cautious-sounding vocabulary. That is a recurring regularity, not a rare fact, so the model has learned it well and will produce it fluently and without hesitation. Compare immunology, where the danger is the opposite kind of fact — a one-off name the model cannot hold. Here the danger is a pattern it holds perfectly. It will hand you a conclusion that reads exactly like the ones in your reading list, and it may still be a claim your study design cannot support.

Writing "associated with" is not a fix

The obvious response is to hedge everything: never write "causes", write "is associated with", hope the marker is satisfied. Haber's data is the argument against it — over half the reviewers read "association" as carrying causal weight anyway, and the papers with cautious linking sentences were the same ones making stronger recommendations two lines later. A verb is not a design.

What licenses a claim is the structure of the comparison and what was controlled. So the order of operations is: name the design, name what could have produced this result other than the exposure, decide what survives, and only then choose the verb. A sentence built in that order is defensible in any wording; one built verb-first is not, however carefully hedged.

One more trap, for the marks at the top

If your course reaches multivariable regression, learn the Table 2 fallacy by name. Westreich and Greenland described it in the American Journal of Epidemiology in 2013: when one model's output is printed as a single table, readers read the confounder coefficients as though they were findings in their own right. That invites "confusion of direct-effect estimates with total-effect estimates for covariates in the model" — and, the sharper point, "these effect estimates may also be confounded even though the effect estimate for the main exposure is not confounded."

It is worth knowing because of how a chatbot behaves: paste a regression table in and ask what it shows, and it will walk down every row interpreting each line as a finding. That is the fallacy, performed on request. Those variables were adjusted for to clean up one estimate, not to measure them.

Split the job

TaskWho does itWhy
Compute a risk ratio or odds ratio from a two-by-two tableYou, by handIt is one minute of division and you can check it by recomputing
Decide which measure the design permitsYou, from the designThe commonest lost mark in the course, and not a calculation
Write the conclusion sentenceYou, alwaysNear-random on inferring causation from correlation across 17 models
Generate candidate confounders to considerThe chatbot, for breadthRecall of plausible common causes is a job where a long list helps
Decide which confounders were handledYou, from the methods sectionOnly the paper knows what it adjusted for
Quiz you on rival explanationsThe chatbot, as examinerAsking relentlessly is the one job where it has no way to be wrong

The workflow, step by step

Four prompts. Paste them into ChatGPT, Claude, Gemini or whatever you already use. Each keeps the interpretation on your side of the desk.

1. Design before anything else. Nothing can be computed until this is settled, so make the tool refuse to move on.

You are my epidemiology design coach. Here is a study I am reading,
or a question I have been set:

[paste the abstract or the question]

Do NOT name the study design, do NOT compute anything, and do NOT
interpret any result.

Ask me one question at a time, waiting for my answer each time:
1. Who was compared with whom? Name the two groups.
2. Did the investigators assign the exposure, or only observe it?
3. Were people selected on their exposure, on their outcome, or on
   neither?
4. Given my answers, which measure of association can be computed
   here - and which one cannot?

If an answer of mine is wrong, say only that it is wrong and ask
again a different way. Do not correct me.

2. Audit your own conclusion sentence. This one turns the Haber finding into a two-minute check you can run on anything you are about to submit.

Here is a conclusion sentence I wrote:

[paste your sentence]

Do not rewrite it, do not grade it, and do not tell me whether it
is correct.

Do this instead:
1. Quote the exact word or phrase I used to link the exposure to
   the outcome.
2. Rate how much causal claim a reader would take from that phrase:
   none, weak, moderate or strong.
3. Ask me one question - what feature of the study design licenses
   that strength? Then stop and wait.
4. Afterwards, quote any recommendation or implication I made, and
   tell me whether it implies MORE causality than my linking phrase
   did.

Do not answer question 3 for me.

Step 4 is the important one: it is the exact disconnect Haber's reviewers found in 44.5% of published papers, and far easier to catch in your own writing when something quotes the two sentences back at you side by side.

3. Confounders, without a verdict. Breadth is genuinely useful here; judgement is not delegable.

I am studying this exposure and outcome:

Exposure: [   ]
Outcome: [   ]
Population: [   ]

List candidate confounders - things that could plausibly cause both
the exposure and the outcome.

Rules:
- Aim for breadth, obvious and non-obvious.
- For each one, give a single clause saying why it could cause BOTH.
- Do NOT tell me whether the study I am reading adjusted for it.
- Do NOT rank them or say which matter most.
- Do NOT draw any conclusion about the association.

Then stop.

Take that list to the methods section and check each one yourself. A candidate the paper never mentions is the interesting kind.

4. Drill the rung it is worst at. Counterfactual reasoning was the weakest category in the 2026 clinical evaluation, so practise it as a question generator rather than an answer source.

Quiz me on rival explanations, one question at a time.

The finding is: [paste it, e.g. "coffee drinkers had 1.4 times the
rate of the outcome"]

Ask me these one at a time, waiting for my answer each time:
1. If the exposure did nothing at all, what else could have produced
   this number?
2. Who was left out of this study, and would including them move the
   result?
3. Which came first, the exposure or the outcome - and how does the
   design tell you?
4. If the outcome is common, what does that do to an odds ratio?

Do not give me the answers. After each answer of mine, ask one
follow-up beginning "what in the study design rules that out?"

Where the line is

The split above sits comfortably inside most academic integrity policies, and the FAQ below covers the ordinary cases. The risk specific to this subject is subtler than copying: a chatbot-written interpretation paragraph reads like competent epidemiology, so it is unusually easy to submit one without noticing you never formed the judgement — and unusually hard to defend when asked a follow-up question. Read your syllabus, and if the wording is vague, our guide to homework help without cheating and the class AI policy checklist cover how to read it and how to ask in writing.

Related reading

FAQ

Can I use ChatGPT to interpret a risk ratio or an odds ratio?

Use it to check your arithmetic, not your sentence. Jin and colleagues built a benchmark of more than 200,000 samples testing exactly this skill — reading correlational statements and determining the causal relationship — evaluated seventeen models, and found they "achieve almost close to random performance on the task". Fine-tuned versions failed as soon as the wording was perturbed. The calculation is a minute of division you can verify by recomputing it; the interpretation has no such check, so it stays with you.

Why can a case-control study not give me a risk ratio?

Because the investigators chose how many ill and how many healthy people to enrol, so the proportion who are ill in the study is a recruitment decision rather than a fact about the population. You cannot get a risk out of it, which is why case-control studies report an odds ratio. The CDC's Principles of Epidemiology course notes when substituting one for the other is defensible: "when the health outcome is uncommon, the odds ratio provides a reasonable approximation of the risk ratio." If the outcome is common, describing an odds ratio as a multiple of risk overstates the effect.

Is it enough to write "associated with" instead of "caused"?

No, and there is evidence on this. In a 2022 American Journal of Epidemiology review of 1,170 articles across 18 high-profile journals, "associate" was the main linking word in 45.7% of abstracts — yet over half the reviewers rated "association" as carrying at least some causal implication. The same review found that for 44.5% of articles the action recommendation implied more causality than the paper's own linking sentence. Hedging vocabulary does not do the work; the study design does. Name the design and the rival explanations first, then pick the verb.

What is the Table 2 fallacy?

It is reading the confounder coefficients in a regression table as though each were a finding about that variable. Westreich and Greenland named it in the American Journal of Epidemiology in 2013, warning that presenting exposure and confounder estimates from one model invites "confusion of direct-effect estimates with total-effect estimates for covariates in the model", and that "these effect estimates may also be confounded even though the effect estimate for the main exposure is not confounded." Watch for it when you paste a results table into a chatbot, because interpreting every row is exactly what it will do.

Is it cheating to use AI for an epidemiology assignment?

It depends on your course policy, but the layers separate cleanly. Having a model write your interpretation or discussion section substitutes its judgement for yours under almost any policy, and pasting an active assignment into a public tool may breach the policy by itself. Being quizzed on rival explanations, having your own conclusion sentence interrogated, or asking why a design permits one measure and not another is ordinary studying. Read the syllabus, and ask your instructor in writing when it is unclear.

Bottom line

Epidemiology is not a hard arithmetic course. It is a course about the discipline of saying only what your comparison group entitles you to say — and a language model is fluent in precisely the register that discipline exists to resist, because it learned that register from the same journals you are being marked against. Do the two-by-two by hand, take the design question seriously before anything else, and treat the tool as an examiner rather than an author. The sentence at the end is the assignment.

Advertisement
Free download: Grab the one-page AI Study Safety Checklist — everything to check before you upload, trust, or submit anything involving AI.
Advertisement