Subject workflows · Updated 19 September 2026
How to Study Developmental Psychology With AI (Safely)
Your textbook prints an age beside every stage and every milestone, and those ages are the easiest thing in the course to ask a chatbot for. They are also the part most likely to have moved since somebody wrote them down.
Use a chatbot for this course, but treat any age it gives you as half an answer until you know two more things: what the number is the age of, and which children it was measured on. A milestone age is not a constant like a boiling point. It is a percentile drawn from a particular sample, and both of those can change without the behaviour changing at all. Models hand back the bare number because that is how the number appears nearly everywhere they have read it.
The same behaviour, two different correct ages
In 2022 the CDC's Learn the Signs. Act Early. programme funded the American Academy of Pediatrics to convene an expert group and rebuild its developmental checklists. The resulting paper in Pediatrics describes the goal as being to "clarify when most children can be expected to reach a milestone (to discourage a wait-and-see approach)." The working group set eleven criteria, the first of which was "using milestones most children (75%) would be expected to achieve by specific health supervision visit ages."
Note what changed there. The older lists were built on the median — and, as the paper puts it, "lists based on median (50th percentile) age milestones might encourage a wait-and-see" approach. Nobody discovered that children had started walking later. The question was swapped. "When does a typical child do this?" and "by when should most children have done this, such that not having done it is worth a second look?" are different questions about the same behaviour, and they have different numbers as answers.
The rebuild was not cosmetic: applying the new criteria "resulted in a 26.4% reduction and 40.9% replacement of previous CDC milestones", and of the milestones retained but shifted to a different age, 67.7% "were moved to older ages." An unrevised pre-2022 source is therefore not randomly wrong. It skews early.
Whose children, though
The second half of every age is the sample, and developmental psychology has an unusually well-documented problem here. Mark Nielsen, Daniel Haun, Joscha Kärtner and Cristine Legare surveyed every article published between 2006 and 2010 in Child Development, Developmental Psychology and Developmental Science — the three highest-impact experimental developmental journals — and published the count in the Journal of Experimental Child Psychology in 2017.
Across 1,582 articles, 90.52% of participant samples were WEIRD: Western, educated, industrialised, rich and democratic. "Less than 3% of the participants contributing to the expansion in our knowledge of children's psychological development came from all of Central and South America, Africa, Asia, and the Middle East and Israel combined" — regions containing roughly 85% of the world's population. Eleven articles featured participants from Central and South America and ten from Africa, which as the authors point out is "fewer than the 23 articles devoted to non-human primates and fewer than the 15 articles that did not specify where their participants were from."
Nor is it history. They re-ran the count on 2008 and on 2015 and got 91.67% and 92.37% — no movement in seven years.
What this costs you in an exam is concrete. The paper's worked example is a 2006 study in which Hai//om children searching for an object hidden among overturned cups used a geocentric strategy, tracking position relative to the surrounding environment, where Western children use an egocentric one. The authors' conclusion is the sentence to carry into the exam hall: "If Haun and colleagues had tested only Western children, then it could have easily been assumed, as is typically written, that children employ egocentric search strategies. But children generally do not do so; only children from specific cultural backgrounds do."
Their recommendation is that where a sample is homogeneous, "consideration should be given to the notion that whatever is being reported may be culturally specific, and hence possibly unrepresentative and not generalizable, and this should be openly acknowledged in print." That acknowledgement is often the sentence your marker is looking for, and the one a chatbot is least likely to volunteer: "children do X by age Y" is simply how this literature writes, and a model reproduces the house style along with the content.
The classic vignettes are the ones it has already read
The second trap in this subject is about the tasks rather than the ages. Developmental psychology is taught through a small set of canonical scenarios — the false-belief task, conservation of liquid, A-not-B, the unexpected-contents test — and those appear verbatim across an enormous amount of text.
That is measurable rather than speculative. The team behind TMBench found that models would reproduce Sally-Anne test scenarios closely matching the originals, so they built a from-scratch inventory "to strictly avoid data leakage" — on which "even the most advanced LLMs like GPT-4 lag behind human performance by over 10% points."
The sharpest demonstration of what memorising a vignette looks like came from Tomer Ullman at Harvard, who took an unexpected-contents task a model had passed and changed one detail at a time: the bag is transparent; a trusted friend has already told Sam what is inside; Sam filled and labelled the bag herself; and the simplest of all — Sam cannot read. "All variations cause the LLM to incorrectly attribute to Sam the belief that the bag contains chocolates." A five-year-old who has understood the task does not get the unreadable-label version wrong.
One caveat, stated plainly: that work tested a 2023 model generation, current models handle the canonical forms far better, and how robust they are to variants is still argued over — a 2025 systematic review concluded only that models show "emerging competence" while "significant gaps persist." What is not in dispute is the part that affects you. The textbook wording of these tasks is in the training data, so asking a model to explain the false-belief task is the easiest question you can ask it, and a fluent answer tells you nothing about whether you understand the concept. Your exam will change the details.
Two free lookups worth more than a prompt
The CDC publishes its revised checklists free, covering 2 months to 5 years, as printable PDFs and in the free Milestone Tracker app. Read them for the framing rather than as an answer key: every item is phrased as what "most children" do, which is that 75%-by-this-visit criterion made visible, and the CDC states plainly that these resources "are not a substitute for standardized, validated developmental screening tools." Your exam is marked against your course's own normative table, not a surveillance checklist, and knowing those are two instruments answering two different questions is itself a chunk of the syllabus.
For the theory, OpenStax publishes a full Lifespan Development textbook free to read online under a Creative Commons licence — a second opinion you can check a chatbot against in thirty seconds, which is faster than arguing with one.
Split the job
| Task | Who does it | Why |
|---|---|---|
| Give you the age for a milestone or stage boundary | Your course's own table | The criterion behind the number varies by source, and your marker is using one specific source |
| Explain the sequence and the mechanism of a stage theory | The chatbot | Order and reasoning recur constantly in the corpus and are checkable against your notes |
| State which sample a finding came from | The paper, or your textbook's citation | Sample details are one-off facts, and 90% of this literature ran on one narrow population |
| Rewrite a classic task with all new surface details | The chatbot, then you verify it still works | Generating a variant is easy for it; you get an untested version to answer yourself |
| Tell you whether a variant task is still a valid test | You | This is the comprehension the exam is actually measuring |
| Flag where your own writing states a finding as universal | The chatbot, but only if instructed | Left alone it writes in the house style of the literature, which is universal by default |
The workflow, step by step
Four prompts for ChatGPT, Claude, Gemini or whatever you already have open. None of them asks for an age on its own.
1. Never accept a bare number. This is the thesis of the whole page turned into a prompt you can reuse all term.
I am going to ask about the timing of a developmental milestone
or stage boundary. Do not give me a single age.
Topic: [the behaviour or stage transition]
Answer in exactly this format:
1. AGE OR RANGE - the figure as commonly reported.
2. CRITERION - what that figure is the age of. Median? The age
by which some percentage of children have done it? If so,
which percentage? Clinical consensus rather than data?
3. SOURCE TYPE - a norming study, a screening tool, a textbook
summary, or a clinical guideline.
4. SAMPLE - which children this was measured on: country,
language, rough year, sample size if known.
5. COMPETING FIGURES - any other age in circulation for the
same behaviour, and what makes it different.
Write NOT STATED for anything you cannot support. Do not
estimate a percentile or a sample to fill a gap.
Rows 2 and 4 are the ones to read. If both come back filled in, you can quote the figure safely. If either comes back NOT STATED, you have learned the useful thing — the number is floating free of its criterion — and you go to your course table instead.
2. Make it rewrite the classic task, not explain it. This is the drill that gets you off the memorised version.
Take this classic developmental task: [e.g. the false-belief
task / conservation of liquid / A-not-B].
Rewrite it as a short scenario in which EVERY surface detail is
different: different names, objects, setting, and container or
material. Keep the logical structure identical.
Rules:
- Do NOT tell me the correct answer.
- Do NOT tell me which stage or age it relates to.
- Do NOT explain what the task is testing.
- End with the single test question and stop.
Then wait. After I answer, tell me only whether my answer
matches the logic of the original task, and if it does not,
name the one feature of the scenario I ignored.
Now do the second half yourself: before you answer, check that the rewrite is still a real test. If it has quietly made the container transparent or given the character a way of knowing, it is no longer the same task — and spotting that is precisely the understanding you are being examined on.
3. Audit your own essay for smuggled universals. Built directly on what Nielsen and colleagues ask authors to do.
Below is a passage from my essay.
Go through it sentence by sentence. For every sentence that
makes a claim about how children develop, output:
- the sentence
- whether it names a population, an age range, or a study, or
states the claim as universal
- if universal, the smallest edit that would make it honest
Do not rewrite my passage. Do not add new claims, studies or
citations. Do not tell me whether the argument is good.
Passage: [paste]
Do not paste its suggested edits back in — write them yourself, in your own sentences. You will find most of the flagged lines are ones where you simply wrote "children" and meant "the children in that study."
4. Drill the order, not the ages. The sequence is the robust part of a stage theory; the ages are the fragile part, and only one of the two survives a rewording.
Examine me on [theory or domain] by sequence only.
Ask me, one at a time, waiting for my answer each time:
1. What comes immediately before this stage, and what marks the
transition?
2. What can a child do at this stage that they could not before?
3. What is the classic evidence for that transition?
4. What would we expect to see if the order were wrong?
Rules: do not mention any ages at any point. Do not give me the
answers, even partially. If I am wrong, say "not yet" and ask
the same thing another way.
At the end, list the transitions I could not explain.
If you can explain why each transition happens and what evidence marks it, the ages become something you look up rather than something you have to trust a chatbot for.
Where the line is
Everything above is studying: being quizzed, reasoning through a fresh scenario, having your own prose challenged. Writing the essay is the graded work and it stays yours. The risk particular to this course is not getting caught — it is a confident paragraph built on an age that turns out to be a pre-2022 median, or a "children do X" that describes one country. Neither reads as cheating; both read as not having checked. If your syllabus is vague, our guide to homework help without cheating and the class AI policy checklist cover how to read it and how to ask in writing.
Related reading
- The rest of psychology. Social psychology, where the question is whether a finding survived rather than whom it was measured on, plus psychology case studies and cognitive theories and memory curves and spaced repetition.
- Reading the numbers properly. Hypothesis testing and p-values, Bayes and conditional probability — percentiles are the whole game on this page — and epidemiology, where what a study design lets you claim is the graded skill.
- Study design and methods. Sociology research methods and survey design covers sampling from the other direction.
- Checking anything a model tells you. The hallucination checklist, verifying AI answers before you study them, and evaluating sources for research papers.
- Writing it up. Literature reviews and the synthesis matrix — add a column for the sample, and the essay half writes itself.
FAQ
Why do different sources give different ages for the same milestone?
Usually because they are answering different questions. An age can be the median, meaning half of children have done it by then, or the age by which some larger share have, or a clinical consensus with no norming study behind it at all. When the CDC and the American Academy of Pediatrics revised their checklists in 2022 they switched from the 50th percentile to the age by which 75% of children would be expected to reach a milestone, because median-based lists "might encourage a wait-and-see" approach. Same behaviour, different criterion, different number. Always ask which one you are looking at.
Are the milestone ages in my older textbook wrong?
Not necessarily wrong, but if they predate 2022 they are likely to skew early relative to the current CDC checklists. That revision cut 26.4% of previous CDC milestones and replaced 40.9% of them, and of the milestones that were retained but moved to a different age, 67.7% moved to an older age. Study from your own course's table, because that is what you are marked against, and treat any age a chatbot offers as a prompt to check that table rather than as a replacement for it.
Can I trust a chatbot to explain the false-belief task?
To recite it, yes — that is the problem. The canonical wording of these tasks appears throughout the training data, and the TMBench researchers found models could reproduce Sally-Anne test samples closely matching the originals, which is why they built a leak-free inventory instead. On that clean set GPT-4 lagged human performance by over ten percentage points. Tomer Ullman also showed that trivial alterations to an unexpected-contents task, such as stating that the character cannot read, flipped a model from passing to failing. That result used a 2023 model and newer ones do better, but the lesson holds: a fluent explanation of the standard version is not evidence that either of you understands it.
Does it matter which children a study used?
In this subject, yes, and it is often worth marks. A survey of every article published over five years in the three leading developmental journals found 90.52% of samples were WEIRD, with under 3% of participants from Central and South America, Africa, Asia and the Middle East combined — regions holding about 85% of the world's population. There were fewer articles featuring participants from Africa than articles about non-human primates. The authors' recommendation is that findings from homogeneous samples should be acknowledged as "possibly unrepresentative and not generalizable", and writing that acknowledgement into your essay is usually the difference between describing a finding and evaluating it.
What is the single best habit for this course?
Refuse bare numbers. Whenever an age turns up — from a chatbot, a revision site or your own memory — ask what the figure is the age of and which children it came from. If you cannot answer both, you do not yet have a fact you can put in an essay. Learning the sequence and the mechanism of each stage is the part that is stable; the ages are the part that moves, and they are the part to look up rather than memorise.
Bottom line
Developmental psychology looks like a subject made of numbers, and that is what makes it easy to study badly with a chatbot. Every age in it is a percentile from a sample, and both halves have moved in the last few years — the CDC changed which percentile it publishes, and the field is still arguing about whose children it has been measuring. So ask for the criterion and the sample, never the number alone. Have the model rebuild the classic tasks with new details instead of explaining the ones it has already memorised. And drill the sequence, which is the part that will still be true when the ages get revised again.