Subject workflows · Updated 17 September 2026
How to Study Social Psychology With AI (Safely)
This is the field that went looking to see whether its own famous results held up, and found that many of them did not. That history is not trivia any more — it is what your exam asks you to reason about, and it is the one place a chatbot will quietly take your side.
Use a chatbot for this course, but ask it one thing before anything else: has this finding replicated? Asked that directly, models are good at flagging a claim the evidence has moved past. Asked the same thing indirectly — with the claim already assumed inside a normal-sounding study question — they measurably tend to go along with it. So the check has to be a separate, deliberate step, not something you hope comes up. Every other habit on this page follows from that one.
The canon got smaller, and your syllabus knows it
In 2015 the Open Science Collaboration ran direct replications of 100 studies from three psychology journals and published the result in Science. Their own summary: "Ninety-seven percent of original studies had significant results (P < .05). Thirty-six percent of replications had significant results." Average effect sizes fell by about half, from r = 0.403 to r = 0.197 — "a substantial decline."
Many Labs 2, published in 2018, went at it with overwhelming force: 28 classic and contemporary findings, protocols peer reviewed in advance, 15,305 participants from 36 countries. Fifteen of the 28 produced a significant effect in the same direction as the original. The median effect size went from d = 0.60 in the originals to d = 0.15 in the replications. And the detail worth carrying into an exam: "9 effects (32%) were in the direction opposite the direction of the original effect."
Not weaker. Backwards. Three of the best-known casualties are ones you will meet by name this term:
- Ego depletion. Twenty-three labs and 2,141 participants ran a preregistered replication of the standard sequential-task design. The pooled effect was d = 0.04, with a confidence interval running from −0.07 to 0.15 — straight through zero.
- Power posing. In 2016 Dana Carney, first author of the original study, posted a statement on her Berkeley faculty page: "I do not believe that 'power pose' effects are real", and "the evidence against the existence of power poses is undeniable."
- The Stanford Prison Experiment. Thibault Le Texier went through the study archives and interviewed 15 participants. His 2019 paper in American Psychologist, free to read as a preprint, reports that the guards "received precise instructions regarding the treatment of the prisoners," were not told they were subjects, and that participants "were almost never completely immersed by the situation."
Why the training data leans toward the confident version
Here is the part that decides how you should use a chatbot. Le Texier's paper opens by noting that the Stanford study "has been criticized on many grounds, and yet a majority of textbook authors have ignored these criticisms in their discussions of the SPE, thereby misleading both students and the general public about the study's questionable scientific validity."
If the textbooks still tell it straight, the rest of the written world tells it straighter. Revision notes, popular articles, talk write-ups, quiz sites — the bulk of text about social psychology is the memorable version of each study, not the correction. A model trained on that text has read the confident story many times over and the retraction once or twice. That is a claim about what has been written rather than a measurement of any particular model, but it sets the odds you are betting against.
The model knows — until you stop asking it straight
The behaviour has been measured, in an adjacent corner of education. Eva Richter, Markus Spitzer and colleagues tested whether chatbots could catch neuromyths — the learning-styles-and-left-brain folklore that circulates among teachers — and published the results in Trends in Neuroscience and Education in 2025. Two findings sit side by side.
First: "LLMs outperformed humans in identifying neuromyth statements as used in previous studies." Put a myth in front of a model as a statement to judge, and it judges it better than most people do.
Second: "when presented with applied user-like questions comprising misconceptions, they struggled to highlight or dispute these" — which the authors attribute to "LLMs' tendency toward sycophantic responses." Bury the same myth inside a practical question and the correction never arrives.
That study is about neuromyths rather than social psychology findings, and the difference is worth being clear about. What it measures is a behaviour — whether a model volunteers a correction nobody asked for — and that depends on the shape of the question, not the subject. "Explain how ego depletion affects revision timetables" has exactly the shape that failed.
The third finding is the useful one, and it is counterintuitive enough that almost nobody does it: "explicitly asking LLMs to correct unsupported assumptions increased the likelihood that misconceptions were flagged considerably, while prompting the models to rely on scientific evidence had only little effects." The instinct — "base your answer on peer-reviewed research" — is close to useless. "Correct any unsupported assumptions in my question" is the one that works.
It will not be your participants either
The same agreeableness shows up in the other job students hand these tools: predicting what people would do. Ziyan Cui, Ning Li and Huaikang Zhou replicated 156 scenario-based experiments from top social science journals using three models, publishing in Nature Computational Science in 2025. The headline number looks encouraging — main effects reproduced 73–81% of the time — but two details are not. Effect sizes came out "approximately 2-3 times higher than human studies." And: "When original studies reported null findings, LLMs produced significant results at remarkably high rates (68-83%)."
Ask a model how participants would respond to your manipulation and it will mostly find your effect, including when the honest answer is that nothing happens. The authors also report "significantly lower replication rates for studies involving socially sensitive topics such as race, gender and ethics" — which is most of the second half of the syllabus.
The tool that settles it in thirty seconds
You do not have to adjudicate any of this yourself. Lukas Röseler and a large team built the Replication Database and described it in the Journal of Open Psychology Data in 2024: a free, searchable platform "hosting 1,239 original findings paired with replication findings," designed to make replications "visible, easily findable via a graphical user interface." Search the effect, see whether anyone has retried it and what happened. It will not cover everything on your reading list, but when it has an entry, that entry beats anything a chatbot will tell you — and it takes less time than writing the prompt.
Split the job
| Task | Who does it | Why |
|---|---|---|
| Decide whether a finding still stands | You, from a database or the paper | A model follows the framing of the question you asked |
| Explain what mechanism a theory proposes | The chatbot | Mechanism is repeated constantly in the corpus and is checkable against your notes |
| Supply a study's sample size, design or effect size | The source paper | One-off numeric details are the least reliable thing to recall from memory |
| Flag a false premise in your own question | The chatbot, but only if ordered to | An explicit correction instruction works; "use good evidence" barely does |
| Predict how real participants would behave | You, then the actual data | Models inflate effects and turn nulls significant 68–83% of the time |
| Argue the opposite of your interpretation | The chatbot, as sparring partner | The one job where agreeableness cannot help it, because you assigned the side |
The workflow, step by step
Four prompts for ChatGPT, Claude, Gemini or whatever you already have open. Each is built to stop the model agreeing with you.
1. Run the replication check cold, before you study the finding. Neutral wording, nothing assumed.
I am going to name a finding from social psychology. Do not
explain it and do not tell me why it matters.
Finding: [name it in the most neutral words you can - no
"the well-known", no "the classic"]
Answer only these, in this order:
1. State whether the original effect has been subject to direct
or multi-lab replication attempts that you are aware of.
2. For each attempt you name, give the year and what it found.
3. Say plainly whether the current evidence supports the effect,
is mixed, or points against it.
4. Mark anything you are not confident about as UNCERTAIN, and
say what I should search to confirm it.
Do not soften the answer to be encouraging.
Check step 2 against the Replication Database or Google Scholar before trusting it. The answer is not authoritative; the point is that asking coldly is the only framing under which you get a straight one.
2. The preamble to paste above every ordinary study question. This is the whole paper turned into four lines.
Before answering anything below, do this first:
List every empirical claim my question assumes to be true.
For each one, say whether it is well supported, contested, or
has failed to replicate. If any assumption in my question is
unsupported, say so explicitly and correct it before you answer.
Only then answer the question.
My question: [paste your normal question]
Keep it in a note on your phone. It costs four lines and it is the single change with measured evidence behind it — and note what it does not say. Asking for peer-reviewed sources is not what moved the numbers. Asking for unsupported assumptions to be corrected is.
3. Make it quiz you on the five facts, without handing them over. This is the exam-answer drill.
You are examining me on one study: [name it].
Ask me, one at a time, waiting for my answer each time:
1. What was the hypothesis, in one sentence?
2. Who were the participants, how many, and from where?
3. What exactly was manipulated, and what was measured?
4. How large was the reported effect?
5. What has happened when others tried to repeat it?
Do NOT give me any of these answers, not even partially, and do
not hint. If I am wrong or vague, say only "not yet" and ask the
same thing a different way. If I say I do not know, tell me which
section of a paper would contain it - not the content.
At the end, list which of the five I could not produce.
Question 5 is the one that separates a first-class answer from a confident one. If you cannot produce it for a study, you do not yet know that study the way this course now grades it.
4. Assign it the other side. Agreement is only a problem while the model is free to choose which side to agree with.
Here is my interpretation of [study or phenomenon]:
[paste your paragraph]
Your job is to argue against it as strongly as the evidence
allows. Specifically:
- Name the alternative explanation my account rules out without
saying so.
- Point out where I treat a correlational result as causal.
- Point out where I generalise beyond the sample that was
actually studied.
- Say which single piece of evidence would most damage my
position.
Do not tell me whether I am right, and do not offer a balanced
summary at the end. Argue one side only.
You are removing the choice. A model told which position to take cannot flatter you by taking yours, and the objections it produces are the ones your marker is holding.
Where the line is
Everything above is studying: being quizzed, having your reasoning attacked, checking whether a claim survived. Writing the interpretation is the graded work and it stays yours. The risk specific to this course is subtler than copying — it is submitting a fluent essay built on a finding the field walked away from, which reads as not having done the reading. If your syllabus is vague about what is allowed, our guide to homework help without cheating and the class AI policy checklist cover how to read it and how to ask in writing.
Related reading
- The rest of psychology. Psychology case studies and cognitive theories, and memory curves and spaced repetition — where some of the sturdiest findings in the discipline live.
- The same problem, other departments. Epidemiology, where the model has learned its causal language from a literature that blurs the same line, and sociology research methods.
- The statistics underneath. Hypothesis testing and p-values explains why 97% of published studies being significant was the warning sign, and Bayes' theorem covers updating on new evidence.
- Checking anything a model tells you. The hallucination checklist, verifying AI answers before you study them, and evaluating sources for research papers.
- When you have to write it up. Literature reviews and the synthesis matrix — a column for replication status is worth adding to both.
FAQ
Will a chatbot tell me a social psychology finding was debunked?
If you ask it that question on its own, usually yes — one 2025 study found models identified misconception statements better than humans did. If the finding is embedded in a broader question you actually wanted answered, often no. The same study found that models "struggled to highlight or dispute" misconceptions inside applied, user-like questions, which the authors put down to a tendency toward sycophantic responses. Treat the replication check as its own separate prompt, asked first and worded neutrally.
Which famous social psychology findings failed to replicate?
The best documented include ego depletion, where 23 labs and 2,141 participants found a pooled effect of d = 0.04 with a confidence interval spanning zero, and power posing, which the first author of the original study publicly disavowed in 2016 with the words "I do not believe that 'power pose' effects are real." The Stanford Prison Experiment was not a failed replication but an archival reappraisal: Le Texier's 2019 paper in American Psychologist documented that the guards received precise instructions and that participants were "almost never completely immersed by the situation." Check any specific finding in the free Replication Database rather than relying on a list.
Can I use AI to simulate participants for a study design assignment?
For generating ideas and spotting obvious problems with your materials, yes. As evidence about what people would actually do, no. The largest test of this — 156 experiments replicated with three models, published in Nature Computational Science in 2025 — found effect sizes roughly two to three times larger than the human studies, and found that when the original study reported a null result the models produced a significant one 68 to 83 percent of the time. A tool that almost always finds an effect cannot tell you whether there is one.
What is the single best prompt habit for this course?
Add one instruction above your question: list the empirical claims my question assumes, and if any is unsupported, say so and correct it before answering. That specific phrasing is the one with evidence behind it. Asking the model to base its answer on scientific evidence, which is what most people try, was found to have "only little effects" on whether misconceptions got flagged.
Is it cheating to use AI for social psychology coursework?
That depends on your course policy, and this subject carries one risk worth naming separately from integrity. An essay that confidently explains an effect the field has moved past reads as poor scholarship whether or not a chatbot was involved, and no disclosure statement repairs that. Being quizzed, having your argument attacked, and checking replication status are ordinary studying. Read your syllabus, and ask your instructor in writing if the wording is unclear.
Bottom line
Social psychology asks you to hold two things at once: what a study claimed, and how well that claim has held up. A chatbot is genuinely useful for the first and will follow your lead on the second, which is the wrong way round for the marks. So split the question. Ask whether it replicated, coldly and on its own. Paste the correct-my-assumptions line above everything else. Then let the model argue against you, which is the one job it cannot do by agreeing.