Advertisement

Subject workflows · Updated 20 September 2026

How to Study Abnormal Psychology With AI (Safely)

A case vignette is the most paste-able thing in your whole degree. It is a short passage of plain text with a single right answer at the end of it, and a chatbot will give you that answer instantly and in a confident tone.

Advertisement

Use a chatbot for this course, but never ask it what the patient has. That single question is the one output the published evaluations say to distrust. When researchers fed clinical case vignettes to current models and asked for diagnoses directly, the models found most of the correct answers — and named a great many incorrect ones alongside them. The thing that fixed it was not a better model. It was forcing the model through the same structured rule-out logic your course already teaches you.

A round wooden-rimmed mesh sieve holding a heap of blue pebbles, tilted above a rectangular wooden tray in which only five small blue pebbles have collected, on a light wooden desk between a blue desk lamp, a terracotta pot with a green plant and a pale blue mug

Two numbers, and only one of them is reassuring

In a study published in Psychiatry Research, Karthik Sarma and colleagues took clinical case vignettes and their correct diagnoses from the DSM-5-TR Clinical Cases book, and diagnostic decision trees from the DSM-5-TR Handbook of Differential Diagnosis, which they "refined for LLM use." Then they ran the vignettes past three models two ways: asking straight out for a diagnosis, or walking the model through the decision trees.

Asked straight out, the best-performing model had a sensitivity of 76.7% and a positive predictive value of 40.4%. In plain words: it found about three-quarters of the diagnoses it should have found, and fewer than half of the diagnoses it named were correct. The authors' summary is the sentence to remember — direct prompting "yielded most true positive diagnoses but had significant overdiagnosis."

Run through the decision trees instead, and precision jumped to 65.3% while sensitivity only slipped to 70.9%. The structure did not make the model cleverer. It stopped it volunteering.

Why this is the worst possible error for a vignette exam. A model that misses a diagnosis leaves a visible hole. A model that adds a plausible extra one leaves nothing at all — the answer looks complete, reads well, and is wrong in a way you cannot see by reading it. On most marking schemes a confident extra diagnosis costs you exactly as much as the one you failed to spot.

It is not just the older models

The obvious objection is that this is a problem of the last model generation. The same group tested that directly, in a follow-up in JMIR AI published in June 2026. They took 106 vignettes from the same casebook, "with permission," and ran them past the newest ordinary model and the newest reasoning model from each of two vendors, with and without an extra self-verification prompt.

Sensitivity ranged from 0.732 to 0.817 and positive predictive value from 0.534 to 0.779. The best configuration — a reasoning model plus the self-verification prompt — reached a PPV of 0.779. So the gap narrowed a lot, and it did not close: in the best setup tested, roughly one named diagnosis in five was still wrong, and in the weakest, nearly half were. Reasoning models improved precision significantly, and so did the self-verification prompt, with neither making a significant difference to sensitivity.

That second finding is the one worth acting on, because a self-verification prompt is not a product you have to buy. It is a paragraph of instructions, and it improved precision even on a model that was already reasoning. Prompt 3 below is that idea written out for a student.

One caveat rather than a buried footnote: both studies evaluate a clinical task on published casebook vignettes, not a student doing homework. What transfers is the shape of the error — these models over-name — not the exact percentages.

Your exam is built out of the rule-outs

Here is why that shape matters so much in this particular subject. The handbook those researchers took their decision trees from is Michael First's DSM-5-TR Handbook of Differential Diagnosis — the same book sitting on your lecturer's shelf, containing 30 symptom-oriented decision trees and 67 differential diagnosis tables. Its six-step framework runs, in the publisher's own description, "from determining if the presenting symptoms are due to a substance/medication or a medical condition, to establishing the boundary between disorder and normality, to determining the primary disorder, and differentiating adjustment disorders from other mental disorders."

Read the order. The first two steps are subtraction, and naming the disorder does not appear until the third. When a vignette drops in a thyroid condition, a new blood-pressure medication, three drinks every evening or a bereavement four months ago, it is testing whether you noticed — and the correct answer is frequently that the criteria are not met, or that a different category applies. A model asked "what does this patient have" begins at step three, because that is the question you asked it.

The boundary step is harder still, and it is contested among specialists rather than merely difficult for students. Writing about prolonged grief disorder in the Australian and New Zealand Journal of Psychiatry, Maarten Eisma identifies exactly this as the weak point: "an important gap in the validity evidence: the distinction of prolonged grief disorder from normal grief." If the people who write the criteria are still arguing about where normal ends, a chatbot's brisk placement of that line tells you nothing.

Which manual is it quoting?

There is a second, quieter problem: a model has read DSM-IV-era, DSM-5-era and DSM-5-TR-era text, and nothing in its answer labels which one it is reproducing. That gap is wider than most students assume. In their World Psychiatry overview of the 2022 text revision, First and colleagues note that unlike DSM-IV-TR, where updates were "confined almost exclusively to the text," DSM-5-TR carries "a number of significant changes and improvements" — more than 70 disorders had criteria or specifier definitions updated. So ask which edition it is using. For a free second opinion, Washington State University publishes Fundamentals of Psychological Disorders under a Creative Commons licence, updated through DSM-5-TR.

Split the job

TaskWho does itWhy
Name the diagnosisYouIt is the graded skill, and the one output the evaluations found least reliable
Pull out every mention of substances, medication and physical illnessThe chatbot, told to quote the vignetteExtraction from a passage is a reading task, and it is step one of the standard framework
Lay a criteria set out as a checklistThe chatbot, then you check it against your own manualStructure is easy for it; the specific thresholds are what drift between editions
Decide whether the duration and impairment thresholds are metYou, against your course's criteria setThese are the elements exam vignettes are built to fail on
Write practice vignettes that nearly qualifyThe chatbotProducing a near-miss is easy for it, and you get a case you have not seen
Supply the answer key to a practice vignetteYour own criteria tableFewer than half the diagnoses a directly prompted model named were correct

The workflow, step by step

Four prompts for ChatGPT, Claude, Gemini or whatever you already have open. None of them asks for a diagnosis.

1. Make it do the subtraction first. This is First's step one and step two, with the naming step withheld.

Here is a case vignette from my abnormal psychology course.
Do NOT name a diagnosis. Not even a likely one.

Work only the ruling-out steps, in this order:

1. SUBSTANCE OR MEDICATION - quote every phrase in the vignette
   about alcohol, drugs, prescriptions, caffeine, withdrawal or
   a recent change of dose. If there are none, write NONE STATED.
2. MEDICAL CONDITION - quote every phrase about physical
   illness, injury, thyroid, sleep or test results. If there are
   none, write NONE STATED.
3. NORMALITY - quote every phrase about how much the symptoms
   interfere with work, study or relationships, and every phrase
   describing a life event the reaction might be expected after.
4. MISSING - list what someone would still need to ask to close
   each of the three steps above, that this vignette does not say.

Quote the vignette word for word. Do not paraphrase, and do not
infer anything that is not written down.

Vignette: [paste]

Section 4 is the one to read twice. Vignettes are written with gaps in them on purpose, and knowing which question is missing is usually worth more marks than a confident guess at the label.

2. Turn the criteria into a ledger. You pick the disorder. Its job is bookkeeping.

I am checking one disorder against one vignette. This is
bookkeeping, not judgement.

Disorder: [you name it - I am not asking you to choose]
Edition: DSM-5-TR. Say so plainly if you are unsure you have it.

Build a table with one row per criterion. Columns:
- the criterion, in plain words
- MET, NOT MET, or NOT STATED
- the exact words from the vignette that decide it, or blank

Give the duration requirement, the distress-or-impairment
requirement, and every exclusion clause their own rows.

Do not total the rows. Do not tell me whether the diagnosis is
met. Do not suggest a different disorder. Stop at the table.

Vignette: [paste]

NOT STATED is the useful column. A criterion the vignette never addresses is not a criterion that was failed, and telling those two apart is most of the skill.

3. Audit it for overdiagnosis. If a model has already given you a list of candidates, this is the self-verification step that improved precision in the 2026 study.

You gave me these candidate diagnoses: [paste its own list]

For each one, in order:
1. Name the single criterion that is hardest to support from
   this vignette.
2. Quote the words in the vignette that support it, or write
   NOTHING IN THE TEXT.
3. If you wrote NOTHING IN THE TEXT, withdraw that diagnosis
   and say "withdrawn".

Then list which withdrawn diagnoses one further question could
restore, and what that question would be.

Do not add any new diagnoses. Do not defend a diagnosis with
what is typical of the disorder - only with what this vignette
actually says.

Expect withdrawals. That is the prompt working, not failing, and the pattern of what gets withdrawn will teach you more about the boundary between two categories than a correct answer would have.

4. Drill the near-misses. The cases you get wrong in an exam are rarely the textbook ones.

Write three short vignettes, about 120 words each.

All three should look like [disorder]. Exactly one should meet
full DSM-5-TR criteria. The other two should each fail on one
different element - for example a duration under the threshold,
an exclusion that applies, or no significant distress or
impairment.

Rules:
- Do not label them.
- Do not say which one qualifies.
- Do not hint in the wording.
- Number them 1, 2, 3 and stop.

After I answer, name the element I missed and nothing else.

Then do the thing this whole page is about: check its answer key against your own criteria table. Writing a near-miss is a creative-writing task and models are good at it; deciding which of the three qualifies is the diagnostic task, and that is the half the numbers at the top of this page apply to.

Where the line is

Everything above is studying — being drilled, working fresh cases, having a list of candidates challenged. The written case formulation is the graded work and it stays yours. One caution is specific to this subject: if your course uses real clinical material, or anything from a placement, do not paste it into a chatbot at all, and our privacy checklist for class materials covers what that means in practice. If your syllabus is vague, homework help without cheating and the class AI policy checklist cover how to read it and how to ask in writing.

This is a study workflow, not a diagnostic tool. Abnormal psychology is a course where the reading gets personal, and a tool that over-names disorders is a bad place to take a worry about yourself or someone you know. Surveying 926 undergraduates in the 2024-2025 Healthy Minds Study, Cindy Liu and Tiffany Yip found that students who had actually had a psychiatric diagnosis trusted AI for mental health information less than those who had not. Your campus counselling service is free, confidential and staffed by people who do step two properly.

Related reading

FAQ

Can I just paste a vignette and ask for the diagnosis?

You can, and you will usually get a fluent answer containing the right diagnosis plus one or two that are not there. In the Psychiatry Research study, the best model asked this way had a positive predictive value of 40.4% — fewer than half of the diagnoses it named were correct, even though it found roughly three-quarters of the ones it should have. The researchers describe it as "significant overdiagnosis." On a vignette exam, an extra confident diagnosis usually costs you as much as a missing one, so this is the single question worth not asking.

Is this still true of the newest reasoning models?

It is better, and it is not solved. The same group repeated the experiment with the latest ordinary and reasoning models from two vendors on 106 casebook vignettes and published it in JMIR AI in June 2026. Positive predictive value ranged from 0.534 to 0.779 across the configurations tested, with the best being a reasoning model given an extra self-verification prompt. So at best roughly one named diagnosis in five was still wrong. Both the reasoning models and the self-verification prompt improved precision significantly, which is the practical takeaway: the structure you add to the question matters more than which tool you opened.

What should I check first in a vignette?

Not the symptoms. Michael First's differential diagnosis framework starts by asking whether the presentation is due to a substance or medication, then whether it is due to a medical condition, then where the boundary with normality sits — naming the disorder is only the third step. Exam vignettes are written to test exactly those first two steps, which is why they so often mention a medication change, a drink count or a recent bereavement in a sentence that looks like background colour. Work the subtraction before you work the label.

Does it matter whether it is quoting DSM-5 or DSM-5-TR?

Often yes. The 2022 text revision was not cosmetic: more than 70 disorders had criteria or specifier definitions updated, prolonged grief disorder was added, and "intellectual disability" was renamed "intellectual developmental disorder". A model has read all three generations of this literature and will not tell you which one an answer came from unless you ask. Ask, then check the criteria against whichever edition your course is marked against, because that is the one your grade depends on.

Should I use a chatbot to work out what is going on with me?

No — and this is the one place on this page where the research is not just about marks. A tool with a documented tendency to over-name disorders is a poor place to take a worry about your own mental health, especially in the term when you are reading about symptoms every week. Among 926 undergraduates surveyed in the Healthy Minds Study, 31% trusted generative AI for mental health information and only 13% trusted it for decisions, with students who had a lifetime psychiatric diagnosis trusting it less than those who did not. Talk to your campus counselling service, which is free and confidential at most institutions.

Bottom line

Abnormal psychology looks like the most chatbot-friendly subject on your timetable, because its exam questions arrive as short passages of text with one right answer. That is precisely why it goes wrong. The evaluations point the same way: these models are reasonably good at finding the diagnosis that is there and unreliable at not naming the ones that are not, and the fix is structure rather than a better model. So make it quote the vignette back at you, make it rule things out before it rules anything in, make it withdraw what it cannot support — and keep the naming step for yourself. That step is the one you are being examined on anyway.

Advertisement
Free download: Grab the one-page AI Study Safety Checklist — everything to check before you upload, trust, or submit anything involving AI.
Advertisement