Show the full module text
Module 14 — ABO — Autistic Burnout in Eight Questions
Why an eight-item burnout screener exists when a twenty-four item one already does, what the research on very short screens shows about the trade they make, and an honest account of what two minutes can and cannot tell you.
Autistic Self-Discovery · Part Five — Burnout · 13 min read, about 24 min with the workbook
The big idea
If you have ever opened a questionnaire, seen the words "24 questions, five to seven minutes" and closed the tab — this module is for that version of you. The screener attached to it is eight questions and takes about two minutes. It is also the least precise instrument in this whole course, and the point of the next few pages is to explain why both of those sentences are true at the same time.
A 29 and a 30 are the same reading. If your score sits within two or three points of a boundary, the band label is the least informative thing on the page and the answers you gave are the most informative. Read the items, not the door they sorted you through.
Step 1 — The lesson
A. The version you can actually take
Module 12 asked a question about your history: what in the shape of your life makes burnout likely over years. Module 13 asked a question about right now, in detail, across four areas at once. This module asks Module 13's question again, in eight items, in about two minutes.
That invites an obvious objection. If the long one exists, why take the short one. It deserves a real answer rather than a shrug, so here is the answer, and then here is what it costs.
Start with what the instrument is trying to detect. The first academic definition of autistic burnout came out of community-based participatory research in 2020: a syndrome resulting from chronic life stress and a mismatch of expectations and abilities without adequate supports, characterised by pervasive, long-term exhaustion — typically three months or more — loss of function, and reduced tolerance to stimulus.1 A year later, twenty-three autistic adults with lived experience of it agreed a definition of their own across three Delphi rounds: exhaustion, withdrawal, executive function problems and generally reduced functioning, with increased manifestation of autistic traits, distinct from depression and from non-autistic burnout.2
Read those two definitions again with the reader in mind rather than the researcher. Executive function problems. Generally reduced functioning. Withdrawal. The state being measured is precisely a state in which sitting down to a long form is hard.
Diagram — A · The version you can actually take. A twenty-four item screener asks better questions and gives you a shape rather than a number. An eight-item screener asks worse questions and gives you something you can do on the worst Tuesday of the month. Neither one replaces the other.
Survey methodology has known the general version of this for a long time, in a duller form. In one experiment the stated length of a web questionnaire was randomly set at ten, twenty or thirty minutes. The longer the stated length, the fewer people started it and the fewer finished. Answers placed later in the instrument came back faster, were skipped more often, and varied less.3 Length does not only lose people. It quietly degrades the answers of the people who stay.
Now apply that to a burnout screen. The people most likely to abandon a twenty-four item form halfway down are exactly the people it was built for. A short screen is not a lazy version of a good one. It is a different design decision with a different set of failure modes, and you are entitled to know both.
B. Two questions that work
The best evidence that a very short screen can carry real information does not come from autism research at all. It comes from depression, anxiety, and from burnout in doctors.
Take burnout first, because the precedent is almost absurdly short. Two single items lifted out of the Maslach Burnout Inventory were tested against the full instrument in 10,951 medical students, residents, faculty and practising surgeons. The one-item measure of emotional exhaustion correlated with the full emotional exhaustion domain at between 0.76 and 0.83. The one-item measure of depersonalization ran weaker, between 0.61 and 0.72.4 One question. Not as good as twenty-two, and not nothing.
The two-item depression screen is the better-studied case. The PHQ-2 asks about interest and about mood over the past two weeks, and that is the entire instrument. In the original study of 6,000 primary care and obstetrics patients, with 580 of them independently interviewed by a mental health professional, a cut-point of 3 gave a sensitivity of 82.9% and a specificity of 90.0% for major depression. Its positive predictive value was 38.4%.5
That last number is the one worth sitting with. Nine out of ten people without depression were correctly passed over, and still, fewer than four in ten of the people it flagged actually had the disorder. That is not a flaw in the PHQ-2. It is what happens when you screen for something uncommon: most positives come out of the large group who do not have it.
Diagram — B · Two questions that work. The same two questions, scored at a different threshold. Drop the cut-point from 3 to 2 and sensitivity climbs to 91 per cent, but a third of everyone without depression now screens positive. You cannot buy one without paying the other.
The scale of that trade-off was pinned down in 2020 by an individual participant data meta-analysis of 100 datasets covering 44,318 people, 4,572 of whom had major depression on a semistructured interview. At a cut-point of 2 or more, the PHQ-2 had a sensitivity of 0.91 and a specificity of 0.67. At 3 or more, sensitivity fell to 0.72 and specificity rose to 0.85.6 Two questions, one threshold moved by a single point, and a completely different instrument in practice.
The anxiety case makes the other half of the argument. A 2025 Cochrane review pooled 48 studies across 19,228 people in 27 countries. At its usual cut-point, the two-item GAD-2 detected generalized anxiety disorder with a sensitivity of 0.68 and a specificity of 0.86. The seven-item GAD-7 managed 0.64 and 0.91.7 Two items were not meaningfully worse than seven, and the review concludes plainly that neither can be used on its own to identify or rule out an anxiety disorder.
That is the honest case for a two-minute screen, and it is a real case. Very short instruments can sort a population about as well as longer ones. What they cannot do is settle anything about one person — and the longer version cannot do that either.
C. What eight items cost
Now the other side, because it is the side that lands on you rather than on a population.
The ABO has eight items, each scored from 1 to 5. Your total lands somewhere between 8 and 40, which is thirty-three possible scores, sorted into three bands. Move one answer by a single step and your total moves by one point. Move one answer from one end of its scale to the other and the total moves by four. On a twenty-four item instrument, a single item is worth a great deal less.
Diagram — C · What eight items cost. Three bands over a thirty-three point range. The wall between the green band and the gold one is a single point wide, and so is the wall between gold and red. Nothing in the instrument makes a 29 different in kind from a 30.
There is published work on exactly this. A simulation study in Psychological Methods asked whether short tests — the ones of at most fifteen items that clinicians actually use — are reliable enough to sort individuals into a treatment group and a non-treatment group. Short tests classified at most 50% of a group consistently. A six-item test built from weakly discriminating items classified 54% of the treatment group inconsistently, rising to between 67% and 83% when the cut-score sat out in the extreme 5% tail. Twenty and forty-item tests did considerably better.8
That is not an attack on any particular questionnaire. It is arithmetic. Fewer items means less information, less information means more measurement error, and more measurement error means that any boundary sitting near your score behaves partly like a coin.
The methodological literature adds a second warning. A much-cited critique of how abbreviated questionnaires get built argues that short forms are routinely written up as though they inherited the reliability and validity of the long instrument they came from, and are routinely derived and then evaluated in the very same sample rather than an independent one.9 Worth noting that the ABO does not even have a parent instrument to inherit from. It was written short.
And a third, which cuts against the best argument for taking a short screen at all. When a brief measure is given twice and the two scores are compared to judge whether someone has changed, shorter tests produce a higher risk of the wrong conclusion about that individual.10 Repeatability is the strongest thing an eight-item screen has, and it is also where the error bites hardest. That tension does not resolve. It tells you how to read the output: watch the direction of several readings over weeks, never the gap between two.
D. What your number does and does not mean
Here is the plain version, ahead of the clinician box rather than buried inside it.
The ABO is New Path's own instrument. We wrote it. It is Tier 4: honest, useful for thinking with, and unvalidated. There is no published validation study, no normative sample, no peer review, no reported internal consistency, no test-retest figure, and no sensitivity or specificity, because nobody has ever compared its output against any independent assessment of anything. Every number in sections B and C above belongs to other instruments. None of it transfers.
Diagram — D · What your number does and does not mean. This is the whole psychometric record. Two lines are honest descriptions of the tool. Five are things that a validated instrument would report and this one cannot, because the studies were never run.
The band boundaries deserve saying out loud. The line between Burnout Not Likely Present and Possible / Emerging Burnout sits between 19 and 20; the line above it sits between 29 and 30. Neither was derived from a sample of people assessed for autistic burnout by any other means, because no such sample was collected. They are reasonable divisions of a range. They are not thresholds anybody crossed in a study.
What is not uncertain is the thing being described. Autistic burnout has a defended definition,1 a lived-experience consensus behind it,2 and a growing measurement literature. What that literature also shows is how hard this is even when it is done properly.
A 2023 study surveyed 141 autistic adults with experience of autistic burnout and tested the 27-item AASPIRE Autistic Burnout Measure. It separated people currently in burnout from people who had been in burnout with poor specificity, an area under the curve of .661. In the same sample, 98% scored above the PHQ-9 cut-off for depression.11 A carefully built twenty-seven item research instrument could not cleanly tell now from then, and the depression overlap was near total.
A 2026 validation of the same measure in 379 autistic adults found excellent internal consistency, omega = 0.98, and then two figures worth more than that one: a twelve-month test-retest correlation of only 0.59, and correlations of 0.52 with a depression questionnaire and 0.54 with an anxiety questionnaire.12 The low retest figure is arguably good news, since burnout is a state and states move. The overlap with depression and anxiety is the problem, and eight items will not solve what twenty-seven could not.
One more thing specific to this screener. The live page offers it to people aged 16 and over, where every other questionnaire in this course says 18 and over. There is a defensible reason: an analysis of 1,127 posts from 683 people described autistic burnout as chronic or recurrent, often first experienced during adolescence, and lasting months or years.13 Sixteen is not an arbitrary place to start looking. But nothing about this instrument has been tested at any age, so if you are under eighteen, treat the result as something to show an adult who can act on it.
INSIDE THE INSTRUMENT
ABO — Autistic Burnout (eight-item self-screener)
Where it came from
Built in-house by New Path as an ultra-brief companion to the ABSI-24 in Module 13, not adapted from any published instrument. Eight items, each scored 1 to 5, total 8 to 40, completion time stated as two to three minutes. The live page describes its content domain as recent depletion, masking fatigue and sensory overwhelm, and offers it to respondents aged 16 and over. There is no development paper, no item pool, no factor analysis and no item-level statistics of any kind.
What it is made of
Eight items yielding a single raw total and three bands: 8-19 Burnout Not Likely Present, 20-29 Possible / Emerging Burnout, 30-40 Burnout Likely Present. No subscales. The construct it targets is well specified in the literature even though the instrument is not: chronic life stress plus a mismatch between expectations and abilities without adequate supports, producing long-term exhaustion, loss of function and reduced tolerance to stimulus,1 with the lived-experience consensus adding withdrawal, executive function problems and increased manifestation of autistic traits.2
How well it performs
Unknown, in the strict sense that no performance data exist. No internal consistency coefficient, no test-retest interval, no ROC analysis, no criterion or convergent validity, no factor structure. For calibration on what an eight-item instrument can and cannot be expected to do, the relevant external evidence is the short-scale literature: short tests of at most fifteen items classified at most 50% of a simulated group consistently, with a six-item low-discrimination test misclassifying 54% and up to 83% at extreme cut-scores,8 and shorter tests raise the risk of wrong conclusions about individual change between two administrations.10
Where it was validated
Nowhere. No validation sample, no normative data, no clinical comparison group, no translation studies, no independent replication. Raw scores are not standardised and are not comparable with any published cut-off. Where a module in this course quotes sensitivity, specificity or reliability, those figures belong to the PHQ-2, the GAD-2, the Maslach items or the AASPIRE measure and are offered as calibration for how brief screens behave in general, never as properties of this one.
What it cannot do
It cannot distinguish autistic burnout from depression, and neither can the validated instruments. In 141 autistic adults with burnout experience, 98% scored above the PHQ-9 cut-off, and the 27-item AASPIRE measure separated current from past burnout with an AUC of only .661.11 In a 379-adult validation of the same measure, correlations with the PHQ-9 and GAD-7 were 0.52 and 0.54.12 An elevated ABO score is a reason to assess mood, sleep, thyroid function and anaemia, not a reason to skip them.
It cannot support a change score. With eight items and a thirty-three point range, the standard error around any single administration is unquantified and almost certainly wide relative to the four-point wall between bands. Simulation work on individual change with short instruments is explicit that shorter tests produce a higher rate of incorrect conclusions about whether a given person has moved.10 A trend across several weekly readings is defensible as a conversation; a two-point difference between two readings is not.
Its band boundaries are conventions, not empirically derived cut-points. The 19/20 and 29/30 lines were not set against any criterion assessment. Compare the PHQ-2, where moving the cut-point from 2 to 3 shifts sensitivity from 0.91 to 0.72 and specificity from 0.67 to 0.85 across 100 datasets and 44,318 participants6 — that is what a threshold looks like when it has been studied, and none of that work has been done here.
It cannot substitute for the longer instruments in Modules 12 and 13. The ABTI-24 addresses standing vulnerability; the ABSI-24 gives a four-part current-state profile. This screener collapses the second into one number. Where short and long administrations disagree, the long one carries more information, and the published advice on brief anxiety screens follows the same logic: a positive brief screen is a trigger for fuller assessment, not a result.7
What a clinician does with it
Use it as a repeated-measure conversation opener rather than an assessment. Its one real advantage over the ABSI-24 is that it can be given weekly without attrition, and repeated brief sampling is better placed than a single retrospective summary to show how a state moves.14 That advantage is only realized if something is done with the reading: in routine outcome monitoring, feedback alone produced d = 0.14, rising to d = 0.49 when it came with guidance on what to change.15 Record the items endorsed rather than the total, and formulate against demand load, masking and recovery access rather than against the band label.
Validity tier: 4 — built in-house by New Path. Honest and useful, not normed, no published validation. The construct has a defensible definition and a real literature; this particular eight-item operationalisation of it has never been tested, and its scores should be treated as structured self-description rather than measurement.
Diagram — E · One reading, many roads. The single most important caveat here. Exhaustion, reduced function and low tolerance for stimulus are the final common path of a great many things, several of which show up on a blood test. A high score means look, and it does not say where.
STRENGTHS LENS
Taking a two-minute check on a bad day is not settling for less.
There is a particular kind of self-criticism that arrives with burnout: that you should be able to do the thorough version, that a shortcut is a cop-out, that if you cannot manage twenty-four questions about your own life then you are not taking it seriously. That reasoning is backwards. Recognizing that your capacity is lower than usual and choosing the tool that fits it is the skill, not the failure.
It is also worth noticing what you have already done by arriving at a page like this. Autistic burnout went undescribed in the academic literature until 2020, and the definition that exists now was built by autistic adults describing their own experience to researchers who listened. Everyone who put words to this before there was a name for it was doing unpaid research. You are reading the results.
F. What helps
An eight-item screen is worth very little as a one-off and quite a lot as a habit. Everything below assumes you will take it more than once.
1. Take it weekly, same day, and read the line rather than the point.
Sunday evening, or whichever slot you will actually keep. Four readings tell you something a single reading cannot, because a summary of the past few weeks written from memory is subject to recall bias in a way that repeated brief sampling is not.14 Write the number somewhere you will still have it in a month. The direction of travel is the finding; the individual number is noise around it.
2. Build your own baseline instead of borrowing ours.
Our three bands were set by judgment. Your own range, once you have five or six readings, is real data about you. If you normally sit at 22 and you are at 29, that is a meaningful move even though both numbers fall in the same band. If you have always sat at 31, the red label is telling you less than the week-to-week variation around it.
3. When the short one and the long one disagree, believe the long one.
If the ABO says one thing and the ABSI-24 in Module 13 says another, the twenty-four item version has more information in it and should win. The same applies to the ABTI-24 in Module 12, which is answering a different question entirely — standing vulnerability rather than current state. A low ABO today does not cancel a high ABTI-24.
4. Rule out the boring physical explanations before you settle on this one.
Deep fatigue, low tolerance for noise and light, and a brain that will not do the things it used to do are also the presentation of anaemia, thyroid dysfunction, sleep apnoea, post-viral illness and untreated depression. A blood test is cheap and this screener is no substitute for one. Excluding a physical cause also makes the burnout conversation easier, because it removes the first thing anyone will ask.
5. Change one demand, then re-screen in two weeks.
Pick the single largest recurring demand you have any control over — a recurring meeting, a social commitment, one masked hour a day — and remove or reduce it. Then take the screener again a fortnight later. Repeated measurement only earns its keep when something is done in response to it; monitoring that changes nothing produces the smallest effects on record.15
Step 2 — Take the screener
Eight questions, two to three minutes, free and confidential. You get one number between 8 and 40 and one of three bands. Take it more than once if you can — a single reading from an eight-item screener is a rough estimate, and four readings over a month tell you something a single reading never will.
Before you start
The ABO was written in-house by New Path. It has no published validation, no normative sample and no peer-reviewed psychometrics, so its score is raw self-description rather than a measurement, and its three band boundaries were set by judgment rather than derived from any study. With only eight items, a single answer moves the total by up to four points, which is enough to cross a band line. Exhaustion, low tolerance for stimulus and reduced functioning are also produced by depression, anaemia, thyroid problems, sleep disorders and post-viral illness, so a high score is a reason to look further rather than an answer. This is a screen, not a test, and only a clinician can diagnose.
Step 3 — Your workbook
Your answers save to this device only — we cannot see a word of what you write. This module's workbook is built for repeated use: somewhere to keep four weekly scores, work out your own baseline instead of borrowing ours, and set the one demand you are going to change before the next reading.
Your ABO result
Took the screener? Put the number in below. Eight items is a small net, so treat the band as a rough door rather than a verdict — and come back and add next week's number underneath. Entirely optional — skip it if you would rather just read.
Score bands: 8–19 = Burnout Not Likely Present; 20–29 = Possible / Emerging Burnout; 30–40 = Burnout Likely Present
Fields: ABO · Autistic Burnout, eight-item self-screener; Total score (8–40); My total (enter 8–40)
Your weekly line
Section F, item 1. One score is a dot. Four scores are a line, and the line is the part worth having.
Fields: Week 1 score and date; Week 2 score and date; Week 3 score and date; Week 4 score and date; Which way is the line going?: Down, clearly / Down, slightly / Flat / Up, slightly / Up, clearly / Too soon to say; What was happening in the week with the highest number
Your own baseline, not our bands
Section F, item 2. Our three bands were set by judgment. Your usual range is real information about you.
Fields: My usual score, roughly; The lowest I have recorded; The highest I have recorded; What a five-point rise usually means for me in practice
Short against long
Section F, item 3. If you have taken the ABSI-24 in Module 13 or the ABTI-24 in Module 12, put them side by side here.
Fields: My ABSI-24 total, if I have one; My ABTI-24 total, if I have one; Do the short and long results agree?: Yes, closely / Roughly / No, the short one reads higher / No, the short one reads lower / I have only taken one; If they disagree, which one matches how the last month actually felt
The boring explanations first
Section F, item 4. Exhaustion, sensitivity and a brain that will not cooperate have several dull physical causes that are worth excluding.
Fields: Tick: I have had bloods done in the last twelve months; Tick: My sleep has been properly looked at; Tick: Mood has been discussed with someone qualified; What I still want ruled out, and who I would ask
One demand, removed
Section F, item 5. Monitoring that changes nothing has the smallest effect on record. Pick one thing and re-screen in a fortnight.
Fields: The demand I am going to reduce or remove; The date I will re-take the screener; What I expect to change, and what actually did
Appendix — Research companion
Peer-reviewed research
7. Akturk Z, Hapfelmeier A, Fomenko A, Dummler D, Eck S, Olm M, Gehrmann J, von Schrottenberg V, Rehder R, Dawson S, Lowe B, Rucker G, Schneider A, Linde K (2025). Generalized Anxiety Disorder 7-item (GAD-7) and 2-item (GAD-2) scales for detecting anxiety disorders in adults. Cochrane Database of Systematic Reviews, 2025(3), CD015455. DOI 10.1002/14651858.CD015455. View the paper Cochrane diagnostic test accuracy review of 48 studies covering 19,228 participants across 27 countries and 24 languages, in non-clinical, mixed clinical and condition-specific settings. At a cut-point of 3 or more the two-item GAD-2 detected generalised anxiety disorder with a pooled sensitivity of 0.68 and a specificity of 0.86, against 0.64 and 0.91 for the seven-item GAD-7 at a cut-point of 10, so the shorter scale was not meaningfully worse at group level. The authors conclude that neither scale can be used on its own to identify or to rule out an anxiety disorder. Limitation: pooling across highly heterogeneous settings and reference standards, and the review addresses single-occasion detection rather than repeated administration over time.
11. Arnold SRC, Higgins JM, Weise J, Desai A, Pellicano E, Trollor JN (2023). Towards the measurement of autistic burnout. Autism, 27(7), 1933-1948. DOI 10.1177/13623613221147401. View the paper Online survey of 141 autistic adults with experience of autistic burnout, 103 of whom reported burnout within the past three months, testing the 27-item pre-publication AASPIRE Autistic Burnout Measure alongside a new 20-item set derived by exploratory factor analysis from 48 candidate items. The new item set reached a Cronbach alpha of .88 overall with subscale alphas of .73 to .86, but the AASPIRE measure distinguished current from past burnout with poor specificity (AUC = .661, n = 136), and 98% of the sample scored above the PHQ-9 cut-off for depression. Limitation: a small, predominantly female, online-recruited and late-diagnosed sample with no comparison group of autistic adults who had never experienced burnout, which the authors describe as too small for ideal assessment tool development.
12. Bougoure M, Zhuang S, Brett JD, Maybery MT, English MC, Tan DW, Magiati I (2026). Measuring autistic burnout: a psychometric validation of the AASPIRE Autistic Burnout Measure in autistic adults. Autism, 30(1), 20-36. DOI 10.1177/13623613251355255. View the paper Psychometric validation of the 27-item AASPIRE Autistic Burnout Measure in 379 autistic adults aged 18 to 77. Internal consistency was very high at omega = 0.98 with a predominantly unidimensional structure, but the 12-month test-retest intraclass correlation was only 0.59, and total scores correlated 0.52 with the PHQ-9 and 0.54 with the GAD-7. Limitation: a predominantly white, cisgender, tertiary-educated and late-diagnosed online sample, with criterion validity resting on a single self-report item and a 12-month retest interval the authors acknowledge may be too long to capture how quickly burnout moves.
8. Emons WHM, Sijtsma K, Meijer RR (2007). On the consistency of individual classification using short scales. Psychological Methods, 12(1), 105-120. DOI 10.1037/1082-989X.12.1.105. View the paper Simulation study asking whether short tests of at most 15 items, of the kind routinely used to make decisions about individual patients in clinical and health psychology, are reliable enough to sort people into a treatment and a non-treatment group. Short tests classified at most 50% of a group consistently at a certainty level of .9; a six-item test built from low-discriminating items classified 54% of the treatment group inconsistently, rising to between 67% and 83% when the cut-score sat in the extreme 5% tail, while 20-item and 40-item tests performed considerably better. Limitation: a simulation conducted under item response theory assumptions rather than an empirical study of any real instrument, so the exact percentages depend on the item parameters chosen and describe the shape of the problem rather than the properties of any named questionnaire.
3. Galesic M, Bosnjak M (2009). Effects of questionnaire length on participation and indicators of response quality in a web survey. Public Opinion Quarterly, 73(2), 349-360. DOI 10.1093/poq/nfp031. View the paper Web survey experiment in which the stated length of the questionnaire was randomly manipulated at 10, 20 and 30 minutes, with thematic question blocks presented in random order. The longer the stated length, the fewer respondents both started and completed the questionnaire, and answers to questions placed later in the instrument showed shorter response times, higher item non-response, shorter open-ended answers and less variability across grid items. Limitation: a general-population web panel study of survey behaviour rather than a clinical sample, and it says nothing about screening accuracy, so it explains why long forms lose people without bearing on what any particular questionnaire measures.
5. Kroenke K, Spitzer RL, Williams JB (2003). The Patient Health Questionnaire-2: validity of a two-item depression screener. Medical Care, 41(11), 1284-1292. DOI 10.1097/01.MLR.0000093487.78664.3C. View the paper Developed the two-item PHQ-2 in 6,000 patients across eight primary care and seven obstetrics-gynaecology clinics, with criterion validity assessed in 580 patients who also received an independent structured interview by a mental health professional. A cut-point of 3 or more gave a sensitivity of 82.9% and a specificity of 90.0% for major depressive disorder, with a positive predictive value of only 38.4% at the 7% prevalence observed in that setting. Limitation: the criterion subsample contained just 41 cases of major depressive disorder, and the operating characteristics are tied to primary care prevalence, so both the optimal cut-point and the predictive value shift in populations where depression is more or less common.
10. Kruyen PM, Emons WHM, Sijtsma K (2014). Assessing individual change using short tests and questionnaires. Applied Psychological Measurement, 38(3), 201-216. DOI 10.1177/0146621613510061. View the paper Simulation study of the situation clinicians face when a brief measure is administered twice and the difference between the two scores is used to judge whether an individual client has benefited. It found that shorter tests produce a higher risk of drawing incorrect conclusions about change in individual clients, and used the results to derive guidelines for the minimum number of items needed to assess individual change reliably. Limitation: a simulation rather than a clinical trial, and its guidelines assume item response theory conditions that an instrument with no published item analysis cannot demonstrate it meets.
6. Levis B, Sun Y, He C, Wu Y, Krishnan A, Bhandari PM, Neupane D, Imran M, Brehaut E, Negeri Z, Benedetti A, Thombs BD (2020). Accuracy of the PHQ-2 alone and in combination with the PHQ-9 for screening to detect major depression: systematic review and meta-analysis. JAMA, 323(22), 2290-2300. DOI 10.1001/jama.2020.6504. View the paper Individual participant data meta-analysis of 100 datasets covering 44,318 participants, of whom 4,572 (10%) met criteria for major depression on a semistructured interview. At a cut-point of 2 or more the PHQ-2 had a sensitivity of 0.91 and a specificity of 0.67; at 3 or more, sensitivity fell to 0.72 and specificity rose to 0.85, with an area under the curve of 0.88. Screening with the PHQ-2 first and administering the PHQ-9 only to those who screened positive reduced full PHQ-9 administration by roughly 57% while giving a sensitivity of 0.82 and a specificity of 0.87. Limitation: substantial heterogeneity across diagnostic interview types, and the authors note limited evidence on how the two-stage approach performs in routine practice rather than in research datasets.
1. Raymaker DM, Teo AR, Steckler NA, Lentz B, Scharer M, Delos Santos A, Kapp SK, Hunter M, Joyce A, Nicolaidis C (2020). Having all of your internal resources exhausted beyond measure and being left with no clean-up crew: defining autistic burnout. Autism in Adulthood, 2(2), 132-143. DOI 10.1089/aut.2019.0079. View the paper Community-based participatory research conducted with the AASPIRE partnership, analysing 19 interviews with autistic adults alongside 19 public internet sources and around 200 posts tagged #AutisticBurnout. It produced the first academic definition of autistic burnout: a syndrome resulting from chronic life stress and a mismatch of expectations and abilities without adequate supports, characterised by pervasive long-term exhaustion of typically three months or more, loss of function, and reduced tolerance to stimulus. Participants described it as distinct from both depression and occupational burnout and linked it directly to masking and to the absence of accommodations. Limitation: a small exploratory qualitative study using a convenience sample that lacked racial diversity and did not include non-speaking people, people with higher support needs or people with low educational attainment, so it defines the construct without estimating how common it is.
14. Shiffman S, Stone AA, Hufford MR (2008). Ecological momentary assessment. Annual Review of Clinical Psychology, 4, 1-32. DOI 10.1146/annurev.clinpsy.3.022806.091415. View the paper Review of ecological momentary assessment, the practice of sampling people's current experience repeatedly and in real time in their own environment rather than asking them to summarise weeks of life at a research or clinic visit. It sets out the case that global retrospective self-report is limited by recall bias and is poorly suited to showing how behaviour changes over time and across contexts, while repeated brief sampling minimises recall bias and improves ecological validity. Limitation: a methodological review rather than a trial, written for research designs using many prompts per day, so it supports frequent brief measurement in principle without establishing that any particular weekly screener is accurate.
9. Smith GT, McCarthy DM, Anderson KG (2000). On the sins of short-form development. Psychological Assessment, 12(1), 102-111. DOI 10.1037/1040-3590.12.1.102. View the paper Methodological critique of how abbreviated questionnaires are constructed and reported in psychological assessment, and one of the most cited papers in the short-form literature. It argues that short forms are commonly treated as inheriting the reliability and validity evidence of the parent instrument when they have not earned it, and criticises the widespread practice of deriving a short form by factor-analysing the long form's own data set and then evaluating it in that same sample rather than in an independent one. Limitation: a conceptual and methodological paper with no new empirical data, addressing forms abridged from longer parent measures rather than instruments written short from the outset.
4. West CP, Dyrbye LN, Sloan JA, Shanafelt TD (2009). Single item measures of emotional exhaustion and depersonalization are useful for assessing burnout in medical professionals. Journal of General Internal Medicine, 24(12), 1318-1321. DOI 10.1007/s11606-009-1129-z. View the paper Tested two single items drawn from the Maslach Burnout Inventory against the full instrument in 10,951 respondents across four samples: 2,248 medical students, 333 internal medicine residents, 465 faculty and 7,905 practising surgeons. The single emotional exhaustion item correlated with the full emotional exhaustion domain at Spearman r = 0.76 to 0.83, and the single depersonalization item with its domain at r = 0.61 to 0.72, with consistent stratification of high-burnout risk across all four groups. Limitation: an occupational burnout study in medical professionals validated against another self-report questionnaire rather than any clinical assessment, and the weaker depersonalization figures show that one item captures some facets of burnout much better than others.
Clinical frameworks and position statements
15. Lambert MJ, Whipple JL, Kleinstauber M (2018). Collecting and delivering progress feedback: a meta-analysis of routine outcome monitoring. Psychotherapy, 55(4), 520-537. DOI 10.1037/pst0000167. View the source Meta-analysis of controlled trials in which brief measures were administered repeatedly during psychotherapy and the results fed back to clinicians. Across 15 studies and 8,649 clients using the Outcome Questionnaire system the overall effect of feedback was d = 0.14, rising to d = 0.33 for clients whose progress was off track and d = 0.49 when feedback was paired with guidance on what could be done about it, while 9 studies covering 2,272 clients using the Partners for Change system gave d = 0.40. Limitation: these are effects of therapist-facing feedback inside active treatment rather than of a person tracking their own score at home, and the smallest effects occur where the measure is collected without any structured response to it.
Lived experience
2. Higgins JM, Arnold SRC, Weise J, Pellicano E, Trollor JN (2021). Defining autistic burnout through experts by lived experience: grounded Delphi method investigating #AutisticBurnout. Autism, 25(8), 2356-2369. DOI 10.1177/13623613211019858. View the source Grounded Delphi study in which 23 autistic adults, all experts by lived experience of autistic burnout, co-produced a definition across three rounds of survey, with a high majority of agreement reached by round 3. The agreed definition describes a highly debilitating condition characterised by exhaustion, withdrawal, executive function problems and generally reduced functioning, with increased manifestation of autistic traits, and explicitly distinct from depression and from non-autistic burnout. Limitation: 23 self-selected participants recruited largely online, and the authors state that further work is still needed to differentiate autistic burnout from other conditions, so this is a consensus statement rather than a validated diagnostic boundary.
13. Mantzalas J, Richdale AL, Adikari A, Lowe J, Dissanayake C (2022). What is autistic burnout? A thematic analysis of posts on two online platforms. Autism in Adulthood, 4(1), 52-65. DOI 10.1089/aut.2021.0021. View the source Thematic analysis of 1,127 posts from 683 users across two online platforms between 2005 and 2019. Contributors described autistic burnout as a chronic or recurrent condition that was often first experienced during adolescence and lasted months or years, with masking named most often as the cause and a pervasive lack of awareness and stigma about autism as the overarching theme. Limitation: publicly posted text from self-selecting online contributors with almost no demographic information and no way to confirm diagnosis, so it describes how the experience is talked about rather than how often it occurs or at what age it typically begins.
Further reading — general background
World Health Organization (2019). Burn-out an occupational phenomenon: International Classification of Diseases. World Health Organization, news release, 28 May 2019. View the source WHO statement clarifying the status of burn-out in ICD-11. It defines burn-out as a syndrome conceptualised as resulting from chronic workplace stress that has not been successfully managed, with three dimensions - feelings of energy depletion or exhaustion, increased mental distance from one's job or cynicism about it, and reduced professional efficacy - and states that burn-out is classified as an occupational phenomenon and not as a medical condition. It adds that the term refers specifically to phenomena in the occupational context and should not be applied to describe experiences in other areas of life. Limitation: this is the occupational construct rather than autistic burnout, and the WHO attaches no diagnostic criteria, no threshold and no instrument to it, so it cannot be used to validate any burnout questionnaire.
Peer-reviewed = checked by independent experts before publication. Clinical model = an established professional framework, not a single study.
Up next
Module 15 - Sensory Profile — Sensory Processing Across Eight Senses
All modules in Autistic Self-Discovery
Two minutes is enough to start a conversation. This course was built by clinicians who are part of the New Path Family. If your weekly numbers have been climbing for a month, that is worth saying out loud to someone rather than tracking alone — and working out which demands can actually be reduced is the part that tends to need another person. Therapy for clients in California and coaching worldwide, all by telehealth, are offered by our sister company New Path Family of Therapy Centers, Inc. A conversation costs nothing and there is no pressure. Saving this for later counts too. Talk with the New Path team
