Show the full module text
Module 3 — AQ-50 — The Best-Known Autistic Traits Questionnaire
Where the fifty questions came from, what the two cut-off scores were built to do, why the five named subscales do not hold up under analysis, and what the AQ's real accuracy figures look like once it leaves the research sample.
Autistic Self-Discovery · Part Two — The Broad Screeners · 14 min read, about 28 min with the workbook
The big idea
If you have ever taken the AQ at eleven at night, landed on 29, and then spent an hour deciding whether you had answered honestly or answered the way you hoped — this is the module for that. The AQ-50 is the most widely taken autistic traits questionnaire in the world, and the one whose numbers are most often read as more than they are.
This is the base-rate problem, and it is the most useful idea in this course. Predictive value is not a property of a questionnaire. It belongs to the questionnaire and the room it is used in. Move the same fifty items from a clinic where three in four people are autistic to an audience where the great majority are not, and an identical score means something different — the pool that can produce a false positive has grown enormously while the pool of true cases has not.
Step 1 — The lesson
A. Fifty questions from Cambridge, 2001
The AQ came out of the Autism Research Centre in Cambridge in 2001. Fifty statements, each answered on a four-point scale from definitely agree to definitely disagree. Every answer in the autistic direction scores one point and every other answer scores nothing, so the total runs from 0 to 50.1
The paper that introduced it tested four groups: fifty-eight adults who already had a diagnosis of Asperger syndrome or high-functioning autism, a hundred and seventy-four randomly selected controls, eight hundred and forty Cambridge students, and sixteen winners of the UK Mathematics Olympiad. The autistic group averaged 35.8, the controls 16.4. Eighty per cent of the autistic adults scored 32 or above, against 2 per cent of controls.1
Hold on to one of those numbers. Fifty-eight. Everything the AQ knows about how an autistic adult answers fifty questions, it learned from fifty-eight British adults around the turn of the millennium, most of them men, all already diagnosed at a point when adult diagnosis was rare and reached mostly one kind of person.
Diagram — A · Fifty questions from Cambridge, 2001. The two findings that matter most for how you read your own score — the collapse in specificity and the failure of the five subscales — arrived long after the questionnaire had become standard. Age and popularity are not accuracy.
Since then the AQ has been given to an enormous number of people, so we know the ordinary range well. A systematic review pooled 73 articles covering 6,934 non-clinical adults and found a mean of 16.94. Men averaged 17.89, women 14.88. Autistic samples in the same review averaged 35.19.2
Those two figures — roughly 17 for men, roughly 15 for women — are the ones quoted on the screener page, and they are the right ones. They also change how the first band reads. That band runs from 0 to 25, and a typical non-autistic man sits at 17. Someone who scores 24 is seven points above that average and still lands in the green. The band is not telling that person they are typical. It is telling them the questionnaire has not seen enough to say anything.
One honest note before you take it. The screener page does not state how individual items are scored, and we will not invent that for you. If the version in front of you does not use the four-point format above, the bands will not mean what the research means by them.
B. Five named parts that will not stay separate
The fifty items were written in five blocks of ten. The published domains are social skill, attention switching, attention to detail, communication and imagination.1 That structure is why the AQ feels well organized while you take it. It is also the part that has held up worst.
The first serious challenge came from a Dutch team in 2008, working with 302 general-population adults and 961 students. The five domains would not separate. What fitted best was two: a broad Social Interaction factor absorbing four of the original five, and a stand-alone Attention to Detail factor that barely related to it, correlating at r = .19.3
Diagram — B · Five named parts that will not stay separate. The five domains are real as blocks of items on a page. They are not five separable dimensions in the data — which is the difference between a profile and a sum.
It has not settled since. In 2020 an Australian group ran confirmatory factor analyses on eleven competing published models of the AQ across 1,702 undergraduates and 1,280 general-population adults. The best-supported model had three factors, not five. More striking is what they concluded about the total: the evidence did not support using total-scale scores at all, and they recommended researchers use subscale scores instead.4
A Rasch analysis in 2017, using 130 autistic adults and 219 students, put it more bluntly. The AQ did not meet the criterion of a unidimensional measure — only 26.2 per cent of the variance was explained by the thing it is supposed to be measuring. Five items misfitted, four from attention to detail, and three of those correlated in the opposite direction to the rest of the scale.5
Here is why that matters to you and not only to psychometricians. Your total is a sum of parts that do not move together. Two people can arrive at 29 by entirely different routes — one through noticing every detail in a room, the other through finding conversation exhausting — and the number will not tell them apart.
C. Two lines, drawn for two different jobs
You will see two thresholds quoted for the AQ, 26 and 32, usually presented as two rungs of one ladder. They are not. They came out of different studies four years apart, chosen to make opposite mistakes.
Thirty-two came from the original paper: the level at which 80 per cent of those fifty-eight autistic adults scored, against 2 per cent of controls.1
Twenty-six came from a 2005 study of a hundred consecutive adults referred to a diagnostic clinic. At 26 or above the AQ was 94.5 per cent sensitive and 51.9 per cent specific, classifying 83 per cent of that sample correctly. At 32 or above it was 76.7 per cent sensitive and 74.1 per cent specific. The area under the curve was 0.78. The authors were explicit about what the two lines were for: use 26 in a clinic, where the expensive mistake is missing someone, and use 32 for screening a general population, where the expensive mistake is flooding the service with false positives.6
Diagram — C · Two lines, drawn for two different jobs. These are the bands you will be given, reproduced exactly. What they do not say is that the lower line was drawn for a room where everybody was already being assessed, and the upper one around fifty-eight people in 2001.
So the two numbers are not possible and likely the way a thermometer is warm and hot. They are two settings on one dial, each picked to guard against a different error in a different room.
There is a short form too. In 2012 the same Cambridge group distilled the AQ down to ten items using 1,000 cases and 3,000 controls across age bands. For the adult AQ-10 the figures came from 449 autistic adults and 838 controls, and at a cut-off of 6 they reported sensitivity 0.88, specificity 0.91 and a positive predictive value of 0.85.7 Those participants were volunteers registered on the research center's website who completed the questionnaire online. Keep that detail.
On the strength of those figures, UK national guidance adopted the short form. Recommendation 1.2.3 of the NICE guideline on autism in adults says to consider the AQ-10 for adults with possible autism who do not have a moderate or severe learning disability, and to offer a comprehensive assessment at 6 or above.8 That recommendation still stands. The next section is what happened when somebody tested it.
D. What the number is actually worth
In 2016 a team at a national specialist clinic in London gave the AQ to 476 adults referred there for assessment, then compared the scores against the outcome each adult received. Three hundred and forty-six of them, 73 per cent, turned out to be autistic.
At a cut-off of 26 the AQ-50 was 0.88 sensitive and 0.20 specific, with a positive predictive value of 0.76. At 32, sensitivity 0.71, specificity 0.35, the same positive predictive value of 0.76. The AQ-10 at a cut-off of 6 came out 0.77 sensitive and 0.29 specific, and it did not predict the outcome better than chance. Of the patients who scored 5 or below on the AQ-10, 70 per cent were autistic anyway. And in the multivariate model, the one variable that predicted an AQ-10 score was generalized anxiety disorder. Being autistic did not.9
Diagram — D · What the number is actually worth. Specificity is the share of people without the condition who correctly score below the line. At 0.20, four in five non-autistic adults in that clinic scored above it. A line almost everyone crosses is not sorting anybody.
The positive predictive value needs unpacking too, because 0.76 sounds respectable. Seventy-three per cent of that room was autistic before anyone filled in a form, so an instrument that simply said yes to everybody would have scored 0.73. The questionnaire added almost nothing to what the referral letter already said.
An online screener sits somewhere between those two rooms, and nobody can tell you where. People who search "am I autistic" late at night are not an unselected population; they are self-selected, often heavily, which lifts the base rate well above the roughly 1 per cent seen in the general population.9 They are not clinic referrals either. The honest answer is that the positive predictive value of your score on a page like this one is unknown, and any site that hands you a confident percentage is making it up.
Two further studies show the number moving with the room. A Dutch study of 210 adults referred for assessment, alongside 63 controls, found sensitivity and specificity well below published figures and concluded that none of the instruments tested had sufficient validity to predict the outcome reliably in outpatient settings.10 A later Dutch study of 92 adults referred to a general mental health service found the AQ at 26 to be 58 per cent sensitive and 86 per cent specific — close to the mirror image of the London result. Predictive values belong to a population and a setting, not to an instrument.11
Then there is who the items were written about. A 2023 study tested seventeen published versions of the AQ, from 6 to 50 items and 1 to 5 factors, in 7,076 UK adults: 5,246 women and 1,830 men. Among the models that fitted acceptably, only two items were free of gender bias. Twenty were biased toward men and twenty-one toward women, with women more likely to endorse the social skill and communication items and men the restricted and repetitive ones.12 In the original 2001 control sample, no woman scored 34 or above, while 4 per cent of men did.1
None of that makes the AQ useless for women. It means a woman's total is assembled from a different mix of items than a man's, so comparing the two directly assumes something the evidence does not support.
The short form carries its own problems. Given to 6,595 non-clinical adults, every internal-consistency coefficient for the AQ-10 fell below 0.7, a single-factor model was rejected outright, and parallel analysis pointed to four factors inside ten items.13 Ten questions drawn from five domains do not add up to one thing.
INSIDE THE INSTRUMENT
AQ-50 — the Autism-Spectrum Quotient
Where it came from
Published by Baron-Cohen, Wheelwright, Skinner, Martin and Clubley in 2001 at the Autism Research Centre, Cambridge. Development groups: 58 adults with Asperger syndrome or high-functioning autism, 174 randomly selected controls, 840 undergraduates and 16 UK Mathematics Olympiad winners. Mean AQ 35.8 (SD 6.5) against 16.4 (SD 6.3) in controls; 80% of the autistic group scored 32 or above against 2% of controls. Test-retest at two weeks r = 0.7 in 17 students; Cronbach's alpha by domain 0.63 to 0.77.1
What it is made of
Fifty forced-choice items with four response options — definitely agree, slightly agree, slightly disagree, definitely disagree — scored 1 for an answer in the autistic direction and 0 otherwise, total 0 to 50. Five ten-item domains: social skill, attention switching, attention to detail, communication and imagination. No age norms, no country norms, no standardised subscale scores. Pooled general-population mean across 73 articles and 6,934 adults, 16.94 (males 17.89, females 14.88); pooled autistic mean 35.19.2
How well it performs
In 100 consecutive clinic referrals: sensitivity 94.5% and specificity 51.9% at 26; sensitivity 76.7% and specificity 74.1% at 32; AUC 0.78.6 The AQ-10 derivation study reported sensitivity 0.88, specificity 0.91 and PPV 0.85 at a cut-off of 6 in 449 autistic adults and 838 controls from an online research register.7 In 476 adults at a national specialist clinic the AQ-50 gave sensitivity 0.88 and specificity 0.20 at 26, sensitivity 0.71 and specificity 0.35 at 32, PPV 0.76 at both against a base rate of 0.73; the AQ-10 performed no better than chance.9 In 92 adults at a general mental health outpatient service, the AQ at 26 gave sensitivity 58%, specificity 86% and PPV 86% against a base rate of 0.68.11
Where it was validated
Originally in Cambridge, UK, in a cognitively able adult sample.1 Independently replicated in Dutch population and patient groups, where alpha for the total was 0.81 in students and 0.71 in the general population and the structure resolved to two factors rather than five.3 A systematic review of 38 articles covering nine instruments for autistic adults of average measured intelligence found the AQ-50, AQ-S, RAADS-R and RAADS-14 to be the only screening tools reaching satisfactory or intermediate psychometric values on strong or moderate evidence, while noting that risk of bias and applicability concerns limit what can be concluded about their performance in practice.14
What it cannot do
It cannot separate autistic traits from anxiety. In the 476-adult clinic sample, generalized anxiety disorder was the only significant predictor of AQ-10 score (p = 0.014) while an autism diagnosis was not (p = 0.537); 21.1% of false positives carried a GAD diagnosis against 3.7% of true negatives, with the self-evaluative phrasing of several items the suggested mechanism.9
Its five named domains do not exist as five factors. Two factors in the Dutch replication, Social Interaction and Attention to Detail correlating at only .19;3 a three-factor model best supported across 2,982 adults, with the same analysis declining to support total-scale scores at all;4 and a failure of unidimensionality under Rasch analysis, 26.2% of variance explained, three attention-to-detail items loading in reverse.5
It is not sex-neutral at item level. Across 17 published AQ models tested in 7,076 UK adults, only two items were invariant by gender; 20 were biased toward men and 21 toward women.12 A total is therefore not directly comparable across sexes.
Its predictive value does not travel. PPV 0.76 with specificity 0.20 in one specialist clinic,9 PPV 0.86 with specificity 86% in an outpatient mental health service,11 figures well below the published ones in a third referred sample.10 It has never been calibrated in a self-selecting online sample, which is where most people now meet it, so no honest PPV can be quoted for this page.
It measures endorsement, not cost. No item about impact, masking or context. Someone who has spent thirty years learning which answers other people give will score what they learned, not what it took.
What a clinician does with it
Use it as a structured opening to a history, not as a triage gate. Read the item pattern rather than the total, in line with the recommendation to prefer subscale over total-scale scores.4 Treat a below-threshold score as uninformative rather than reassuring: at the specialist clinic, 70% of those scoring 5 or below on the AQ-10 were autistic. NICE 1.2.3 already allows clinical judgment with informant history to override a low score, and that clause is doing most of the work.8 Do not use the total as evidence in a formulation.
Validity tier: 1 — validated and published, calibrated by independent labs. The AQ has an original peer-reviewed validation, independent replication in several countries and one of the better evidence bases of any adult screening instrument. The tier describes that evidence base, not the accuracy of any one score — and it is the evidence base itself that tells you the specificity is poor.
Diagram — E · What the fifty items report. The blank rows are why a score cannot settle anything alone. Everything that decides whether a trait is a difficulty or simply a feature of you — the job, the household, the masking, the cost afterwards — sits outside the fifty items.
STRENGTHS LENS
You went looking for a pattern instead of accepting a character note.
Most people who take this questionnaire have been carrying a private hypothesis for years, assembled from evidence nobody else noticed: the way one kind of noise ends an evening, the fact that friendships have always needed a structure to happen inside, the decades of being told you are too literal or too much or too quiet. Turning that into a question you can test is what a clinician does, with fewer resources and much more at stake.
The attention-to-detail items deserve a second look for the same reason. In the original study, science and mathematics students scored higher than humanities students, and mathematicians highest of all. That gets quoted badly, as though autistic traits were a qualification. They are not. But the capacity to hold a pattern steady after everyone else has moved on is real, and it has been doing quiet work in your life for a long time.
F. What helps
None of this is about improving your score. The AQ is a set of questions, not an exam, and the useful work starts once the number is on the page.
1. Read the five blocks, not the total.
The clearest psychometric recommendation in this literature is to use subscale scores rather than the total, because the total sums parts that do not move together. If your screener returns only a number, go back through the fifty items and mark which blocks your agreements clustered in. That shape describes you. The total describes a group you were compared against.
2. Take it twice, a fortnight apart.
Test-retest reliability in the original study was r = 0.7 over two weeks in seventeen students — respectable, not immovable. Anxiety in particular pushes scores up, so a form filled in at the end of a bad week is a different form. Two readings show which of your answers are stable and which are weather.
3. Write down every item you could not honestly answer.
In a study of 117 autistic adults completing standard research questionnaires, participants reported that questions lacked context, that many were unclear, and that the response scale forced choices they did not want to make.15 Your list of unanswerable items is not a failure to finish the task. It is often the most revealing page you will produce, and the one to take to an assessment.
4. Treat a low score as unfinished rather than settled.
At the specialist clinic, seventy per cent of the adults who scored 5 or below on the AQ-10 turned out to be autistic. If the questions did not fit, the instrument registers the fit and not the reason. A result below the line is a reason to try a questionnaire built on a different premise, not a reason to close the question.
5. Turn the score into three sentences, not a number.
No clinician is moved by a screener total, and the guidance already tells them to weigh history above it. What moves an assessment forward is specific: the school reports, the friendship that ended and exactly why, the job you left three months after the office went open-plan. Pick the three items you agreed with most strongly and write one real example under each.
Step 2 — Take the screener
Fifty questions, seven to ten minutes, free and confidential. You get a total between 0 and 50 and a band that goes with it. Read section C before you look at the band if you can: the two thresholds in common use were drawn by different research teams for different purposes, and neither of them was ever meant to settle anything.
Before you start
The AQ-50 is a published, peer-reviewed questionnaire that measures how far your answers resemble those of a research group of 58 already-diagnosed autistic adults tested in Cambridge in 2001. It is a screen, not an assessment, and it cannot tell you whether you are autistic. Its accuracy varies enormously with setting: in one specialist clinic of 476 adults, specificity at the standard cut-off of 26 was 0.20, meaning four in five non-autistic people there still scored above the line, and generalized anxiety predicted the score better than autism did. Its five named subscales have repeatedly failed to replicate, and its items are not neutral between men and women. Only a clinician can diagnose, and this screener is best used as a way of organizing what you already suspect.
Step 3 — Your workbook
Your answers save to this device only — we cannot see a word of what you write. This module's workbook records your total, then does the more useful work: finding which of the five blocks carried your score, listing the items you could not honestly answer, and turning the number into three concrete examples you could take to an assessment.
Your AQ-50 result
Took the screener? Put your total in below. Section D is the one to read first — the band is real but it is narrower than it looks. Entirely optional — skip it if you would rather just read.
Score bands: 0-25 = Fewer autistic traits; 26-31 = Possible - worth a closer look; 32-50 = Strong likelihood
Fields: AQ-50 · Autism-Spectrum Quotient · 50 items; Total score (0-50); My total (enter 0-50)
Where your agreements clustered
Section F, item 1. The five blocks are worth more than the sum of them. Guess if you have to — you will usually be right about which one carried your score.
Fields: The block I agreed with most: Social skill / Attention switching / Attention to detail / Communication / Imagination / Two or more are level; The block I agreed with least: Social skill / Attention switching / Attention to detail / Communication / Imagination / Two or more are level; What the highest block looks like on an ordinary Tuesday
The items you could not honestly answer
Section F, item 3. The ones where you wanted to write "it depends who, and on what day". This is the page that travels best into an assessment.
Fields: Items I could not answer as written; What I would have answered if I could have added a sentence; Tick: I want to bring this list to an assessment
Two readings, two different weeks
Section F, item 2. Test-retest was r = 0.7 over a fortnight, and anxiety pushes scores up. Come back and fill the second row in later.
Fields: First score, and the date; Second score, and the date; How the second week compared: Much calmer week / Slightly calmer / About the same / Slightly harder / Much harder week; Which answers moved, and why I think they did
If your score came in low
Section F, item 4. Seventy per cent of the adults who scored 5 or below on the AQ-10 at a specialist clinic were autistic anyway. A low score is unfinished business, not an answer.
Fields: What the questionnaire did not ask that I would have said yes to; How much I was answering as the version of me other people see: Not at all / A little / Quite a lot / Almost entirely / I cannot tell; What I still want an answer to
Three sentences, not a number
Section F, item 5. Three concrete examples beat any total. Pick the items you agreed with most strongly and put a real memory under each.
Fields: Example one; Example two; Example three; Who I would show this to first; The opening sentence I would use
Appendix — Research companion
Peer-reviewed research
7. Allison C, Auyeung B, Baron-Cohen S (2012). Toward brief red flags for autism screening: the short autism spectrum quotient and the short quantitative checklist in 1,000 cases and 3,000 controls. Journal of the American Academy of Child and Adolescent Psychiatry, 51(2), 202-212. DOI 10.1016/j.jaac.2011.11.003. View the paper Derivation and validation of the ten-item short forms across age bands. For the adult AQ-10 the analysis used 449 autistic adults and 838 controls, split into derivation and validation halves, and at a cut-off of 6 reported sensitivity 0.88, specificity 0.91 and positive predictive value 0.85. These are the figures on which UK national guidance rests. Limitation: participants were volunteers registered on the research centre's website who completed the questionnaire online, so the sample is self-selected rather than clinical or population-based, and the reported accuracy was not reproduced when the AQ-10 was later tested in a real referral clinic.
9. Ashwood KL, Gillan N, Horder J, Hayward H, Woodhouse E, McEwen FS, Findon J, Eklund H, Spain D, Wilson CE, Charman T, Murphy DG (2016). Predicting the diagnosis of autism in adults using the Autism-Spectrum Quotient (AQ) questionnaire. Psychological Medicine, 46(12), 2595-2604. DOI 10.1017/S0033291716001082. View the paper 476 adults assessed at a UK national specialist diagnostic clinic, 355 male and 121 female, median age 29; 346 of them (73%) received an autism diagnosis. At a cut-off of 26 the AQ-50 gave sensitivity 0.88, specificity 0.20 and PPV 0.76; at 32, sensitivity 0.71, specificity 0.35 and PPV 0.76; the AQ-10 at a cut-off of 6 gave sensitivity 0.77, specificity 0.29 and PPV 0.76 and did not predict diagnosis better than chance. 70% of those scoring 5 or below on the AQ-10 were autistic anyway, and generalised anxiety disorder was the only significant predictor of AQ-10 score in the multivariate model (p = 0.014) while an autism diagnosis was not (p = 0.537). The authors state that NICE recommendations supporting the AQ may need to be reconsidered. Limitation: a tertiary specialist clinic with a 73% base rate is an extreme setting, so the poor specificity partly reflects how heavily pre-selected the sample was and does not transfer directly to primary care or online use.
14. Baghdadli A, Russet F, Mottron L (2017). Measurement properties of screening and diagnostic tools for autism spectrum adults of mean normal intelligence: a systematic review. European Psychiatry, 44, 104-124. DOI 10.1016/j.eurpsy.2017.04.009. View the paper Systematic review, appraised with COSMIN and QUADAS-2, of 38 articles comprising 32 studies, five reviews and one book chapter, covering nine instruments used to screen or assess autistic adults without intellectual disability. Among screening tools, only the AQ-50, AQ-S, RAADS-R and RAADS-14 reached satisfactory or intermediate values on their psychometric properties with strong or moderate evidence behind them. The authors nonetheless conclude that risk of bias and applicability concerns limit the evidence for all of them, and that self-reported information and clinical expertise must be used alongside these instruments. Limitation: the review closed its search in 2016, so it predates the largest gender-invariance and factor-model studies of the AQ.
1. Baron-Cohen S, Wheelwright S, Skinner R, Martin J, Clubley E (2001). The autism-spectrum quotient (AQ): evidence from Asperger syndrome/high-functioning autism, males and females, scientists and mathematicians. Journal of Autism and Developmental Disorders, 31(1), 5-17. DOI 10.1023/A:1005653411471. View the paper The originating paper. Four groups: 58 adults with Asperger syndrome or high-functioning autism, 174 randomly selected controls, 840 Cambridge students and 16 UK Mathematics Olympiad winners. The autistic group scored a mean of 35.8 (SD 6.5) against 16.4 (SD 6.3) in controls, with 80% of the autistic group scoring 32 or above versus 2% of controls; no control woman scored 34 or above while 4% of control men did. Test-retest at two weeks was r = 0.7 in 17 students, and Cronbach's alpha by domain ran from 0.63 to 0.77. Limitation: the entire autistic validation group was 58 already-diagnosed, cognitively able British adults recruited around 2000, so the 80% figure and the 32 threshold describe that sample rather than autistic adults generally, and the authors state plainly that the AQ is not diagnostic.
12. Belcher HL, Uglik-Marucha N, Vitoratou S, Ford RM, Morein-Zamir S (2023). Gender bias in autism screening: measurement invariance of different model frameworks of the Autism Spectrum Quotient. BJPsych Open, 9(5), e173. DOI 10.1192/bjo.2023.562. View the paper 7,076 UK general-population adults, 5,246 women and 1,830 men, mean age 32.22, used to test measurement invariance across 17 published AQ models ranging from 6 to 50 items and 1 to 5 factors. Eleven models had satisfactory fit; among those, only two items were invariant by gender, with 20 items biased towards men and 21 towards women. Women were more likely to endorse items on social skills and communication, men more likely to endorse items on restricted and repetitive behaviours. The authors call for items to be rephrased to remove gender-related bias. Limitation: the sample was recruited from the general population without diagnostic confirmation, so the analysis establishes item-level bias but cannot quantify how many autistic women are missed by a given cut-off.
11. Bezemer ML, Blijd-Hoogewys EMA, Meek-Heekelaar M (2021). The predictive value of the AQ and the SRS-A in the diagnosis of ASD in adults in clinical practice. Journal of Autism and Developmental Disorders, 51(7), 2402-2415. DOI 10.1007/s10803-020-04699-7. View the paper 92 adults aged 18 to 62 referred for autism assessment at a Dutch outpatient general mental healthcare service, of whom 68% received a diagnosis. At the standard cut-off of 26 the AQ gave sensitivity 58%, specificity 86%, positive predictive value 86% and negative predictive value 58% - close to the mirror image of the specialist-clinic findings. The authors conclude that predictive values are specific to a population and a setting rather than being fixed properties of an instrument. Limitation: 92 participants is a small sample for estimating predictive values, and the confidence intervals around each figure are correspondingly wide.
4. English M, Gignac GE, Visser T, Whitehouse AJO, Maybery M (2020). A comprehensive psychometric analysis of autism-spectrum quotient factor models using two large samples: model recommendations and the influence of divergent traits on total-scale scores. Autism Research, 13(1), 45-60. DOI 10.1002/aur.2198. View the paper Confirmatory factor analyses of 11 competing published AQ models in two large samples, 1,702 undergraduates and 1,280 general-population adults, using polychoric rather than Pearson correlations. The three-factor model described by Russell-Smith and colleagues was best supported; the original five-factor structure was not. Critically, the analysis did not support the use of total-scale scores, because some proposed factors are uncorrelated or negatively correlated, and the authors recommend that researchers use subscale scores instead. Limitation: both samples were non-clinical, so the recommendation against total scores is established for research use in general populations rather than for screening in clinical settings.
3. Hoekstra RA, Bartels M, Cath DC, Boomsma DI (2008). Factor structure, reliability and criterion validity of the Autism-Spectrum Quotient (AQ): a study in Dutch population and patient groups. Journal of Autism and Developmental Disorders, 38(8), 1555-1566. DOI 10.1007/s10803-008-0538-x. View the paper Independent Dutch replication using 302 general-population adults, 961 students and small clinical groups with autism spectrum conditions, social anxiety disorder and obsessive-compulsive disorder. Cronbach's alpha for the total AQ was 0.81 in students and 0.71 in the general population. The original five subscales did not hold: confirmatory analysis supported two factors, a higher-order Social Interaction factor absorbing social skill, communication, attention switching and imagination, and a separate Attention to Detail factor, correlated at only r = .19. Limitation: the clinical comparison groups contained just 12 participants each, far too few to support the criterion-validity claims on their own.
5. Lundqvist LO, Lindner H (2017). Is the Autism-Spectrum Quotient a valid measure of traits associated with the autism spectrum? A Rasch validation in adults with and without autism spectrum disorders. Journal of Autism and Developmental Disorders, 47(7), 2080-2091. DOI 10.1007/s10803-017-3128-y. View the paper Rasch analysis of the full AQ in 349 adults: 130 with an autism spectrum diagnosis (mean age 29.3) and 219 university students without one. Item and person separation were adequate and targeting was good for the autistic group, but the AQ failed the criterion of unidimensionality, with only 26.2% of variance explained by the primary measure. Five items misfitted, one from imagination and four from attention to detail, and three of those showed negative point-measure correlations. The authors conclude a 12-item subset would carry almost the same information. Limitation: the comparison group was university students rather than a clinical or general-population sample, which inflates the contrast between groups.
2. Ruzich E, Allison C, Smith P, Watson P, Auyeung B, Ring H, Baron-Cohen S (2015). Measuring autistic traits in the general population: a systematic review of the Autism-Spectrum Quotient (AQ) in a nonclinical population sample of 6,900 typical adult males and females. Molecular Autism, 6, 2. DOI 10.1186/2040-2392-6-2. View the paper Systematic review of 73 articles reporting 78 independent studies, pooling 6,934 non-clinical adults alongside 1,963 matched autism spectrum cases. The pooled non-clinical mean AQ was 16.94 (95% CI 16.4 to 17.4), with males at 17.89 and females at 14.88; the pooled autistic mean was 35.19. These are the figures behind the typical-score claims quoted on most screener pages. Limitation: the pooled samples are dominated by university students and other convenience samples rather than representative population sampling, so the means are best read as a broad reference range and not as norms.
10. Sizoo BB, Horwitz EH, Teunisse JP, Kan CC, Vissers CTWM, Forceville EJM, Van Voorst AJP, Geurts HM (2015). Predictive validity of self-report questionnaires in the assessment of autism spectrum disorders in adults. Autism, 19(7), 842-849. DOI 10.1177/1362361315589869. View the paper 210 adults referred for autism assessment in the Netherlands, of whom 139 received an autism spectrum diagnosis and 71 another psychiatric diagnosis, plus 63 controls, tested with the RAADS-R, AQ-28 and AQ-10. Sensitivity and specificity were considerably lower than the values reported in the original validation literature; the RAADS-R had the highest sensitivity at 73% and the AQ short forms the highest specificity at 70% and 72%. Negative predictive values indicated that only about half of the referred non-autistic patients were correctly identified. The authors conclude that none of these instruments has sufficient validity to reliably predict an autism spectrum diagnosis in outpatient settings. Limitation: the study used the AQ-28 and AQ-10 rather than the full 50-item AQ, so its figures constrain the short forms directly and the full instrument only by inference.
13. Taylor EC, Livingston LA, Clutterbuck RA, Shah P (2020). Psychometric concerns with the 10-item Autism-Spectrum Quotient (AQ10) as a measure of trait autism in the general population. Experimental Results, 1, e3. DOI 10.1017/exp.2019.3. View the paper Psychometric evaluation of the AQ-10 in 6,595 non-clinical adults. Every Cronbach's alpha coefficient fell below 0.7, indicating poor internal reliability; confirmatory factor analysis rejected a single-factor model on every fit index, and parallel analysis identified four factors within the ten items. The authors attribute the multifactorial structure to the items having been drawn from five different subscales of the full AQ, and caution against using the AQ-10 as a measure of trait autism in general-population research. Limitation: the sample was entirely non-clinical, so the findings concern the AQ-10 as a trait measure in research and do not by themselves overturn its clinical screening use.
6. Woodbury-Smith MR, Robinson J, Wheelwright S, Baron-Cohen S (2005). Screening adults for Asperger syndrome using the AQ: a preliminary study of its diagnostic validity in clinical practice. Journal of Autism and Developmental Disorders, 35(3), 331-335. DOI 10.1007/s10803-005-3300-7. View the paper One hundred consecutive adult referrals to a UK diagnostic clinic, median age 32, male to female ratio 4:1. At a cut-off of 26 or above the AQ was 94.52% sensitive and 51.85% specific, correctly classifying 83% of the sample; at 32 or above it was 76.71% sensitive and 74.07% specific. Area under the ROC curve was 0.78. The authors recommend 26 in clinical settings to limit false negatives and the higher threshold of 32 for general-population screening to limit false positives. This is the study that put the 26 cut-off into circulation. Limitation: 100 people, all self-referred or clinician-referred with autism already suspected, so the specificity figure applies only to that pre-selected population.
Clinical frameworks and position statements
8. National Institute for Health and Care Excellence (2021). Autism spectrum disorder in adults: diagnosis and management, recommendation 1.2.3. NICE Clinical Guideline CG142, updated 14 June 2021. View the source UK national guidance for adult autism assessment. Recommendation 1.2.3 states that for adults with possible autism who do not have a moderate or severe learning disability, clinicians should consider using the Autism-Spectrum Quotient - 10 items, and that a score of 6 or above, or clinical suspicion taking account of informant history, should lead to a comprehensive assessment. The guideline was amended in June 2021 and the recommendation remains current. Limitation: the recommendation rests on the AQ-10 derivation study rather than on independent clinic data, and a large specialist-clinic evaluation published in 2016 concluded that this recommendation may need to be reconsidered.
Lived experience
15. Stacey R, Cage E (2023). Simultaneously vague and oddly specific: understanding autistic people's experiences of decision making and research questionnaires. Autism in Adulthood, 5(3), 263-274. DOI 10.1089/aut.2022.0039. View the source 117 autistic adults completed an online survey in which they answered four standard research questionnaires and then gave open-ended feedback on each, analysed by content analysis. Participants reported that questions lacked context, that many items were unclear or hard to interpret, and that Likert response formats forced choices that did not represent them, leading the authors to describe the measures as of questionable validity for autistic people. Decision-making itself was reported as shaped by internal state, energy, perceived pressure, environment and time. Limitation: an online self-selected sample of autistic adults able to complete a long written survey, and the feedback concerns research questionnaires generally rather than the AQ specifically.
Peer-reviewed = checked by independent experts before publication. Clinical model = an established professional framework, not a single study.
Up next
Module 4 - GQ-ASC — The Autism Questionnaire Written for Women
All modules in Autistic Self-Discovery
A number is a starting point, not an answer. This course was built by clinicians who are part of the New Path Family. If you have landed in the middle band and cannot tell whether that means something or nothing, that ambiguity is exactly what a conversation is for — and it is the most common place people get stuck with this questionnaire. Therapy for clients in California and coaching worldwide, all by telehealth, are offered by our sister company New Path Family of Therapy Centers, Inc. A conversation costs nothing and there is no pressure. Saving this for later counts too. Talk with the New Path team
