Show the full module text
Module 1 — What an Assessment Is For
What an assessment produces, what it cannot produce, and why the answer at the end is a clinical judgement rather than a measurement.
The Assessment · Part One — Before you book
The big idea
Short on capacity today? The big idea: Most people arrive at an assessment expecting a test. Something will be administered, a number will come out, and the number will settle it.
That is not what happens. An assessment produces a judgement — formed by a trained person, from several different kinds of evidence, none of which decides it alone. That is a weaker promise than a test, and a much more useful one. This module is about why, and about what follows from taking it seriously.
Step 1 — The lesson
A. Who this course is for
You have decided to seek an in-depth assessment, or you are close enough to deciding that you want to know what you are walking into.
That is the only assumption made here. This course is a walkthrough, not a decision aid. It does not argue you into an assessment and it does not argue you out of one. If you are still weighing it up, the Self-Identification course has a module built for exactly that question, and it is a better use of your time than this one.
One thing carried across from there, said once and then left alone: self-identification and an in-depth assessment are a fork, not a ladder. This course is one branch of that fork. It is not the top of anything, and nobody who took the other branch is standing on a lower rung. You are here because this branch is the one that gets you what you need.
Over eight modules the course walks the process end to end: who assesses and why it matters, what you will be asked and how to prepare, which instruments you are likely to meet, what happens in the outcome conversation, how to read the report, and what a diagnosis does and does not change afterwards. It describes assessment as it is practised generally, not any particular service.
B. What an assessment actually produces
The screener courses are built on a deliberately narrow claim. A questionnaire, however good, returns one of two things: low or no indication, or possible or elevated indication. Two lamps. Nothing else. A screener cannot tell you that you are autistic, and it cannot tell you that you are not.
An assessment gives you a third thing, and it is worth being precise about what that third thing is. It is a clinical judgement: a trained person’s conclusion about whether your presentation meets diagnostic criteria, and about what else might account for it.
That judgement is assembled from several sources, and the single most common misconception about assessment is that one of them decides it. None of them does.
- Self-report. What you say about your own experience — the inner account nobody else has access to. Necessary, and on its own insufficient, because how you experience yourself and how you present to an observer are genuinely different measurements.
- Developmental history. What was happening at five, at nine, at fourteen. Both autism and ADHD are developmental conditions, so evidence of early presence is not a formality — it is part of the definition.
- Informant accounts. A parent, a sibling, a long-standing partner. Somebody who watched you from the outside, ideally over years. This is often the part people most want to skip and the part that most changes an assessor’s confidence.
- Structured instruments. Interviews and rating scales with defined questions, defined scoring, and published data behind them. Module 5 goes through the ones you are likely to meet.
- Direct observation. How you actually are in the room, on the call, over an hour or three — including things you would never think to report about yourself.
An assessor’s job is to hold all five against each other and reach a defensible conclusion about the whole picture. When the sources agree, the judgement is easy. When they disagree — and they disagree often — the judgement is where the skill lives. That is the work you are waiting for.
Diagram — A · Five strands, one rope. A questionnaire, however good, returns one of two things: low or no indication, or possible or elevated indication. An assessment gives you a third thing, and it is worth being precise about what it is: a clinical judgement, a trained person's conclusion about whether your presentation meets diagnostic criteria and about what else might account for it. That judgement is assembled from five sources, and none of them decides it alone. Self-report is what you say about your own experience, the inner account nobody else has access to — necessary, and on its own insufficient, because how you experience yourself and how you present to an observer are genuinely different measurements. Developmental history is what was happening at five, at nine, at fourteen; both autism and ADHD are developmental conditions, so evidence of early presence is not a formality, it is part of the definition. Informant accounts come from a parent, a sibling or a long-standing partner — somebody who watched you from the outside, ideally over years — and this is often the part people most want to skip and the part that most changes an assessor's confidence. Structured instruments are interviews and rating scales with defined questions, defined scoring and published data behind them. Direct observation is how you actually are in the room, on the call, over an hour or three, including things you would never think to report about yourself. An assessor's job is to hold all five against each other and reach a defensible conclusion about the whole picture.
C. Two assessors, one person
Now the part that the rest of this course rests on, and the part most descriptions of assessment leave out.
A judgement made by a person can differ from a judgement made by another person. This has been measured, repeatedly, and the honest answer is: agreement is good but it is a long way from perfect, and it depends heavily on how the assessment is done.
The standard measure is Cohen’s kappa — agreement between two raters after subtracting the agreement you would expect from chance alone. A kappa of 1.0 is perfect; 0 is chance-level.
The DSM-5 field trials. When the current diagnostic manual was being finalised, its authors ran a test-retest study across multiple clinical sites: two clinicians, independently, assessing the same patient. For autism spectrum disorder, the pooled kappa was 0.69 (95 per cent confidence interval 0.58–0.79). For ADHD it was 0.61 (0.51–0.71). Both fell into the trial’s “very good” band, and both were among the strongest results in the whole exercise — only five diagnoses reached that band at all, while three came out below 0.20.
One qualification, said plainly rather than buried in a footnote: both figures come from the trials’ child and adolescent sites. Neither autism nor ADHD was among the fifteen diagnoses tested in the adult arm, so no adult kappa exists for either. These are the closest measurements available, not a measurement of what happens in an adult clinic.
With that said, the first thing the evidence says is reassuring: as psychiatric categories go, these two are at the reliable end.
The second thing is less comfortable. Look inside those pooled figures. The ADHD result was 0.71 at one site and 0.45 at another — the same criteria, the same manual, the same study, and a difference big enough to matter to an individual person sitting in front of an individual clinician. Reliability is not a property of the diagnosis. It is a property of the diagnosis as practised in a particular place, by particular people.
Diagram — B · The same target, two clinics. A judgement made by a person can differ from a judgement made by another person. This has been measured, repeatedly, and the honest answer is that agreement is good but a long way from perfect, and that it depends heavily on how the assessment is done. The standard measure is Cohen's kappa: agreement between two raters after subtracting the agreement you would expect from chance alone, where 1.0 is perfect and 0 is chance-level. When the current diagnostic manual was being finalised, its authors ran a test-retest study across multiple clinical sites, with two clinicians independently assessing the same patient. For autism spectrum disorder the pooled kappa was 0.69, with a 95 per cent confidence interval of 0.58 to 0.79; for ADHD it was 0.61, interval 0.51 to 0.71. Both fell into the trial's very good band and both were among the strongest results in the whole exercise — only five diagnoses reached that band at all, while three came out below 0.20. One qualification belongs in the same breath rather than in a footnote: both figures come from the trials' child and adolescent sites. Neither autism nor ADHD was among the fifteen diagnoses tested in the adult arm, so no adult kappa exists for either, and these are the closest measurements available rather than a measurement of what happens in an adult clinic. The first thing the evidence says is reassuring: as psychiatric categories go, these two are at the reliable end. The second thing is less comfortable. Inside the pooled ADHD figure, the result was 0.71 at one site and 0.45 at another — the same criteria, the same manual, the same study, and a difference big enough to matter to an individual person sitting in front of an individual clinician. Reliability is not a property of the diagnosis; it is a property of the diagnosis as practised in a particular place, by particular people.
And when the structure is removed, agreement falls further. In a 2017 Australian study, 27 health professionals each watched two of nine video-recorded assessments and rated the person against the diagnostic criteria. There was 100 per cent agreement on the classification for only three of the nine cases. Only 24 per cent of the clinicians reached good or excellent agreement — kappa above 0.6 — with the original assessment team’s conclusion.
The mirror image matters just as much. In the same year, a multi-site study had ten recorded administrations of a structured diagnostic interview rated by five raters each, drawn from eleven clinicians across eight clinical sites. Agreement on the diagnostic classification was kappa 0.83 — high, and comparable to what specially trained research interviewers achieve.
Read those two studies together and you get the most useful sentence in this module:
Structure is what produces agreement. Where clinicians work through the same defined interview, they largely land in the same place. Where they are left to form a global impression, they come apart — on the same person, from the same recording.
This is not an argument against assessment. It is an argument for a well-structured one, and it is why asking what instruments a service uses is not pedantry — it is the difference between the 0.83 study and the 24 per cent study.
Diagram — C · With a straight-edge, and without. When the structure is removed, agreement falls. In a 2017 Australian study, 27 health professionals each watched two of nine video-recorded assessments and rated the person against the diagnostic criteria. There was 100 per cent agreement on the classification for only three of the nine cases, and only 24 per cent of the clinicians reached good or excellent agreement — kappa above 0.6 — with the original assessment team's conclusion. The mirror image matters just as much. In the same year, a multi-site study had ten recorded administrations of a structured diagnostic interview rated by five raters each, drawn from eleven clinicians across eight clinical sites. Agreement on the diagnostic classification was kappa 0.83 — high, and comparable to what specially trained research interviewers achieve. Read those two studies together and you get the most useful sentence in this module. Structure is what produces agreement: where clinicians work through the same defined interview they largely land in the same place, and where they are left to form a global impression they come apart, on the same person, from the same recording. This is not an argument against assessment. It is an argument for a well-structured one, and it is why asking what instruments a service uses is not pedantry — it is the difference between the 0.83 study and the 24 per cent study. It also means something for you personally: an answer you did not expect is information, not a verdict, and a second opinion is a normal thing to seek.
It also means something for you personally. An answer you did not expect is information, not a verdict. A second opinion is a normal thing to seek, and the evidence above is why.
D. The differential, and what “gold standard” really means
If a judgement can vary, why buy one at all?
Because the assessor is not only asking is this autism or is this ADHD. They are asking what else could produce this exact picture — anxiety, trauma, depression, chronic sleep debt, a thyroid problem, a specific learning difference, or two of those at once. That is the differential, and it is the one thing here you genuinely cannot do for yourself: not because you are not clever enough, but because you have exactly one case to compare against and no way to see yourself from outside. The value is not the label at the end. It is everything that got ruled in or out on the way there.
Which brings us to the instruments. You will hear certain tools described as the gold standard — structured diagnostic interviews, standardised observation schedules, normed rating scales completed by you and by someone who knows you. Module 5 takes the specific ones apart properly. For now, one correction is enough: a gold-standard instrument is a tool that improves agreement between assessors. It is not an oracle, and it does not return the answer by itself.
Two findings make the point concretely, and both are about adults.
In a study of 88 adults assessed at a specialist service, the standard observation schedule used with adults showed 92 per cent sensitivity — it picked up nearly everyone who went on to receive a diagnosis — but only 57 per cent specificity. Of the people who scored above its cut-off, about half received a diagnosis. Scoring above the threshold on the best-regarded observation tool in adult autism assessment was, in that sample, close to a coin toss.
And in a study of 2,310 people aged 4 to 72, specificity was lower for adolescents and adults than for children on the same instruments — and for the adult group, a classifier built on observation items alone performed as well as one that also used the developmental interview.
Neither finding means the instruments are bad. They mean the instruments are inputs. The judgement is made by the person reading them alongside your history, your account, and everything else in section B. That is the whole architecture of an assessment, and now you know why it is built that way.
Diagram — D · A wide net, and a loose sieve. If a judgement can vary, why buy one at all? Because the assessor is not only asking is this autism or is this ADHD. They are asking what else could produce this exact picture — anxiety, trauma, depression, chronic sleep debt, a thyroid problem, a specific learning difference, or two of those at once. That is the differential, and it is the one thing here you genuinely cannot do for yourself: not because you are not clever enough, but because you have exactly one case to compare against and no way to see yourself from outside. The value is not the label at the end; it is everything that got ruled in or out on the way there. Which brings us to the instruments. You will hear certain tools described as the gold standard — structured diagnostic interviews, standardised observation schedules, normed rating scales completed by you and by someone who knows you. One correction is enough for now: a gold-standard instrument is a tool that improves agreement between assessors. It is not an oracle, and it does not return the answer by itself. Two findings make the point concretely, and both are about adults. In a study of 88 adults assessed at a specialist service, the standard observation schedule used with adults showed 92 per cent sensitivity — it picked up nearly everyone who went on to receive a diagnosis — but only 57 per cent specificity, and of the people who scored above its cut-off about half received a diagnosis. Scoring above the threshold on the best-regarded observation tool in adult autism assessment was, in that sample, close to a coin toss. And in a study of 2,310 people aged 4 to 72, specificity was lower for adolescents and adults than for children on the same instruments, and for the adult group a classifier built on observation items alone performed as well as one that also used the developmental interview. Neither finding means the instruments are bad. They mean the instruments are inputs, and the judgement is made by the person reading them alongside your history, your account, and everything else.
E. Four things an assessment cannot do
It cannot tell you who you are. It can tell you whether your presentation meets a set of criteria. Those are not the same question, and the second one is not a clinical question at all. People who go in hoping to be handed an identity often come out holding a document instead.
It cannot retroactively fix the years before it. The school you struggled through, the job you lost, the relationship that ended — a diagnosis explains them. It does not repair them, and the grief that arrives with an explanation is a real and well-documented part of this. It is not a sign that something went wrong.
It cannot guarantee access to anything by itself. A diagnosis usually opens doors that a self-description does not, but the specific door you need has its own rules about what evidence it takes, from whom, and how recently. Find that out before you book, not after.
It cannot be relied on to be quick. In England, in June 2026, 294,792 people had an open referral for suspected autism, and 86.8 per cent of them had been waiting at least thirteen weeks. Independent routes are faster, and speed is much of what is being paid for. Either way, plan for this to be a process, not an appointment.
→ The module promised in section A — the one built for the question of whether to seek an assessment at all — is Self-Identification, Module 11.
Diagram — E · A blank key and a cut key. One: write down, before anything else, what you want the assessment to change. One sentence. If it names something concrete — an adjustment at work, a prescription, a document somebody has asked you for — the rest of this course is straightforwardly useful to you. If it names certainty, read section C again, because certainty is the thing an assessment is least able to deliver, and knowing that now is better than discovering it in the outcome conversation. Two: start gathering the history early, because it is the input you control. Structure and corroboration are what make a judgement reliable, so school reports, old work appraisals, a parent or partner willing to answer questions and concrete examples of where things went wrong are the single highest-value preparation you can do. Three: ask what the process is before you commit to it — what instruments are used, who assesses and what qualification they hold, how many appointments, what document you receive, and what happens if the answer is no. A service that answers those clearly is describing the structured end of the evidence. Four: decide now how you will hold a no. Not to rehearse disappointment, but to make it usable: given kappa values between 0.45 and 0.83 depending on how the work is done, a no is a considered clinical judgement and not a fact about the universe. If it does not fit what you know about yourself, a second opinion is a reasonable next step; if it comes with a different explanation that does fit, that is the differential doing its job, and it may be worth more than the answer you went in for. Five: bring the question you actually want answered, not the one you think is allowed. Assessors are used to am I autistic and they hear it constantly; what produces a more useful hour is the specific version, such as I cannot hold down a job past eighteen months and I want to know why. That is the question the differential is built to attack.
F. What helps
1. Write down, before anything else, what you want the assessment to change.
One sentence. If it names something concrete — an adjustment at work, a prescription, a document somebody has asked you for — the rest of this course is straightforwardly useful to you. If it names certainty, read section C again. Certainty is the thing an assessment is least able to deliver, and knowing that now is better than discovering it in the outcome conversation.
2. Start gathering the history early — it is the input you control.
Sections B and C both point the same way: structure and corroboration are what make a judgement reliable. School reports, old work appraisals, a parent or partner willing to answer questions, concrete examples of where things went wrong. Module 3 is entirely about assembling this, and it is the single highest-value preparation you can do.
3. Ask what the process is before you commit to it.
What instruments are used, who assesses and what qualification they hold, how many appointments, what document you receive, and what happens if the answer is no. A service that answers those clearly is describing the structured end of the evidence in section C.
4. Decide now how you will hold a “no”.
Not to rehearse disappointment — to make it usable. Given kappa values between 0.45 and 0.83 depending on how the work is done, a no is a considered clinical judgement and not a fact about the universe. If it does not fit what you know about yourself, a second opinion is a reasonable next step. If it comes with a different explanation that does fit, that is the differential doing its job, and it may be worth more than the answer you went in for.
5. Bring the question you actually want answered, not the one you think is allowed.
Assessors are used to am I autistic. They hear it constantly. What produces a more useful hour is the specific version: I cannot hold down a job past eighteen months and I want to know why. That is the question the differential is built to attack.
Module 2 goes through who assesses, what their training actually covers, and how the professions differ — because the answer to “who is doing this” turned out, in section C, to matter as much as the instruments do.
Up next
Module 2 — Who Assesses You, and Why It Matters
