What a Test Paper Can’t Tell You
Why question level analysis is not the diagnosis it appears to be
It is the Tuesday after the autumn assessment, and Priya has her spreadsheet open. Thirty children in her Year 5 class, each row colour-coded question by question. Three of the questions on the maths paper involved fractions, and twelve of her children got at least one of them wrong. The conclusion seems to write itself: this class is shaky on fractions. She blocks out a week to reteach the topic.
It is one of the most reasonable things a teacher can do. The paper has been sat and marked, the data is sitting right there, so why not read it for clues about what to teach next? This is the promise of question level analysis, or QLA, and it is why, despite years of people like me grumbling about it, teachers keep coming back to it. The instinct is sound. The problem is that the test paper was never built to answer the question Priya is now asking of it.
A test and a diagnosis are different tools
Almost every set-piece assessment a school runs is designed to produce a summary. A score, a grade, a rank, a sense of how much a pupil has learned across a domain. It answers the question how much has this pupil learned overall?
A diagnosis answers a different question: what, precisely, does this pupil understand and misunderstand, and what should they work on next? The two sound like cousins, but they pull a test’s design in opposite directions. A summary test wants questions that spread pupils out and cover a lot of ground quickly. A diagnostic wants questions that isolate one idea at a time and reveal the reason behind a wrong answer. You cannot usually get both from the same paper, any more than a thermometer can tell you what is causing the fever. A test score is a temperature reading. It tells you something is off without telling you what.
This is why QLA so often disappoints. We take an instrument built to summarise and ask it to diagnose. There are three reasons it struggles.
The wrong answer doesn’t tell you why
Consider one of Priya’s fraction questions: a wordy problem that asks pupils to work out, in context, the result of adding 3/8 and 1/4. To get it right, a child has to read and interpret the problem, recognise that the denominators differ, find a common denominator, carry out the addition, and not trip over any arithmetic along the way. That is at least four distinct things bundled into one mark.
So when Jack gets it wrong, what have we learned? Almost nothing we can act on. Was it the reading? The common denominator? A careless slip in the arithmetic? “Wrong” is mute about the cause. Exam-style questions are deliberately built this way, because for summarising attainment we want questions that draw on several things at once. For diagnosis, that same richness is fatal. A question that tests four things cannot tell you which of the four has gone wrong.
One or two questions is mostly noise
Even setting the bundling aside, three questions is a tiny amount of evidence. We have written before about reliability and about how much noise sits inside any single response. A child who gets a fractions question right might have guessed, or used a trick that won’t generalise. A child who gets one wrong might have known it perfectly well and misread the question. On a different day, with a slightly different set of three questions, the same child could easily land on a different score.
This is why a verdict on an individual child from two or three questions is so fragile. The signal is real but faint, and it is swamped by the random business of guessing and slipping.
Three questions are a keyhole, not a window
Finally, there is the matter of what was asked. The fractions strand of the curriculum is large: equivalence, comparing, adding, subtracting, multiplying, fractions of quantities, and more. Priya’s paper sampled three points in that landscape. If a child stumbles on adding fractions, that tells us little about whether they can compare or simplify them because those weren’t on the paper.
This is the trap of reading a QLA grid as though each question stands in for a topic. The questions are a handful of footholds, not a map of the terrain.
What would actually diagnose
It is worth knowing what a true diagnostic looks like, because it throws the test paper’s limits into relief. A diagnostic question is stripped back so that it tests the one idea that you want to know if a student can do. If you want to know whether a child can add fractions with unlike denominators, you ask them exactly that, in the plainest possible form, with nothing else in the way.
In the case of multiple choice questions, a good diagnostic question is built so that each tempting wrong option corresponds to a particular misconception. The child who adds numerators and denominators together to make 4/12 has told you something specific and useful.
A diagnostic is not trying to produce a precise measure of attainment. It is trying to point a child’s next bit of study in the right direction. This is the heart of it: diagnosis is about working out what to do next, not about ranking.
So what should we do on Tuesday morning?
None of this means Priya should close the spreadsheet. It means she should read it for what it is.
Treat the pattern as a question about the class, not a verdict on individuals. Twelve children out of thirty struggling with a fractions question is a hypothesis worth following up, not a diagnosis to act on blindly. It is a good place to point your attention next.
Then follow up cheaply. The fastest diagnostic instrument in the room might be a quick set of questions to isolate what problems with fractions the class has (compare these two fractions, now add these two with the same denominator, now these with different ones). Within ten minutes you will know far more than the paper told you. Cheaper still, you can simply ask: was that a slip, or were you stuck? Children in a low-stakes setting are surprisingly honest, and surprisingly accurate, about which of their errors were careless and which were real.
The relative QLA reports that test providers supply are just a starting point. They do some useful work in adjusting for question difficulty, but they are sensitive to how recently you taught a topic. Doing brilliantly on fractions may mean your class is strong, or merely that you finished the unit last week while other schools did it a term ago.
Teachers generally know more than the test paper. In deciding what interventions and remediation is needed, lean on everything you already know about them - classwork, quizzes, earlier tests - rather than three questions on one paper. The paper’s evidence has value, especially because it reflects what pupils can recall now rather than during the original teaching, but it should never be the sole basis for a decision.
The questions our pupils get wrong are informative. They just begin an enquiry rather than settle one.



Yes.
Priya probably has a sense of the issue (particularly since she is teaching at KS2 and not KS4) but she doesn't have the set of multiple-choice questions easily available that would help with the diagnosis. We're creeping towards it - various maths schemes are getting there slowly, I think - but it's quite likely Priya has to choose between taking an educated guess and spending yet another bit of her weekend crafting the questions.
Moving away from maths, the problem is worse. When it's a 3 mark "explain" question in science, it's a fair bit harder. The low mark is often shaky recall, of incomplete understanding, coupled with an inability to turn what's in your head into words on the page. Those things can be somewhat separated, although it's not easy and it's not something we're good at, so it's pretty common to just end up teaching it all again, but faster. Even if you do better than that, fixing any one of those things might not make any difference.
Excellent piece. QLA sounds like a sensible way of analysing components of knowledge. The issue is often it is the prior knowledge that is the issue, not just the nugget we are focused on.