An email lands with a link and a deadline. It calls itself a “short online assessment.” It does not tell you which test it is, how long you get, or what happens when the clock beats you. So you search, and every result is either the vendor telling you not to worry or a prep company telling you to worry a lot and buy something.
I have taken four of the five tests below. Not read about them. Sat them, watched the timer, and lost questions I could have answered. So this is the ranking I wanted when I was the one staring at the invite: five cognitive aptitude tests, scored on how brutal they actually are, with the reasons spelled out rather than asserted.
One thing before the list, because it is the whole reason to trust the rest of it. I have not sat the Watson-Glaser. I have sat the CCAT, the Wonderlic, the PI Cognitive Assessment, and an SHL numerical reasoning test. The Watson-Glaser I rank on its documented format and on what it demands of a candidate, and I flag exactly where that ranking is inference rather than experience. A first-person ranking that quietly pretends to five sittings is worth nothing, so I would rather show you the seam.
The quick answer
The hardest of the five, taken as a whole, is the CCAT. Not because it has the tightest clock, because it does not, but because it is the only one that stacks three difficulties at once: a heavy question load, a difficulty curve that rises toward the end, and five answer options instead of four.
Worth separating two questions that usually get mashed together. “Hardest across a full sitting” and “hardest to see coming” have different answers. The CCAT wins the first, for the reasons above. The PI wins the second, and I have said so elsewhere: the CCAT is hard in a way you can anticipate, and the PI is hard in a way that surprises you. If you are asking which one will ambush an unprepared candidate, bet on the PI. If you are asking which one grinds down a prepared one, it is the CCAT.
For the pure-clock answer, the PI Cognitive Assessment and the Wonderlic tie for worst at roughly 14 seconds a question. And for “the one where doing the maths correctly still loses you the point,” that is the SHL.
How I am scoring “brutal”
Difficulty rankings usually collapse into one number, which is why they are useless. Three separate things make an aptitude test hard, and most people are only weak at one.
Time pressure. Seconds per question if you tried to answer every one. The number everyone quotes, and the least interesting of the three, because it is also the most trainable.
Question density. How much you have to decode before you can start answering. A number series is low density: read eight characters and you are working. A six-division revenue table with a footnote about thousands is high density: half your budget goes on finding the number, not using it.
Recoverability. What one bad question costs you. Nobody scores this axis and it is the one that decides your result. A test whose hard questions sit at the end punishes a slow start twice: once when you lose the time, again when you reach the nastiest material with nothing left.
Here is the ranking on those three axes.
| Rank | Test | Load | Seconds per question | Time pressure | Question density | Recoverability | I sat it |
|---|---|---|---|---|---|---|---|
| 1 | CCAT | 50 Q / 15 min | ~18 sec | High | Medium | Worst | Yes |
| 2 | PI Cognitive | 50 Q / 12 min | ~14 sec | Worst | Low | Medium | Yes |
| 3 | SHL numerical | 10 to 18 Q / 18 to 25 min | ~75 to 108 sec | Low | Worst | Poor | Yes |
| 4 | Wonderlic | 50 Q / 12 min | ~14 sec | Worst | Low | Good | Yes |
| 5 | Watson-Glaser | 40 Q / 30 min | ~45 sec | Medium | Medium | Good | No |
Question counts and time limits are the published specs from the test makers: Criteria Corp for the CCAT, The Predictive Index for the PI, Wonderlic for the Wonderlic, SHL for the Verify range, and TalentLens for the Watson-Glaser. The three difficulty columns are my scoring.

1. CCAT: the only one where a slow start compounds
The CCAT gives you 50 questions in 15 minutes, about 18 seconds each. On paper that is three and a half seconds more room per question than the PI, and that extra room is exactly the trap.
The first four minutes felt comfortable when I took it. Comfortable is the wrong feeling on a test where roughly one candidate in a hundred reaches the end. I answered about twelve questions in those four minutes and thought I was doing well, which meant I was burning around 25 seconds each against an 18-second budget without noticing. Then a number series arrived that I could not immediately see, and I did the thing you must not do: I stared at it. Thirty seconds. On any other test that is a rounding error. Here it was nearly two questions I would never reach.
The CCAT does not ask whether you are smart. It asks whether you stay calm while a clock takes something away from you.
What puts it at number one is the ramp. The early questions are gentle and the last ten are designed to be slow to parse, so the material gets harder precisely as your remaining time gets shorter. Add five answer options rather than four, which drops a blind guess from a 25 percent shot to 20 percent, and a bad first five minutes is close to unrecoverable. On the PI a slow start costs you questions. On the CCAT it costs you the questions and the ones you would have got right.
It also leans harder on vocabulary than the other four. I hit sentence-completion and analogy items where I either knew the word or I did not, and no reasoning bridges that in 18 seconds. If English is your second language, that lean matters more than any difficulty score captures. Seeing how the question types are structured and scored before you meet them cold removes the one cost on this test that is entirely avoidable.
The full sitting, minute by minute, is in how hard the CCAT actually felt from the chair.
2. PI Cognitive Assessment: the worst clock, the kindest structure
Fifty questions, twelve minutes, about 14 seconds each. The tightest clock of the five, and it feels like it. My twelve minutes evaporated. I caught myself re-reading a word problem a second time and could physically feel two or three questions slide away while I did it.
So why second and not first? Because everything except the clock is merciful. The difficulty is flat rather than ramped, with numerical, verbal, and abstract items interleaved, so there is no back third where the test turns on you. Four answer options rather than five makes a blind guess a one-in-four shot. There is no guessing penalty. And you can move backward and forward between screens, so a skipped question is parked rather than lost.
That is what makes a bad start survivable here. I lost the plot on an abstract pattern early, decided around halfway to stop trying to be right and start being efficient, and still finished somewhere in the mid-thirties for questions attempted. On the CCAT the same recovery would have landed me in the hardest material with two minutes left.
The interface does one genuinely cruel thing: it shows time remaining in whole minutes rather than seconds. You do not get a countdown, you get a blunt notice that another minute is gone, which removes any way to pace inside the minute.
I put the two head to head in PI Cognitive Assessment vs CCAT, having taken both, and the longer account is in what 50 questions in 12 minutes actually feels like.
3. SHL numerical reasoning: correct arithmetic, wrong box
The SHL is the odd one out and the reason this ranking needed three axes instead of one.
On time pressure it is by far the gentlest here: roughly 75 to 108 seconds per question depending on the version. Verify Interactive numerical runs up to 10 questions in 18 minutes, Verify numerical reasoning up to 18 in 25 minutes, and Verify numerical ability up to 16 in 20 minutes. Against 14 seconds on the PI, that is luxurious. It is still brutal, because the difficulty moved somewhere else entirely.
The SHL does not punish you for being bad at maths. It punishes you for reading the chart one row too fast. The arithmetic is percentages, ratios, and the occasional currency conversion. With no clock you would be near perfect. What eats the time is the hunting: six divisions, four quarters, two currencies, and a footnote telling you the figures are in thousands. You are not calculating, you are locating.
And the wrong answers are engineered. If the right answer is a 12.5 percent increase, one option is what you get dividing by the end value instead of the start, another is the adjacent quarter’s jump, another is the raw difference with “percentage” ignored. Every distractor is a mistake someone has actually made. I got a question shaped almost exactly like that wrong on a practice run, not because I could not do the sum but because I divided by the wrong base.
That is why recoverability is poor despite the loose clock. On the CCAT a wasted question costs you time. Here a confidently wrong answer costs you a mark and gives you no signal that anything went wrong, so the misreading habit rides into the next question. Full account: the SHL numerical test and the data questions nobody warns you about.
4. Wonderlic: the same clock as the PI, none of the teeth
Fifty questions in twelve minutes, roughly 14 seconds each, identical to the PI. On raw time it is joint worst. On everything else it is the most forgiving test here, which is why it lands at four. That three-place gap between the CCAT and the Wonderlic is the single most useful thing in this ranking: the two look almost identical on a spec sheet, both 50 questions against a tight clock, but the CCAT ramps its difficulty and gives you five answer options while the Wonderlic stays flat and forgives a guess, so the same candidate will usually score better on the Wonderlic relative to the norm group.
The content is the gentlest of the three speeded tests: verbal, numerical, and logic items that are individually not hard, with no ramp and no vocabulary cliff. There is no penalty for a wrong answer, so a blank is strictly worse than a guess. And it is the most trainable of the five, because so much of the score is pacing and format familiarity rather than reasoning horsepower.
The moment that stuck with me was hitting a word problem I knew I could solve, working out that it would cost me around 40 seconds, and making myself skip it. Forty seconds is three or four easier questions. That is the whole test in one decision: not “can you solve this,” but “can you tell in two seconds which questions to abandon.”
Where the Wonderlic gets genuinely confusing is afterwards. It hands you a raw number out of 50 with no key, and the average sits somewhere around 20 to 22 depending on whose sample you trust, so a score that sounds like a failing grade in school terms is often perfectly good. I untangled the raw-score-versus-percentile mess, and why the “multiply by five for your IQ” rule is a trap, in what a good Wonderlic score actually is.
5. Watson-Glaser: the one I have not sat
This is the call I am least certain about, so here is exactly why.
I have not taken the Watson-Glaser Critical Thinking Appraisal. Everything below comes from its documented format and what that format demands, not from the chair. If you have sat it and would rank it higher, you are probably right and I am probably underweighting it.
On the published specs it is 40 questions in 30 minutes, about 45 seconds each, across five sections: inference, recognition of assumptions, deduction, interpretation, and evaluation of arguments. It is published by TalentLens, a Pearson division, was first written in 1925, and it is the standard screen for UK trainee solicitor pipelines at the Magic Circle firms.
It goes last because the clock is loose by this list’s standards and because a bad answer is contained: each question stands alone against its own decision rule, so one misread does not propagate.
The argument against my own ranking is that it measures something the other four do not touch. They ask how fast you reason. This one asks whether you can switch off your opinion and evaluate an argument on its own logic, judging whether a conclusion follows from the passage rather than from what you know about the world. That is a genuinely unnatural skill, and because each section applies a different decision rule, skimming the instructions costs real percentile points. If your invite says Watson-Glaser, the highest-value hour you can spend is on the five decision rules it scores you against, section by section, rather than on generic reasoning practice.
The ranking is wrong for you, and that is the point
The hardest test on this list is the one that lands on your specific weakness. If you freeze under a racing clock, the PI is your worst outcome and my number-one ranking is irrelevant to you. If you are precise but read too fast, the SHL will take your marks with a smile. If you are used to being the person in the room with the strongest opinion, the Watson-Glaser is built to make that a liability.
So do not memorise a difficulty league table. Take one timed run of whichever test is actually in your inbox, see which of the three axes broke you, and fix that one.
Common questions
Is 33 a good CCAT score? Yes, comfortably. The average sits around 24 out of 50, so a 33 is well above the middle and in the range competitive technical and senior roles target. Employers set their own cutoff per role rather than a universal pass mark, so a 33 opens most doors and a few of the most competitive will still want more.
Is the CCAT like an IQ test, or equivalent to an IQ? It overlaps with what an IQ test measures, since both tap general cognitive ability, but it is not equivalent to one and does not convert to an IQ number. The CCAT is normed against job applicants rather than the general population and is built to predict how quickly you pick up a new role. The practical difference is speed, and speed is trainable in a way raw ability is not.
Is there a cheat sheet for the CCAT? Not in the sense people mean. There is no leaked question set that is both real and current. What genuinely functions like one is format familiarity: if the matrix and number-series layouts are automatic on test day, you save the seconds everyone else spends decoding, and those seconds are the score.
How difficult is the CCAT? Individually the questions are not hard. Handed any one with unlimited time, you would very likely get it right. The difficulty is 18 seconds each, a rising difficulty curve, and the fact that almost nobody finishes.
Do most people finish the CCAT? No. Only about 1 in 100 candidates answers all 50 questions. Knowing that in advance is worth real points, because the moment you realise you will not finish is the moment most scores collapse.
Which of these is easiest? Of the four I sat, the Wonderlic. Same clock as the PI, gentler content, no ramp, and the most responsive to practice.
Why I bothered ranking these at all
I did not sit four cognitive aptitude tests for fun. I sat them because I build the practice platform people use to prepare for them, and I did not want to be another founder writing about an experience he had only read about.
That mattered more than I expected. Sitting the CCAT taught me the useful drill is not “more questions,” it is timed full-length runs where you practise abandoning a question at the 20-second mark. Sitting the SHL taught me that reviewing why a distractor was tempting beats grinding another fifty items. Neither insight is available from a spec sheet.
That is the thinking behind PrepClubs, where the CCAT, Wonderlic, PI, SHL, and Watson-Glaser banks each start with a free full-length diagnostic. I am biased about the platform. I am not biased about the order of operations: take one free timed run first, find out which axis broke you, and only then decide whether the paid drilling is worth your money.
If your invite has landed and you know which test you are facing, start with the lived account of that one: the CCAT, the PI Cognitive Assessment, the Wonderlic, or SHL numerical reasoning.
