Does Question Count Predict DILR Set Difficulty Explained

During a sweep, candidates look for anything that separates the sets quickly, and the number of questions attached to each is right there on the screen. The theory that a larger cluster signals a harder set, or an easier one, is appealing because it costs nothing to apply.
It is also unreliable, and the reason is worth understanding rather than just accepting: question count tells you about the value of a set, not about its cost, and those are separate things. Once you separate them, the count becomes genuinely useful for a different purpose. This piece explains which purpose.
Cost estimation comes from volume within a type. Data tables practice is one place it is built.
- Question count measures value, not difficulty, and the two are independent.
- Nothing is published about how counts relate to difficulty, so claims are inferences.
- Count does matter for return per minute, which is a real selection input.
- A large cluster raises both the upside and the downside of a wrong structure.
- Judge cost by whether a first placement is forced, then use count to rank.
Count Measures Value, Not Cost
Separate the two and most of the confusion goes.
The number of questions tells you how many marks are available once you have built the structure. That is the value of the set, and it is known before you start.
The cost is how long the structure takes to build, which depends on whether constraints force placements early, how much branching there is, and whether the questions need one structure or several. None of that is visible in the question count.
Two sets with the same count can differ by eight minutes in cost, and two sets with very different counts can cost the same. The variables are independent, which is why the count on its own predicts nothing about difficulty.
Using count as a difficulty proxy because it is visible instantly. It is the same error as using length: the properties that decide cost take slightly longer to read, and acting on the fast visible signal is worse than spending ten more seconds on the right one.
Nothing Is Published About This
Worth stating, because confident claims circulate in both directions.
CAT publishes a general syllabus. It does not publish how sets are constructed, how question counts are assigned, or any relationship between cluster size and difficulty. There is no rule being followed that anyone could reverse engineer.
What circulates is recollection after the exam, and that is a weak instrument. Candidates remember the sets that went badly, the sample across a few papers is small, and DILR is more variable between years than Quant, which makes pattern-spotting from a handful of papers particularly unreliable.
Every visible-at-a-glance signal candidates use for selection has this property: it is available instantly and it is uncorrelated with what matters. Length, tidiness, position, question count. The useful signals all cost a few more seconds, which is exactly why they get skipped.
Where Count Genuinely Matters
Having separated value from cost, the count becomes useful for ranking rather than for predicting.
| Situation | How count matters | What to do |
|---|---|---|
| Two sets you assess as equally cheap | Higher count is more marks for the same work | Take the larger cluster first |
| A cheap set with few questions | Low return per minute of building | Still worth it, but later |
| An expensive set with many questions | High upside, high risk | Only with time and confidence |
| An expensive set with few questions | Poor on both axes | Leave it |
| Late in the section | Building time is scarce | Favour smaller clusters you can finish |
That is the real use: the count ranks sets you have already assessed for cost, rather than replacing the cost assessment. A candidate who estimates cost first and then uses count to break ties is using both variables correctly.
Return Per Minute Is the Right Frame
Thinking in return per minute rather than in difficulty makes the selection decision cleaner.
A set costing six minutes with four questions returns more per minute than one costing six minutes with three, assuming you can answer them. That is a genuine argument for the larger cluster when the cost is comparable.
The qualifier matters. Larger clusters frequently include at least one question requiring full resolution, while the smaller questions in a set may be answerable from a partial structure. So the realistic return is what you can actually answer rather than the cluster size.
Read all the questions in a set during the sweep, not just the count. The count tells you how many; the questions tell you how much structure each needs, and that is the difference between a four-question set worth entering and one worth avoiding.
A Larger Cluster Raises Both Tails
One property of big clusters deserves specific attention because it is asymmetric in a dangerous way.
Questions come off one structure. If the structure is right, the whole cluster lands and a large set is a very good outcome. If the structure is wrong, every answer derived from it is wrong together, and on MCQs that is several penalties at once rather than one.
Under the marking scheme, a correct answer is plus three and an incorrect MCQ is minus one, so a five-question cluster answered off a bad structure is a large negative swing. That is a reason to be more confident before committing to a big cluster, not less.
What Actually Predicts Cost
- Is a first placement forced? A condition, or a pair, that pins something down with no choice. Sets offering this tend to cascade.
- How much branching? If every condition still permits several arrangements, the bookkeeping will cost more than the reasoning.
- Do the questions want one structure or several? Different quantities off the same base means rebuilding.
- Is the information gathered or scattered? Conditions spread across preamble, table and questions have an assembly cost before any solving.
Those four take about twenty seconds per set and they predict cost far better than any visible-at-a-glance property. The count then ranks the sets those twenty seconds identified as affordable.
Calibrate It On Your Own Data
You can settle the question empirically rather than arguing about it, and it takes a couple of weeks.
For every practised set, record the question count, your predicted cost in minutes, and the actual cost. Then check whether count correlates with anything in your own numbers. Most candidates find it does not, which is more convincing than being told.
The same log answers the more useful question, which is whether your cost predictions are any good. That is the estimate selection actually runs on, and it only improves if predictions are checked against outcomes.
Late in the Section, Count Changes Sides
There is one point where the count becomes a stronger input than usual, and it is worth planning for because it arrives every time.
Early in a section, building time is plentiful relative to what remains, so an expensive set with a large cluster can be worth entering. Late in the section the calculation inverts: you may have time to build one more structure, and a set you cannot finish returns nothing regardless of how many questions were attached to it.
So in the last ten minutes, prefer sets you are confident of completing over sets with the highest potential return. A three-question set you will certainly finish beats a five-question set you probably will not, because the five-question set's value is contingent on a build you may not complete.
That also affects which questions to answer within a set you have entered. With time short, take the cheap forms first, the ones answerable from a partial structure, rather than working towards the question that needs full resolution. Banking two answers from an incomplete grid is a better use of six minutes than building towards a fifth question you will not reach.
DILR Is More Variable Than Quant
One reason to be cautious about any pattern claim in this section.
Set types and counts swing between years more than Quant's areas do, so the frame of a stable relationship between structure and difficulty applies less well here. What transfers across papers is recognition of the recurring families and the judgement to assess a set quickly, not a rule about how sets are built.
That also means a bad DILR section in one mock carries less information than a bad Quant section. Judge your DILR on a trend across several attempts rather than on any single paper, and judge selection heuristics the same way.
The Wider Lesson About Fast Signals
This question belongs to a family, and noticing the family saves you from the next three versions of it.
Candidates reach for signals that are visible in a second because a sweep is time-constrained and anything requiring real reading feels unaffordable. Length, position, tidiness, question count, how familiar the scenario sounds: each is instant, and each is close to uncorrelated with what a set will cost.
The properties that do predict cost all require a few seconds of actual reading. Whether a placement is forced cannot be seen without looking at a condition. Whether the questions want one structure or several cannot be known without reading them. That is an uncomfortable trade during a sweep and it is the right one, because a sweep that reads nothing is not selection, it is sorting by appearance.
The practical resolution is to budget for it explicitly. Under a minute per set, spent on the four cost signals rather than on absorbing the scenario, is affordable across a section and it produces a ranking based on the variables that matter. Candidates who feel they cannot spare that minute are usually spending several on a set the minute would have told them to avoid.
The Summary
Question count measures the value of a set rather than its cost, and the two are independent. Two sets with the same count can differ by eight minutes in build time, which is why the count on its own predicts nothing about difficulty.
Nothing is published about how counts relate to difficulty, and DILR's year-to-year variability makes pattern claims from a few papers particularly unreliable. The count joins length, tidiness and position as a signal that is visible instantly and uncorrelated with what matters.
Assess cost first, using whether a first placement is forced, how much branching there is, whether the questions want one structure or several, and whether the information is gathered or scattered. Then use the count to rank the affordable sets by return per minute, remembering that a large cluster raises both the upside and the cost of a wrong structure.
- Do you assess cost before looking at the question count?
- Do you read all the questions during the sweep, not just count them?
- Have you checked whether count correlates with cost in your own logs?
- Are you more careful about the structure before a large cluster?
If your selection runs on signals visible in a second, it is running on the ones uncorrelated with cost. A CAT preparation strategy review will show what that is costing across your mocks, and a personalised CAT preparation plan builds the estimation drill into your practice.
Count Is Value. Cost Is Separate.
Assess what a set will take, then let the count rank the ones you can afford.
Build My Weekly PlanFrequently Asked Questions About Set Question Counts
Does the number of questions in a DILR set predict its difficulty?
No. Count measures how many marks are available once the structure is built, while cost depends on whether constraints force placements early and how much branching there is. The two variables are independent.
Is question count useless for selection then?
It is useful for ranking rather than predicting. Between two sets you have assessed as equally cheap, the larger cluster returns more marks for the same building work, so the count breaks ties after the cost assessment.
Are larger clusters riskier?
They raise both tails. Questions come off one structure, so a right structure lands the whole cluster while a wrong one produces several penalties at once under minus one for each incorrect MCQ. That argues for more confidence before committing.
What should I judge cost by?
Whether a first placement is forced, how much branching the conditions permit, whether the questions want one structure or several, and whether the information is gathered or scattered. Those take about twenty seconds and predict cost far better.
Practice DILR sets, chapter by chapter
CAT DILR practice across LR puzzles and DI sets, each with a worked solution.
More from DILR
Continue reading

Why CAT RC Answer Options Often Feel Equally Correct

How To Study For CAT When You Are Mentally Exhausted

How To Stop Procrastinating On Your Weakest Section

How To Recover From A Wrong Start On A DILR Set Explained
Put it into practice