Likert scales: 5 or 7 points and what to do with the middle
A Likert item is a statement plus a symmetric, labelled agreement scale. The three decisions that matter are how many points it has, whether the middle exists, and whether every point is labelled — and each one changes how you are allowed to analyze the results.
What it is, precisely
A Likert item presents a statement — "The checkout process was straightforward" — and asks how strongly the respondent agrees, on a scale with the same number of positive and negative options around a middle. Strictly speaking, a Likert scale is several such items about one underlying attitude, summed into a single score; the individual row is a Likert item. Most people use "Likert scale" for both, which is fine in conversation and matters in analysis.
It matters because a single item is a noisy, ordinal measurement of an attitude, while a set of items about the same thing averages out some of that noise. If you intend to combine items into a score, they need to be about the same construct. Averaging "the checkout was easy" with "delivery was on time" produces a number that describes nothing.
One more distinction worth keeping straight: a Likert item asks for agreement with a statement, while a rating scale asks for a judgement on a dimension. Where you can, prefer the dimension — "How easy or difficult was checkout?" rather than "Do you agree checkout was easy?" Agreement wording invites acquiescence, the well-documented tendency for people to agree by default when they are unsure or in a hurry.
The five decisions
How many points
Five and seven are the common choices. More points only help if respondents can genuinely discriminate at that resolution — a customer rating their broadband is not resolving seven levels of agreement, while a trained assessor rating performance might. On a phone, seven labelled options either wrap awkwardly or shrink to unreadable.
Whether the middle exists
An odd number of points gives a defined middle; an even number forces respondents to lean one way. Removing the middle does not remove ambivalence — it records it as a mild opinion, which moves the noise into your data instead of labelling it.
Which points get labels
Fully labelled scales make each point mean roughly the same thing to everyone. Endpoint-only labels are faster to read and work well with numbers, but the middle points then mean whatever the respondent decides they mean.
Which direction it runs
Negative on the left and positive on the right, or the reverse — either is defensible. What is not defensible is changing direction partway through a survey, which reliably produces wrong answers from people who had learned the pattern.
Where "does not apply" lives
If an item might not apply to some respondents, give them an explicit option outside the scale. A "not applicable" hidden inside the neutral midpoint is indistinguishable from genuine ambivalence once it is in your results, and the two mean opposite things.
The defaults we would pick
None of these is a law. They are the choices that go wrong least often, and consistency matters more than any single one of them.
| Decision | The trade-off | A sensible default |
|---|---|---|
| 5 points or 7 | 7 gives finer resolution to respondents who can use it; 5 is easier to read and fits a phone | 5 for general audiences, 7 for expert or repeat raters |
| Neutral midpoint | Keeping it invites satisficers to park; removing it forces ambivalent people to pick a side at random | Keep it, and label it "Neither agree nor disagree" |
| Label every point | Full labels standardise meaning; endpoint labels are cleaner and read faster | Label every point on agreement scales, endpoints only on 0–10 numeric scales |
| Scale direction | Either direction works; only inconsistency does damage | Negative on the left, unchanged for the whole survey |
| Not applicable | On-scale is tidier; off-scale keeps the data interpretable | A separate option outside the scale, placed last |
| Changing scale between waves | A better scale next quarter breaks the comparison with last quarter | Keep the old scale unless it was actively broken |
The midpoint argument, settled as far as it can be
The case against the midpoint is that it is a parking spot. Someone who has not thought about the question can select the middle and move on, so removing it forces a real judgement and increases the spread of your data.
The case for it is stronger: ambivalence and mild opinion are genuinely different states, and a forced choice makes them look the same. If you take the middle away, the person with no real view still has to answer, and they answer somewhere — you have not gained information, you have added noise you can no longer identify.
The practical resolution is to keep the midpoint and be precise about what it says. "Neither agree nor disagree" describes a balanced view. "Neutral" is vaguer and gets used as a shrug. "No opinion" and "Don't know" are different again, and both belong off the scale next to "not applicable" if your item might not apply. Distinguishing balanced view, no view, and no basis for a view costs you one extra option and saves you from reporting three different things as one.
Analysing it without kidding yourself
Likert data is ordinal. You know that "strongly agree" is more than "agree," but not that the gap between them is the same size as the gap between "disagree" and "strongly disagree." Taking a mean assumes those gaps are equal, which is an assumption, not a fact.
In practice most teams report means anyway, and for tracking the same question over time that is defensible — the assumption is wrong in the same way every quarter, so the direction of change is still informative. What is not defensible is showing the mean on its own. A mean of 3 out of 5 can be everyone choosing the middle or half the sample at each extreme, and those two situations call for completely different responses.
The reporting summary that survives scrutiny is top-two-box: the percentage of respondents choosing the two most positive points. It does not assume equal spacing, it is easy for a non-specialist to interpret, and it is stable enough to track. Pick one summary, show the full distribution beside it, and use the same summary every time.
Sample size deserves the same caution here as anywhere else. A five-point item from 30 respondents has enough sampling variation that a shift of a tenth of a point means nothing. Report the response count next to the score, always.
Building one in Zunoform
For a single item with numbers, use a rating question. It runs from 5 to 10 points, optionally starting at zero, with labels on the two ends — which makes it a good fit for numeric scales and a poor fit if you need each point named.
When every point needs a word, use a choice question with the five or seven labels as options. You lose the compact scale look and gain unambiguous meaning, which is usually the better trade on an agreement scale.
When several statements share one scale, use a matrix question: each row is a statement, and the columns carry the scale labels, which you set yourself. Keep matrices short. A grid of fifteen statements is where people abandon a survey on a phone — split by topic and put the most important block first.
All three question types are on the free plan, along with 500 responses a month and conditional logic, so an entire attitude survey can be built and run without paying anything.
Questions people ask
Should a Likert scale have a neutral option?
Usually yes. Removing it does not remove ambivalence, it just relabels ambivalent respondents as mildly positive or negative and makes that noise invisible. Keep the midpoint, label it "Neither agree nor disagree," and add a separate "not applicable" option if the item might not apply.
Is a 5-point or 7-point Likert scale better?
For general audiences on phones, 5. For expert or trained raters who can genuinely distinguish finer degrees, 7. If you are continuing an existing survey, keep whatever the previous wave used — a changed scale breaks your comparison more than a slightly suboptimal scale hurts your data.
Can I calculate an average from Likert responses?
You can, and most teams do, but it assumes the gaps between points are equal, which is not guaranteed. It is reasonable for tracking the same question over time, provided you show the distribution alongside the mean. Top-two-box percentages avoid the assumption entirely.
What's the difference between a Likert scale and a rating scale?
A Likert item asks how strongly you agree with a statement. A rating scale asks you to judge something directly on a dimension, like easy to difficult. Where both would work, the rating version is usually better, because agreement questions invite people to agree by default.
How many statements should a matrix question have?
Enough to fit on one phone screen without side-scrolling — roughly five to eight rows. Longer grids are a common abandonment point, and splitting them by topic costs you nothing but a page break.
Keep reading
Better forms.
Better data.
Build your first form in under 60 seconds. Free forever for personal use, no credit card required.