Likert Scales: What They Are, When They Break, and How to Write One
A Likert scale measures agreement on an ordered set of options. The scale is easy; almost everything that goes wrong happens before anyone answers. A practical guide to writing, administering and analysing them.
In short
A Likert scale measures how strongly someone agrees with a statement, usually on five or seven points running from strong disagreement to strong agreement. It is named after Rensis Likert, who introduced it in 1932, and it is the most widely used attitude measure in survey research.
The scale itself is easy. Almost everything that goes wrong with it happens before anyone answers: a statement that contains two ideas, a midpoint that collects people who mean three different things, labels that are not evenly spaced, or an analysis that averages responses the scale never licensed you to average.
This guide covers what a Likert scale is, how to write one that holds up, how many points to use, what you may legitimately do with the results, and the specific mistakes that turn a clean instrument into an unusable one.
What a Likert scale actually is
A Likert item is a statement, not a question. The respondent is shown something like "The government is handling the cost of living well" and asked how strongly they agree, choosing from a set of ordered options.
The classic five-point form runs:
| Point | Label |
|---|---|
| 1 | Strongly disagree |
| 2 | Disagree |
| 3 | Neither agree nor disagree |
| 4 | Agree |
| 5 | Strongly agree |
That ordering is the whole idea. Each option sits further along a single dimension than the one before it, which is what separates a Likert scale from a list of unordered choices.
Likert item or Likert scale?
These get used interchangeably and they are not the same thing.
A Likert item is one statement with one set of agree-disagree options.
A Likert scale, in the original and stricter sense, is several related items combined into one score. Likert's 1932 proposal was that you ask about an attitude from several angles and sum the responses, because any single statement carries wording quirks that a combined score averages out.
The distinction matters for analysis, and we come back to it below, because it decides whether taking a mean is defensible.
How many points should a Likert scale have?
The honest answer is that it depends on what you are measuring and how, and that the differences between five and seven are smaller than most arguments about them suggest.
Five points is the default for good reason. It is easy to read aloud, it works on a phone call, and respondents hold it in their heads without strain.
Seven points gives more room for people to discriminate and tends to produce slightly better reliability for attitudes people hold strongly. It costs you something in administration: seven labels read out over the phone is a lot to retain, and respondents start collapsing them anyway.
Four or six points, with no midpoint, force a direction. That is a legitimate choice when you believe the midpoint is being used as an escape hatch rather than a genuine position, and a bad choice when some people genuinely have no view, because you have now made them pick one.
Ten or eleven points are common in satisfaction and recommendation measures, where the scale is closer to a rating than an agreement. Net Promoter Score uses eleven, zero to ten.
A practical rule: use five if the survey is spoken, seven if it is read, and reach for more points only when you have a specific reason you can articulate.
The midpoint problem
"Neither agree nor disagree" is the most misunderstood option on any survey. It collects at least three different kinds of person: those who genuinely sit in the middle, those who have no opinion, and those who do not want to say what they think.
Those are not the same thing and they should not share a box. If the distinction matters to your decision, add an explicit "Don't know" or "Prefer not to say" option alongside the midpoint rather than forcing all three groups into one.
In politically sensitive research this is not a technicality. A midpoint inflated by people avoiding the question looks exactly like genuine moderation, and the difference changes what the number means.
How to write a Likert item that works
Most Likert problems are writing problems.
One idea per statement
The commonest fault is the double-barrelled item. "The government is handling inflation and insecurity well" cannot be answered by someone who thinks one is going well and the other badly. Split it.
Avoid negatives, and never use two
"The government is not failing to address inflation" is a sentence people have to unpick before they can answer. Reverse-worded items are sometimes used deliberately to catch inattentive respondents, but a double negative is just a trap for everyone.
Keep the labels evenly spaced
The points should feel like equal steps. "Excellent, very good, good, fair, poor" is not even: three of those five are positive, so results drift upward before anyone has formed a view.
Match the scale to the statement
If a statement is about frequency, agree-disagree is the wrong scale and you want never through always. If it is about importance, use an importance scale. Forcing everything into agree-disagree produces awkward items like "I agree that I often buy this brand".
Watch acquiescence
People agree more than they disagree, across cultures and especially where an interviewer is present and politeness matters. You cannot eliminate it, but you can reduce it by balancing the direction of statements across a block so agreement is not always the positive answer.
Write for the mode and for the language
A scale that works on paper may not survive being read aloud. And a respondent answering in their second language gives shorter, flatter, more agreeable answers. In Nigerian fieldwork that is a live issue: an item that is crisp in English can become muddy in Pidgin, Hausa, Yoruba or Igbo unless it is adapted rather than translated word for word. Pilot it in every language it will be administered in.
What you may legitimately do with the results
This is where the arguments happen, and where a lot of published analysis quietly overreaches.
The ordinal problem
Likert responses are ordinal. You know "strongly agree" is further along than "agree", but you do not know that the gap between them equals the gap between "agree" and "neutral". The numbers 1 to 5 are labels for ordered categories, not measured quantities.
Strictly, that means the mean is not defined, because averaging assumes equal intervals the scale never established.
What to do instead
For a single item, report the distribution. The percentage who agree or strongly agree, the percentage who disagree, the percentage in the middle. This is almost always more informative than a mean anyway: a mean of 3.0 could be everyone sitting on the fence or the country split violently in half, and those are different findings.
The median and the mode are both defensible for ordinal data.
Top-two box scoring, combining "agree" and "strongly agree", is the standard reporting convention in commercial research and it communicates well.
When a mean is defensible
Where a Likert scale proper is used, several items combined into one score, treating the summed score as interval data is widely accepted and well supported in the methodological literature. The combination of multiple items smooths the unevenness of any single one.
So the practical rule: report distributions for single items, and means only for multi-item scales you have constructed deliberately and whose internal consistency you have checked.
Significance testing
For ordinal single items, use tests that do not assume intervals: Mann-Whitney for two groups, Kruskal-Wallis for more, chi-square for the full distribution.
The mistakes that cost you a study
- A double-barrelled statement. Nobody can answer it and you cannot interpret what they did answer.
- Treating the midpoint as moderation. It may be avoidance, and in sensitive research it usually partly is.
- Averaging single items and reporting one decimal place. The precision implied is not in the data.
- Changing the scale mid-tracker. Moving from five points to seven breaks your trend line, and the break will look like a real change in opinion.
- Too many items in one block. Twenty agree-disagree statements in a row produces straight-lining, where respondents pick the same option down the list. Break the block up, and flag straight-liners in quality checks.
- Not piloting in every administration language. A translated item is a different item until you have tested it.
A worked example
Suppose you want to measure confidence in the electoral process ahead of an election.
A weak single item: "The elections will be free, fair and credible." Three concepts, one answer.
A better multi-item scale, each answered on the same five-point agreement scale:
- My vote will be counted accurately.
- The results announced will reflect how people voted.
- I can vote without fear of intimidation.
- The body running the election will act impartially.
Four items, each on one idea, measuring the same underlying attitude from different angles. Report each distribution, and build a combined confidence score from the four where you want one number to track over time.
That is a Likert scale in the sense Likert meant, rather than a single item wearing the name.
Related reading
- Writing survey questions that work
- How many people you need to survey
- How we design and run studies
- The standards we publish against
Frequently asked questions
What is a Likert scale?
A Likert scale measures agreement with a statement on an ordered set of options, usually five or seven points running from strong disagreement to strong agreement. It was introduced by Rensis Likert in 1932 and is the most widely used attitude measure in survey research. Strictly, a Likert scale is several related items combined into a single score, while one statement with agree-disagree options is a Likert item.
How many points should a Likert scale have?
Five is the sensible default and works well when a survey is read aloud or administered by phone. Seven gives respondents more room to discriminate and often improves reliability slightly, at the cost of being harder to hold in mind during a spoken interview. Four or six points remove the midpoint and force a direction, which suits some research questions and distorts others. Use more points only when you can articulate why.
Should a Likert scale have a neutral midpoint?
It depends on whether a genuine middle position exists for your statement. The risk is that the midpoint collects three different groups: people who truly sit in the middle, people with no opinion, and people unwilling to say. If that distinction matters, keep the midpoint and add a separate "Don't know" or "Prefer not to say" option rather than forcing everyone into one box.
Can you calculate the mean of a Likert scale?
For a single Likert item, strictly no, because the data is ordinal and the gaps between points are not known to be equal. Report the distribution, the median, or a top-two-box percentage instead. For a Likert scale proper, where several related items are combined into one score, treating the summed score as interval data and taking a mean is widely accepted and well supported in the literature.
What is the difference between a Likert item and a Likert scale?
A Likert item is one statement with one set of ordered agree-disagree responses. A Likert scale is several related items combined into a single score, which is what Likert originally proposed. The distinction decides what analysis is defensible: distributions for single items, means for properly constructed multi-item scales.
What is a double-barrelled question and why does it matter?
A double-barrelled item asks about two things at once, such as "The government is handling inflation and insecurity well". A respondent who thinks one is going well and the other badly has no honest answer available, and whatever they choose cannot be interpreted. It is the single most common fault in Likert items and the easiest to fix: split it into two statements.
What is straight-lining and how do you catch it?
Straight-lining is when a respondent selects the same option down a long block of items without reading them. It is a quality problem rather than an opinion. Shortening blocks reduces it, varying the direction of statements helps, and response-pattern checks should flag it during data processing so affected interviews can be reviewed.
Do Likert scales work in multilingual surveys?
They work, but only if the scale and the statements are adapted into each language rather than translated literally. Agreement labels in particular do not map evenly between languages, and a respondent answering in a second language tends to give shorter and more agreeable answers. Any multilingual instrument should be piloted in every language it will be administered in before full fieldwork.
Tags
Cite this article (CC BY 4.0)
NigeriaPolls Research Desk. (1 October 2026). "Likert Scales: What They Are, When They Break, and How to Write One." NigeriaPolls. CC BY 4.0. https://nigeriapolls.com/blog/likert-scale-guide
Free to share, remix, and republish with attribution. See terms.
Live data behind this story
Keep reading
Running Survey Research Under the Nigeria Data Protection Act
Survey research sits squarely inside the Nigeria Data Protection Act 2023. The six obligations that matter to fieldwork, why political opinion counts as sensitive data, and the two traps specific to Nigeria.
Writing Survey Questions That Survive Contact With a Respondent
Most bad survey data comes from questions that were clear to the person who wrote them and ambiguous to everyone else. A practical guide to questionnaire design, wording, ordering, length and piloting.
How Many People Do You Need to Survey? Sample Size, Worked Through
About 1,000 completed interviews gives a national result at roughly plus or minus 3 points, and population size barely matters. What actually drives sample size is the breakdowns you intend to report separately.
