Synthetic respondents will fail Nigeria before they fail anywhere else
Pew asked a language model to stand in for American survey respondents and missed a real result by 21 points. The structural reason that happened is far worse here than it is there.
This is argued opinion from the NigeriaPolls Research Desk. It is not a poll finding. Our measured readings carry a sample size, field dates, mode and a margin of error, and are published at polls and indices. We argue about policy, markets and method, and we do not argue about who should win an election.
Somebody is going to offer you Nigerian public opinion without interviewing any Nigerians. The pitch will be good. It will be fast, it will be cheap, it will arrive as a clean deck on a Friday, and it will carry percentages to one decimal place. Before anyone buys that, it is worth reading what happened when a serious research organisation tested the idea properly on the easiest possible population.
What Pew actually did
Pew Research Center assigned demographic personas to a large language model and asked it to answer questions drawn from three of its own American Trends Panel surveys, then compared the model's answers against what real people had said. The headline is in their title: not really.
The detail is more useful than the headline. On whether it is acceptable for immigration officers to wear face coverings, 38 per cent of American adults said yes and the synthetic respondents came back at 17. That is a 21 point miss on a question with a clean, recently measured answer. On presidential job approval the model went the other way, reading 46 against a real 34. A survey that was wrong by 12 points in one direction on one question and 21 points in the other direction on another is not a survey with a margin of error. It is a guess with a confident voice.
Two of their other findings matter more than the misses themselves. Synthetic respondents scored far higher than humans on factual knowledge questions, which means the model does not simulate a public, it simulates a well informed version of a public. And the answers moved when the underlying model changed, with everything else about the exercise held constant. A measurement instrument whose readings depend on which vendor you bought it from this quarter is not a measurement instrument.
Why the Nigerian case is worse, structurally
Here is the part that nobody selling this will raise with a Nigerian buyer.
The United States is the single best represented country in the text these models are trained on. American opinion surveys, American news, American forum arguments, decades of American polling writeups. If a model has a fair shot at any population on earth, it is that one. It still missed by 21 points.
Nigeria is represented in that same training text in a way that is not just thinner but differently shaped. The Nigerians who write at volume in English on the open internet are younger, more urban, better off and more online than Nigerians as a whole. They are, almost exactly, the slice that a properly drawn Nigerian sample spends real money and effort trying not to over-represent. So a synthetic Nigerian respondent is not a blurry copy of the Nigerian public. It is a sharp copy of Nigerian Twitter.
The error that produces does not look like noise. Noise is survivable, because it widens honestly and you can report it. This error has a direction. It will systematically read the country as more connected, more literate in English, more urban and more comfortable than it is, and it will do so while reporting no uncertainty at all, because the model has no sampling frame to be uncertain about.
Ask what that does to a decision. A consumer goods company sizing demand outside Lagos. A funder allocating between states on perceived need. An agency deciding where a health message has landed. Every one of those decisions is made worse in the same direction by the same bias, and nothing in the output tells the buyer it happened.
What this costs to do properly, and why we do it anyway
The honest comparison is not synthetic against perfect. It is synthetic against the real thing, with the real thing's costs stated.
Our own method is slower and it is published so it can be argued with. We field by live telephone interview, SMS, WhatsApp and in-person interviewing across all 36 states and the Federal Capital Territory, with Nigerian interviewers based in the states where the work happens. Outbound calling is screened against the NCC do-not-disturb register. Back-checks re-contact a share of completed interviews. Every published result carries its field dates, total sample size, mode breakdown, weighting variables and a margin of error at 95 per cent confidence.
That last sentence is the whole argument. Each of those disclosures is a place where we can be caught being wrong. A synthetic sample offers none of them, not because its vendors are hiding something, but because none of them exist to disclose. There is no field date. There is no mode. There is no margin of error, because there was no sample.
Where the models are genuinely useful
None of this makes the technology useless to a research desk, and pretending otherwise would be its own kind of dishonesty.
Language models are good at drafting a questionnaire before a human cuts it, at flagging a leading question, at coding open-ended responses that a human then audits, at translating an instrument into Hausa, Yoruba or Igbo for a human to check, and at finding the three transcripts worth reading out of four hundred. In each of those the model is working on data that came from people, under review by someone accountable for the result.
The line is simple enough to hold. Use the model on the data. Do not use the model instead of the data.
What to ask when the deck arrives
Four questions, and they are not hostile ones. Any supplier doing real work will have the answers ready.
When was this fielded, and over what dates. How many people were interviewed, and how were they reached. What is the margin of error at 95 per cent confidence. And if the answer to any of those is that no one was interviewed, what is the evidence that this model reproduces Nigerian opinion specifically, measured against a real Nigerian sample.
That last question is the one that matters, and at the moment nobody can answer it, because the benchmark does not exist. Pew had an American benchmark to test against and the test failed. For Nigeria there is not yet a public benchmark to fail against, which is not the same thing as passing.
We would rather be the slow, expensive, checkable option than the fast one that cannot be checked. That is a commercial position as well as a methodological one, and we would rather state it plainly than have it discovered.
What this is based on
Respond to this argument at [email protected]. We correct errors of fact and we say so on the page when we do.
