Automation has quietly become the default answer to every fieldwork constraint. Screening is faster because a model scores it. Fraud detection is better because a model flags it. Open ends are coded overnight because a model read them. Most of these claims are true in the narrow sense and misleading in the way that matters, which is what happens when the model is wrong.
The useful questions are not about whether a supplier uses AI. They are about what happens at the boundary between the model and a person.
What the model actually decides
There is a large difference between a model that ranks responses for human review and a model that removes them. The first is a productivity tool. The second is a sampling decision, and it needs to be documented as one.
Ask which decisions are made without a human in the loop, and ask for the removal rate. If a supplier cannot tell you what share of completes their automated checks reject, they are not monitoring the thing that most affects your data.
- Which decisions are fully automated, and which are model-assisted?
- What is the automated rejection rate, and how has it moved over the last year?
- What happens to a respondent the model rejects — are they told, and can they appeal?
- Is the same model applied to every market, and was it validated in each one?
Open-end coding is where it shows first
Automated coding is now good enough that the output looks plausible regardless of whether it is right. That is precisely the failure mode to worry about: a code frame that has been applied consistently and wrongly produces clean tables and a false finding.
The mitigation is boring and effective. A human codes a sample of the same verbatims independently, the two are compared, and the agreement rate is reported to the client rather than kept internally.
Validation per market, not in aggregate
A fraud model validated on English-language responses from three markets will behave differently on a fourth in another language. It usually behaves worse, and the degradation shows up as a higher rejection rate among respondents whose writing is being scored by a model that was not trained on how they write.
That is not a hypothetical fairness concern. It skews your sample towards the respondents the model recognises, which in multi-country work is a systematic bias with a plausible-looking dataset on the other side of it.
The nine questions
- Which fieldwork decisions are made by a model without human review?
- What is the automated rejection rate, by market?
- How was the model validated in each language you field in?
- What is the human-model agreement rate on open-end coding?
- Who reviews rejections, and how often are they overturned?
- Is model output retained so a finding can be re-examined later?
- What is the fallback when the model is unavailable mid-field?
- How is respondent data used in training, and was consent obtained for that?
- Can you supply the documentation a client audit would need?
None of this is an argument against automation in fieldwork. It is an argument for being able to say what the automation did, which is the same standard we have always applied to an interviewer.