AI for Market Research: A Practical Guide

AI for Market Research: A Practical Guide

2026-08-05 · Tommaso Maria Ricci

The insights industry crossed 153 billion dollars in 2024, and the fastest growing part of it was not research at all. According to ESOMAR data published in Research World, the research software sector reached 62 billion dollars and grew 11.5%, while traditional market research services reached 56 billion and grew 4.8%. Software is now the larger half of the industry, and it is pulling away. That single comparison tells you more about where AI for market research is going than any vendor deck will.

Here is what it means in practice. The part of research that involves fielding, coding, cleaning, tabulating and summarizing is being absorbed into tooling. The part that involves deciding what question is worth asking, and what a company should do about the answer, is not. Most research teams are staffed as if the first part were still the job.

I run companies, I do not sell research panels, and my interest in this topic is narrow: I want to know whether a business decision is being made on evidence or on the loudest opinion in the room. Over the last few years I have watched AI move research from something a company commissioned twice a year to something it can run continuously. That change is real, and it is also badly oversold. This guide covers what actually works, what does not, what it costs, how long it takes to pay back, and how to build the capability without buying a platform you will abandon in nine months.

Why market research is unusually well suited to AI

Four properties make research a high-return, moderate-difficulty domain.

The raw material is mostly unstructured text. Open-ended survey responses, interview transcripts, support tickets, reviews, sales call recordings, community threads. This is precisely the material that language models handle far better than the previous generation of tools, and precisely the material that most companies collect and never read.

The bottleneck was always analysis capacity, not data collection. Ask any insights team what happens to the open-ended questions on a large tracker. In most organizations they get skimmed, a handful of quotes get lifted for the deck, and the rest is never touched. The constraint was human reading time, and that constraint just moved.

Output quality is checkable. You can hold back a portion of the data, code it manually, and compare. Few business functions allow this kind of clean validation, and it removes most of the political argument that stalls AI projects elsewhere.

The work repeats with a stable structure. Screener, questionnaire, fielding, coding, weighting, analysis, report. Repetition with known structure is where automation compounds.

There is also a market context worth naming. The number of decisions companies want evidence for is rising much faster than research budgets. AI does not solve that by making research cheaper in a straight line. It solves it by making a class of questions worth asking that previously were not worth a study.

The six jobs a research function actually does

Before talking about tools, be precise about which decision you are automating, because they differ enormously in difficulty.

  1. Framing. Turning a vague business question into a researchable one.
  2. Designing. Building instruments that do not bias the answer.
  3. Collecting. Recruiting the right people and getting real responses from them.
  4. Processing. Coding, cleaning, weighting, quality control.
  5. Analyzing. Finding the pattern that matters and separating it from noise.
  6. Landing. Getting the organization to act on the finding.

AI touches all six. It is genuinely transformative on three, useful on two, and close to useless on one. Knowing which is which is the whole game.

The seven applications of AI in market research

Ordered by value-to-difficulty ratio. For most companies under a few hundred million in revenue, the right starting point is in the first three.

1. Open-end coding and thematic analysis

This is the highest-return, lowest-risk application, and the one most teams still do manually or skip entirely. A model reads every open-ended response, assigns it to themes, flags responses that do not fit any existing theme, and produces frequency counts with example verbatims attached to each theme.

The gain is not only speed. Manual coding on a large dataset is done by several people working from a shared codebook, and inter-coder consistency degrades as fatigue sets in. A model applies the same criteria to response one and response fifteen thousand. That consistency is a genuine quality improvement, not just a cost saving.

There is one non-negotiable requirement: every theme must be traceable to the specific responses that produced it. If the system gives you a theme and a percentage without letting you click through to the underlying verbatims, you have replaced a slow process with an unauditable one. When you evaluate vendors, ask to see this. It separates the serious tools from the demos.

2. Qualitative synthesis across interviews

A study with twenty in-depth interviews produces roughly twenty hours of transcript. The traditional approach is that one senior researcher reads everything, holds it in their head, and writes the synthesis. It works, it is expensive, and it does not scale past a certain number of interviews.

Models change the economics here. They extract mentions of a given concept across every transcript, compare how different segments describe the same problem, surface contradictions between what a participant said early and late in a session, and identify the point at which new interviews stop producing new themes.

That last capability deserves attention. Deciding when you have reached saturation is normally a judgment call made under budget pressure. Making it measurable is a real methodological improvement, and it often reveals that studies were being stopped too early rather than too late.

What models do not do well is notice the thing nobody said. The absence that a skilled qualitative researcher registers, the discomfort in a pause, the question the participant deflected, is not in the transcript. Treat synthesis as a first pass that a human interrogates, not as the deliverable.

3. Continuous listening across owned and public sources

Most companies already hold a large body of unsolicited customer feedback: support tickets, reviews, chat logs, sales call notes, community posts, churn reasons. It is the most honest research data a company owns, and it is almost never analyzed systematically because it arrives in a form nobody has time to read.

A pipeline that classifies this stream by topic, sentiment, product area, customer segment and severity converts it into something you can trend. The value shows up in a specific way: you stop discovering problems in the quarterly tracker and start seeing them the week they emerge.

The practical warning is selection bias, and it is severe. People who write reviews and open tickets are not a representative sample of your customers. Continuous listening tells you what is going wrong for the vocal minority. It does not tell you what the silent majority thinks, and treating it as if it does is the most common analytical error I see in companies that adopt this first. It complements survey research, it does not replace it. The same discipline of using existing data before collecting new data applies across the business, as I set out in the guide on AI ROI for business.

4. Instrument design and questionnaire review

Bad questions produce confident wrong answers, and most questionnaires are written under deadline by people reusing last year's file. A model reviews a draft instrument for the standard failure modes: double-barreled questions, leading phrasing, unbalanced scales, missing response options, order effects, questions that assume knowledge the respondent does not have, and items that duplicate each other.

This is a modest-sounding application with an outsized effect on data quality, because it catches errors at the only point where fixing them is free. Once the study is in field the money is spent.

It also helps on the other side. Given a business decision, a model can propose the set of questions that would actually inform it, which is useful precisely for the teams that most need it: product and marketing groups running their own research without a trained researcher in the room.

5. Analysis assistance and segmentation

Once data is clean, models help in two distinct ways that get confused with each other.

The first is conversational analysis: asking questions of a dataset in natural language instead of building crosstabs. This is genuinely useful for exploration and genuinely dangerous for conclusions, because the model will answer a badly specified question rather than tell you it is badly specified. Use it to find where to look, then verify the finding with a proper test.

The second is segmentation. Clustering respondents on attitudes and behaviors is a mature statistical technique that predates AI. What is new is speed of iteration: you can run twenty segmentation variants in an afternoon and evaluate each for stability and business usability, instead of running two and picking the better one. The risk is finding structure that is not there. Any segmentation you intend to act on must be validated on a holdout sample, and the ones that fail replication are more common than most agencies admit.

6. Synthetic respondents and simulated audiences

This is the most hyped application in the category and the one that requires the most careful handling.

The idea is to use models to simulate how a defined audience would respond, either to pretest an instrument, to fill in hard-to-reach segments, or to run cheap directional reads before committing to fieldwork. Adoption in the industry is real but narrow: broad AI use among researchers is now close to universal, while regular use of synthetic respondents remains a small minority of practitioners, and skepticism inside the profession is high.

The defensible uses, in my view, are three. Pretesting a questionnaire before it goes to real people. Generating hypotheses to test properly. Stress-testing a concept internally before spending on fieldwork.

The indefensible use is treating simulated output as evidence about the market. A model trained on text reproduces how people write about their preferences, which is not the same as what they do. It is systematically weakest exactly where research matters most: novel products, unfamiliar categories, and populations underrepresented in training data. Any organization adopting this needs a written rule about which decisions may and may not rest on synthetic data, agreed before the first study, not after the first uncomfortable result.

7. Report generation and knowledge retrieval

The final application, and the most mature, is turning the archive of past studies into something you can query. Most companies have years of research sitting in slide decks on a shared drive, unfindable and effectively lost. New studies get commissioned to answer questions that were answered two years ago.

A retrieval layer over that archive changes the default. Before commissioning, you ask what is already known. In several organizations I have seen this alone cut the number of new studies by a meaningful margin, which is a strange kind of return: the tool pays for itself by preventing work.

Report drafting also compresses well. The model produces the descriptive layer, what the data says, and the researcher writes the interpretive layer, what it means and what to do. That division is the correct one, and inverting it is where output starts sounding fluent and saying nothing. This connects directly to the broader pattern described in the guide to generative AI for business.

What AI does not do in market research

This section exists because nearly every failed project I have reviewed failed on wrong expectations, not wrong technology.

It does not decide what is worth researching. The highest-leverage decision in the whole process is which question to ask. That requires knowing the business, its economics, and what decision is pending. No model has that context.

It does not fix a bad sample. Analysis quality is bounded by data quality. Fast, cheap analysis of a badly recruited sample produces wrong answers more efficiently. If anything, cheaper analysis increases the temptation to cut corners on recruitment, which is the wrong place to save money.

It does not establish causation. Models are pattern finders and will happily present correlation with a confident explanation attached. Causal claims still require experimental design or careful quasi-experimental methods.

It does not replace the interview. Rapport, follow-up, the willingness to abandon the discussion guide when a participant says something unexpected: this is where qualitative value is created, and it is human work.

It does not survive without maintenance. A coding scheme built in March degrades as the product changes and new terminology enters customer vocabulary. Without someone monitoring theme drift, the system keeps producing clean-looking output that is quietly measuring last year's categories.

Self-assessment: is your organization ready

Score each question 0 for no, 1 for partial, 2 for yes, and total.

Block A, data foundations

  1. Is all customer feedback from the past two years, including tickets and reviews, stored somewhere a single query can reach?
  2. Are past research reports catalogued so that someone can find what is already known before commissioning a new study?
  3. Can survey responses be linked to actual customer behavior, or do they sit in an isolated survey tool?
  4. Is there a consistent taxonomy for product areas and issue types used across support, product and research?

Block B, method maturity

  1. Is there a written standard for how open-ended responses are coded, or does each project invent its own?
  2. Do you know your current cost per completed interview and per study, fully loaded including internal time?
  3. Has any study in the past two years been designed to test a specific decision rather than to produce a general read?
  4. Is there a documented rule for which claims require statistical testing before they go into a deck?

Block C, organizational capacity

  1. Is there at least one person accountable for research quality, as opposed to research delivery?
  2. Do the teams that commission research act on it, or does the deck get presented and filed?
  3. Is research involved before a decision is framed, or asked to validate one already made?
  4. In the past three years, has any research finding actually reversed a planned decision?

Reading the score.

Zero to eight means your constraint is not AI, it is data foundations and process. The next ninety days should go into consolidating feedback sources and building a searchable archive of past work. Buying a platform now automates disorder.

Nine to fifteen is where most mid-sized companies land. Start with open-end coding and continuous listening. Both produce visible results in weeks without integration work.

Sixteen to twenty means you can take on qualitative synthesis at scale, questionnaire review, and a retrieval layer over the research archive.

Twenty-one to twenty-four puts you in a small minority. Your issue is governance: model documentation, validation protocols, and clear rules on synthetic data. The NIST AI Risk Management Framework is voluntary but gives a well-structured way to map risks, define roles and set measurement practices, and it costs a research organization less to adopt than most functions because documentation discipline is already cultural.

Question 12 is the one that matters most. If no research finding has changed a decision in three years, the problem is not analysis speed, and no tool will fix it. That is a conversation worth having with someone who has run the commercial side of a business, before signing any software contract.

A 30, 60, 90 day roadmap

This is the sequence I use with companies that want to bring AI into their research function. It is deliberately slow in the first thirty days, because early haste is the most common cause of failure at day ninety.

Days 1 to 30: measure what you have

Weeks 1 and 2, inventory. List every study run in the past twenty-four months: business question, method, cost, elapsed time, and whether the finding changed anything. In parallel, inventory unsolicited feedback sources and volumes. Most teams have never seen these two lists side by side, and the comparison is usually uncomfortable.

Week 3, baseline three numbers. Fully loaded cost per study. Elapsed time from brief to delivered finding. Percentage of collected open-ended data that was actually analyzed. The third number is typically far below what anyone expects and becomes the baseline you will report against.

Week 4, pick one use case and define the metric. One. With a number and a date: "cut time from field close to coded results from three weeks to four days by November 30," not "make research more efficient."

Output after thirty days is a two-page document. Software spend: zero.

Days 31 to 60: build the pilot on your own data

Weeks 5 and 6, assemble the validation set. Take a completed study where open-ends were manually coded. That existing human coding is your ground truth, and having it is what separates a real pilot from a demo.

Weeks 7 and 8, run in parallel and compare. The model codes the same dataset. Compare against the human coding on two axes: agreement rate, and the nature of disagreements. Read a sample of the disagreements yourself. In practice a portion of them turn out to be human coding errors, which is informative in its own right and usually changes the internal conversation about accuracy standards.

One rule I do not negotiate: the pilot runs on your data, in your categories, with your messy verbatims and your industry jargon. A vendor demonstrating only on their own curated dataset is concealing the integration cost.

Days 61 to 90: validate, decide, industrialize

Weeks 9 and 10, validation against baseline. Two metrics together: agreement with human coding, and stability when the same data is processed twice. A system that produces different themes on repeated runs cannot be used for tracking, whatever its accuracy on a single pass.

If the pilot does not beat the current process, stop. Having found that out in ninety days rather than two years is a good outcome.

Week 11, design the steady-state process. Who reviews model output, at what sampling rate, who owns the codebook, how new themes get approved, what happens when the model and the analyst disagree. A model without a human owner drifts within six months and nobody notices.

Week 12, decide. Three outcomes are all legitimate: industrialize, iterate for another thirty days, or stop and document why. Healthy organizations can take the third.

By day ninety you should have one use case in production with a measured benefit, not five pilots in permanent limbo. The entire difference between companies that get value from AI and companies that accumulate proofs of concept lives in that discipline, a pattern I unpack further in the enterprise AI adoption framework.

What it costs and how the return is calculated

The figures below reflect what I have seen companies between roughly ten and five hundred million in revenue actually pay. They are orders of magnitude, not price lists.

Entry cost by application

Open-end coding. Model consumption on typical volumes runs a few hundred dollars per month. Setup, validation and integration into the existing workflow runs 10,000 to 30,000 dollars. Time to production: four to eight weeks. This is the cheapest meaningful entry point in the category.

Continuous listening pipeline. 20,000 to 60,000 dollars depending on how many sources need connecting and how much taxonomy work is required. Time: two to four months. Most of the cost is plumbing, not modeling.

Qualitative synthesis tooling. 15,000 to 45,000 dollars, plus per-seat platform costs if you buy rather than build. Time: six to twelve weeks.

Research archive retrieval. 15,000 to 50,000 dollars, almost entirely document preparation and metadata work. Time: two to four months. Boring, and one of the highest-return projects on this list.

Segmentation and analysis tooling. Highly variable. If your data is already in a warehouse, 10,000 to 30,000 dollars. If it is not, the warehouse is the project and this is a later phase.

The line item nobody budgets

Data preparation absorbs between 50% and 70% of total effort on any project of this kind. In research specifically there is a second underestimated item: building the validation set. Someone who knows the categories has to hand-code a sample so you can measure whether the model is right. That is tens of hours of a senior person, and it is not delegable to someone unfamiliar with the domain.

The practical rule: budget an amount for data preparation and validation comparable to what you budget for technology. If that seems excessive, consider the alternative, which is a system nobody can prove is accurate reporting themes into a board deck.

Three components of return

Analyst hours released, high certainty. This is the number that survives scrutiny in a budget meeting. A team that recovers 60 hours a month across coding, transcript review and report assembly, at a fully loaded 60 dollars per hour, releases 43,200 dollars a year of capacity.

Studies avoided, medium certainty. Count how many commissioned studies in the past two years asked a question the existing archive could have answered. Multiply by fully loaded cost per study. In organizations with a decent research history this figure is frequently larger than the hours saving.

Better decisions, low certainty. Real, and unquantifiable in advance. State it as an expected benefit and exclude it from the payback calculation. A business case resting on unquantifiable benefits gets dismantled at the first review.

Built this way, payback typically lands between eight and sixteen months. The full calculation method applies to any initiative and is set out in the AI readiness assessment guide.

Real cases: what I have measured

These come from my own work. I report the numbers that were measured, not the ones that sound best. None of them started as a research project, which is exactly why they are worth reading: the mechanism that produced the result was the same one described above.

WSB Sport, 30% sales increase

The engagement started in marketing. The multiplier was reconstructing actual profitability by channel and by product line, which revealed that a share of advertising spend was feeding low-margin items. Reallocating budget toward the healthy lines, combined with automating creative production, produced a 30% increase in sales.

The research lesson is direct. The data required to reach that conclusion already existed inside the company. Nobody had assembled it because assembling it was too slow to be worth doing, which is precisely the constraint that has now moved. The marketing side of this pattern is covered in the AI marketing strategy guide.

Hotel, revenue from nine to ten million

The work involved building a demand forecast by segment and booking window, tied to weekly rather than seasonal rate revision. Revenue moved from nine to ten million at the same occupancy, on a higher average rate.

The relevant part for research: the same techniques applied to reviews and complaints showed that a large share of negative ratings concentrated on two repeated operational causes, not the twenty the team believed it was managing. That finding came from text nobody had been reading.

Medical center, 20% capacity increase

No new equipment, no new hires. The bottleneck was scheduling and no-shows. A model predicting per-patient no-show risk enabled selective overbooking and differentiated reminders, raising delivered appointments by 20% with the same staff and facility.

The point for this guide is that the predictive signal came from operational records, not from a study. A great deal of what companies commission research to learn is already latent in data they hold.

Mistakes I see repeatedly

Buying a platform before measuring the current process. The correct order is data, diagnosis, process, tool. Inverting it produces active subscriptions and usage that quietly stops by the second quarter.

Skipping the validation set. Without human-coded ground truth you cannot state accuracy, which means you cannot defend a finding when someone senior disputes it. That moment always arrives.

Letting cheap research inflate study volume. When unit cost falls, the temptation is to run more studies rather than better ones. The result is more decks and no more decisions changed, and the credibility of the whole program suffers.

Using continuous listening as if it were representative. It is a biased sample by construction. Useful for early warning, unreliable for sizing anything.

Treating synthetic respondents as data. Useful for pretesting and hypothesis generation. Not evidence about a market. The rule needs to be written down before anyone is under deadline pressure.

Confusing adoption with impact. According to the Stanford HAI AI Index Report 2025, the share of organizations reporting AI use rose to 78% in 2024, up from 55% the year before. McKinsey's 2025 State of AI survey puts organizational adoption higher still, while finding that only a minority can point to measurable impact on enterprise earnings. Using AI and getting value from it remain two different conditions.

Ignoring change management. A senior researcher whose standing rests on being the person who has read everything will not welcome a system that makes that knowledge explicit and transferable. Involve them as the owner of the new process rather than the recipient of a decision made elsewhere. In my experience this is the single factor most correlated with project success.

How the researcher's role changes

Worth being explicit, because it is the real concern of people doing this work.

The researcher of 2020 spent most of their time collecting, processing and summarizing. The researcher of 2028 will have that ratio inverted. Three capabilities become decisive, and none of them is technical.

  1. Question framing. Translating a business decision into a researchable question, and knowing when the honest answer is that research will not resolve it. This requires commercial understanding, which is why it stays human.
  2. Methodological judgment. Knowing which claims need an experiment, which need a representative sample, and which can rest on directional evidence. As the cost of producing analysis falls, the value of knowing which analysis to trust rises.
  3. Making findings land. Getting an organization to act on evidence that contradicts what it wanted to hear. This is a political and narrative skill, and it becomes more important as the volume of available findings increases.

Where training is needed, the realistic horizon is forty to sixty hours to bring an experienced researcher to independent use of these tools and, more importantly, to correct skepticism about their output. The return is faster than any platform purchase, and it scales in the same way described in the guide for small businesses adopting AI.

Governance and the regulatory picture

Two governance items matter for research specifically, and both are frequently missed.

The first is privacy and consent. Feeding customer verbatims, interview transcripts or support tickets into an external model is a data processing activity, and it needs the same scrutiny as any other transfer of personal data to a third party. The practical checks: what the vendor does with submitted data, whether it is used for training, where processing happens, and whether your existing consents cover it. This is a contract question, and it belongs in the vendor selection stage rather than after the pilot.

The second is transparency. Under the EU AI Act, whose official text is published on EUR-Lex, most research tooling falls into the limited-risk category. The obligations that apply are principally transparency and AI literacy for staff using these systems, and the simplification package adopted in 2026 did not defer either of those. If you interact directly with respondents using an AI system, for example an automated interviewer, disclosure is not optional.

The third item is not regulatory but professional: if a finding presented to a client or a board was produced with substantial model involvement, say so in the methodology note. Industry norms on this are still forming, and the organizations that will look worst in three years are the ones that were quiet about it.

The path for smaller companies

Smaller companies have three constraints that large ones do not, and three advantages they almost never use.

The constraints are familiar: no dedicated insights function, limited budget, and customer feedback scattered across tools nobody has connected.

The advantages are less obvious. Decisions are fast because there is one decision maker. Processes are unlayered and can be redesigned in weeks. And the starting point is low enough that the first intervention produces percentage gains that would be impossible in a mature function.

The sequence that works has a specific shape. Start by analyzing feedback you already own, which costs almost nothing and usually produces the first surprise within a month. Then add lightweight, frequent surveys to your own customer base, analyzed automatically, to fill the gaps that unsolicited feedback cannot cover. Only at the third step commission representative external research, now with a much sharper question than you would have asked at the start. Starting with expensive external research when you have not read your own support tickets is the most common way small companies waste a research budget. The commercial framing of this sequencing is covered in the go-to-market strategy framework.

On vendor selection, a simple defensive rule: be wary of anyone who proposes a tool before seeing your data, who will not put the pilot success metric in writing, and who cannot explain what happens when your product terminology changes.

Checklist before signing anything

Use this as a final filter. If you cannot answer yes to every line, you are not ready to sign.

  • I know my fully loaded cost per study and my elapsed time from brief to finding.
  • I know what percentage of my open-ended data currently gets analyzed.
  • I have one use case defined, with a numeric metric and a measured baseline.
  • I have a human-coded validation set to test the system against.
  • The pilot will run on my own messy data, in my categories, for at least one month.
  • Every theme or finding the system produces can be traced back to source responses.
  • I know who owns the codebook and who arbitrates when model and analyst disagree.
  • I have budgeted data preparation and validation at a level comparable to the technology.
  • I have checked what the vendor does with submitted data and whether it trains on it.
  • I have a written rule on which decisions may not rest on synthetic or simulated respondents.

Where to start tomorrow morning

If you want one concrete action in the next forty-eight hours, this is it: export every open-ended response, support ticket and review from the last twelve months, run it through a single thematic pass, and rank the resulting themes by volume.

In most companies the outcome is the same. Three or four themes account for the majority of the volume, and at least one of them is not what the leadership team believes the main problem to be. That single output tells you both where the evidence points and which part of your research program is currently answering questions nobody needed answered.

AI for market research is not a technology decision. It is the decision to stop treating evidence as something you commission twice a year and start treating it as something the business runs on continuously. The companies that make that shift over the next eighteen months will be deciding with evidence at a cost their competitors cannot match, in an industry where the software half is already growing at more than twice the rate of the services half. If you are weighing where to start and want a direct conversation about the right first step for your organization, with no platform to sell you, that is exactly the kind of discussion I begin with the leadership teams I work with.

FAQ

What is AI for market research?

It is the use of language models and machine learning to automate the repetitive parts of the research process. In practice that means five things: coding open-ended survey responses into themes with traceable verbatims, synthesizing findings across many interview transcripts, classifying continuous streams of unsolicited feedback such as tickets and reviews, reviewing questionnaires for design flaws before they go to field, and making the archive of past studies searchable so teams stop commissioning research that already exists. It does not decide what is worth researching, and it does not fix a badly recruited sample. It removes the analysis bottleneck, which historically meant most collected qualitative data was never read.

How much does it cost to introduce AI into a research function?

The cheapest meaningful entry point is automated open-end coding, which typically runs 10,000 to 30,000 dollars for setup, validation and workflow integration, plus a few hundred dollars a month in model consumption. A continuous listening pipeline across support tickets, reviews and chat logs runs 20,000 to 60,000 dollars depending on how many sources need connecting. A retrieval layer over the past research archive lands between 15,000 and 50,000 dollars, almost all of it document preparation. The rule that matters most: budget for data preparation and validation at a level comparable to the technology itself, because that work absorbs 50% to 70% of real effort.

How long until the investment pays back?

Typical payback is eight to sixteen months, calculated on two defensible components. The first is analyst hours released: a team recovering 60 hours a month at a fully loaded 60 dollars per hour frees 43,200 dollars of capacity a year. The second, often larger, is studies avoided, counted as the number of commissioned studies in the past two years whose question the existing archive could have answered, multiplied by fully loaded cost per study. Better decision quality is real but should be excluded from the payback math, because a business case built on unquantifiable benefits gets dismantled at the first review.

Can synthetic respondents replace real ones?

No, and treating them as evidence about a market is the most consequential mistake in this category. Simulated respondents are defensible for three narrow uses: pretesting a questionnaire before it reaches real people, generating hypotheses you will then test properly, and stress-testing a concept internally before committing to fieldwork. They are systematically weakest exactly where research matters most, meaning novel products, unfamiliar categories, and populations underrepresented in training data. Adoption data from the industry reflects this: general AI use among researchers is now near universal, while regular use of synthetic respondents remains a small minority and professional skepticism is high. Write down which decisions may not rest on synthetic data before your first study, not after an uncomfortable result.

Which application should we start with?

Automated coding of open-ended responses. It has the best value-to-difficulty ratio: the data already exists, no system integration is required, results arrive within weeks, and you can validate accuracy against work your team has already coded by hand. The second step is a continuous listening pipeline over support tickets and reviews, which shifts you from discovering problems quarterly to seeing them weekly. Qualitative synthesis at scale and research archive retrieval come later, because they require document preparation that most organizations have not done.

How do we know the model is coding responses correctly?

By building a validation set before you start. Take a completed study where humans coded the open-ends, treat that coding as ground truth, run the model over the same data, and measure both the agreement rate and the character of the disagreements. Read a sample of disagreements yourself, because a meaningful share of them usually turn out to be human coding errors rather than model errors. Then check stability by running the same dataset twice: a system that produces different themes on repeat runs cannot be used for tracking, regardless of its accuracy on any single pass. Without a validation set you cannot state accuracy, and you cannot defend a finding when a senior stakeholder disputes it.

Is it safe to put customer verbatims and transcripts into an AI system?

It requires the same scrutiny as any transfer of personal data to a third party, and it should be settled during vendor selection rather than after the pilot. The four checks that matter: what the vendor does with submitted data, whether that data is used to train their models, in which jurisdiction processing happens, and whether the consents you already collected from respondents cover this use. Under the EU AI Act most research tooling sits in the limited-risk category, where the live obligations are transparency and staff AI literacy, neither of which was deferred by the 2026 simplification package. If an AI system interacts directly with respondents, disclosure is mandatory.

Will AI replace market researchers?

It replaces a large share of the collecting, processing and summarizing that has historically filled the role, which is most of the working week for junior positions. What becomes more valuable is framing the business question, exercising methodological judgment about which claims need an experiment versus a directional read, and getting an organization to act on findings it did not want to hear. As the cost of producing analysis falls, the ability to judge which analysis to trust becomes the scarce skill. The real risk to a researcher is not being replaced by a model: it is remaining the person who spends three weeks coding verbatims while a colleague spends those weeks changing a decision.