How do you monitor what AI assistants say about your company?

How do you monitor what AI assistants say about your company?

How do you monitor what AI assistants say about your company?

THE SHORT ANSWER

Prompt monitoring means running a fixed set of buyer questions against each assistant on a schedule and logging what comes back. Because generation is probabilistic, a single run tells you almost nothing, so repeat each prompt three times and record the proportion of runs where you appear. Log four things per run: were you named, were you linked, was the description accurate, and which competitors appeared. Keep the prompt set frozen or the trend becomes unreadable.

The first attempt at this usually looks like someone typing their company name into an assistant, screenshotting a flattering answer and pasting it into a channel. That is not monitoring, it is a mood. Assistants generate different text for identical prompts, personalise by account and location, and change behaviour when the underlying model is updated, so a single observation carries almost no information.

What works is ordinary sampling discipline borrowed from any other noisy measurement problem. Fix the questions, control the conditions, repeat enough times to see through the variance, and record structured outcomes rather than impressions. It is dull, it takes an hour a month once built, and it is the only way to know whether your visibility work did anything.

The numbers, at a glance

  • Prompt set size: twenty to forty questions is the workable range; below twenty the sample is too small, above forty nobody keeps it up

  • Repeats per prompt: three runs per prompt per assistant, because a single run has roughly coin-flip reliability on contested questions

  • Session hygiene: each run from a logged-out or fresh session with no memory or custom instructions, otherwise you are measuring your own history

  • What to log: four separate fields per run: named, linked, described accurately, and which competitors appeared alongside you

Designing a prompt set that measures something

Prompts are not keywords. Write them the way a buyer speaks, including the hedges and the context, because that phrasing changes which search the assistant runs underneath. Then balance the set across four types so one category of movement cannot dominate the score.

  • Category prompts. Who supplies exclusive solar leads in the Netherlands. These test whether you exist in the assistant's picture of the market at all.

  • Comparison prompts. What are the alternatives to a named competitor. These are where challenger brands win or lose, and where absence is most expensive.

  • Attribute prompts. Which lead suppliers do not resell the same enquiry. These test whether your actual differentiator is legible to the machine.

  • Direct prompts. What does this company charge, where does it operate. These measure accuracy rather than visibility, and they are the ones that catch commercial damage.

Write the set once, then freeze it. The single most common way teams destroy their own baseline is adding prompts in month three because someone thought of a good one. Park new ideas in a separate list and start a second baseline if you must.

Running it so the numbers mean something

Same day each month, same order, fresh session, no personalisation, and a recorded location if the assistant asks for one. Three repeats per prompt per assistant, which for a thirty-prompt set across four assistants is three hundred and sixty runs. That sounds heavy until you realise it is a couple of hours with a browser and a spreadsheet, or minutes with a script if the assistant exposes an interface you are permitted to use.

Record mention rate as the fraction of runs where you appear, not a yes or no. A prompt where you show up in one run out of three is a genuinely different position from one where you show up in three out of three, and treating both as a mention discards the most useful signal you have. Movement from one-in-three to two-in-three is real progress and would be invisible under a binary.

The accuracy log, which matters more than the mention rate

For every run where you are named, score the description as correct, partially correct or wrong, and write down the specific error in plain language. Over a few months this log becomes the most operationally useful artefact in the programme, because errors cluster and each cluster points to a fixable source.

The recurring causes are predictable: a price that only exists inside a downloadable document, a services list on a page that was superseded but never redirected, a country you exited two years ago still listed in a directory, or a competitor's comparison page describing your offer inaccurately and ranking well enough to be retrieved. Each has a different fix and none of them is discoverable from a mention-rate chart.

What to do with a month of results

Three questions, answered in order. Did mention rate move by more than the noise floor, which for three repeats is roughly a third of a point per prompt, so ignore anything smaller. Did any accuracy error appear in more than one assistant, which means it is a source problem rather than a model quirk. Did a new competitor appear across several prompts, which usually means they published something newly retrievable.

Then act on exactly one thing. The failure mode of monitoring programmes is generating a report nobody acts on, and the discipline that prevents it is committing to a single fix per cycle, applied to whichever error appeared most often. Slow, but it compounds, and it produces a causal record you can actually read a year later.

Standing up a monitoring programme this week

  1. Write twenty to forty prompts balanced across category, comparison, attribute and direct types.

  2. Freeze the list and store it where nobody can quietly extend it.

  3. Build a sheet with one row per prompt per assistant per run and four outcome columns.

  4. Run three repeats per prompt from clean sessions on a fixed date each month.

  5. Close each cycle by fixing the single most frequent accuracy error, then log what you changed.

Want leads like this in your pipeline?

Flock runs the campaigns, screens the enquiries and hands you only the ones that match your service area, job size and capacity. You pay per lead, not per month.

Book a 15-minute fit check  |  See lead package pricing

Related answers

Frequently asked questions

How many prompts is enough?

Twenty to forty. Fewer than twenty and a single volatile question swings the whole score; more than forty and the monthly run becomes a job nobody wants, which is how programmes quietly die in month four. Balance across question types matters more than raw count.

Should I buy a monitoring tool?

The tools automate the repetitive runs, which is genuinely worth paying for above a certain scale. Be sceptical of any dashboard reporting a rank or a share-of-voice score, because neither exists inside a generated answer. The underlying data is always just did-you-appear, counted and averaged.

Why do I get a different answer every time I ask?

Generation samples from a probability distribution, retrieval runs live against a changing web, and personalisation varies by account, memory and location. That is precisely why the method is repeated sampling from clean sessions rather than a single check, and why one screenshot proves nothing.

How quickly should I expect the numbers to move?

Assistants that browse live can reflect a new page within days to weeks. Anything depending on what a model absorbed during training moves on the model release cycle, so allow two to three quarters before judging whether a change of approach worked.

NEED A CLEARER PLAN?

Let’s turn your next move into momentum.

Talk to us →

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet

contact

hello@flockleads.com

Reply within 24 hours

REMOTE

Remote-first

Serving clients worldwide

All meetings via Teams or Google Meet