ChatGPT brand monitoring: how to track brand mentions in AI answers
When a buyer asks ChatGPT for the best tool in your category, does your brand come up? ChatGPT brand monitoring is how you answer that with evidence instead of anecdote: ask the engines the questions your buyers ask, record what comes back, and read the pattern over weeks.
This page covers what you can actually learn, a manual method you can run this afternoon without buying anything, what changes when you automate it, and exactly what Boldmention does. It also answers the question a vendor should answer first: why not just use the report Microsoft already gives publishers about their own site.
Published 12 September 2026. Every third-party statement below carries its source and the date we read it.
What you can find out, and what the engines do not tell you
You can measure outputs. Ask a question, record the answer, repeat. That is the whole evidence base, and it is enough to act on: which questions name you, which name somebody else, and which pages get cited.
You cannot find out why, at least not from the engines themselves. OpenAI’s public crawler documentation covers access, not selection — it explains that “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links” (OpenAI crawler docs, read 12 September 2026). That is a permission rule. It says nothing about which of the permitted sources an answer ends up using, and we have found no OpenAI document that does.
One engine family is an exception, and only about its own surfaces. Google documents that AI Overviews and AI Mode use “retrieval-augmented generation (RAG)”, relying on core Search ranking systems to retrieve pages, plus “query fan-out”, where the model issues related queries alongside yours (Google, AI features and your website, read 12 September 2026). That is Google describing Google Search. It is not a description of ChatGPT, Perplexity or Claude, and it should never be read as one.
In the same guide Google writes: “No third-party tool has access to our internal ranking or AI systems.” Take that as the frame for this entire category, ours included. A monitoring tool is an instrument pointed at the output of a system, not a window into it. That is worth holding every tool to, this one included.
The manual method: run it today, buy nothing
Do this before you evaluate any tool. It takes an afternoon and it teaches you what the numbers in any dashboard are made of.
- Write 20–30 questions your buyers actually ask. Not your brand name. “Best AI visibility tool for a B2B SaaS marketing team” is the kind of question that decides whether you exist; “what is [your brand]” only tells you how you are described once someone already knows you. Write both, mostly the first kind.
- Ask each one in a fresh chat, with the wording fixed. Anything else in the conversation is an uncontrolled variable, so a thread of follow-ups is not a repeatable measurement. Start a new chat and paste the same string every time.
- Ask each question three times. Expect different answers. Treat any single answer as one sample, never as the state of the world. This is the step most people skip, and it is the one that stops you rebuilding your homepage because of a single bad reply.
- Record seven columns. Date, question, engine, brand named (yes/no), position in the list, the other brands named, and the URLs the answer cited. The last two are the ones that change what you do next.
- Repeat weekly and do not touch the wording. Editing a question restarts its series. Keep a changelog if you must change one.
- Read it as rates. Named in 4 of 30 questions this week and 7 of 30 next week is signal. One answer flipping is not.
The arithmetic is what eventually pushes people to automate: 30 questions, three repeats, one engine is 90 answers to read and log. Four engines is 360, every week, with the wording identical each time.
What automating it actually requires
- Volume and cadence. Questions multiplied by engines multiplied by repeats, performed on the same schedule whether or not anyone is watching.
- Stable inputs stored with the outputs. If you cannot see the exact string that produced an answer, the series is not comparable across weeks.
- Matching that does not lie. Naive substring matching finds short brand names inside unrelated words and inflates every metric built on top.
- Somewhere to put the links. The URLs an answer cites are the part you can act on. They need extracting and attributing, not just reading.
How Boldmention runs it
This is the pipeline, in the order it executes. Where a number appears, it is a number in the code.
Setup: questions and competitors are derived, not typed
You give us a website URL. We read its sitemap, crawl a bounded set of pages from it, build a description of the business from what we find, and ask a model for six to eight topics and 30–40 candidate questions. You pick up to six to start with, and you can add more later from the dashboard.
There is no field for competitors, because you do not supply them. The same analysis runs a two-pass discovery: one model produces a broad candidate list, a second filters it down to direct alternatives, and the validated list is stored against your company. It then keeps growing on its own — roughly one tracked answer in ten triggers another discovery pass over that answer’s text, and any name not already stored is added.
Each run: five models, different grounding by design
Every active question goes out in parallel to the models your plan allows, drawn from a set of five: GPT-5 Nano (OpenAI), Gemini 2.5 Flash, Grok 3 Mini (xAI), Perplexity Sonar and DeepSeek V3.2 Chat.
How each one gets web results differs, and the difference is deliberate rather than incidental:
- For GPT-5 Nano, Gemini 2.5 Flash and Grok 3 Mini we call an external search API (Serper, with Brave as the fallback), take up to seven organic results for the question, and prepend them to the system message. Gemini’s own native search is switched off when we do this, so all three see the same results.
- Perplexity Sonar searches on its own, so we inject nothing.
- DeepSeek V3.2 Chat runs with no web results attached at all. It answers from the model.
Detection, scoring and what each number means
- Mention detection is a word-boundary regular expression on your exact brand string, so a brand name sitting inside a longer word is not counted.
- If the brand is absent, we store the answer with a visibility score of 0 and make no further model call.
- If the brand is present, a second call (GPT-5 Nano) returns three fields: position in the list of recommendations, sentiment, and a 0–100 prominence score. Be clear about what those are: one model judging another model’s answer. We store them because they are consistent and useful, not because they are ground truth.
- References. Every URL in the answer text is extracted with its domain and matched against your brand and the stored competitor names.
- Share of Voice for a question is your mentions divided by all brand mentions in its answers — yours plus every competitor counted. Per answer we count at most ten competitor names, so an answer naming more than ten is truncated there, and that moves the number in our favour. Worth knowing before you quote it.
Repeat runs are driven by a scheduled batch job, and that job filters free-trial accounts out explicitly. A free trial gives you one snapshot at onboarding, not a series. That is worth knowing before you judge the product on a trial.
What Boldmention does not do
- It does not open chatgpt.com. Our calls go to the providers’ APIs. An API model with search results attached is close to what a logged-in user sees in the app, and it is not the same product. It is a distinction worth asking any vendor about, us included.
- It does not cover engines outside the five models above. No Claude, no Copilot, no Google AI Overviews.
- It does not tell a paid placement from an organic mention. OpenAI’s own news feed lists “Testing ads in ChatGPT” (11 August 2026) and “ChatGPT Ads expands across Europe” (18 August 2026) (OpenAI news feed, read 12 September 2026). Our pipeline records the text of an answer. Nothing in it identifies how a brand came to be there.
- It does not explain why an engine answered as it did, and it cannot promise you a place in an answer. See Google’s line about third-party tools above.
Why not just use Bing Webmaster Tools instead?
For part of this job, you should. Microsoft gives publishers first-party citation data about their own site, and no third-party measurement can match first-party data on that specific question. Check it before you pay anyone, including us.
Its AI Performance report, launched in public preview on 10 February 2026, reports “Total Citations,” “Average Cited Pages,” “Grounding queries” and “Page-level citation activity” across “Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations” (Bing Webmaster blog, read 12 September 2026). Grounding queries are the phrases the AI used when retrieving. For contrast, Google’s generative-AI report in Search Console documents impressions broken down by page, country, date and device (Search Console Help, read 12 September 2026); the query itself is not among the dimensions that page documents.
Since 16 June 2026 it also has Citation Share, which Microsoft describes as available in preview globally and defines as “the percentage of citations attributed to your site out of all citations shown across all sites for that same grounding query”, alongside Intents (grounding queries classified by intent), Topics (clustered into themes) and Compare (overlay a previous period) (Bing Search blog, read 12 September 2026). If you have read anywhere that Bing Webmaster Tools has no share metric, that stopped being true on that date.
Microsoft states the bounds itself, and they are the honest answer to the heading above. Citation Share “is designed as an observational metric – not a ranking system or a competitive scoreboard. It does not expose competitor domains, represent traffic share, or assign quality scores to content.” (Bing Search blog, read 12 September 2026). And of the grounding-queries data specifically: “The data shown represents a sample of overall citation activity.” (Bing Webmaster blog, read 12 September 2026).
So the split is clean. If your question is “how often do Microsoft’s AI surfaces cite my pages, and for which grounding queries”, go and get that report; it is first-party and better than any inference. If your question is “when someone asks for a recommendation in my category, which brands get named, on which engines, and how does that move month to month”, it is a different measurement: it spans engines Microsoft does not report on, it is about questions where nobody may have cited you at all, and by Microsoft’s own description it is not what Citation Share was built to show.
Start with the questions, not the tool
Whatever you use, the question list decides what you learn. A tool run against 30 brand-name queries will report healthy numbers and tell you nothing, because the answer to “what is [your brand]” was never in doubt. The uncomfortable, useful questions are the category ones, where the engine picks somebody.
If you would rather have it run for you, Boldmention starts on a free trial, and the trial is exactly the one-off run described above: request an AI visibility report and you get the questions, the answers and the brands named in them. For which models each plan covers, how the metrics are computed and how often prompts run, see the frequently asked questions. For longer write-ups on Share of Voice, LLM drift and AI search visibility, see our blog.