LLM Visibility: I Audited My 10 Product Sites and Found One Rule Blocking ChatGPT

You have probably asked ChatGPT about your own category and watched it name a competitor, or nobody at all. Then you open an LLM visibility dashboard, get a share-of-voice percentage, and still have no idea what to fix on Monday.

LLM visibility is how often AI assistants such as ChatGPT, Google AI Overviews, Perplexity, Claude and Copilot mention or cite your site when people ask the questions your buyers ask. It is not one score. It is three checks in order: can the AI crawlers reach your pages, do you show up in the search indexes those assistants retrieve from, and do they actually cite you. On 26 September 2026 I ran all three checks on my own 10 product sites. The results: zero tracked LLM mentions on any of them, Google’s AI Overview cited me on 2 of 10 test queries, ChatGPT searched the web on only 4 of 10, and one Cloudflare rule was returning a 403 to every AI crawler on my biggest site. Below are the exact numbers, the fix, the curl test, and the 10-query audit sheet I now use.

What is LLM visibility, and why is it three numbers instead of one?

LLM visibility is the share of relevant AI answers that mention or link to you, and it only exists if three earlier things are true. An assistant cannot cite a page its crawler was refused. It rarely cites a page that does not surface in the index it searched. And even when both are fine, it may answer from memory and cite nobody.

That is why I stopped reading any single “AI visibility score”. Kevin Indig’s analysis of Omnia data found that only 2.35% to 2.45% of cited URLs appeared in ChatGPT, Perplexity and Google AI Overviews for the same prompt, and 91% of citations showed up in only one engine. A page can be a regular in AI Overviews and invisible in ChatGPT. So I measure per engine, per layer:

Layer Question How I measure it
1. Access Can the AI search crawlers fetch the page and robots.txt? curl with each bot’s user agent, then the firewall log
2. Retrieval Does the page rank where the assistant searches (Google, Bing)? Search Console and Bing Webmaster data
3. Citation When the assistant answers, does it cite or name you? Live checks per engine on a fixed query list

If layer 1 fails, layers 2 and 3 do not matter. That is exactly what I found.

What did my 10-site LLM visibility audit actually find?

My 10 sites are close to invisible in AI answers today, and the reasons are mostly boring: thin organic reach on six of them, and a firewall rule on the seventh. The domains were murkuz.com, ertiqah.com, mjunaidkhalid.com, contextli.com, locul.ai, ligosocial.com, prepclubs.com, ccattests.com, testimonials.ltd and meisa.io. I ran everything on one day, 26 September 2026, US English, using DataForSEO’s LLM Mentions API, its ChatGPT scraper, a live Google SERP pull for AI Overviews, and plain curl.

Check Result
Tracked LLM mentions (ChatGPT and Google AI) 0 rows for all 10 domains, on both platforms
Domains with any Google US top-10 keyword 4 of 10; 6 had none at any volume
Google AI Overview present on my 10 test queries 10 of 10
AI Overview cited one of my domains 2 of 10, both ligosocial.com blog posts
ChatGPT searched the live web 4 of 10 queries; the other 6 were answered from memory with no citations
ChatGPT cited one of my domains 1 of 10, and it picked the wrong one of mine (below)
Domains that block AI crawlers 1 of 10: ligosocial.com

Three findings changed what I work on.

First, “zero mentions” in a tracker is a finding, not a bug. The LLM Mentions index returned a clean, empty result for every domain. At my current scale none of these sites are inside the tracked corpus yet, so a paid dashboard would have shown me a row of zeros too.

Second, ChatGPT only browsed on 4 of my 10 queries. For the other 6 it answered from training data and cited nobody. No on-page edit changes that. For those queries the lever is being mentioned widely enough to be in the model’s memory, which is a much slower game.

Third, the engines do not agree even inside my own portfolio. For “ccat score range”, my personal blog was the page ranking on Google. ChatGPT searched, skipped it, and cited a CCAT scoring guide on ccattests.com instead, a different site of mine that DataForSEO did not even show in the top 10. Ranking on Google did not decide the ChatGPT citation.

Which misconfiguration was silently blocking ChatGPT?

A Cloudflare security challenge on ligosocial.com was serving a 403 “Just a moment…” page to every AI crawler, including OpenAI’s search crawler, while letting Googlebot and Bingbot through. ligosocial.com is my strongest site on Google by a distance: 145 top-10 keywords against single digits everywhere else. It was also the only site where Google’s AI Overview cited me. And ChatGPT had never cited it once.

When I tried to score six ligosocial.com posts, every fetch came back 403. So I tested the same URLs with different user agents. I re-ran the test across all 10 domains on 26 September:

User agent ligosocial.com The other 9 domains
Googlebot 200 200
bingbot 200 200
OAI-SearchBot 403 200
ChatGPT-User 403 200
GPTBot 403 200
Claude-SearchBot 403 200
ClaudeBot 403 200
PerplexityBot 403 200

Even /robots.txt returned the challenge page. That detail matters: my robots.txt could have said “welcome” in capital letters and it would not have mattered, because the bots never got to read it. robots.txt is a request; a firewall rule is a wall.

OpenAI is blunt about what that costs. Its crawler documentation says: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers”. A firewall challenge is an opt-out you never chose. The pattern also explains the split I saw: Google’s AI Overview could cite ligosocial.com because Googlebot was let through, and ChatGPT could not because its crawler was not.

The lesson I took is not technical. The rule did its job: it kept junk traffic off the site. Nobody went back to ask who else it was keeping out once the list of crawlers that matter changed. The fix is to let the AI search crawlers past the challenge, then confirm they get a 200. I have not measured the after yet. I will update this section with the ChatGPT citation result once the crawlers have had a few weeks to recrawl.

How do you check whether AI crawlers can reach your site?

Run one curl loop against your homepage and one key article, look for anything that is not a 200, then confirm in your firewall log. Here is the exact loop I used. Replace the URL:

for ua in Googlebot bingbot OAI-SearchBot ChatGPT-User GPTBot Claude-SearchBot ClaudeBot Claude-User PerplexityBot; do printf '%-18s %sn' "$ua" "$(curl -s -o /dev/null -w '%{http_code}' -A "Mozilla/5.0 (compatible; $ua)" https://yoursite.com/)"; done

Read it like this. All 200s: layer 1 is probably fine. A 403 or 503 only on the AI crawlers, as on ligosocial.com: you have a bot rule to fix. A 403 on everything, including Googlebot: your challenge is blocking all non-browser traffic, which is a bigger problem.

One honest caveat. curl only pretends to be each bot. Cloudflare and other firewalls can treat the real crawler differently, because they verify it by IP. So a 403 here is a strong signal, not proof. Confirm it in Cloudflare’s security events log by filtering on the bot names, and check two settings: the Block AI bots toggle and AI Crawl Control, which let you allow search crawlers while still refusing training crawlers.

Then make robots.txt say what you mean. This is the block I now use. It allows the crawlers that decide live citations and leaves training crawlers as a separate choice:

User-agent: OAI-SearchBot
Allow: /

User-agent: ChatGPT-User
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: Claude-User
Allow: /

User-agent: Bingbot
Allow: /

If you want to opt out of model training but stay citable, add a separate User-agent: GPTBot group with Disallow: /. OpenAI states that each setting is independent. The token names come from the vendors’ own pages: OpenAI’s crawlers, Anthropic’s ClaudeBot, Claude-User and Claude-SearchBot, Perplexity’s crawlers and Bing’s crawlers. Perplexity says PerplexityBot “is not used to crawl content for AI foundation models”, so blocking it only removes you from its results.

Why does Bing matter for LLM visibility?

Bing matters because ChatGPT search and Copilot are widely reported to lean on Bing’s index, and a site Bing will not serve has one less door into those answers. I am careful with that claim. Most of the “Bing powers ChatGPT” material I read comes from vendors, and I found no independent test that proves the size of the effect. But checking Bing takes minutes, so I check it.

I learned this the expensive way. From March 2026, when I added contextli.com to Bing Webmaster Tools, until late September, Bing showed the site for nothing: zero impressions on every day of roughly six months, even though 579 pages sat in its index and its own URL inspection said “Indexed successfully”. No crawl errors, no blocked URLs, and Cloudflare was ruled out. The only fix was a support ticket. On 26 September Bing replied that the issue was resolved and that the site should start serving again within two to three weeks. The recovery is in progress, and I do not have an after number yet.

Two habits came out of it. I submit new URLs through IndexNow instead of waiting for a crawl. And I look at Bing numbers next to Google’s every week, not once a quarter. I pull both into Claude through Murkuz’s Search Console connector, which also exposes Bing search performance (tool list), so “is Bing showing this site at all” is one question in a chat instead of two dashboards. I wrote up that setup in I connected Google Search Console to Claude.

Which LLM visibility tools should you use?

For a small or young site, a fixed list of 10 queries checked by hand, plus the curl test above, tells you more than a tracker subscription. Trackers measure layer 3. Most small sites are failing at layer 1 or 2, and a tracker cannot see that; it just shows zeros. Here is what I actually used as an LLM visibility checker, and what I skipped:

Tool What it measures What I found with it My verdict
curl loop (above) Crawler access, layer 1 The ligosocial.com 403 Run first, every site
Bing Webmaster Tools Bing index and impressions Six months of zero serving on contextli.com Mandatory, and it is where support tickets go
Search Console plus Bing in Claude via Murkuz Retrieval, layer 2 Which queries each site really ranks for What I use weekly
DataForSEO LLM Mentions API Tracked mentions, layer 3 0 rows on all 10 domains Useful once you are big enough to appear
DataForSEO ChatGPT scraper Does ChatGPT browse, and whom it cites Browsed on 4 of 10; cited ccattests.com once Best value for a per-query check
Dedicated trackers (Profound, Ahrefs Brand Radar, Peec AI, Otterly) Share of voice across prompts Not tested; I have not paid for any of them Worth it once layers 1 and 2 pass and you have mentions to track

My pick: if your site is under a year old or has little organic reach, start with the curl loop, Bing Webmaster Tools and the 10-query sheet below, and skip the subscription. The tracker becomes worth its price when there is something to track, which for 6 of my 10 sites there is not yet.

The 10-query LLM visibility audit sheet I now use

Copy this table, fill one row per query, and rerun it monthly. Pick queries where you already rank on Google, because that is where a missing citation is most likely a fixable problem rather than a reach problem. My real rows from 26 September are the example:

Query Your Google rank AI Overview cites you? ChatGPT browsed? ChatGPT cites you? Cited instead Bot access
linkedin comment 6 Yes No No none 403
linkedin post character limit 5 Yes No No none 403
comment examples 7 No No No none 403
linkedin about me template 6 No No No none 403
invites sent 7 No Yes No linkedin.com 403
linkedin account name 9 No Yes No linkedin.com 403
linkedin premium career pricing 10 No Yes No premium.linkedin.com 403
what is a good ccat score 11 No No No none 200
ccat percentiles 9 No No No none 200
ccat score range 10 No Yes My other site ccattests.com 200

Then act on each row with this checklist, in this order:

  1. Bot access is not 200: fix the firewall rule before touching content.
  2. ChatGPT did not browse: no page edit will help this query. Work on being mentioned elsewhere, or pick a query where it does browse.
  3. ChatGPT browsed and cited someone else: read the cited page. Compare its opening to yours, and look for a short, direct answer near the top. That comparison is the subject of my next piece, on what separated cited pages from skipped ones.
  4. AI Overview cites you but ChatGPT does not: suspect access or Bing first.
  5. Rank is outside the top 10: this is a reach problem. Fix it with normal SEO before you spend on GEO.

What does not move LLM visibility?

Three things sold as LLM visibility fixes had no measurable effect in the best evidence I could find, so I have stopped spending time on them.

A word on sources. A lot of the loudest LLM visibility advice, including Press Ranger’s long “How to Rank Anything on ChatGPT” email series, is vendor marketing for a paid product, and some of its own numbers contradict each other between issues. I read it for ideas and test nothing on its say-so. The strongest controlled evidence is still the Princeton and Georgia Tech GEO paper, which found that rewriting pages with added statistics, quotations and cited sources “can boost visibility by up to 40%” in generative engine responses. That is a content lever, and it only works after the crawler gets through.

FAQ

What are the best LLM visibility tools?

For small sites, the best LLM visibility tools are the ones that test access and retrieval first: a curl test, Bing Webmaster Tools and Search Console. For citation checks, a per-query ChatGPT scraper such as DataForSEO’s is enough. Dedicated share-of-voice trackers like Profound, Ahrefs Brand Radar, Peec AI and Otterly make sense once you have mentions to track.

How do I improve LLM visibility?

Fix it in layer order. Make sure OAI-SearchBot, PerplexityBot, Claude-SearchBot and Bingbot get a 200. Make sure Bing and Google actually serve your pages. Then improve the pages: a direct answer near the top, specific numbers with sources, attributed quotes and first-hand data.

Aren’t LLM visibility scores just vanity share-of-voice numbers?

Often, yes, for a small site. A share-of-voice number does not tell you why you are missing. The 10-query sheet does, because each row points to a layer: blocked, not ranking, not browsed, or out-written.

Is being cited inaccurately worse than not being cited?

I have not hit this myself yet, so treat this as my view. A wrong answer about your product is fixable only if the assistant can read your correct page, which brings you back to access. Make sure the crawlers get your real pricing, feature and policy pages, and state those facts plainly near the top.

Is GEO just SEO renamed?

Mostly it sits on top of SEO. Retrieval runs on search indexes, so a page that does not rank or is not served by Bing has little chance. But access rules for AI crawlers, per-engine citation checks and answer-first writing are real additions that classic SEO audits never covered.

Why does ChatGPT cite my competitor when I rank higher on Google?

In my audit, ChatGPT skipped the page ranking on Google and cited a different page with a clearer, more direct answer. Google rank is not the deciding factor. Ahrefs found that only 37.9% of URLs cited in AI Overviews also appeared in the top 10 blocks, and the overlap is lower for ChatGPT.

How often should I run an LLM visibility audit?

Monthly for the 10-query sheet, and the curl test after any change to your firewall, CDN or bot settings. Citation behaviour shifts often, so a single snapshot can mislead.

Does ChatGPT cite the same sources every time?

No, and you should not expect it to. My audit is one snapshot per query on one day, and ChatGPT scraper results can vary from run to run, including whether it searches at all. That is why the audit sheet is a monthly habit. Read the pattern across runs, not a single answer.

Where this comes from

I build and run my products solo, and my back office runs on Claude Code schedules. This audit was not a client project; it was my own ten sites, measured on 26 September 2026, with the raw responses saved. The most useful finding was the least glamorous one: my best-ranking site was invisible to ChatGPT because of a security rule, not because of anything I had written. If you only do one thing from this article, run the curl loop on your own site today. For the SEO stack behind these numbers, see why I built mine on DataForSEO instead of Ahrefs, and for how the schedules run, Claude Code best practices from running it unattended.