How to Get Cited by ChatGPT: What I Learned Scoring 45 Real Pages

You rank on Google, your article is better than the one ChatGPT quotes, and ChatGPT still links someone else. Most advice on how to get your content cited by AI is a list of nine tips with no way to tell which ones matter. I wanted to know which ones matter, so I measured it.

To get cited by ChatGPT, three things must be true: ChatGPT searches the web for the question, its crawler OAI-SearchBot can read your page, and your page opens with a 40 to 60 word passage that answers the question on its own, backed by specific numbers and first-hand evidence. In a scorer I built and calibrated on 45 real pages, 23 that ChatGPT or Google’s AI Overview cited and 22 they skipped, the strongest single signal was a quotable answer capsule in the first 300 words. Then came specificity, scannable structure, attributed quotes and first-hand evidence. Outbound links and freshness made no difference in my sample. Below are the numbers, the test that failed, the capsule template, and the command I run before I publish.

Does ChatGPT even search the web for your question?

Often it does not, and when it answers from memory it cites nobody. On 26 September 2026 I ran 10 informational queries where my own sites already rank on Google through DataForSEO’s ChatGPT scraper. ChatGPT browsed the live web on only 4 of the 10. On the other 6 it answered from training data with zero citations, whoever ranked first.

So the first step is to check whether your query is one ChatGPT searches for. Ask it the question in a fresh chat and look for a sources list. If there is none, rewriting your page will not change that answer. You either pick a query where it does search, or play the slower game of being mentioned across many independent sites so the model knows you. I am doing the first and not pretending the second is quick.

Can ChatGPT’s crawler read your page?

If OAI-SearchBot gets a 403, you cannot be cited in ChatGPT search, no matter how good the page is. OpenAI puts it plainly in its crawler documentation: “Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers.”

I found this out on my own biggest site. ligosocial.com ranks for 145 keywords in Google’s top 10 and had been cited by Google’s AI Overview, yet a Cloudflare challenge was returning a 403 to OAI-SearchBot, ChatGPT-User, GPTBot, ClaudeBot, Claude-SearchBot and PerplexityBot, while Googlebot and Bingbot got through. Even robots.txt was behind the challenge. None of my other nine sites had the problem.

Test yours in ten seconds. Replace the URL with one of your articles:

curl -s -o /dev/null -w '%{http_code}n' -A "Mozilla/5.0 (compatible; OAI-SearchBot/1.0)" https://yoursite.com/your-article/

Anything other than 200 is your first fix. curl only imitates the bot, so confirm a 403 in your firewall’s event log, then allow the search crawler. In robots.txt that is a User-agent: OAI-SearchBot group with Allow: /. OpenAI notes a robots.txt change can take about 24 hours to reach its systems.

What separated the pages AI search cited from the ones it skipped?

The cited pages answered the query in a tight, quotable passage early, made more specific claims, and were easier to scan. I measured this with a scorer I wrote in Python. The fuzzy judgments (“is there a passage that answers the query?”, “is this first-hand?”) are made by Jev, TypeSafe’s fast judge model, which returns a probability rather than prose. The mechanical ones (numbers per 300 words, headings, lists) are plain code.

For the test set I took 12 queries across 12 niches, from “how to improve email deliverability” and “best password manager” to “how does compound interest work” and “what is a ccat test”. A page counted as cited if it appeared in ChatGPT’s sources or in Google’s AI Overview references, and as not cited if it ranked in Google’s top 10 but appeared in neither. That gave 45 pages: 23 cited, 22 not. Because every query contributes both kinds, the comparison is mostly like-for-like on topic.

The number that matters is AUC: the chance that a random cited page scores higher than a random skipped one. 0.5 is a coin flip.

Signal AUC Avg score, cited Avg score, skipped
Answer capsule in first 300 words 0.74 0.67 0.44
Specific claims (names, numbers, versions) 0.72 0.87 0.71
Scannable structure (H2s, question headings, lists, tables) 0.69 0.75 0.61
Attributed quotations 0.67 0.26 0.19
First-hand evidence or original data 0.67 0.44 0.27
Statistics density 0.61 0.87 0.72
Outbound links to sources 0.51 0.59 0.60
Freshness (visible, recent date) 0.52 0.54 0.49
Full score 0.74 0.67 0.55

Two more results. The full score with the judge model reached an AUC of 0.74, against 0.62 with the code checks alone, so most of the signal lives in the questions a regex cannot answer. And when I compared a cited page directly with a skipped page for the same query, the cited one scored higher in 67% of 39 pairs.

Be clear about the limits. 45 pages is small, all US English, one snapshot on one day. I did not measure internal links, images or paragraph counts. This is a strong edit guide, not a citation predictor. It lines up with the one controlled study in the field, Aggarwal et al.’s GEO paper, whose authors found that GEO “can boost visibility by up to 40% in generative engine responses”, with added statistics, quotations and cited sources among the strongest rewrites.

How does this compare with the big “cited page” studies?

The best-known numbers describe what cited pages look like, not what separates them from skipped ones. Jaclyn Ranere’s summary of Evertune’s analysis of 33,000 ChatGPT-cited URLs gives the medians: about 941 words, 18 paragraphs, 4 H2s, 28 internal links, 15 external links and 10 images. Search Engine Land built its article on the content traits LLMs quote most on the same research.

Those medians are a useful shape to aim for. But a median of cited pages cannot tell you whether skipped pages look the same, and in my sample, on outbound links, they did. That is why I compared cited and skipped pages for the same query. I did not measure internal links or images, so I cannot confirm or contradict those two numbers. Treat them as a companion check alongside the capsule test.

Why did my first “answer-first” test fail?

Because “answer first” does not mean “the first sentence answers it”. My first version of the question asked the judge, “do the first two sentences answer the query?” On the 45 pages it scored an AUC between 0.42 and 0.47 across my test runs, worse than a coin flip. Cited pages often open with a sentence or two of context, just like this article does.

What actually separated them was whether a short, self-contained passage anywhere in the first 300 words answered the query completely, so an assistant could lift it as the answer. Rewording the question to that took the signal from below 0.5 to 0.74. One caveat I owe you: I picked the better question on the same 45 pages, so 0.74 is optimistic.

The lesson I took for my own writing: stop obsessing over sentence one. Put a bold, complete, 40 to 60 word answer somewhere in the opening screen, and make sure it still makes sense if it is the only thing a reader sees.

How do you write an answer capsule ChatGPT can quote?

Write one short paragraph that answers the exact query on its own, with the query’s words in its first sentence, a concrete verdict or steps, and no links inside it. Here is the template I now use:

[Query restated as a statement]: [the direct answer in one sentence].
[The 2-3 conditions, steps or numbers that make it true, as one sentence].
[One specific proof point: a number, a date, a named source or your own result].

And a filled example, for a query like “what is a good ccat score”:

A good CCAT score is one that clears the cutoff for the job you applied to, which employers set per role rather than Criteria setting one pass mark. The raw score is out of 50 questions, and hiring teams usually read it as a percentile against other candidates. Check the role’s benchmark before you decide what “good” means for you.

Rules I follow:

  1. 40 to 60 words. Long enough to be complete, short enough to quote whole.
  2. No links inside the capsule. Put sources in the next paragraph, so the quotable part stands alone.
  3. No “as mentioned above”, no unexplained “it” or “this”. The capsule must survive being cut out of the page.
  4. Every H2 gets a mini capsule. Open each section with one or two sentences that answer its heading. My scorer checks this separately, because an assistant can quote a section rather than the intro.
  5. Keep it true to the body. A capsule that promises more than the article delivers is the fastest way to lose a reader who clicks through.

If you write with Claude, the SEO bundle in the Locul skills library includes the SERP-first article writer I use for first drafts. I still write the capsule myself.

What made no difference in my sample?

Outbound links and freshness were flat: AUCs of 0.51 and 0.52, a coin flip. Cited and skipped pages linked out at almost the same rate, and a recent date did not separate them. Crawl access showed no variance at all, because every page in the set was crawlable. I chose them that way, and that is exactly how ligosocial.com dropped out.

I still link sources and keep dates current, for two reasons. Bigger studies than mine back them, and they help readers. But I no longer treat “add more outbound links” as a citation fix.

Two things sold as GEO fixes did not earn my time either. Ahrefs tracked 1,885 pages adding schema and saw AI citations barely move. And SE Ranking’s look at roughly 300,000 domains found no relationship between llms.txt and how often a domain is cited. Schema scored 0.60 in my set, but I read that as bigger publishers having both better schema and more citations, not one causing the other.

The edits I am making to my own articles

The clearest case in my own portfolio is a pair of CCAT pages. For “ccat score range”, my personal blog was the page ranking on Google. ChatGPT searched, skipped it, and cited a CCAT scoring guide on ccattests.com, another site of mine that did not even show in DataForSEO’s Google top 10. That cited page scored 65.8 on my scorer: 2,129 words, 8 H2s, a named byline, updated two days before the test, and a direct scoring answer near the top.

So these are the edits I am working through on my older articles, starting with the CCAT percentile guide and the other posts that rank but are not cited:

  • [ ] Add a bold 40 to 60 word capsule for the page’s main query in the first 300 words.
  • [ ] Turn vague H2s into the questions people ask, and open each with a one or two sentence answer.
  • [ ] Replace generic claims with named tools, versions and numbers, each with a source link.
  • [ ] Add one or two attributed quotes from primary sources, such as the vendor’s own documentation.
  • [ ] Add what I actually did or measured, with dates, where I have it.
  • [ ] Re-score, and only republish at 80 or above.

This is in progress. I have not measured a before and after on ChatGPT citations yet, and I will not guess at one. I will add the result here once the edited pages have been live for a few weeks. What I can say is that this article went through the same gate before I published it.

To choose which pages to edit first, I pull Search Console into Claude through Murkuz’s Search Console connector and ask for queries where a page sits in positions 5 to 15 with real impressions. Those are the pages close enough to be retrieved but not yet the obvious answer. The setup is in I connected Google Search Console to Claude.

How can you score your own page before you publish?

Score the draft against the signals in the table above, and fix the capsule before anything else. This is the exact command I run on every draft. The script is my own internal tool, so the command is here for completeness:

python3 geo_score.py draft.md --url https://yoursite.com/your-slug/ --keyword "your focus keyword"

It prints a 0 to 100 score, a line per criterion, and a fix for each failure. The weights I settled on after calibration give the most points to the capsule (12), statistics density (11), specific claims (10), then attributed quotes, structure and first-hand evidence (9 each). Schema gets 2, because I found no evidence that it moves citations.

You do not need my script to use the idea. The portable part is the judge questions. Give any capable model the first 300 words of your draft and ask these one at a time, each expecting a yes or no:

  1. Within this text, is there a short self-contained passage (one to three sentences, or a tight list) that directly and completely answers “[your query]”, such that an AI assistant could quote it as the answer?
  2. Does each section’s opening answer its heading in one or two sentences that make sense without the rest of the article?
  3. Does the article report the author’s own test, measurement, dataset or dated real-world result with specifics?
  4. Do most claims name specific tools, versions, numbers, settings or worked examples?

If question 1 is a no, fix that and nothing else first. It was the single biggest gap between cited and skipped pages in my sample.

FAQ

Does ChatGPT crawl my site directly, or does it use Bing?

ChatGPT search fetches pages with OAI-SearchBot, and live user requests come from ChatGPT-User. It is widely reported to lean on Bing’s index to find candidates, but I have seen no independent test of how much. In practice, make sure both OAI-SearchBot and Bingbot get a 200, and that Bing actually serves your pages in its results.

Is getting cited by ChatGPT just SEO renamed?

Partly. Retrieval runs on search indexes, so a page nobody can find will not be cited. But in my 45 pages, Google rank alone did not decide it. Ahrefs found that only 37.9% of URLs cited in AI Overviews also appeared in the top 10 blocks. The capsule and specificity signals are extra work that normal SEO checklists do not ask for.

Why does ChatGPT cite me in one chat and not the next?

Because it does not browse on every run, and the sources it pulls vary between runs. My test was one snapshot per query. Check the same queries monthly and read the pattern, not a single result.

Is there proof beyond anecdotes?

The strongest controlled evidence is the Princeton and Georgia Tech GEO study. My own 45-page calibration is small but real: cited and skipped pages from the same queries, scored by the same rubric, with the numbers in the table above. Treat any single “5 steps” post, including this one, as a hypothesis to test on your own pages.

Isn’t this just another “5 steps to get cited” list?

The steps are similar to other lists. The difference is that each one here comes with a measured result from cited versus skipped pages, including the ones that made no difference. If a tip is not in my table, I have no evidence for it either way.

Do I need to rank number one on Google to get cited?

No. In my own portfolio ChatGPT cited a page that was not in Google’s top 10, and skipped the one that was. Ranking helps you get retrieved; the passage quality decides who gets quoted.

Do schema markup or llms.txt help you get cited by ChatGPT?

Not measurably, in the best evidence I found. Keep schema for normal SEO. Skip llms.txt unless it costs you nothing.

Can you pay to get cited by ChatGPT?

Not that I have found. What vendors sell is press releases and listicle placements that might be picked up as sources. Press Ranger’s long email series on ranking in ChatGPT is a good example: interesting ideas, but vendor marketing for its own distribution product, with numbers that change between issues.

Where this comes from

I build and run my products solo, and most of my back office, including the blog pipelines that publish this site, runs on Claude Code schedules. I built the scorer because I wanted to know which of the usual tips actually separate cited pages from skipped ones before I rewrote my own articles. The biggest surprise was my own first question scoring worse than a coin flip, which is why I now trust measured signals over anyone’s list of tips, mine included. The capsule template and the four judge questions above are the parts I use on every draft. For how the scheduled pipelines run, see how I use Claude Code to run a business.