ChatGPT and Perplexity cite pages that a retrieval system can find, parse and quote without extra context. In practice that means crawlable HTML, paragraphs that answer one question each, specific numbers with sources, and a brand name that reads the same everywhere. Nobody can guarantee a citation. You can raise the odds and measure the change with a fixed prompt set
We work on this for clients in Dubai and on our own site, and the evidence base is thinner than most GEO marketing suggests: one peer-reviewable study, Google's public documentation, and whatever you log yourself
How do ChatGPT and Perplexity choose sources?
Both engines run a search step before they write. A model rewrites your question into one or more queries, a search index returns candidate pages, the system pulls short passages from those pages, and the model writes an answer that links back to the passages it used
Every stage of that pipeline can drop you:
- Retrieval. If a crawler cannot fetch the page, or the index never ranked it for the rewritten query, the page is not in the candidate set. Classic SEO decides this stage
- Passage extraction. The system works with chunks of a few hundred words. A chunk that opens with "as mentioned above" or depends on a chart image gives the model nothing to quote
- Selection. Among the passages it holds, the model prefers ones that state a fact plainly, carry a number or a named source, and agree with other passages it trusts
- Attribution. The answer links the passage it leaned on. If your brand appears under three different names across the web, the model may credit the claim to someone else or merge you with a competitor
Google describes the same structure for its own products. AI Overviews and AI Mode may use a "query fan-out" technique, issuing several related searches across subtopics before composing a response, according to Google Search Central's guide to AI features. OpenAI and Perplexity publish less detail about ranking, so anything specific you read about their weighting is inference from outside testing. We treat it that way too.

What does the research say?
The most cited study is "GEO: Generative Engine Optimization" by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande (arXiv:2311.09735). The authors built GEO-bench, a benchmark of queries across many domains, rewrote source pages with different methods, and measured how much of each rewritten page showed up in generated answers.
Three edits did most of the work: adding statistics, adding quotations from credible people, and adding citations to sources. Those methods lifted visibility by roughly 30 to 40 percent on the benchmark, and the abstract reports gains of up to 40 percent. Keyword stuffing, the reflex from old SEO, did not help. The authors also found that results varied by domain, so a tactic that wins for history questions may do little for legal ones
Two caveats keep us careful with this paper. It measured a research setup built on generative engines in 2023, and the commercial engines have changed since. It also measured visibility inside answers for pages already in the candidate set, so it says nothing about getting retrieved in the first place. We use it as evidence for what to write, and we keep the SEO basics for what gets fetched
In the GEO benchmark, adding a statistic, a named quote and a citation lifted visibility by roughly 30 to 40 percent, and keyword stuffing changed nothing.
Google's position is the other half of the research picture. The same Search Central page says a page must be indexed and eligible to show a snippet to appear as a supporting link, and that "there are no additional technical requirements." It also states that AI Overviews are shown only when Google's systems judge them additive, "and as such, often don't trigger." Read together, those lines rule out any promise of inclusion: no special markup, no special file, no guaranteed slot
What should you change on your site?
Start with the page types buyers ask engines about: pricing, comparisons, how a process works, and anything regulated. Then work through the factors below. We built this table from the GEO paper, Google's documentation and our own client work, and we order it by how often a fix is both cheap and missing
| Factor | Why engines care | What to do this week |
|---|---|---|
| Self-contained passages | Extraction works on chunks; a chunk with no referent cannot be quoted | Rewrite each section so its first 1 to 3 sentences answer the heading alone. Remove "as mentioned above" and "the former" |
| Statistics with sources | GEO paper: statistics and citations lifted visibility roughly 30 to 40% | Put real numbers in the copy, such as AED price ranges, timelines in weeks, and response times. Link the primary source for each external figure |
| Quotations | GEO paper: quotation addition was one of the strongest methods | Quote a named person from your team with a role, saying something specific enough to be wrong |
| Entity consistency | Attribution fails when one company has several names | Use one brand name and one descriptor in copy, schema, footer and social profiles. Match the legal entity everywhere it appears |
| Structured data that matches the page | Google asks that structured data match visible text | Add Organization, Person, BlogPosting, FAQPage and BreadcrumbList as JSON-LD. Never mark up content the reader cannot see |
| Original data | A table nobody else has is a passage only you can supply | Publish one artifact per key page: a price table, a field map, a benchmark with method and date |
| Freshness | Engines answer fast-moving questions from recent pages | Show the publish date, and change the modified date only when the content changed |
| Crawl access | No fetch, no candidate | Keep robots.txt permissive for GPTBot, ClaudeBot and PerplexityBot unless you have a reason to opt out |
A few of these need more than a table row
Entity consistency matters more in the UAE than most owners expect. A Dubai brokerage often trades under a marketing name, holds a RERA-registered legal name, lists under a third variant on Property Finder and Bayut, and has a fourth spelling in Arabic. When an engine sees four names, it may treat them as four companies. We pick one canonical name and one descriptor, then make the site, the schema Organization node, the portal profiles and the founders' LinkedIn pages agree. Our own footer states PHII Labs, the legal entity PHII LABS (FZC) and the SRTIP free zone licence in the same words our schema uses
Original data is where a small company can beat a large one. Global agencies own the definitions of GEO and AEO, and you will not out-explain them. They do not have an Ejari, SPA and NOC field map from a Dubai brokerage pipeline, or AED budget ranges from builds delivered in the UAE. Our document automation guide and our checklist for picking an agency each carry one artifact for that reason. An engine answering a Dubai-specific question has few other places to find those passages
Structured data deserves a clear expectation. FAQPage and Organization markup make your content machine-readable and keep entities unambiguous. They do not move a weak page into an answer. Google's guidance lists "making sure your structured data matches the visible text on the page" alongside internal links and page experience as ordinary SEO hygiene, and the schema vocabulary itself is documented at schema.org. We add it because it is cheap and correct, and we never sell it as the lever.
llms.txt is an experiment. We publish one at /llms.txt: a plain-text file listing our canonical facts, key pages and crawler policy. It took under an hour to write. Some AI crawlers may read it. Google does not use it for AI Overviews, since its guidance requires nothing beyond normal indexing. We will report whether it changes anything once our prompt logs show a difference. Until then it stays in the "cheap, unproven" column
How do you measure whether it works?
Run the same prompts against the same engines on a fixed schedule, log which domains each answer cites, and compare the logs over time. Rankings tools do not cover this, and screenshots of one lucky answer prove nothing because answers vary between runs
Our method, which you can copy:
- Write 20 to 40 prompts in the words buyers use. Pull them from sales calls, WhatsApp enquiries and Search Console queries. For a Dubai brokerage that means prompts like "best property management company in Dubai Marina" and "how long does Ejari registration take", not your brand name
- Freeze the set. Changing prompts between runs breaks comparison. Add new prompts as a separate batch with its own start date
- Pick engines and settings. We use ChatGPT with search on, Perplexity, Gemini and Google AI Overviews where they trigger. We log in from a clean session and note the location, because results from a UAE IP differ from results in Europe
- Run each prompt three times per engine per date. Generative answers are not deterministic, so a single run tells you very little. Three runs show whether a citation is stable or a coin flip
- Log every cited URL, including competitors'. The competitor list is often more useful than your own score
- Change one thing at a time on the site, note the date, and wait. Recrawls take days to months, so we compare monthly runs, not daily ones
The log is a flat table. One row per citation:
run_date | engine | prompt_id | run | position | cited_domain | cited_url_path
2026-10-01 | perplexity | P07 | 1 | 2 | example-brokerage.ae | /guides/ejari-registration
2026-10-01 | perplexity | P07 | 2 | 1 | example-brokerage.ae | /guides/ejari-registration
2026-10-01 | perplexity | P07 | 3 | 4 | competitor.com | /blog/ejari
2026-10-01 | chatgpt | P07 | 1 | 3 | example-brokerage.ae | /guides/ejari-registration
From that table we report two numbers per engine: the share of prompts where the client is cited in at least two of three runs, and the median position when cited. We also list the five domains cited most often. The rows above are illustrative; real client logs stay private
Alongside the prompt log we watch three cheaper signals. Search Console includes AI Overview and AI Mode appearances in the regular Performance report under the "Web" search type, per Google's documentation. Server logs show GPTBot, ClaudeBot and PerplexityBot fetches, which tells you whether the crawlers reach the pages you changed. Analytics referrers from chatgpt.com and perplexity.ai show whether citations turn into visits
What results can you honestly expect?
Expect slow, partial movement on the queries where you published something specific, and little movement on generic ones. Nobody controls the engines, so the claim we make to clients is narrow: we will apply changes the evidence supports, and we will show you the before and after logs
The closest result we can point to is classic SEO. On the Trusted Real Estate website we built a chatbot that pre-qualifies buyers and generates PDF contracts, and organic SEO traffic grew 21 percent while lead-to-call time fell 17 percent. That was search traffic, not AI citations, and we report it as such. It matters here because retrieval is stage one of every AI answer, and pages that Google already trusts are the pages AI engines find first
What we will not do is sell "guaranteed ChatGPT placement", in AED or any other currency. If an agency offers it, ask for their prompt set, their run dates and their raw logs. A method you can audit is worth more than any promise, and the differences between SEO, AEO and GEO are mostly differences in what you measure
The takeaway
Engines cite passages they can retrieve, extract and attribute. Make each section answer its heading in the first two sentences, put sourced numbers and named quotes in the copy, keep one brand name everywhere, and publish at least one table nobody else has. Then measure with a frozen prompt set, three runs per engine, logged monthly. Our AI search visibility service runs that process
If you want to know where your brand stands today across ChatGPT, Perplexity, Gemini and AI Overviews, book a free audit. We send the prompt set, the logs and a list of the first five changes worth making
FAQ
Does FAQ schema help with answer engine optimization?
It makes Q&A machine-readable and eligible for rich results. It is one signal among many; passage quality and source authority matter more
What structured data should I use for AI search?
Organization, Person, BlogPosting, FAQPage and BreadcrumbList, consistent with visible content. Schema clarifies entities; it does not guarantee citation
Does Google use llms.txt?
Google states that AI Overviews need no special files and does not use llms.txt today. Publishing one is a cheap experiment, not a ranking lever
How do I allow AI search crawlers in robots.txt?
A permissive wildcard allow rule already covers GPTBot, ClaudeBot and PerplexityBot. Explicit allow rules are only needed if you block by default
How do you measure AI-search visibility without promising inclusion?
Run a fixed prompt set monthly across engines, log cited domains and positions, and report the table. We report methods and measured results and never guarantee inclusion
Sources
- Google Search Central's guide to AI featuresdevelopers.google.com
- arXiv:2311.09735arxiv.org
- schema.orgschema.org
