This page is the study design, not the findings. The results tables below are deliberately empty and every one is marked pending. Nothing here should be read, quoted or cited as a finding yet.
It is set to noindex until real data replaces the placeholders. The methodology is published first on purpose — pre-registering how you will measure, before you look, is what stops a study quietly becoming whatever the data needed to say.
120 commercial prompts across ChatGPT, Gemini, Perplexity and Google AI Overviews — measuring who gets mentioned, who gets cited with a link, and what those sources have in common.
Share the methodology while collection is still open.
Short on time? Open this article in an answer engine and have it summarised for you.
There is a great deal of published advice about AI search visibility and remarkably little published measurement — particularly for the markets I work in. Almost every benchmark I can find is US-centric, and almost none of them publish their prompt set, which makes the results impossible to check or reproduce.
So the niche here is deliberate: India and the UAE, commercial intent, open methodology. Narrower than a global study, but genuinely uncontested, and far more useful to the businesses actually buying these services.
When a real buyer in India or the UAE asks an AI assistant a commercial question, which sources get cited, and what do those sources have in common? Not "how do I rank in ChatGPT" in the abstract — what is observably true across a fixed prompt set, tested identically on each platform.
| Platform | Surface | Mode |
|---|---|---|
| ChatGPT | ChatGPT Search | Web search enabled, logged out, memory off |
| Google Gemini | Gemini app | Default model, fresh session |
| Perplexity | Perplexity Search | Default, not Pro/focus mode |
| Google Search | AI Overviews / AI Mode | Incognito, location set per market |
Anything that lets results drift is a threat to the finding, so these are fixed in advance:
Twenty prompts per vertical, split across three intent stages so the study can separate discovery from decision:
| Vertical | Discovery (8) | Comparison (6) | Decision (6) |
|---|---|---|---|
| SEO / digital marketing | "how do I improve my website's Google ranking" | "SEO agency vs freelance SEO consultant" | "best SEO consultant in Indore" |
| Website design & development | "how much does a business website cost in India" | "Webflow vs custom website for a small business" | "who can build my website in a week" |
| Local services | "how do I get my business on Google Maps" | "local SEO vs Google Ads for a service business" | "Google Business Profile expert in Dubai" |
| eCommerce | "why is my Shopify store not getting traffic" | "Shopify SEO vs WooCommerce SEO" | "eCommerce SEO agency for a 10,000 SKU catalogue" |
| SaaS | "how do SaaS companies get organic signups" | "content marketing vs paid ads for SaaS" | "SaaS SEO consultant India" |
| Marketing automation | "how do I automate marketing reporting" | "Make.com vs n8n vs Zapier" | "who can build a GA4 reporting automation" |
The full 120-prompt list ships with the results as a CSV, so anyone can reproduce or challenge the study. A benchmark whose inputs are secret is marketing, not research.
| Field | Values | Why it is measured |
|---|---|---|
| Brands mentioned | List | Mention without a link still shapes the buyer |
| Sources cited | URLs | The actual citation, the thing being competed for |
| Domain type | Brand / Reddit / YouTube / publisher / directory / forum | Tests whether owned sites win at all, or aggregators dominate |
| Ranks in Google top 10? | Yes / No / Position | Tests how tightly citation tracks classic ranking |
| Content format | Guide / listicle / product / comparison / forum thread | Which formats get lifted |
| Freshness | Published / modified date | Whether recency correlates with citation |
| Entity signals | Org schema, sameAs, consistent NAP | Whether entity clarity correlates with citation |
| Word count | Integer | Depth vs brevity |
| Has original data? | Yes / No | Tests the information-gain hypothesis directly |
| Third-party mentions | Count | Off-site authority |
One row per prompt per platform — 480 rows at full coverage:
market,vertical,intent,prompt_id,prompt,platform,run_date,
brands_mentioned,sources_cited,domain_type,google_rank,
content_format,published_date,modified_date,has_org_schema,
has_sameas,nap_consistent,word_count,has_original_data,
third_party_mentions,screenshot_ref,notes
Written down in advance so they can be wrong:
Stated up front because a benchmark that hides them is not worth citing. Single-run sampling means model non-determinism is uncontrolled — the same prompt can return different sources. Results are a snapshot of one collection window, not a trend. 120 prompts across 6 verticals is a small n per vertical. Logged-out testing removes personalisation, which real users have. None of this invalidates the study; it bounds what it can claim.
Once populated, this becomes the thing almost no AI-search article currently has: observed data from a stated market, with a published prompt set anyone can rerun.
Whether it confirms the hypotheses or contradicts them is genuinely open — and a study that can only confirm what its author already sells is not a study. If H1 fails and cited pages routinely do not rank, that is a more interesting finding than if it holds.
When results publish, the prompt CSV, the scoring rubric and the aggregated dataset publish with them. Rerun it, disagree with it, or run it for your own market — the design is deliberately reusable.