Method
How AfterLaunch measures AI search visibility
Every number AfterLaunch publishes about AI search comes off one corpus of answers, collected the same way every time. This page says exactly how, so that anyone reading a figure from us can decide for themselves whether it means anything.
It is versioned. When the method changes, the version changes, and figures stay attached to the version that produced them.
1. The engines
Each one is queried directly. Nothing is scraped from a search results page.
| Engine | What is queried |
|---|---|
| ChatGPT | the model, with browsing available |
| Google Gemini | the model, with Google grounding |
| Google AI Overviews | the overview block returned for a query |
| Perplexity | the answer endpoint |
Every engine is asked the same question in the same wording on the same day. Answers are stored whole, with their citations, at the moment they are given. The corpus is not re-derived later from a cache.
2. The questions
Fifteen questions per company per wave. They are generated from what the company actually sells, not from its brand name, because a buyer who already knows your name is not the buyer worth measuring. The questions read like a prospect's: what tool does X, what is the best Y for Z, how do people solve W.
A brand-name question tests recall. A category question tests discovery. We measure discovery.
Where the question comes from, and why it is versioned
The source text behind a question changes the answer more than any other choice in this method, so it is versioned and stated on every figure.
Questions are built from the company's live homepage text: title, meta description, headings and lead paragraphs as they read on the day of the measurement. The brand name is removed before generation. The generator names the concrete category a buyer shops for and ignores mission statements and slogans.
Measured 2026-08-22: same 90 companies, same engines, same detection rule, only the question source changed.
| Question built from | Young named | Established named |
|---|---|---|
| A directory one-line description | 2.7% | 27.8% |
| The live homepage | 18.9% | 62.5% |
20 companies moved from unnamed to named, 2 the other way, 68 unchanged. The directory descriptions were slogans, and a slogan does not name a category. One firm describing itself as "people infrastructure for the future of work" sells background checks.
3. The cohort
Figures come from a fixed, published list of companies, re-measured on a schedule. A fixed cohort is what makes change over time measurable at all: a rolling sample cannot tell you whether the engines moved or the sample did.
The cohort has two legs, and the second one is the point.
- The young leg. Recent-batch startups, shipped within roughly the last three years.
- The established leg. Companies the engines have known for years.
The established leg is not the sample. It is the control. Almost every published statistic about AI search visibility is computed over well-known brands, which makes it an upper bound presented as an average. Running both legs through the identical instrument on the identical day turns that bias into the measurement: the gap between the legs is the finding.
Cohort v1 is 102 companies, 80 young and 22 control, drawn by a seeded random draw from Y Combinator's public directory and stratified by YC's own industry fields and by batch year. Both legs come from the same institution, which holds founder quality and funding access roughly constant. The seed and the full list are published alongside the cohort.
4. What counts as a mention
A mention is the company being named in the answer text. Detection is deliberately recall-first: where a judgement is close, the answer is counted as a mention.
The reason is asymmetric cost. An undercount tells a company it is invisible when it is not, which is a false alarm it can act on and disprove in a minute. An overcount tells a company it is visible when it is not, which is a false comfort it cannot detect at all. Between the two errors, we take the loud one.
5. What counts as a citation
A citation is a source the engine attached to its answer. Citations are read from the structured response each provider returns, never by looking for links in the answer prose.
This distinction is not pedantry, it decides the number. Some engines write sources inline in the text and some return them only as structured data alongside it. Counting links in the prose makes an engine that cites on every single answer look like it never cites at all. Inline versus structured is a formatting convention. It says nothing about engine behaviour.
6. Repeatability
Ask an engine the same question twice and you will not always get the same answer. Measured 2026-08-23: ten companies, identical query, three repeats, both engines.
| Engine | Same answer all three times |
|---|---|
| Perplexity | 10 of 10 |
| Google AI Overviews | 7 of 10 |
| Combined verdict | 9 of 10 |
All three disagreements ran the same way: named on the first ask, absent on the second and third, with a full Overview rendered each time.
What follows regardless: one question is not a reliable verdict on one company. Fifteen questions turns a coin flip into a rate, and a rate is stable where a flip is not.
7. What is excluded, and why
Exclusions are part of the method, so they are stated here rather than in a footnote.
- Gemini is included, read from the right field. Gemini returns citations as links to vertexaisearch.cloud.google.com, Google's own redirect wrapper, rather than as publisher URLs. Read the URL alone and 8,308 of Gemini's 8,311 citations look like they point at Google. They do not: every one of those citations also carries a title, and on 8,308 of 8,311 that title is the publisher's bare domain. We take the domain from there.
- Page-level Gemini data is captured at scan time or not at all. The redirect resolves to the full article URL for roughly thirty days and then returns 404. Figures published from answers older than that window are domain-level for Gemini.
- Answers without stored text are excluded. The corpus holds rows for calls that failed and for a prompt version that did not persist answer text. Only rows carrying an answer count. As of 2026-08-22 that is 2,759 answers out of 5,001 stored rows, and the usable corpus begins on 2026-07-19.
- Answers not attached to a cohort company are excluded from per-company figures. They remain in aggregate source analysis, where the question is what the engines cite rather than who they name.
8. Reporting rules
- Every figure names the corpus it came from, its size, and its date.
- Any figure computed over established brands is labelled a ceiling, not an average.
- Row counts are never reported as answer counts.
- Nothing is published from a doc or a cached summary. Figures are recomputed against the corpus at the time of writing.
9. Checking us
The cohort list and the question set are published. The engines are named, the dates are on every figure, and the exclusions are above. Anyone with access to the same engines can ask the same questions on the same companies and see whether they get what we got.
If a figure of ours does not reproduce, that is worth knowing, and we would rather hear it than not.
10. Versions
- 1.2, 2026-08-23. Question construction versioned in (built from live homepage text, brand redacted, category-anchored) after a paired test showed the source text moves the result more than any other choice. Repeatability added as a stated property.
- 1.1, 2026-08-22. Cohort rebuilt as a seeded draw from YC's public directory. Gemini moved from excluded to included in source analysis.
- 1.0, 2026-08-22. First published method. The engine set named above, 15 questions per company, two-leg cohort of 20, recall-first mention detection, structured citation extraction, Gemini excluded from source analysis.
AfterLaunch runs this measurement continuously for the companies it works with. The method above is the same one behind every figure on this site.