schedule a call
← All posts

The AEO Measurement Framework: Tracking Mentions in ChatGPT, Gemini, Perplexity

June 24, 2026by Marco CoronadoASO & SEO
Marketing analyst comparing brand mention reports across ChatGPT, Gemini, and Perplexity on a multi-monitor setup.

Answer Engine Optimization is the new layer above SEO, and most companies measuring it are measuring the wrong things. They track aggregate share-of-voice numbers from new AEO platforms that show pretty dashboards but cannot tell you whether your brand actually gets mentioned when a buyer in your category asks a real question. The result is AEO investment that produces no measurable lift because nobody knew what lift to look for.

This article is the AEO measurement framework Semnexus uses with clients in 2026. It covers what to measure, how to sample without overspending on API tokens, how to track changes over time, and how to act on the data when a brand mention drops or rises.

What AEO measurement actually looks like

AEO measurement has three layers. Most companies stop at layer one.

Layer What it measures What it tells you
Mention presence Are you cited? Are you visible at all?
Mention quality Are you cited correctly and positively? Is the visibility valuable?
Mention drivers Which sources feed the mention? What to do next?

Tracking only layer one — presence — produces a number that goes up and down with no actionable lever to pull. The measurement work is in layers two and three.

The three answer engines that matter in 2026

ChatGPT, Gemini, and Perplexity dominate consumer and prosumer answer engine usage in 2026. Each has different retrieval behavior and produces different mention patterns.

  • ChatGPT retrieves from a curated index plus live search. Mention behavior is consistent run-to-run but varies by prompt phrasing. ChatGPT mentions skew toward authoritative sources and named brands.
  • Gemini integrates with Google Search and Google's knowledge graph. Mention behavior favors entities that have strong Google entity presence, fresh content, and structured data.
  • Perplexity retrieves heavily from live search and shows citations explicitly. Easiest to measure because citations are visible per response.

Tracking all three is the right baseline. Anthropic's Claude is increasingly used in B2B contexts and is worth tracking as a fourth where the audience justifies it.

What to measure

The minimum AEO scorecard has five metrics.

1. Mention rate by category prompt

The share of category-defining prompts that mention your brand. The category prompts are the 20 to 50 questions a typical buyer in your space might ask an answer engine. Examples for a mobile-app marketing agency:

  • "Who are the best mobile app marketing agencies?"
  • "How do I market a new mobile app?"
  • "What does an ASO agency do?"
  • "Best app store optimization companies for SaaS"

Your mention rate is the share of these prompts that name your brand at least once.

2. Citation rate by category prompt

How often the engine cites your domain as a source, even when it does not mention you by name. Especially relevant for Perplexity, where citations are explicit. A high citation rate with low mention rate means your content is informing the answer but not getting brand credit.

3. Mention position

When mentioned, where in the response do you appear? First brand mentioned is meaningfully more valuable than fourth. Track position as a numeric average.

4. Mention sentiment

When mentioned, is the description positive, neutral, or negative? An engine that mentions you as "outdated" or "expensive" is doing damage, not value.

5. Source attribution

Which sources feed the mention? A mention sourced from your own pillar page is repeatable; a mention sourced from a single Reddit thread can disappear without warning.

How to sample without overspending

Running every category prompt against three engines every day is overkill and expensive. The right cadence:

  • Weekly: Run all category prompts against all three engines, capturing the full response.
  • Daily: Sample a rotating subset of 5 prompts per engine to catch sudden drops.
  • Per content update: Re-run the affected category prompts within 7 days of any pillar content change.

For a 30-prompt category set, weekly sampling produces 90 responses to analyze per week (30 prompts × 3 engines). That is manageable manually and trivial through API.

API approach: use the official OpenAI, Google AI, and Perplexity APIs where available. Token cost for the weekly run for 30 prompts is typically $20 to $80 depending on model selection. For prompts that need browse or live-search behavior, use the engines' search-enabled modes; this raises cost modestly.

For Gemini and ChatGPT, the public consumer experience differs slightly from the API. Cross-check at least once per month by running a sample manually through the consumer interface.

How to act on the data

Measurement without action is theater. The data should drive one of three responses.

Response A: A category prompt has 0% mention rate

Your brand is not in the engine's retrieval index for this question. The fix is content. Publish a pillar page that directly answers the question, get external citations to that page (PR, podcasts, niche communities), and wait 30 to 60 days for the engines to recompute. Recheck.

Response B: A category prompt has 50%+ citation rate but low mention rate

You are informing the answer but not getting brand credit. The fix is in the content style. The engines extract attributable claims from sources. If your content phrases insights in attribution-friendly ways ("Semnexus reports that..." or "according to a Semnexus analysis...") the engines are more likely to surface your name with the citation.

Response C: Mention sentiment is negative

This is the most damaging signal in AEO. Investigate the source. Often a single outdated review, comparison page, or forum thread is feeding the negative description. Address the source: get the page updated, push fresher content with the correct framing, or engage with the community.

Tools and stack for 2026

The AEO measurement stack is still maturing. As of 2026:

  • Direct API approach. OpenAI, Google AI, Anthropic, and Perplexity APIs. Cheapest and most flexible. Requires a small custom analyzer for mention extraction.
  • Vertical AEO platforms. Several 2026 vendors (Profound, Otterly, Athena AI, and emerging competitors) handle prompt orchestration, mention parsing, and reporting. Useful at scale but verify their methodology matches yours.
  • Hybrid. Most mature teams in 2026 run a custom prompt set through the APIs and use a vertical platform for trend visualization. The custom set captures category-specific prompts the platform may not surface.

Avoid platforms that only show "AI share of voice" with no per-prompt breakdown. Share of voice without prompt-level resolution is not actionable.

What good AEO measurement looks like at maturity

A mature AEO measurement function inside a company has:

  • A defined category prompt set updated quarterly
  • Weekly automated runs against all three engines
  • A dashboard showing mention rate, citation rate, position, and sentiment per prompt
  • A monthly review meeting with content, PR, and product to action the data
  • A budget line for AEO content investments tied to specific prompt gaps

This is not a heavy lift. A single analyst plus an engineer can stand up the function in 4 to 6 weeks and run it indefinitely.

Frequently asked questions

How long does it take for AEO measurement to show actionable signal? The first weekly run is a baseline; the first month is calibration; meaningful signal emerges in months 2 and 3. Below that timeframe, the data is too noisy to act on.

Does ranking in Google still matter for AEO? Yes. Google search ranking and Google entity presence are major inputs to Gemini's retrieval. Strong traditional SEO compounds with AEO; weak SEO caps AEO.

Should I measure AEO across non-English languages? For most companies, English is enough. If your business has significant non-English customer base, sample at least one major language per quarter.

How does AEO measurement integrate with SEO measurement? They share a small overlap (organic discovery, brand mentions) but the metrics differ. AEO measures mention quality in generative answers; SEO measures ranking position in search results. Both inform a unified visibility scorecard.

What is the right team to own AEO measurement? Usually content marketing or SEO, with engineering support for the measurement plumbing. PR and product should be consumers of the data, not owners.


If you are setting up AEO measurement for the first time or want a second opinion on your current setup, the AEO marketing team at Semnexus runs framework and measurement engagements as part of every AEO project. The website marketing team handles the content side once the measurement layer shows which prompts to address.

lets connect

SEM Nexus is ready to help you find unique solutions for your app. Get in touch to learn more about your project and receive the full SEM Nexus treatment.

By partnering with SEM Nexus, you can confidently launch your app and get your product into the hands of customers, achieving unparalleled mobile growth.

get in touch now!
breaker
logo 98 Cuttermill Road STE 223N,
Great Neck, New York, 11024
follow us
facebookinstagramlinkedin
our newsletter
subscribe!