AI summary: Prompt monitoring measures how AI answers respond to configured prompts over repeated runs. In the Geolify snapshot, 1,323 configured prompts ran across 6,438 monitoring records and produced 25,965 executions — three different denominators you must never merge.
Keyword tracking asks where a page ranks in a results list. Prompt monitoring asks what an AI engine actually says when a user asks a question — whether your brand is mentioned, whether you are cited with a link, and how that changes run to run.
In the Geolify snapshot dated 2026-07-10: 1,323 configured prompt records, 6,438 monitoring runs, and 25,965 prompt executions (Geolify internal snapshot, MG-PROMPT-01, MG-MON-01, MG-MON-02).
The core lesson: prompt, run, and execution are three different units. Your metric is only meaningful when you name which one is the denominator.
Everything downstream depends on these definitions.
Configured prompt — the definition. The text you decide to monitor, plus its settings. The snapshot holds 1,323 configured prompt records (Geolify internal snapshot, MG-PROMPT-01). This is your test plan, not your data volume.
Monitoring run — one scheduled execution cycle. The snapshot holds 6,438 monitoring records (Geolify internal snapshot, MG-MON-01). A single run can carry many prompts. A single prompt runs across many cycles.
Prompt execution — one issued prompt within one run. The snapshot holds 25,965 executions: 20,679 completed, 4,125 failed, the rest still processing (Geolify internal snapshot, MG-MON-02). Executions are the atomic events you analyze.
The common mistakes are now obvious:
Keyword tracking evolved for ten blue links — a stable, ranked, deterministic list. AI answers break that model in three ways:
A page can rank well in classic Search and still be absent from an answer. Google frames AI Overviews and AI Mode as built on the same Search systems, with no additional technical requirements (Google Search Central). Ranking matters as an input — but it feeds a synthesis step you must measure separately.
Prompt Monitor is how you measure that synthesis step directly.
These metrics are a proposed framework, not vendor-standard definitions. Label them that way in your reporting.
Answer presence rate. Of completed executions for a given prompt, the share where your brand is mentioned in any form. Denominator: completed executions.
Citation presence rate. Of completed executions, the share where your domain is cited with a link. Keep this separate from answer presence — a mention without a link and a linked citation have different value.
Competitor share. Of completed executions where any competitor is mentioned, the share per named competitor. This is relative within your tracked set. Never a whole-market claim.
Sentiment tilt. Of completed executions mentioning your brand, the distribution across positive, neutral, and negative. Report the rubric and sample size — sentiment on small samples is fragile.
Volatility. Run-to-run variance of any metric above, for a fixed prompt over a fixed window. This is the metric keyword tracking never needed. AI monitoring cannot live without it.
Every definition shares one trait: a stated denominator and explicit failure handling.
The snapshot makes the problem concrete: 4,125 of 25,965 executions failed (Geolify internal snapshot, MG-MON-02). Compute any rate over all 25,965 and you are mixing real answers with timeouts, provider errors, and unfinished work.
Fix it before you compute anything:
This is standard event-pipeline hygiene. It is what keeps a monitoring trend honest.
Walk through what naive math does to the snapshot data. There were 25,965 executions: 20,679 completed and 4,125 failed, with the remainder in processing states (Geolify internal snapshot, MG-MON-02).
A naive "success rate" of 20,679 over 25,965 quietly treats in-flight work as failure. Re-run the same report an hour later, after processing finishes, and the rate moves — with zero change in actual AI visibility. Anyone comparing the two reports would see a trend that never happened.
The defensible version separates three statements. Completed executions: 20,679 — this is the analysis population. Failed executions: 4,125 — this is an operational metric with its own trend line. Processing: excluded from both until terminal. Now every downstream rate is stable under re-computation, and a rising failure line is visible as a data-quality problem instead of masquerading as a citation drop.
This is the single most common error we see in AI-visibility reporting. It is also the cheapest to prevent: one completion-state filter, applied before any percentage is computed.
Your prompt library is a test plan. Its quality determines everything downstream. Three design rules:
Write prompts as real user intents. "What is the best way to validate whether AI crawlers can read a JavaScript-heavy page?" is monitorable. A bare keyword is not.
Group prompts into intent clusters. Report at the cluster level to avoid overreacting to a single volatile prompt.
Keep a stable core set. Your volatility and trend metrics need a fixed baseline. Add experimental prompts in a separate cohort. If you constantly rewrite core prompts, you lose the ability to measure change — which is the whole point.
Size the library to your capacity, not your ambition. The snapshot's 1,323 configured prompts represent a mature library (Geolify internal snapshot, MG-PROMPT-01). A team starting out does better with 20 well-written intents it can review monthly than 400 it cannot. Every prompt you add multiplies executions, storage, and review time across every future run.
None of this retires your rank tracker. GEO is not SEO — but it is not anti-SEO either.
Keyword tracking remains the right instrument for three jobs. It measures the classic results surface, which still carries most B2B discovery traffic. It feeds the input side of AI answers: Google states its AI features build on the same Search systems (Google Search Central), so a page that cannot rank is fighting uphill to be synthesized. And it stays comparable across years of history you already own.
The division of labor is clean. Keyword tracking tells you whether the inputs to AI answers include you. Prompt monitoring tells you whether the outputs mention and cite you. A visibility program in 2026 runs both, with separate metrics and separate denominators — you use a rank tracker for SEO and Prompt Monitor for GEO, and neither number substitutes for the other.
AI answers move. Your brand can appear in four of five runs one week and two of five the next — without any change on your site. The model sampled differently, or the index shifted.
Single-run screenshots are the weakest possible evidence. Volatility deserves a first-class metric.
Practical rule: Report ranges and trends, not point values. For a fixed prompt over a window, report the answer presence rate with its run count and observed spread. Only call a change real when it exceeds the historical spread for that prompt.
Correlation with a deploy is a hypothesis. Test it with more runs before calling it a finding.
Instrument the three units at the point of capture — not by reconstructing them later. When a run executes, write one record per execution with:
That schema lets you compute every metric by filtering on completion state and grouping by prompt or cluster. Failures become a first-class field, not a silent gap.
Cadence is a design decision, not a default. Daily cadence gives tight volatility resolution but multiplies cost. Weekly cadence smooths noise but misses short-lived swings. Choose per cluster based on how fast the topic moves. Keep it stable long enough to build a baseline. When you change cadence, annotate the timeline — a sampling change will itself change apparent volatility.
Lead with cluster-level answer presence and citation presence. Show each with its completed-execution denominator and run count. Then show volatility as a range.
Drill into individual prompts only after the cluster view. Reserve competitor share and sentiment for a labeled secondary section — they rest on smaller samples and softer rubrics.
A report built this way is defensible line by line. That matters when a stakeholder challenges a number in a QBR.
If you present AI visibility to executives or clients, one page is enough when the units are clean:
The order matters: trend, evidence, noise floor, quality, action. A stakeholder who can follow that chain will approve the next quarter of monitoring without a fight.
Prompt monitoring tells you what tracked answers said for your configured prompts at your cadence. It does not:
A brand mention is not a click. A citation is not a conversion. Report every number with its conditions: prompt library, run cadence, and failure handling.
The snapshot also warns against inventing benchmarks. There is no supported figure for a universal citation-loss rate. If a stakeholder asks for a single AI visibility percentage, give them a defined metric with its denominator — not a headline number stripped of context.
This week: Separate your data into three tables — configured prompts, monitoring runs, and executions. Define your analysis population as completed executions. Track failures separately.
This month: Publish four metrics: answer presence, citation presence, competitor share, and volatility. Print the denominator next to each number. Keep a fixed core prompt set for trend continuity.
This quarter: Connect this measurement to the broader GEO strategy in the GEO optimization checklist and What Is GEO.
One closing pitfall list, because these recur in almost every failed monitoring program: blending the three units into one count, computing rates over all executions instead of completed ones, rewriting core prompts mid-quarter and calling the resulting movement a trend, treating one run's screenshot as evidence, and changing cadence without annotating the timeline. Each is invisible in the moment and expensive at the quarterly review. The fix for all five is the same habit — write the unit, the denominator, and the window next to every number the day you first compute it.
Primary-source checks were performed on 2026-07-10 against Google Search Central.
| Layer | Evidence ID | Count | Correct Language |
|---|---|---|---|
| Configured prompt definitions | MG-PROMPT-01 | 1,323 | Prompt records / configured prompts |
| Monitoring records/runs | MG-MON-01 | 6,438 | Monitoring runs (a prompt can run many times) |
| Embedded executions | MG-MON-02 | 25,965 | Executions inside runs (totalQueries aggregate) |

Google AI Mode vs. AI Overviews: A Measurement and Budget Framework for B2B Teams AI summary: AI Overviews and AI Mode both build on standard Search systems. Measurement starts with Search Console and GA4 — not special tooling. This framework sets instrumentation limits and a 30-day experiment that ends in a budget decision. The practical […]

Building a Brand Monitoring System for AI Search AI summary: An AI-search brand monitoring system combines two signal families — prompt executions that show what answers say, and crawler traffic that shows who fetches your pages. In the Geolify snapshot, 25,965 executions and single-digit-percent AI traffic show why both signals, and clear denominators, matter. A […]

We Analyzed 4,704 AI Visibility Audit Runs: What Static-Content and Rendering Checks Reveal AI summary: Among 4,413 completed audit records, 2,283 (51.73%) scored below 50 for static content and 2,877 (65.19%) carried a Client-Side Rendering warning. These results describe Geolify audit records — not all websites — and do not prove citation loss. Half of […]