BLOG · METHOD

Prompt generation for GEO: what an ideal tracking set looks like

By Alexandros Palma · Published · Last updated

An ideal prompt set for AI-visibility tracking is about 30 questions: brand-neutral, written the way a buyer actually talks, in the buyer's language, and mapped across the whole customer journey (roughly 10 awareness, 14 consideration, 4 purchase and a few geographic prompts), measured on a fixed weekly cadence. Everything else in this article explains why each of those properties exists, with the data behind it.

Prompts are the measurement instrument

In generative engine optimization (GEO), you don't rank for keywords. You either appear inside an AI-written answer or you don't. That makes the prompt set the instrument that defines what your visibility numbers mean. A skewed instrument produces confident, useless numbers: track only flattering prompts and you'll “win” a market you don't actually reach. Our own monitoring runs show how fragmented that market is: in a single full pass for one Romanian lighting retailer, the engines named 67 distinct brands across just 14 answers. Share of voice in AI answers is sliced thin, and which slice you see depends entirely on which questions you ask.

The funnel taxonomy and why every stage earns its place

The structure below follows the prompt-selection methodology published by Peec AI, whose guide recommends starting with “10-20 awareness prompts (concerns people have about your category), 20-30 consideration prompts (generic and segment-specific)” and tracking for 30 days before refining. Their core warning: don't only track “best [category]” prompts, because “one buying decision involves multiple prompts at different stages.”

Awareness

~10 prompts

The concerns people have about the category before they want any product: durability, running costs, how to choose, common failure modes. Deliberately NOT “best X”. For an LED shop: “how long do LED strips last?”, “do LED panels use much electricity?”, “why do LED bulbs flicker?”.

Why: Engines answer these with explainers, and the brands cited inside those explainers get the first recommendation slot later in the conversation.

Consideration

~14 prompts

Shortlist-building questions: the broad “what are the best X” plus segment-specific variants for each buyer persona, room, budget or use-case visible on the site: “best LED strip for a kitchen”, “waterproof spotlights for a terrace”.

Why: This is where recommendation lists live. Segment variants matter because a brand can rank #1 for “magnetic track lighting” and be invisible for “bathroom spotlights”. One number can't show that.

Purchase

~4 prompts (D2C only)

Ready-to-buy moments: “where can I buy dimmable GU10 bulbs online?”, delivery and payment questions. Skipped entirely for B2B businesses, where nobody asks an assistant for a checkout link.

Why: The last click before money changes hands. If a marketplace is always the answer here, that's a strategic fact worth knowing early.

Geographic

0-3 prompts

Only when a physical location matters: “best LED showroom in Iași”. Added when the site signals a showroom, service area or local pickup.

Why: Peec AI's guide is blunt about this: “Don't assume your brand performs equally across different regions. AI search results can vary dramatically by location, even for identical prompts.”

Rule one: brand-neutral, always

The single most common mistake is putting your own brand in tracking prompts. “Is [brand] good?” guarantees the brand appears in the answer, so visibility reads 100% and means nothing. Peec AI's guide draws the same line: track brand-evaluation prompts “separately from your main tracking to keep your visibility metrics clean.” The ideal set measures one thing only: when a buyer asks a generic question, does the engine bring you up on its own? That's the number that predicts new customers, because those buyers haven't heard of you yet.

Rule two: the buyer's language, the buyer's location

A Romanian buyer asks ChatGPT in Romanian, so a Romanian shop must be tracked in Romanian, since translating the prompts to English measures a different market. Location is just as physical: engines localize answers, so the collection itself has to run from the buyer's market, not through prompt wording like “answer as if in Romania” but as a real location parameter on the request. The differences this surfaces are not subtle. In one of our July 2026 runs, the same brand had 50% visibility on Gemini and 16.7% on ChatGPT for the identical prompt set, and Google's AI Overviews appeared for 0 of 8 commercial Romanian queries, twice in a row. If you only tracked one engine, you'd know a third of the story.

Rule three: conversational phrasing, not keywords

Assistants are asked questions, not fed keyword strings. Peec AI's example: the Google query “best iPhone for photography” becomes the AI prompt “Which iPhone is best for photography?”. Keyword-style prompts produce answers no real user sees. The practical test we apply when generating a set: would a person actually type this sentence into a chat box? If not, it doesn't belong in the instrument.

What the data says happens when you get this right

Two kinds of evidence. From research: the foundational GEO study ( Aggarwal et al., 2023) measured that content optimized for generative engines (citations, quotations, statistics, answer-first structure) can improve visibility in AI answers by up to 40%. From measurement: visibility gaps are real and large. Querying a corpus of 350,394 Romanian AI responses, the lighting retailer we monitor had 0 brand mentions while a national DIY chain had 1,404. You cannot see, let alone close, a gap like that without a prompt set that covers the journey where those mentions happen.

A well-built prompt set also tells you what to fix, because each stage maps to a GEO pillar. Losing awareness prompts usually means the site has no answer-first explainer content for engines to cite. Losing consideration prompts with competitors present points to missing fact density, comparison content or schema (FAQ, Product) on category pages. Losing purchase prompts to marketplaces is a crawlability and structured-data fight. The prompt set is the diagnosis; the GEO score is the treatment plan.

How many, and how often

Around 30 prompts is the working sweet spot: enough to segment by stage and persona, small enough that every prompt earns its place and week-over-week movement is legible. Peec AI's guidance lands in the same range and adds the discipline that matters more than the count: pick the set, track it unchanged for ~30 days, then refine based on where you're invisible. As Peec puts it, “prompt volume alone won't cut it.” Weekly cadence is frequent enough to catch real shifts without chasing the day-to-day randomness inherent in generative answers. Collection cost is no excuse anymore: tracking 30 prompts weekly across three engines costs on the order of a euro per month in API fees.

The checklist

An ideal tracking set, compressed:

  • ~30 prompts, mapped: ~10 awareness · ~14 consideration · ~4 purchase (D2C only) · 0-3 geographic
  • Brand-neutral: no brand names, no “X vs Y” (track those separately if at all)
  • Written as real questions, in the buyer's language
  • Collected from the buyer's location, across multiple engines
  • Held stable for ~30 days, reviewed against visibility gaps, then refined

This is exactly the methodology our project-setup wizard automates: it reads your site, infers the niche, the buyer and the business type, and drafts a funnel-mapped, brand-neutral set in your market's language, which you then edit, because you know your buyers better than any model does.

See which questions your buyers ask, and whether AI answers with you.Start monitoring →