How to Choose Which Prompts to Track in a GEO Tool

Buying a GEO tool takes an afternoon. Deciding which prompts actually represent buying intent is the part that takes weeks. Here is the process I use.

The GEO tool demo is the easy part. Someone shows you a dashboard with a presence score, a citation breakdown by domain, a sentiment chart, and a competitor comparison that puts a red bar next to your name. You sign. Two weeks later you are sitting in front of an empty prompt field with a plan limit next to it, and the tool is politely waiting for you to tell it what to track.

That field is where the actual work is. Everything downstream, the score you report to leadership, the trend line you defend next quarter, the content brief you hand to a writer, is a function of what you type into it. The tool measures whatever you point it at. Point it at the wrong thing and it will measure the wrong thing very accurately, on a schedule, with a nice chart.

I have spent a while now building prompt sets, running them, watching some of them go stale, and retiring the ones that were never worth a tracking slot. Most of what follows I learned by getting it wrong first. This is the process I have ended up with.

The part nobody warns you about

Every GEO vendor’s onboarding assumes you already know what your buyers ask. Some of them will offer to generate a starter set for you, which is worth doing once so you can see the shape of it, and worth throwing away immediately after. The generated sets are fluent and generic. They read like a category overview page. They will happily produce forty prompts about “data integration” that no human has ever typed into anything.

The reason this problem is unglamorous is that it has no tooling. There is no volume metric to sort by. There is no Ahrefs for prompts. Nobody can tell you that 4,400 people asked ChatGPT which vendor to shortlist last month, because the platforms do not publish that, and the third party estimates are modelled guesses layered on search data. You are working without the crutch that SEO gave you for fifteen years.

What you have instead is judgement, plus access to people who talk to buyers. That combination is why the good prompt sets inside companies look nothing like the ones you see in vendor case studies.

A prompt is not a keyword

The first version of my prompt list was a keyword list wearing a costume. I took high intent commercial keywords, added a question word to the front, and shipped it. “Best data integration tools.” “What is reverse ETL.” “Data warehouse pricing.”

They tracked fine. They told me almost nothing, and it took me a while to work out why.

A keyword is a fragment. It is what someone types when the interface punishes them for typing more. A prompt is a complete request with context attached, because the interface rewards context. Real people do not ask “best data integration tools.” They ask something closer to “we are a retail company pulling from about forty source systems, customer records are duplicated across three of them, and our last analytics project stalled because nobody trusted the numbers. What tooling should we be looking at and what will actually fix the trust problem?”

Those two inputs produce different answers, cite different sources, and name different vendors. The second one is the one that decides whether you make the shortlist. The first one gets you a listicle summary where being mentioned means very little.

Once you internalise that, the whole exercise changes. You are not trying to cover a keyword space. You are trying to build a representative sample of the conversations that happen before someone buys.

Stop inventing prompts and go find them

The most useful thing I did was stop writing prompts at my desk and start harvesting them from places where buyers were already asking questions in their own words.

Sales call recordings are the best source by a wide margin. Whatever your team uses, Gong, Chorus, Fathom, the transcripts contain buyers asking questions in full sentences with all their constraints attached. That phrasing is the raw material. You do not have to guess what a buyer sounds like when you have three hundred hours of them talking. I pull questions from discovery calls specifically, because that is the stage where people are still deciding what they need rather than negotiating a contract.

Support tickets and community threads are the second source. These skew toward existing customers, so the intent is different, but they surface the objections and the failure modes. Objection prompts matter more than people expect, and I will come back to that.

RFP and procurement questionnaire documents are the third. They are dry and they are gold, because they contain the exact evaluation criteria a committee wrote down. Anything that appears in three different RFPs is something buyers compare vendors on, which means it is something an AI answer will be asked to compare vendors on.

Then the smaller sources. Search Console filtered to question shaped queries, which still tells you something about how people phrase problems. Sales enablement battlecards, which name the competitors you are actually losing to rather than the ones marketing likes to benchmark against. Reddit and Slack communities in your category. Your own SDRs, who will give you fifteen real questions in a ten minute conversation if you ask them for the questions they hear most and promise not to turn it into a project.

Harvesting takes a week or two. It is boring. It is also the difference between a prompt set that survives its first review with sales and one that gets quietly ignored.

The two questions I ask before a prompt makes the list

After harvesting I usually have far more candidates than I can track. Cutting them down is where most of the value gets created or destroyed, and I run every candidate through two filters.

The first is a commercial filter. If an AI assistant gave a completely wrong answer to this prompt, would it cost us a deal? Not “would it be inaccurate,” but would it change who gets on a shortlist, who gets a meeting, who gets ruled out. A surprising number of prompts fail this. “What is a data warehouse” is a fine question and we should be cited in the answer, but nobody chooses a vendor based on it. Definitional prompts belong in your set as a small fixed allocation, not as the backbone of it.

The second is an information filter, and this one gets ignored almost universally. Does the answer to this prompt actually vary?

If every model names you every time, and has for six months, the prompt has no information content left. You have a flat line at 100 percent that you cannot improve and that will not warn you about anything. The same is true in reverse for prompts where you are structurally never going to appear, like a category you do not compete in. Both are dead weight in the set. They inflate or deflate your aggregate score and they cost you tracking slots.

The prompts worth their slot are the ones near the decision boundary, where you appear sometimes, where a competitor appears sometimes, where the answer shifts when the underlying content shifts. Those are the ones where doing work actually moves the number. Volatility is the signal that a prompt is contested, and contested is where GEO effort pays.

I did not design the set this way at first. I got there by noticing which rows in my dashboard had never moved and asking what I was paying for them.

How I split the set

Once the filters are applied, I allocate rather than just accumulate. Without an explicit split, prompt sets drift toward whatever the loudest stakeholder cares about, which is almost always branded prompts and almost always the product they own.

The rough shape I use looks like this:

BandShare of the setWhat it answers
Problem framing15%Does the assistant describe the problem in terms where our approach makes sense?
Solution and category20%When someone asks how to solve this class of problem, do we come up as an approach?
Vendor comparison and shortlisting35%When someone asks who to evaluate, are we named, and next to whom?
Objection and validation20%When someone asks whether we are any good, what comes back?
Branded10%Is the assistant describing us accurately at all?

The comparison band is the largest because that is where money changes hands. The objection band is the one most teams skip and the one that has surprised me most, because it surfaces things like pricing complaints, migration horror stories, and old incidents that a model will happily repeat from a forum thread written in 2019. You cannot fix what you do not track.

Branded prompts get a floor and a ceiling. You need a few, because factual drift about your own product is real and you want an alarm on it. You do not need fifty, because you will always show up for your own name and the resulting number flatters everyone and informs nobody.

The percentages are mine and they are a starting position, not a law. If you sell into a category where nobody knows the problem exists, weight problem framing higher. If you are the incumbent being displaced, weight objection handling higher.

Same question, three phrasings, three different answers

The thing that took me longest to accept is that a prompt is not a stable unit of measurement.

Ask the same question three ways and you get three different citation sets. Not slightly different. Different vendors named, different sources cited, sometimes a completely different framing of what the question is even about. Add a persona to the front of it, “as the CISO of a mid size European bank,” and the answer shifts again, usually more than any rewording of the verb does.

This has two practical consequences.

Track intents, not strings. Every intent that survives the filters gets two or three phrasings in the set, and I read them as a group. A single phrasing is one sample from a noisy distribution, and drawing a trend line through one sample is how you end up presenting a fluctuation as a result.

Be deliberate about modifiers. Persona, company size, region, and regulatory context all change answers meaningfully, and each one you add multiplies your prompt count. The rule I use is that a modifier earns its place only if it maps to a segment you actually sell to differently. If your enterprise segment in one region has its own messaging, its own competitors, and its own objections, then it deserves its own prompts. If it does not, you are just tripling your bill.

Multi language sets are where this gets awkward, and there is no clean answer. You can translate your existing prompts so the wording stays parallel, or you can research fresh prompts in each language so the phrasing is authentic to how people there actually ask. Translation keeps your sets comparable, which matters if you want to say a market is behind or ahead. Native phrasing is closer to reality, but a difference between two markets then becomes a difference in the question as much as a difference in the answer, and you lose the ability to read them side by side. Pick based on what the reporting is for. Comparison across markets pushes you toward translation. Local content research pushes you the other way.

Deduplicate before you upload, not after

Two prompts that mean the same thing return the same answer and count twice. If a chunk of your set is near duplicates, your presence score is measuring your own redundancy.

I run a semantic deduplication pass before anything gets uploaded. Embed every candidate prompt, compute pairwise similarity, and review anything above a threshold by hand. I have found that setting the threshold high and reviewing manually beats trusting an automatic cut, because some intents that read as different sit close together in embedding space, and some near identical phrasings sit further apart than you would expect.

Two things I look for in the review. First, whether the near duplicates are the deliberate multiple phrasings of one intent, in which case they stay and get read as a group. Second, whether they came from different stakeholders who described the same buyer question in their own team’s vocabulary, which happens constantly and is the main source of accidental duplication.

Doing this after upload is worse in every way. You have already paid for the runs, you have already reported the inflated number, and correcting it means breaking your own trend line.

Freeze the core set, rotate everything else

A prompt set that keeps changing cannot produce a trend. Every edit resets your baseline, and if you edit continuously you will never have a comparable quarter over quarter number, which is the only number leadership actually wants.

Two tiers fixes it.

The core panel is frozen. It is dated, versioned, and changed only at a scheduled review, and when it changes the change is documented so anyone reading a chart knows where the discontinuity is. This is the set that produces the number in the deck.

The rotating set is expected to change. Launches, campaigns, announcements, competitive responses, anything with a short shelf life. It answers “did this thing we just did land,” and it is read as a snapshot rather than a trend.

The split matters most for the fast moving content, the kind with no stable page or topic behind it. Forcing that into a fixed panel gives you a set that is obsolete in three weeks and a trend line nobody can interpret. Keeping the durable reputation questions in the frozen panel and running everything short lived as a dated snapshot alongside it is cleaner, and it stops one noisy quarter of launches from distorting a number you want to read over a year.

When to retire a prompt

Prompt sets rot. Nobody schedules a review, so nobody notices.

I mark a prompt for retirement when it has been pinned at the same result for two consecutive quarters with no competitor movement, because it has stopped carrying information. Also when the product it maps to has been repositioned or sunset, when the phrasing has aged out of how people talk about the category, or when three months of review show nobody has ever used it to make a decision.

That last one is the strongest test and the least comfortable. For every prompt in the set, someone should be able to name the decision it informs and the person who cares about the answer. If a prompt cannot survive that question, it is costing you a tracking slot and diluting your aggregate score in exchange for nothing.

Retirement does not have to mean deletion. Moving a prompt to a quarterly spot check instead of continuous tracking preserves the safety net without paying for daily runs.

If I were starting from zero on Monday

The sequence I would follow, compressed.

Spend the first week harvesting only. No writing, no editing, no filtering. Pull questions verbatim from discovery calls, support tickets, RFPs, and a conversation with two SDRs. Expect a messy list of several hundred.

Cluster that list into intents, ignoring phrasing. You will find that a few hundred raw questions collapse into far fewer real intents, and the collapse ratio itself tells you something about how focused your market is.

Run both filters on the intents. Would a wrong answer cost a deal. Does the answer plausibly vary. Cut hard here and keep the cut list, because it becomes your rotating set candidate pool later.

Allocate the survivors against the bands so no single stakeholder’s interests dominate. Write two or three phrasings per intent. Semantic dedupe. Record where each surviving prompt came from, because in six months when someone challenges one you will want to be able to say it came from four discovery calls rather than from your imagination. Then date the set and freeze the core.

Book the review before you finish, because a review that is not in a calendar does not happen.

What a good prompt set still cannot tell you

I want to close honestly, because this space has an overclaiming problem.

A well designed prompt set tells you whether you are present in the conversations that precede a purchase, and how that presence is changing. It does not tell you that presence caused revenue. There is no click for most of these interactions, the attribution chain is broken by design, and anyone selling you a clean line from citation to closed deal is selling you a model, not a measurement.

It is a leading indicator. Treated as one, it is useful, and it is the closest thing we have to visibility into a surface that is growing fast and reporting almost nothing back to us. Treated as attribution, it will eventually embarrass you in a meeting.

The prompts you choose are the entire experiment design. The tool just runs it.