How to Choose an SEO Agency: Seven Questions and a Way to Score the Answers
· Published August 27, 2026 · 12 min read
Choose an SEO agency by scoring its method, not its pitch. The seven questions below are scored 0, 1 or 2, so two people at your company can interview the same agency and arrive at the same number. That is the whole point: turning a gut call about a sales conversation into something you can compare, defend and revisit.
Most guides to this stop at a list of things to ask. A list is not a decision. What follows is the list, what a good answer actually sounds like, the red flag for each, a scorecard, real prices, and one test you can run in a minute that separates an agency measuring AI answers from one saying the words.
What an SEO agency actually does in 2026
An SEO agency earns your pages a position in search results and, increasingly, earns your brand an accurate mention inside AI generated answers. Those used to be one job. They are now two, and most agencies still sell one.
The first is the work you recognise: technical fixes so pages can be crawled and rendered, content built to answer the questions your buyers ask, and links earned from places that carry weight.
The second goes by generative engine optimization, answer engine optimization, or just AI visibility. When someone asks ChatGPT, Perplexity, Claude, Gemini or Google’s AI Overviews which product to use, the engine assembles an answer from pages it has read. Being in that answer, named and described correctly, is a different outcome from ranking third for a keyword.
The two are connected. Engines mostly draw on pages that already rank, so a strong SEO foundation feeds both. But they are measured differently, and this is where agency selection gets hard. A ranking report tells you nothing about whether an engine names you, gets your pricing right, or cites a competitor’s page when describing your product.
Here is the uncomfortable part, and it comes from our own measurement rather than from opinion. Across 43 domains we scored on site factors against how often each domain was actually cited by AI engines. The correlation was minus 0.068, which is nothing. Brand size correlated at plus 0.878. Tidying a website does not earn citations. It removes reasons you cannot be cited, which is worth doing and is not the same claim.
Any agency telling you that a technical cleanup will get you into AI answers is selling something the data does not support. That is the first thing the questions below are designed to catch.
The seven questions to ask an SEO agency
Score each answer 0, 1 or 2 as you go.
1. Which questions will you measure us against, and how did you pick them?
Rankings for 500 keywords tell you very little about revenue. Your buyers ask a much smaller set of questions that actually decide a signup, a deposit or a demo. The agency should build that set from your sales calls, your support tickets and your search data, hand it to you to strike and add, and then hold it constant so every report compares like with like.
Holding it constant matters more than it sounds. A number that moves because somebody changed the questions is a story, not a measurement.
A good answer sounds like: “We draft the set in week one from your call recordings and Search Console, you sign it off, and we run the same set every month. If we need to add a question later it starts a new series and we label it as not comparable to what came before.”
Red flag: a keyword list sorted by search volume, with no column explaining what a buyer asking that question is trying to do.
2. Which engines do you check, how many times, and can we read the raw answers?
AI answers are not deterministic. The same question on the same engine returns different answers on different runs. One prompt, one engine, one day is an anecdote.
We learned the size of this the hard way. In one early scan of ours, 7 of 18 contradictions appeared in only one run out of three. A single pass tool would have sold all seven as findings, at a noise rate of 39%. Five runs is our floor now because of that scan.
There is a second reason to ask. Cited sources turn over fast: measured across 82,619 prompts over 17 weeks, week to week source replacement ran at 56% on Google’s AI Mode and 74% on ChatGPT. If an agency reports a change without telling you how many runs produced it, you cannot tell progress from churn.
A good answer sounds like: “Five engines, each question run at least five times, share of voice reported per engine and never blended into one score, and the raw answers stored where you can read them.”
Red flag: “We optimise for AI” with no engine list, no run count, and nothing you can read. In a scan of 17 agencies selling AI visibility work in August 2026, not one published how many times it runs a prompt. It is a fair question and almost nobody is answering it.
3. When an engine says something wrong about us, what happens?
This is the question that separates monitoring from marketing, and it matters more than absence. Being missing from an answer costs you a lead. Being described wrongly, on pricing, licensing, availability or what your product does, costs you the buyer and sometimes more than that.
There is a related failure most agencies cannot name. Roughly 61.7% of citations are what we call ghost citations: the engine reads your page, uses it, links it, and never says your name. Your content is doing the work and a competitor is getting the credit. The fix for a ghost citation is different from the fix for absence, and an agency that does not distinguish the two will apply the wrong one.
A good answer sounds like: “We flag it the day we find it rather than holding it for a report, we trace it back to the page the engine cited, we fix or replace that page, and we rerun the exact prompt at the same run count until the answer changes or we tell you it has not.”
Red flag: “Engines are unpredictable, we focus on brand mentions.” Also a red flag: promising the engine will correct itself by a date. Nobody publishes a refresh timeline and nobody controls it.
4. Who does the work, and what did that person run before?
Ask for the name of the person who will touch your account each week, and what they were responsible for before this. Operating history inside your category counts for more than a certification, because it is the difference between knowing what passes a compliance review and finding out in month three.
A good answer sounds like: a named lead, a named prior role you can look up, a description of what they built there, and a commitment that the person on this call is the person on the account.
Red flag: a wall of client logos with no humans attached. Or a senior seller on the pitch and a different team after signature. Ask directly: will you be doing this work?
5. What can we verify without taking your word for it?
Evidence comes in three forms and any one of them is enough. Named client case studies with dates and numbers. A founder or lead whose track record you can check independently. Read only access to a live dashboard for a reference client.
The strongest signal costs the agency nothing to offer: does the agency rank, and get cited, for the thing it sells? An agency selling AI visibility that no engine names when you ask about AI visibility agencies has answered the question.
A good answer sounds like: “Here is the role and the dates, here is a client you can call, and here is our own site ranking and getting cited for the terms we sell you.”
Red flag: results with no dates, screenshots with no URLs, or a case study that names no company and no year.
6. What ships in the first 90 days, and what do you need from us?
Ask for phases with exit criteria, and ask what the agency needs from you. Search Console and analytics access, a named person who can confirm facts about your product, sales call recordings, and a realistic legal review turnaround.
That last part is diagnostic. An agency that needs nothing from your team is planning to ship work that could have been written for anyone.
A good answer sounds like: weeks one to three establishing a baseline and getting a question set and a fact sheet signed, then a short correction phase clearing whatever blocks access, then the long stretch of building and earning. With a named exit condition for each.
Red flag: “We will start publishing eight articles a month” with no baseline. If nothing was measured first, nobody can tell you later whether any of it worked.
7. What do you report, how often, and which number should we act on?
There is a trap in this question and it is worth knowing before you ask it. Weekly reporting sounds like diligence and usually is not. Given how fast cited sources turn over, a weekly report mostly documents the engines changing their minds. The thing that genuinely needs to be fast is escalation, not reporting.
So the answer you want separates the two: check often, report on a cadence long enough for the numbers to mean something, and escalate anything wrong immediately rather than saving it for the next send.
A good answer sounds like: “We check every engine weekly and escalate anything wrong the day we find it. You get what shipped every two weeks with links, a full read monthly, and a rerun of the whole set every quarter against the same fact sheet.”
Red flag: a monthly ranking PDF and silence in between. Also a red flag: a single blended visibility score. Engines disagree with each other, and averaging them away hides the only useful fact, which is which engine and which question.
How to score the answers
| Question | 0 points | 1 point | 2 points |
|---|---|---|---|
| 1. What gets measured | Keyword list by volume | Keywords tagged by intent | A fixed question set you sign off, held constant |
| 2. Engines and runs | No method described | One or two engines, single runs | Five engines, five or more runs, raw answers you can read |
| 3. Wrong answers | No process | Noted in the monthly | Traced to the cited page, fixed, and rerun until it clears |
| 4. Who does the work | Unnamed team | Named lead, no relevant history | Named operator with history in your category |
| 5. What you can verify | Nothing checkable | Screenshots without dates | Named roles, dated results, or a reference you can call |
| 6. First 90 days | A content calendar | Milestones | Phases with exit criteria and a list of what they need from you |
| 7. Reporting | Monthly PDF | Monthly plus a call | Weekly checks, same day escalation, biweekly log, quarterly rerun |
Fourteen is the ceiling.
Ten and above: shortlist. Seven to nine: buy a paid diagnostic before committing to a retainer, and use it to test the answers they gave you. Below seven: the method is not there, and a good sales conversation will not put it there.
One caution on using this. The scorecard measures whether an agency has a defensible method. It does not measure whether they are right for your category, your budget or your team. A 13 in the wrong specialism still loses to an 11 in the right one.
The red flags, in one list
- Guaranteed rankings, or guaranteed AI mentions. Nobody controls the engine. An agency can promise the work and should never promise the outcome.
- Reports with no line to signups, revenue or whatever you actually count.
- Links from networks the agency owns, or any offer that quotes a number of links per month.
- Traffic rising while conversions stay flat. Bots convert at zero.
- Directory badges presented as awards when the directory sells placement.
- An AI claim with no engine list, no run count and no raw output.
- A proposal that needs nothing from you.
- A blended score with no per engine breakdown underneath it.
What an SEO agency should cost
Published retainers in this market commonly run from around $5,000 to $30,000 a month. That range is wide because it is measuring scope, not quality.
The more useful signal is whether the agency publishes a number at all. When we checked 17 agencies selling AI visibility work in August 2026, two published a real retainer figure. The rest asked you to book a call. That is a choice about the sales process, and it tells you something about how much of the rest will be explained to you.
For what it is worth, since arguing for published pricing while hiding our own would be incoherent: Zion Labs runs a three week audit at $7,500 which credits in full against the first month of any retainer, and retainers at $6,500, $7,500, $15,000 and $30,000 a month depending on how many markets are read and how often. The pricing page has the full breakdown.
Use price to check scope. A retainer priced below the cost of one senior person is buying you volume produced by somebody junior. A retainer with no diagnostic in front of it is buying you guesses.
How long it takes
Judge a timeline by what ships in each phase, never by a promised date for a result.
- Weeks one to three: a crawl, a signed question set, an agreed fact sheet, and a baseline reading across every engine. Nothing is optimised yet, and an agency that admits this is being straight with you.
- The following month or so: whatever is stopping engines reading you gets cleared, and every wrong answer gets traced back to the page that taught it.
- From there: the slow work of getting into the sources engines actually quote. This is the part that takes months and the part that matters.
- Every quarter: the same questions, the same fact sheet, rerun, so the comparison is real.
Ninety days is a fair minimum before anyone should expect a readable signal. A year is a fair expectation for owning a category of questions. Anyone offering faster is either targeting terms too easy to matter or describing something they do not control.
Agency or in house?
Build in house when you have a senior operator who has done this before and engineering time to ship their tickets. The advantage is control and continuity, and it is real.
Hire an agency when you need the method and the monitoring on day one, or when the work needs relationships with publications you do not have.
Most teams that get this right run both. One internal owner who controls priorities and knows the product. One external operator who brings the process and does the measuring. One shared log they both read, so nobody is reporting to the other.
The one minute test for an AI visibility claim
Do this before the second call, not after.
- Open ChatGPT, Perplexity, Claude and Gemini.
- Ask each the question your buyers actually ask. Not your brand name. The question someone asks when they do not yet know you exist.
- Note who gets named, whose pages get cited, and whether anything said about you is wrong.
- Then ask the agency to show you those same four answers from their own monitoring, with the run counts.
If they cannot produce them, they are not monitoring. If they produce them and the numbers match what you just saw, run the seven questions properly. This takes about a minute and it is the single most efficient filter in this entire article.
The short version
Ask the seven questions. Score them out of 14. Shortlist at 10.
Every agency will tell you they are different. Very few can show you the questions they measure, the engines they check, how many times they check, and what they do the day something is wrong. That gap is the whole decision.