The fourteen questions, grouped
Score each answer 0 to 2. Zero means the agency failed the question, one means a partial answer, two means it cleared. Questions 4, 6 and 9 count double, which puts thirty-four points on the table. Under 20, keep looking.
Method
| Question | A strong answer sounds like | Red flag | How to verify |
| 1. Walk me through one recent client piece, from topic selection to published page. | A specific page, why that topic, who wrote it, what it was meant to win, and what happened. | A process diagram instead of a page. No named example. | Ask for the live URL and read it. |
| 2. What would you refuse to do for us in the first 90 days, and why? | A named tactic they have decided against, with the reasoning. | "We'd do everything." An agency right for everyone is right for no one. | Compare their answer against what is actually in the proposal scope. |
| 3. Which buying-intent topics would you deliberately skip for us? | Named keywords they would not chase, and the commercial reason. | An unprioritised keyword export. | Ask what the skipped list has in common. |
Delivery
| Question | A strong answer sounds like | Red flag | How to verify |
| 4. Who writes the content, and do you ship published pages or hand us briefs to fill in? (Counts double.) | Named writers, and pages that go live under their process rather than a brief handed back to your team. | "We have a team" that stays vague under one follow-up. A deliverable defined as a brief, an outline, or a recommendation. | Ask for three pages published in the last 60 days and the name of the person who wrote each one. |
| 5. How much of our citation volume should come from pages we do not own, and how do you earn those? | A split between owned pages and third-party pages, with a named method for the second and a clear way to report it. | Off-site work described only as "digital PR" with no target and no measurement. | Ask which third-party pages they got a current client named on, then open those pages. |
Owned pages and third-party pages behave like different assets. Ask how each is measured, who controls the third-party placement, and what happens to it when the engagement ends. Put the answer in the contract rather than accepting a general digital-PR promise.
Measurement
| Question | A strong answer sounds like | Red flag | How to verify |
| 6. Which AI engines do you track, how often, and on which prompt set? (Counts double.) | Named engines, a stated cadence, and a defined buyer-prompt list they will agree with you. | "We do AI SEO" with no metric attached. A single blended number. | Ask to see one real citation map for a current client. |
| 7. How will you measure where we stand on day one, before any work ships? | A per-engine baseline across a fixed prompt set, captured and handed to you before the first change. | No baseline, or a baseline captured after work starts, which makes every later number unfalsifiable. | Ask for the baseline as a file, with the date and the prompt list in it. |
| 8. Do you distinguish visibility, share of voice, and position, and do you report retrieval separately from citation? | Yes, with the definitions. Visibility is the percentage of AI responses where your brand appears. Share of voice is your share of all tracked-brand mentions in those responses. Position is your average rank when you do appear. Retrieved and cited are a separate pair: a URL of yours can be read as a source without being referenced in the visible answer. | Visibility and share of voice used interchangeably, or one blended figure standing in for all three. | Ask which of the three their headline case-study number is. |
| 9. How do you tie any of these numbers to pipeline and revenue? (Counts double.) | Demos booked, qualified leads, and pipeline influenced, reported next to the visibility numbers in the same monthly document. | A report that leads with traffic and rankings and reaches revenue only when you ask. | Ask for a sample monthly report, redacted, before you sign. |
A page can be read by an engine and never quoted in the answer. Those are two separate events, and an agency that reports only one of them is reporting half the job. Our own reading of this is in how to get named in AI search, and the log-file method behind it is in the AEO log-file playbook.
Proof
| Question | A strong answer sounds like | Red flag | How to verify |
| 10. Can you guarantee we will be cited in ChatGPT or appear in AI Overviews? | No, and here is what we can commit to instead. | Any yes. Any timeline attached to a guarantee. | Read what Google publishes about guarantees and about AI Overviews eligibility. |
| 11. Can I speak to a current client at a similar ARR that you did not pick for me? | Yes, with an introduction. | A curated reference only, or a testimonial page offered as a substitute. | Ask that client what went wrong and how the agency handled it. |
| 12. Show me a result you expected that did not happen, and what you changed. | A specific miss, the diagnosis, and the correction. | A portfolio with no losses in it. | Ask what the next client got that this one did not. |
Question 10 is the cheapest disqualifier on the list, because the vendors themselves have published the answer. Google's guidance for anyone hiring an SEO states that "No one can guarantee a #1 ranking on Google", and that "Google never accepts money to include or rank sites in our search results". For the AI surfaces specifically, Google's documentation says a page must be "indexed and eligible to be shown in Google Search with a snippet" and that "There are no additional technical requirements", then states plainly that "Indexing and serving isn't guaranteed".
Google also rejects most of the technical folklore sold as AI optimisation. On llms.txt and similar files: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn't use them." On breaking pages into fragments: "There's no requirement to break your content into tiny pieces for AI to better understand it." If an agency's method rests on those, ask what the method is without them.
OpenAI documents a narrower thing, and the distinction is worth holding. What a site controls is crawler access rather than placement: "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links." Allowing the crawler permits search access. It does not make you cited.
The honest version of the timeline question is that AI citations move at three different speeds. First pickup on an uncontested prompt can happen within a day. Holding a slot in the cited-source set across repeated re-evaluations takes weeks. Leading a competitive prompt cluster takes months. An agency that quotes the first speed while billing for the third is selling the wrong number. More on that in what an AI search agency should deliver in the first 90 days.
Commercials
| Question | A strong answer sounds like | Red flag | How to verify |
| 13. What is the price, what is included, what is extra, and who does the work by name? | A number, a scope, a named list of what falls outside it, and the people who will touch your account with their roles. | Price withheld until a second call. Scope that grows after signature. A team that stays anonymous under one follow-up. | Ask for the scope in writing before the contract, and ask to meet the writer and the strategist. |
| 14. What is the minimum term, the notice period, and what triggers a renewal? | A short initial term, a stated notice period, and renewal you have to opt into. | A twelve-month lock-in presented before any proof point. Automatic renewal you have to hunt for in the terms. | Read the termination clause yourself, not the summary of it. |
We could not find a non-vendor benchmark for what agency contract terms usually look like in this category. The figures we found all trace back to agency blogs citing unnamed surveys. So treat the market-norm argument as unavailable to both sides, and negotiate the terms on their merits instead. What a capable agency costs is a separate question, and the tiers are broken down in AEO agency pricing for B2B SaaS.
Five AI-search contract terms to put in the agreement
Put these five terms in the agreement before signature. They cover the prompt set, tracking records, baseline, off-site placements, and crawler and log access. No agency-evaluation guide we reviewed covers all five together, so ask for each one directly.
| Term | What to ask | Why it matters |
| The prompt set | Who decides which prompts we are tracked against, and can I see and change the list? | The prompt set is the denominator of every percentage the agency will ever report to you. An agency that controls it privately controls the score. |
| The tracking data | Do we get the raw per-engine, per-prompt records, or only the dashboard view? | A summary cannot be re-analysed or handed to the next agency. The raw records can. |
| The day-one baseline | Do we keep the baseline dataset on exit, including the comparison records against competitors? | Without the baseline you cannot prove what the engagement changed, to your board or to yourself. |
| Off-site placements | What happens to third-party pages that name us, and are any of them on properties the agency owns? | Placements on an agency's own network can be removed when you leave. Genuinely earned ones cannot. |
| Crawler and log access | Who holds the log-file and crawler-control access, and does it revert to us cleanly? | Server logs are the highest-fidelity record of which AI crawlers actually reached your pages. Losing that access loses the evidence. |
The off-site row deserves the most attention, for the reason question 5 exists. Visibility built on pages you own behaves like an asset you keep. Visibility built on pages someone else controls behaves like a rental, and the contract should say which one you are buying.
Score the answers
Fourteen questions, 0 to 2 each, is 28 points. Questions 4, 6 and 9 count double, which adds 6 and puts the total at thirty-four.
Twenty-eight or above, strong fit. Between 20 and 27, workable if the gaps sit in the single-weight questions. Under 20, keep looking.
Those three carry double weight because each one covers a failure you cannot fix later from your side of the table. Content that ships is where retainers quietly fail, since a brief-only shop hands the hardest part back to the team you hired it to replace. Per-engine measurement is what separates a 2026 organic-growth partner from a 2023 SEO shop, since an agency that cannot measure citations per engine cannot improve them. Pipeline reporting decides what the agency optimises for, because a team that never reports revenue influence will report traffic instead.
Layer price on top of the score rather than underneath it. The cheapest option that clears the bar usually beats a more expensive one that does not.
Red flags that should end the conversation
Some answers are not a low score. They are a no.
- A guarantee of rankings or citations on a fixed clock. Google says no one can guarantee a number-one ranking, and its AI features do not guarantee indexing or serving.
- Vanity metrics as the headline. A dashboard that leads with sessions and rankings rather than pipeline is selling activity.
- No clear answer on who writes the content. Vagueness under one follow-up usually means a content mill.
- B2C or local case studies presented as SaaS proof. Category-specific prompt fluency does not transfer from a roofing account.
- Setup billing before any deliverable. Months of onboarding fees before a single asset ships protects their revenue ahead of your outcome.
- A long lock-in before any proof. A twelve-month term signed before a 90-day checkpoint protects them rather than you.
- One blended AI number. A single percentage across all engines hides which engine is losing.
A useful closing question for any finalist: who is a bad fit for you, and what would you refuse to do? The full list is in red flags when hiring an AI search agency, and the method for checking a claimed result is in how to verify an AEO agency's results.
Run the selection in three weeks
Week one: build a shortlist of three or four, no more. Pull two from a ranked list you trust, one from a peer referral, and one from the agencies already cited when you run your own category questions through ChatGPT and Perplexity. Treat an agency's presence in those answers as a screening signal, then ask for the prompt, source, and work behind it. Our audit of which AI search agencies publish a named client's pipeline is one place to start that list.
Week two: send all of them the same fourteen questions, and ask for price, scope, and one relevant case study in writing. Standardising the ask is the only honest way to compare. Score each response as it lands.
Week three: take your top two into a reference call with a current client they did not pre-select. Ask that client the one thing an agency cannot coach: what went wrong, and how did they handle it. Then decide.
Before any of this, settle whether you need an agency at all. The build-versus-buy maths is in in-house SEO vs agency for B2B SaaS.
How LoudFace answers these fourteen questions
Use the same fourteen on us.
- 1. One piece, end to end. The methodology page explains our eight-stage Answer Chain. Before you sign, ask us to name a recent page, the topic decision behind it, the writer, and the result. A process diagram alone is not enough.
- 2. What we refuse. We do not build llms.txt files, fragment pages into chunks, or sell a link package as AI search work. Google's own guidance says the first two do nothing for it.
- 3. What we skip. We name the buying-intent topics we would not chase for you, and we say which ones you should not pay us for.
- 4. Who writes, and what ships. Your proposal should name the writer and strategist, state whether the deliverable is a published page or a brief, and show how the work reaches your site. Do not accept a generic team label.
- 5. Off-site. Your proposal should state any third-party placement scope, how it is measured, who controls the placements, and what happens on exit. We do not treat a general off-site promise as proof.
- 6. Engines and cadence. The tracked panel covers ChatGPT, Perplexity, and Google AI Overviews on a defined buyer-prompt set. Monitoring runs daily and reporting is weekly. Gemini, Claude, and Copilot receive periodic snapshots rather than the tracked panel, and every reported number should name which group it came from.
- 7. Day-one baseline. Recorded per engine and per prompt in week one, before the first page change ships, and handed over as a file with the date and prompt list in it. The technical fixes and the first articles follow inside that same week.
- 8. Metric labels. We label visibility, share of answer, and position separately, with a measurement window, and report retrieval separately from citation. For Toku, AI visibility on the category's top stablecoin payroll prompt read 97.8% at average cited position 3.1 on a 30-day reading ending in August 2026, over an 18-month engagement that has since ended. That is a visibility reading.
- 9. Pipeline. Ask to see how demos booked, pipeline influenced, and visibility appear together in the monthly report before you sign. A traffic-only report is not enough.
- 10. Guarantees. None, on placement. We commit to the work, the cadence, and the reporting.
- 11. References. Ask us to discuss a relevant reference process and permission before you sign. Do not accept a testimonial as a substitute for a conversation.
- 12. A result that missed. Ask us for a result that fell short and the record used to diagnose it. We will not use a broad claim about off-site placements as proof.
- 13. Price and scope. The pricing article publishes retainers from $5K to $18K per month with no setup fees across three tiers. Your proposal should state what falls outside the scope.
- 14. Term and exit. A fixed scope runs inside a retainer with a 3-month minimum. Notice, renewal, and exit rights belong in your proposal and contract. Read those terms before you sign.
Our own current reading is 21.5% visibility on non-branded prompts across the three tracked engines over 18–28 September 2026, measured on the 141 non-branded prompts we track after a 17 September prompt cut that lifts the reading. Per engine that is 38.8% on ChatGPT, 11.2% on Perplexity, and 14.3% on Google AI Overviews, at an average position of 3.4 when named, 3rd of 53 tracked brands. Check the Toku case study and our own case study for dated outcomes, the methodology page for measurement scope, and the pricing article for the published band. Writer assignment, off-site scope, and exit terms belong in the proposal.
Send the fourteen, then shortlist
Send the same fourteen questions to every agency on your list, in writing, and score the replies the day they arrive. Questions 4, 6 and 9 will separate the field faster than anything else, because shipping published pages, measuring each engine separately, and reporting pipeline are all cheap to promise in a meeting and expensive to sustain for a quarter. When you have your shortlist, the ranked roster of AEO agencies for B2B SaaS is a reasonable place to fill the last slot.

