TRINZIK.AI

Blog · AI visibility

What an AI recommendation score is worth: how to audit the number before you trust it

  • A published AI recommendation score is worth nothing you cannot audit. Ask how it was built before you trust it.
  • Four checks separate a benchmark from a chart: locked prompts, a constant competitor field, verification against your real domain, and a fixed monthly rerun.
  • A score that matches on your name alone can hand your win to a same-named company. Grounding on your real domain fixes that.
  • Published indexes describe a category from outside. A locked head-to-head tells you whether you are being chosen.
  • Trinzik counts a win only as first place for your verified domain on a locked prompt, rerun on schedule.

By John Michaels

For twenty-five years before Trinzik I sat in forecast reviews, on both sides of the table. You learn one rule fast: a number you cannot verify is a number that will cost you. A rep says the deal closes in the third quarter, and there is exactly one honest response. Show me. Where is it in writing, who signed it, what is the date. A pipeline figure without those answers is not a forecast. It is a hope with a dollar sign on it.

AI recommendation scores are that same problem in new clothing. An index publishes a number, your brand rates a 72, or ranks eighth in its category this month, and it looks like measurement. The forecast question still applies. Show me. What was asked, of which AI platforms, on what date, and how do you know the answer credited your company and not a firm that shares your name? Researchers keep finding the gap between the score and the truth. A 2024 study presented at NAACL found "a disparity between the inflated benchmark scores and the actual performance of LLMs." The lesson is not that scores lie. It is that a score is only as good as the method you can check behind it.

What is an AI recommendation score?

An AI recommendation score is a single number that sums up how often and how favorably AI platforms like ChatGPT, Gemini, Perplexity, Claude, and Grok recommend a brand when people ask them category questions. Industry benchmarks and published indexes turn that number into rankings: who leads a category this month, who moved. There is no standard behind any of it. Each publisher decides which questions to ask, which platforms to count, and how to add it all up. So two scores that carry the same name can measure completely different things.

The scores exist because the stakes are real. When someone asks an AI who to hire, it tends to name one or two brands and leave the rest invisible. A number that claims to track your place in that answer is worth having. It is only worth trusting once you know how it was built.

What a score is actually made of

Most AI recommendation scores are built from three ingredients, and the differences between them decide whether the number means anything.

Mention frequency counts how often a brand shows up in AI answers to a set of questions. It is the easiest input to produce and the easiest to misread, because a mention can be a recommendation, a passing aside, or a warning, and a raw count treats them all the same.

Share of voice, your slice of every brand mention in the same set of answers, is a better single indicator, because it is measured against a field of competitors instead of floating on its own.

Verification is the ingredient most scores skip. It asks whether the tool checked that the winning answer actually pointed to your company's real web domain, or just matched your name. That gap is where a score quietly breaks, and it is the whole difference between a benchmark and a decoration.

A published index

A number handed to you, computed somewhere you cannot see, on questions you did not set. You can read it. You cannot rerun it. When it moves, you are told that it moved and asked to believe why.

A locked benchmark

The same questions, the same rivals, put to the platforms on a set schedule, every result checked against real domains. You can rerun it and get a comparable number. When it moves, you can see for yourself what changed.

Four questions that audit any AI recommendation score

You do not need to be technical to audit a score. You need four questions, and the nerve to keep asking until you get real answers.

1. Are the prompts locked? A benchmark is only comparable to itself when the questions never change. Lock the prompt set at the start and rerun it untouched, so this month and last month are the same test. There is a second reason to lock them, and measurement has a name for it. Goodhart's law, in Marilyn Strathern's 1997 phrasing, holds that "when a measure becomes a target, it ceases to be a good measure." A score whose questions get quietly reworded to look better each month has stopped measuring anything. Locked prompts are what keep it honest.

2. Is the competitor field constant? A win only means something against the same opponents. If the rival list shifts between runs, a jump in your score can be a weaker field rather than a stronger you. Ask who you were compared against, and whether that set is held constant from month to month. In a head-to-head comparison the roster matters as much as the result.

3. Is every result verified against your real domain? This is the check that catches the expensive mistake. AI platforms confuse companies that share a name, and a score that matches on name alone can credit your win to someone else, or someone else's win to you. Real verification, what we call grounding, means the platform had to return your company's own web address, and the result counts only when that address is yours. We have watched this single check rewrite a client's true standing, because more than a dozen firms shared their name inside the engines. The full story is in the case study.

4. Does it rerun on a fixed cadence? One reading is a snapshot that starts aging the moment it is taken, because the platforms change their models and their sources constantly. Monthly tracking on the same locked test is what turns snapshots into a trend you can trust. It also guards against a quieter failure. A static, widely published benchmark degrades over time as its questions leak into what the models learn from. The same NAACL research measured this directly, reporting that on one popular benchmark ChatGPT and GPT-4 could guess hidden test answers 52% and 57% of the time. A number that never changes its test is not stable. It is going stale.

The research community keeps landing on the same standard: a score is trustworthy when you can see how it was made. The Stanford team behind HELM, one of the most cited efforts to evaluate AI models fairly, built the project "to improve the transparency of language models" and chose to "release all raw model prompts and completions publicly" so anyone can check the work. Hold any AI recommendation score to that same bar. If the method is a secret, the number is a leap of faith.

Ask a vendor these four questions and watch what happens. The ones running a real benchmark answer plainly. The ones selling you a chart change the subject.

A number you read, or a race you run: what an index actually buys you

Here the honest comparison is not feature against feature. It is two different kinds of thing. A published index is a product you read. A locked benchmark run as a service is a race someone runs for you and can prove. Comparing them head to head is apples to oranges, so the real question is which one your situation calls for.

An independent index publisher does something genuinely useful. It stands outside every vendor, applies one yardstick to a whole category, and gives the market a shared reference point. That neutrality has real value. A number nobody in the race controls is worth reading, and a common standard is how any young field grows up. If what you want is a rough, third-party read on where a category sits, an index earns its place, and we would not tell you to ignore one.

What an index cannot do is tell you whether you are winning your own comparisons, because it was never built around you. It did not ask your buyers' exact questions. It did not hold your real competitors constant. It cannot verify each answer against your domain, and it will not write the pages that change the result. It describes the race from the grandstand. It does not run it in your lane.

Trinzik sits on the other side of that line. We are a boutique service, white glove by design, with our own technology underneath. We run a locked head-to-head benchmark: your buyers' real questions, put to the major AI platforms on a fixed schedule, against the exact competitors taking your deals, every result checked against real domains. That benchmark is one instrument of several. We pair it with third-party SEO and competitive data, so the read reflects what is actually happening on the ground and not one dashboard's opinion of it. Then our editorial team writes the evidence the platforms said was missing, a named person approves every word, and the next run tells us whether the recommendation moved.

The questionA published indexTrinzik
What it isA number you readA benchmark run for you, and proven
Who sets the questionsThe publisherYour buyers' real queries, locked
The competitor fieldFixed by the publisher, category-wideYour actual rivals, held constant
VerificationUsually name-level, if stated at allEvery result checked against your real domain
CadenceWhenever the publisher updatesThe same locked test, every month
After the scoreYou are told the numberWe write the content that changes it, human-approved
Who does the workYou read and interpretOur team runs the whole loop
A pipeline number you cannot verify is a number that costs you your forecast. An AI recommendation score is no different.

We keep our own definition of a win deliberately narrow, because a loose one flatters everybody. A win is a first-place recommendation for your verified domain, on one locked buyer-intent question and platform, rerun on the same schedule. Not a mention, not your name in a list. First place, your real domain, the same question next month. That is how we count, and it is the number we are willing to be judged on.

The method is not guesswork. The original GEO study by Aggarwal and colleagues found that optimizing content can "boost visibility by up to 40% in generative engine responses." Our practice is that lever, aimed at the exact comparisons you are losing.

480%

verified AI recommendation wins in 11 weeks

Soapbox Bulletin, 5 to 29 of 80 locked queries

300%

more wins, May to June

wealth-management client, from a zero April baseline

40%

visibility gain shown in the research

original GEO study, Aggarwal et al., 2023

Based in Austin, Texas, our team runs that loop end to end. The honest read: if you want a neutral, third-party snapshot of a whole category, a published index is a fair tool. If you want to know whether the AI recommends you over the specific rivals taking your buyers, and you want someone to own that number and move it, that is what a locked, verified benchmark run as a service is for. Our published case studies hold the method to numbers you can check.

Where this is heading

Scores are about to be everywhere. A year ago, a number for your standing in AI answers was a novelty. Soon every vendor will hand you one. When that happens, the score stops being the differentiator and the method behind it becomes the whole game. The durable question is not what did I score. It is can I audit it, and did it end in a verified change to what the AI recommends, with a person in control of everything published to get there. Apply that test to every number you are handed, ours included. A score that cannot survive the four questions was never worth trusting. One that can is worth building a strategy on.

Questions this raises

What is an AI recommendation score?

An AI recommendation score is a single number that sums up how often and how favorably AI platforms like ChatGPT, Gemini, Perplexity, Claude, and Grok recommend a brand when people ask them category questions. Industry benchmarks turn that number into rankings. There is no standard behind it: each publisher chooses which questions to ask, which platforms to count, and how to add it up, so two scores with the same name can measure very different things.

How do you audit an AI recommendation score?

Ask four questions. Are the prompts locked, so every run is the same test? Is the competitor field held constant, so a higher score means a stronger you and not a weaker field? Is every result verified against your company's real web domain, not just your name? And does the test rerun on a fixed schedule? A vendor running a real benchmark answers all four plainly. A vendor selling a chart changes the subject.

Are published AI recommendation benchmarks reliable?

A neutral, third-party benchmark has real value as a shared yardstick for a whole category, and it is worth reading for a rough outside view. It cannot tell you whether you are winning your own comparisons, because it was not built around your buyers' questions or your real competitors, and a static published index can drift as its questions age. For a decision about your brand, verify against your own domain on a locked, rerun test.

Sources

  1. Deng, Zhao, Tang, Gerstein, Cohan, Investigating Data Contamination in Modern Benchmarks for Large Language Models, NAACL 2024
  2. Liang, Bommasani, Lee, Tsipras et al., Holistic Evaluation of Language Models (HELM), arXiv 2211.09110
  3. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, GEO: Generative Engine Optimization, arXiv 2311.09735
  4. Strathern, M. (1997), 'Improving ratings: audit in the British University system', European Review 5(3): 305-321 (Goodhart's law phrasing)

Read your business the way an agent will.

Book a consult and we'll run our measurement live against your own site, so you can see what the AI engines see.