For twenty-five years before Trinzik I sat in forecast reviews, on both sides of the table. You learn one rule fast: a number you cannot verify is a number that will cost you. A rep says the deal closes in the third quarter, and there is exactly one honest response. Show me. Where is it in writing, who signed it, what is the date. A pipeline figure without those answers is not a forecast. It is a hope with a dollar sign on it.
AI recommendation scores are that same problem in new clothing. An index publishes a number, your brand rates a 72, or ranks eighth in its category this month, and it looks like measurement. The forecast question still applies. Show me. What was asked, of which AI platforms, on what date, and how do you know the answer credited your company and not a firm that shares your name? Researchers keep finding the gap between the score and the truth. A 2024 study presented at NAACL found "a disparity between the inflated benchmark scores and the actual performance of LLMs." The lesson is not that scores lie. It is that a score is only as good as the method you can check behind it.
What is an AI recommendation score?
An AI recommendation score is a single number that sums up how often and how favorably AI platforms like ChatGPT, Gemini, Perplexity, Claude, and Grok recommend a brand when people ask them category questions. Industry benchmarks and published indexes turn that number into rankings: who leads a category this month, who moved. There is no standard behind any of it. Each publisher decides which questions to ask, which platforms to count, and how to add it all up. So two scores that carry the same name can measure completely different things.
The scores exist because the stakes are real. When someone asks an AI who to hire, it tends to name one or two brands and leave the rest invisible. A number that claims to track your place in that answer is worth having. It is only worth trusting once you know how it was built.
What a score is actually made of
Most AI recommendation scores are built from three ingredients, and the differences between them decide whether the number means anything.
Mention frequency counts how often a brand shows up in AI answers to a set of questions. It is the easiest input to produce and the easiest to misread, because a mention can be a recommendation, a passing aside, or a warning, and a raw count treats them all the same.
Share of voice, your slice of every brand mention in the same set of answers, is a better single indicator, because it is measured against a field of competitors instead of floating on its own.
Verification is the ingredient most scores skip. It asks whether the tool checked that the winning answer actually pointed to your company's real web domain, or just matched your name. That gap is where a score quietly breaks, and it is the whole difference between a benchmark and a decoration.
A published index
A number handed to you, computed somewhere you cannot see, on questions you did not set. You can read it. You cannot rerun it. When it moves, you are told that it moved and asked to believe why.
A locked benchmark
The same questions, the same rivals, put to the platforms on a set schedule, every result checked against real domains. You can rerun it and get a comparable number. When it moves, you can see for yourself what changed.
Four questions that audit any AI recommendation score
You do not need to be technical to audit a score. You need four questions, and the nerve to keep asking until you get real answers.
1. Are the prompts locked? A benchmark is only comparable to itself when the questions never change. Lock the prompt set at the start and rerun it untouched, so this month and last month are the same test. There is a second reason to lock them, and measurement has a name for it. Goodhart's law, in Marilyn Strathern's 1997 phrasing, holds that "when a measure becomes a target, it ceases to be a good measure." A score whose questions get quietly reworded to look better each month has stopped measuring anything. Locked prompts are what keep it honest.
2. Is the competitor field constant? A win only means something against the same opponents. If the rival list shifts between runs, a jump in your score can be a weaker field rather than a stronger you. Ask who you were compared against, and whether that set is held constant from month to month. In a head-to-head comparison the roster matters as much as the result.
3. Is every result verified against your real domain? This is the check that catches the expensive mistake. AI platforms confuse companies that share a name, and a score that matches on name alone can credit your win to someone else, or someone else's win to you. Real verification, what we call grounding, means the platform had to return your company's own web address, and the result counts only when that address is yours. We have watched this single check rewrite a client's true standing, because more than a dozen firms shared their name inside the engines. The full story is in the case study.
4. Does it rerun on a fixed cadence? One reading is a snapshot that starts aging the moment it is taken, because the platforms change their models and their sources constantly. Monthly tracking on the same locked test is what turns snapshots into a trend you can trust. It also guards against a quieter failure. A static, widely published benchmark degrades over time as its questions leak into what the models learn from. The same NAACL research measured this directly, reporting that on one popular benchmark ChatGPT and GPT-4 could guess hidden test answers 52% and 57% of the time. A number that never changes its test is not stable. It is going stale.
The research community keeps landing on the same standard: a score is trustworthy when you can see how it was made. The Stanford team behind HELM, one of the most cited efforts to evaluate AI models fairly, built the project "to improve the transparency of language models" and chose to "release all raw model prompts and completions publicly" so anyone can check the work. Hold any AI recommendation score to that same bar. If the method is a secret, the number is a leap of faith.
Ask a vendor these four questions and watch what happens. The ones running a real benchmark answer plainly. The ones selling you a chart change the subject.
A number you read, or a race you run: what an index actually buys you
Here the honest comparison is not feature against feature. It is two different kinds of thing. A published index is a product you read. A locked benchmark run as a service is a race someone runs for you and can prove. Comparing them head to head is apples to oranges, so the real question is which one your situation calls for.
An independent index publisher does something genuinely useful. It stands outside every vendor, applies one yardstick to a whole category, and gives the market a shared reference point. That neutrality has real value. A number nobody in the race controls is worth reading, and a common standard is how any young field grows up. If what you want is a rough, third-party read on where a category sits, an index earns its place, and we would not tell you to ignore one.
What an index cannot do is tell you whether you are winning your own comparisons, because it was never built around you. It did not ask your buyers' exact questions. It did not hold your real competitors constant. It cannot verify each answer against your domain, and it will not write the pages that change the result. It describes the race from the grandstand. It does not run it in your lane.
Trinzik sits on the other side of that line. We are a boutique service, white glove by design, with our own technology underneath. We run a locked head-to-head benchmark: your buyers' real questions, put to the major AI platforms on a fixed schedule, against the exact competitors taking your deals, every result checked against real domains. That benchmark is one instrument of several. We pair it with third-party SEO and competitive data, so the read reflects what is actually happening on the ground and not one dashboard's opinion of it. Then our editorial team writes the evidence the platforms said was missing, a named person approves every word, and the next run tells us whether the recommendation moved.
| The question | A published index | Trinzik |
|---|---|---|
| What it is | A number you read | A benchmark run for you, and proven |
| Who sets the questions | The publisher | Your buyers' real queries, locked |
| The competitor field | Fixed by the publisher, category-wide | Your actual rivals, held constant |
| Verification | Usually name-level, if stated at all | Every result checked against your real domain |
| Cadence | Whenever the publisher updates | The same locked test, every month |
| After the score | You are told the number | We write the content that changes it, human-approved |
| Who does the work | You read and interpret | Our team runs the whole loop |
A pipeline number you cannot verify is a number that costs you your forecast. An AI recommendation score is no different.
We keep our own definition of a win deliberately narrow, because a loose one flatters everybody. A win is a first-place recommendation for your verified domain, on one locked buyer-intent question and platform, rerun on the same schedule. Not a mention, not your name in a list. First place, your real domain, the same question next month. That is how we count, and it is the number we are willing to be judged on.
The method is not guesswork. The original GEO study by Aggarwal and colleagues found that optimizing content can "boost visibility by up to 40% in generative engine responses." Our practice is that lever, aimed at the exact comparisons you are losing.
480%
verified AI recommendation wins in 11 weeks
Soapbox Bulletin, 5 to 29 of 80 locked queries
300%
more wins, May to June
wealth-management client, from a zero April baseline
40%
visibility gain shown in the research
original GEO study, Aggarwal et al., 2023
Based in Austin, Texas, our team runs that loop end to end. The honest read: if you want a neutral, third-party snapshot of a whole category, a published index is a fair tool. If you want to know whether the AI recommends you over the specific rivals taking your buyers, and you want someone to own that number and move it, that is what a locked, verified benchmark run as a service is for. Our published case studies hold the method to numbers you can check.
Where this is heading
Scores are about to be everywhere. A year ago, a number for your standing in AI answers was a novelty. Soon every vendor will hand you one. When that happens, the score stops being the differentiator and the method behind it becomes the whole game. The durable question is not what did I score. It is can I audit it, and did it end in a verified change to what the AI recommends, with a person in control of everything published to get there. Apply that test to every number you are handed, ours included. A score that cannot survive the four questions was never worth trusting. One that can is worth building a strategy on.
Questions this raises
What is an AI recommendation score?
An AI recommendation score is a single number that sums up how often and how favorably AI platforms like ChatGPT, Gemini, Perplexity, Claude, and Grok recommend a brand when people ask them category questions. Industry benchmarks turn that number into rankings. There is no standard behind it: each publisher chooses which questions to ask, which platforms to count, and how to add it up, so two scores with the same name can measure very different things.
How do you audit an AI recommendation score?
Ask four questions. Are the prompts locked, so every run is the same test? Is the competitor field held constant, so a higher score means a stronger you and not a weaker field? Is every result verified against your company's real web domain, not just your name? And does the test rerun on a fixed schedule? A vendor running a real benchmark answers all four plainly. A vendor selling a chart changes the subject.
Are published AI recommendation benchmarks reliable?
A neutral, third-party benchmark has real value as a shared yardstick for a whole category, and it is worth reading for a rough outside view. It cannot tell you whether you are winning your own comparisons, because it was not built around your buyers' questions or your real competitors, and a static published index can drift as its questions age. For a decision about your brand, verify against your own domain on a locked, rerun test.
Sources
- Deng, Zhao, Tang, Gerstein, Cohan, Investigating Data Contamination in Modern Benchmarks for Large Language Models, NAACL 2024
- Liang, Bommasani, Lee, Tsipras et al., Holistic Evaluation of Language Models (HELM), arXiv 2211.09110
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, GEO: Generative Engine Optimization, arXiv 2311.09735
- Strathern, M. (1997), 'Improving ratings: audit in the British University system', European Review 5(3): 305-321 (Goodhart's law phrasing)