Every vendor in the AI visibility category sells a number. Audit your site, get a score, raise the score, get recommended by ChatGPT. I sell one too: our WordPress plugin grades every page on the structural things assistants are supposed to reward, answer-first openings, question-form headings, schema, crawler access.
So I tested the claim the category is built on. Not once, because one study in one industry proves very little, but twice, in two industries chosen to be as different as possible, with the hypothesis, the statistical test, and the rules for reading the result written down and dated before a single query ran.
The score did not predict who got recommended in either industry. Company size did, by a wide margin, in the one study where I measured it. And the firms that did beat bigger names all did it the same way, which is the useful part.
What was I testing, and on whom?
The claim under test: sites that score higher on a structural AI-visibility audit get recommended by AI assistants more often than sites that score lower.
Study one used the 64 agencies in WordPress VIP’s service partner directory as of August 26. Same platform, same gatekeeper, all vetted for enterprise delivery, publicly listed. My own firm is not in that directory; my site was audited alongside them as a benchmark and held out of the comparison.
Study two used the IRS list of certified professional employer organizations, 60 brands consolidated from 122 legal entities as of August 27. I picked it because it fixes the biggest weakness of study one: the IRS certifies PEOs on audits and bonding, not marketing. Nobody curated this list for how well they present themselves online, so if a curated directory was compressing study one’s result, this population would show it.
In both studies I audited every site from the outside with the same instrument, then asked ChatGPT and Perplexity ten questions a real buyer would ask, in fresh sessions with memory off, and wrote down every firm they named and every source they cited. Twenty-two sessions in study one, twenty-six in study two, 285 named mentions across 106 scored firms.
How do you test this without fooling yourself?
You write everything down first. The queries, verbatim. The statistic, a permutation test on the difference in mean score between named and never-named firms. The rule for what counts as a null. The rule for what to do if the confound explains the data better than the hypothesis. Then you run it and publish whatever comes out, including the parts where you were wrong.
I was wrong several times. Study one’s analysis script claimed in its own comments that a fixed random seed made the result reproducible. It did not; three runs on identical data returned three different p-values, because the groups were built from a Python set and string hashing is randomized per process. Separately, my own site had ended up inside the never-named comparison group at a score of 83. Both fixed, both disclosed in dated amendments, and the correction moved the result toward the hypothesis rather than away from it, which is the direction that costs you the right to hide it.
Study two got pre-registration done properly because of that. Feasibility gates were frozen before the audit. The protocol was frozen before the first query. The size covariate was collected before capture so it could not be shaped by knowing who got named. The minimum detectable effect was computed in advance instead of after.
Does a structural AI visibility score predict getting recommended?
No large effect, in either industry.
Study one, enterprise WordPress agencies: firms the assistants named averaged 73.4 out of 100. Firms never named averaged 70.5. A 2.9 point gap, p = 0.13, 95% interval from minus 1 to plus 6.6.
Study two, certified PEOs: named firms averaged 73.6. Never-named averaged 72.3. A 1.3 point gap, p = 0.56, interval from minus 3.1 to plus 5.8.
Here is the part almost nobody publishes. These designs detect a gap of about six points in study one and eight in study two, at 80% power. They would have caught the gaps I actually measured about a third of the time. So I cannot tell you the score means nothing. I can tell you that if it mattered a lot, both tests would have caught it, and neither did. The claim is no large effect, not no effect, and the difference between those is the whole argument.
Removing the curated directory in study two did not change the picture. The score estimate got smaller, not larger, in a population with visibly wider score spread.
Two rows make it concrete. In study one, the most-recommended agency in the whole study scored 65, below the median. In study two, the never-named group contains two of the three highest-scoring sites in the corpus, and eleven of its twenty-two members score above the mean. The scorer is not measuring something these firms lack. It is measuring something the assistants do not use to choose.

So what does predict getting recommended?
Size matters.
Study one could not answer this because I had not collected company size, and I said so. Study two collected it for all 45 scored firms before capture, as an employee-count bracket from public company profiles, and ran the exact same test on it.
Named firms averaged a size bracket in the 201 to 1,000 employee range. Never-named firms averaged around 51 to 200. p = 0.0004. Restrict “named” to firms named in an actual recommendation rather than a recital of the certified list, and it sharpens: the eight PEOs that were ever recommended average a bracket in the thousands, the 37 that were not average about 51 to 200, p = 0.0001. Score and size are uncorrelated in this corpus, so size is not working through the score. It is working instead of it.
The rule I wrote down in advance said that if size separated the groups better than score did, size was the headline. It did, by three orders of magnitude.

Those eight, for the record: ADP TotalSource, CoAdvantage, Engage PEO, GMS, Justworks, Paychex, Questco, and TriNet. Four are national brands, and on the general questions, best PEO, which PEO for a 40-person company, who should we switch to, both assistants returned those same names plus a few non-certified platforms, from listicles, from competitors’ own comparison pages, or from memory with no sources at all, run after run. That layer is closed. Three more, CoAdvantage, Engage PEO, and Questco, arrived only through third-party pages on the construction and healthcare questions. One, GMS, got in through its own page.
How did the small ones get in?
Through one page, answering one specific question, on an assistant willing to go read it.
GMS scores 70, below the corpus mean. Perplexity recommended it for a trucking company because GMS owns a page whose URL is, almost word for word, the question: groupmgmt.com/about-us/industries/peo-for-transportation-companies/. Nothing about GMS’s site overall put it there. That page did.
Precise PEO is not even on the IRS list. Its medical and dental industry page won a citation on the healthcare question in two of three draws, while the rest of the names in the answer changed almost completely between draws. The stable unit was the page, not the firm.

Study one had the same pattern before I knew to look for it. A small studio called CODY was recommended first by ChatGPT for enterprise healthcare WordPress work, in an answer that named no directory member at all, off a page at wearecody.com/healthcare that says exactly what an IT-review buyer wants to hear. rtCamp owned the evaluation guide for its category and was recommended off it. Human Made owned the headless-at-scale guide.
The cleanest case is TriNet, which is not small but did something deliberate. Perplexity cited trinet.com/ai-instructions, a page TriNet wrote specifically for AI assistants to read, as its source for TriNet’s healthcare expertise. That is the first time in either study I watched an intentional AEO artifact win a citation in a controlled test.
And Justworks, the most-recommended PEO in the study, named first in twelve of the twenty answers that named anyone, scores 69. Its price is public: 79 dollars and 124 dollars per employee per month. That number was quoted in five separate answers, reached three different ways, from a listicle, from Justworks’ pricing page, and from its help-center FAQ. No other certified PEO’s price appeared in any answer. Justworks is recommended first because it is the only one that gave the assistant a number to build an answer around.
Across both studies the pattern in the counts is the same. On vertical questions in the PEO study, “PEO experienced with healthcare practices,” “PEO for a trucking company,” assistants cited providers’ own pages 61% of the time. On superlative questions, 17%. On switching questions, never. The more a question describes work to be done, the more the assistant reads the firms’ own pages, and the more a firm without a national brand can win.
What are the limits of all this?
They are real and they belong next to the result. Each study is a snapshot: two assistants, one location, one day. Both assistants quietly localized answers to my state and county without being asked, so this is what a buyer in Northern California sees. The tests are underpowered for small effects, and I have quoted the detectable size rather than pretending otherwise. The size covariate is a bracket from public profiles, rough, and it still separated the groups decisively. The most-named brand in the PEO study could not be audited at all because its site blocks every crawler I have, so it is missing from every test. The scorer is a heuristic on public HTML and was corrected twice during study one. And the page-ownership finding is a tally of instances, not a tested effect; nobody ran the counterfactual.
One more limit that is really a finding. Asked to recommend on healthcare, ChatGPT named nobody in three of three attempts and returned a checklist instead, while Perplexity recommended providers and cited their healthcare pages every time. The two assistants agreed on 27% of names in study one and 37% in study two. A visibility number from one assistant describes that assistant.
What should you actually do?
Keep the structural floor. It is hygiene, and pages that are not quotable do not get quoted. Just stop treating the score as a rank you climb, and be suspicious of anyone, including me, who sells it as one.
Then do what the winners did. Pick the specific question your buyers ask right before they ask who to hire. Not “best,” the real one, with the industry or the constraint in it. Publish the definitive page on it, with the question in the URL. Put the concrete facts on it that an assistant can build an answer around, price if you dare. Accept that you will not win the superlative queries against national brands and stop paying for that fight. And measure on more than one assistant, because one of them may have decided not to answer your question at all.
Both studies, protocols, amendments, data, and every place I had to correct myself: Both studies, protocols, amendments, data, and every place I had to correct myself are on the methods and data page.
Key takeaways
- Across two industries and 106 firms, a structural AI-visibility score had no large effect on which firms AI assistants recommended (p = 0.13 and p = 0.56). Effects under six to eight points were not detectable, so the claim is no large effect, not no effect.
- Company size separated recommended from ignored firms decisively (p = 0.0004) in the study that measured it. Score and size were uncorrelated.
- 62% of the 106 scored firms were never named once by either assistant.
- Every small firm that beat bigger names did it with one page answering one specific question, often with the question in the URL, on an assistant that retrieves.
- The two assistants agreed on roughly a quarter to a third of the names.
FAQs
No. It means they measure readiness, not results. A quotable page is a floor. Two pre-registered tests found no large effect of site-level score on being recommended, and found a strong effect of company size instead.
One industry proves little. The second population, IRS-certified PEOs, was chosen because nobody curated it for marketing quality, which was the main objection to the first. The result did not change.
The queries, the statistic, the interpretation rules, the confound rule, and, in study two, the feasibility gates and the minimum detectable effect, all frozen and dated before the relevant data existed. Changes were made only by dated amendment.
Own the page that answers a specific buyer question, with the question in the URL and concrete facts on the page. That is the only mechanism by which a firm without a national brand got recommended in either study.
ChatGPT and Perplexity, in fresh sessions with memory and personalization off. Claude and Gemini were not run.


