Systems · August 09, 2026
Your business is not missing from AI. It was never a candidate.
We asked AI about ten real businesses 257 times. Nine of them got both answers.
Not both opinions. Both answers to the question "do you know this business."
Here is a test you can run before you finish this article. Open your AI assistant of choice, start a fresh chat, and ask it the question your best customer would ask on the day they decided to go looking. Not your brand name. The category and the city. "Who are the best custom home builders in Melbourne." "Who should I use for a business valuation in Brisbane." Whatever the buyer types when they have a problem and no shortlist.
Now do it seven more times, each one in a new chat so nothing carries over. Write down who gets named. Count how many times you appear.
Most people who run that test are surprised twice. The first surprise is how often they are absent. The second surprise is bigger, and it is the one this article is about: the list changes every time. Different names, different order, sometimes ten firms and sometimes two. If you had run it once and taken a screenshot, you would have walked away with a fact that was never a fact.
What we ran
Over the last few months we have been measuring what AI assistants say about real Australian businesses. Not brands with marketing departments. Trading businesses in Melbourne with staff, an ABN, a website and a phone that rings.
The run I want to talk about covered ten of them across two categories, custom home builders and residential architects. We put the same set of buyer questions to more than one frontier model, with live web access turned on and turned off, and repeated each question so that we had a denominator rather than an anecdote. What the machine said was recorded by a separate process with no memory of the conversation, because a model asked to grade its own output is a witness with an interest in the verdict. Two hundred and sixty four conversations were attempted. Two hundred and fifty seven came back clean.
One thing to say up front, because it is the sort of thing that gets left out. That run was on the developer surface, not the consumer product your customers use. It is not a substitute for what a buyer sees in the app. We have consumer surface work too and I will come to it, but where the numbers below come from a developer surface, I am telling you so.
What came back
Across the ten businesses there were 570 separate opportunities to be named in answer to a category question, cold, with no prompting and no brand name supplied. They were named 18 times.
Four of the ten were never named once.
These are not new businesses. They have addresses, awards, published directors, project galleries and years of trading behind them. Asked by name, the machines described them well. Asked the question a buyer asks, they were not in the room.
Then the finding that changed how I think about this whole category. As part of the run we asked each model, straight out, whether it knew the specific business. Nine of the ten got both answers. The same business, the same model, some runs saying it recognised the firm and others saying it had never heard of it. One business got four of each.
Nine out of ten got both answers. Not both opinions. Both answers.
Sit with what that does to a report. If a system contradicts itself about whether a business exists in its knowledge, then any single check of that system is a coin toss with a logo on it. You can buy that coin toss right now for a few hundred dollars a month from about forty different vendors, and it will arrive as a dashboard with a number on it and no denominator anywhere in the document.
This is not only our finding
I want to be careful here, because a lone claim from a company selling the fix is worth what you would expect it to be worth.
In January, Rand Fishkin at SparkToro published a study on the same underlying problem. Six hundred volunteers ran 2,961 prompts across ChatGPT, Claude and Google's AI. The odds of getting the same list of brands twice were under 1 in 100. The odds of getting the same list in the same order were closer to 1 in 1,000. Independent of us, different method, same conclusion: the output is not stable, so a single observation of it is not data.
What we measured is one layer beneath that. Fishkin measured the list moving around. We measured the machine changing its mind about whether the business is in its head at all. Those are different failures and the second one is worse, because a business that flickers in and out of existence cannot be optimised into a better position. It has to be made solid first.
Knowable is not reachable
The clearest case we have is our own, which is convenient for the argument and inconvenient for the ego, so here is the disclosure. Lumen & Lever is a business I am involved in. It sells AI readiness and governance work to mid-sized firms. In August we ran a consumer surface session, ChatGPT with search, and asked the question a buyer would ask: who are the best AI consultants in Melbourne, outside the big players, for small to medium business.
The machine searched 27 websites and returned 94 sources on that first question. Not one of the 27 mentioned Lumen & Lever. Not one of the 94 sources was a Lumen & Lever page. Eleven competitors were named, and three of them got the full "here is who I would call" treatment.
Then we asked about the business by name. The machine described it with accuracy and detail, straight off the site. When we pushed back and asked why it had not been mentioned, it said it would now rank the firm at or near the top, because it fit the brief well.
The machine invented nothing. It read the site, described the business well, and had still never considered it.
That is the distinction the whole industry is missing. The information was public, indexed, accurate and reachable. It was never a candidate. Everything sold under the AEO and GEO banners is aimed at the answer: schema markup, question and answer formatting, entity clean-up, getting quoted. All of that operates on a race you were not entered in. The candidate set is assembled before your content is read, and if you are not in it, the quality of your page is a matter of no consequence.
Being knowable is not the same as being reachable.
Who is deciding the category
There was a second thing in that session worth more than the first. We looked at which pages the machine leaned on to decide who was best. Three of the four were consultancies ranking their own market. A firm publishes "top ten AI consulting firms in Melbourne", puts itself in it, and that page becomes source material for the machine's answer.
We have now seen this pattern in four unrelated categories: buyers advocates, business brokers, custom home builders and AI consultants. It is structural, not a quirk of one vertical. The listicle layer has become the ranking layer, and a good part of it is written by the contestants.
Three of the four pages the machine trusted for "who is best" were consultancies ranking their own market.
The rarer failure, and the expensive one
Most of what we find is absence. Absence is cheap to miss because there is nothing to see. No error, no alert, no angry customer. You are just not there, and the enquiry goes to someone else, and nobody involved learns anything.
The other failure is smaller in volume and much sharper in consequence. Every one of the ten businesses in that run had credentials attributed to it by a machine. ABNs, ACNs, builder licence numbers, registration years, directors, office locations. Between ten and twenty distinct credential strings per business, offered with the same even tone as everything else.
Some of them are right. Some of them are not. In an earlier run we chased fourteen candidate errors and five of them survived verification against the business's own website or a public register. Not model confusion, not a bad paraphrase. A machine stating a specific licence number, a director, or an address that the register contradicts, to a buyer who has no reason to doubt it.
BrightLocal's 2026 consumer survey, a panel of just over a thousand US adults, found 45% had used AI to find a local business in the past year, up from 6% the year before. Sixty three per cent of those users say they trust what it tells them. That survey is American and Australia will not be identical, but the direction is not in doubt, and a wrong licence number attached to your name in a trusted answer is not a marketing problem.
Why the advice you are being sold does not touch any of this
The market for this is loud right now and most of it is empty. Three letters get rearranged, a deck gets built, and the offer is visibility in AI answers with no statement of how visibility will be measured, no baseline, no repeat count and no date when anything will be checked again.
Some of that is cynical. A lot of it is not. A lot of it is people who watched the ground move under their SEO business, read a few things, and are trying to sell forward into a discipline that does not have its instruments yet. I have some sympathy for that position. It stops being sympathetic when the client pays for twelve months and nobody can say whether anything moved.
Here is the reason it cannot be said. Half of these products take one reading. The output they are reading is not stable, and everyone who has measured it says so, in public, with sample sizes. One reading of an unstable system is not a measurement. It is a screenshot.
A screenshot has no denominator. It cannot tell you whether you moved or the machine sneezed.
The four questions
If you are buying anything in this space, ask these before you ask about price.
Out of how many? If they show you an answer where you appear, ask how many times they ran it and how many of those you were in. If the reply is a screenshot, you have your answer about the vendor.
On what surface? The developer API and the consumer app are different systems with different retrieval behaviour. Work measured on one and sold as the other is not being honest with you, and you can ask which one it was.
Against what truth? When a machine states your ABN or your director or your service area, someone has to check it against your own site or a public register. If nobody checked, it is not a finding, it is a transcript.
Measured again when? A number with no repeat date is a souvenir. The whole point is the before and the after, and the after has to be taken the same way as the before or the comparison means nothing.
Ask the vendor four questions: out of how many, on what surface, against what truth, measured again when.
The same failure, in a different suit
None of this is confined to marketing. It is the same failure wearing different clothes all over the AI economy right now, and once you see the shape you cannot unsee it.
Someone builds a prototype with an AI coding tool and it works on the demo path with one user and clean data. It goes to production and falls over under real load, real edge cases and real users who do things nobody typed into a prompt. MIT's Project NANDA looked at enterprise generative AI last year and reported that 95% of pilots produced no measurable impact on the P&L, off the back of interviews, surveys and a review of 300 public deployments. That figure gets thrown around without its context, so treat it as directional. The direction is the point. The build was never the hard part.
An agent makes one convincing sales call and gets a video. Nobody publishes the conversion rate across a hundred calls, or what it said to the customer who had a complaint, or what happened the week the pricing page changed.
It is the same mistake every time: one successful instance treated as a property of the system. Production is where that assumption goes to die, and the AI visibility market is running the same play with a shorter fuse, because the system it is measuring changes its answer while you watch.
What we do about it, and what we got wrong
I have been building the measurement instead of the dashboard, which is slower and less fun. Videt exists to run this with a denominator: a fixed question set, repeats, both surfaces, evidence kept as it was recorded, and a re-measure that can be compared to the first one. Observations do not get edited. When an interpretation changes, that creates a new record and the old one stays where it is. If we tell you a machine misrepresents your business, you get to see what was observed, when, and how we decided.
Which brings me to the part I would rather leave out.
When we published our own case study on Lumen & Lever, two of the findings in it were wrong. We reported that two domains were splitting authority, and they were not, because the redirects had been correct the whole time and had already been consolidated. We reported the firm was absent from a directory ranking when it was in it, near the bottom, because whoever checked read the first screen of the page and inferred the rest. That second one is a sampling error, which is the exact mistake this product exists to stop other people making.
We published two findings about our own business that turned out to be wrong, and we left the corrections in.
Both are still in the document, marked as corrections, with the reason and the evidence. They stay there because a measurement business that edits its own record without saying so is not a measurement business. You should hold us to that, and you should hold everyone selling you a number in this category to the same standard, starting with whether they can tell you the denominator.
Go and run the eight chats. It takes ten minutes and it will tell you more about your position than any report you have been sent this year.
Lee Powell founded Lumen & Lever, an AI readiness and governance advisory in Melbourne, Australia. Lumen & Lever works with mid-sized firms on document-heavy processes, and its client record is listed on Clutch.
Lee Powell created Scrivener for Windows, the writing application used by more than a million writers, and built and maintained it from 2011 until leaving the project in 2022. Behind that sit thirty years of commercial software across IBM, banking and pharmaceuticals, and a Masters in Computer Science from Oxford on a full scholarship.
Lee Powell is also the founder of Videt, the measurement platform that produced the figures in this article. Videt measures how AI assistants describe and recommend a business, using repeated runs across more than one assistant rather than a single check, and keeps every observation as evidence a customer can inspect.
Sources. Zero-click and search behaviour: SparkToro analysis of Similarweb clickstream data, January to April 2026, Rand Fishkin. AI recommendation instability: SparkToro, 28 January 2026, 2,961 prompt runs. Consumer adoption and trust: BrightLocal Local Consumer Review Survey 2026, 1,002 US adults, published March 2026. Enterprise pilot outcomes: MIT Project NANDA, "The GenAI Divide: State of AI in Business", 2025. Business measurements: Videt probe runs, Melbourne, August 2026. Lumen & Lever is a business the author is involved in and is disclosed as own work wherever it is used.