If you have repositioned your company, published new work, or changed what you sell, the AI assistants your buyers ask are probably still describing the version of you from three years ago. Not because they are wrong, but because they are working from an index, and an index lags. Here is what that looked like on my own site, why it happens, and the four changes that fixed it in about a day.
Key takeaways
- AI assistants describe your company from a stored index, not from reading your site fresh. Old pages carry more authority than new ones because authority accrues to URLs over time.
- Publishing new content does not replace the old picture. It adds a candidate to a competition your outdated pages have been winning for years.
- The fastest fixes are subtractive: delete pages telling the retired story, then link your best current evidence from the homepage so authority transfers to it.
- Check by asking each assistant about your domain in a clean session with memory off, and by watching whether AI crawlers reach your new pages at all.
What did the AI actually get wrong?
In July I repositioned Freestyle Web Solutions around AI consulting and automation, rewrote the site, and published two case studies about AI work delivered in 2025 and 2026. In August I asked ChatGPT what it could find about freestylewebsolutions.com.
It got the identity right. Founded 2011, San Francisco Bay Area, four services, founder named correctly. It even praised the specificity of who we say we help.
Then it graded the AI track record as “plausible but needs verification,” and told a prospective buyer they should ask me for evidence of recent AI projects.
That evidence was already on the site. Two case studies, both live, both describing exactly the work it said it could not confirm. What the assistant surfaced instead was client work from years earlier: AppDynamics, CaptivateIQ, BetterCloud, InCountry. Real projects, accurate, and completely wrong about what the company does now.
It also found a page I had forgotten existed: a team page listing six people who do not exist, with placeholder biographies, left over from a theme installation. The assistant quoted their names back to me and called it sloppy. It was right.

So the AI was not hallucinating. Every single thing it said was true at some point. It was describing a company that existed, accurately, using the most established information it could find. The problem was that the most established information was the oldest.
Why does old content beat new content?
Because authority accrues to URLs, and it accrues over time.
A page that has existed for four years has been crawled hundreds of times, linked to, referenced, and cached in more places than you can enumerate. A page you published last week has been crawled once, maybe, and is linked from one menu. When an assistant assembles a picture of your company, both pages are candidates, and the old one carries more weight for the same reason a well-worn path is easier to walk than a new one.
This is not how most people imagine AI search works. The intuition is that these systems read your website the way a person would: open the homepage, follow the navigation, form an impression from what is there now. What actually happens is closer to memory. The assistant combines what it has stored about you, gathered over months or years, with at most a handful of pages it fetches live during the conversation.
I can be specific about the live part, because I was watching. Running a crawler log on my own site, I could see an assistant fetch the homepage during one of these sessions. One page. Everything else in its answer came from what it already had stored, some of it years old. One fresh page, blended with a stale picture of everything else, presented in a single confident paragraph with no seam showing.
That is the mechanism worth internalizing. Publishing new content does not overwrite the old picture. It adds a candidate to a competition the old content has been winning for years.
How do you know if this is happening to you?
Two checks, and you can run both right now.
The first is to ask. Open ChatGPT, Claude, Perplexity, and Gemini and ask each one what it can find about your domain. Do it in a fresh session with memory turned off, because personalization will show you a flattering version a stranger will never see. Read the answer for what is missing and what is stale, not for whether it is complimentary. If the assistant describes services you stopped offering, cites clients from a previous era, or tells a buyer to ask you for evidence that is already published, you have index lag.
The second is to watch the crawlers. AI crawlers identify themselves when they visit, and they come in two kinds: training crawlers building the stored picture, and retrieval crawlers fetching pages mid-conversation to answer someone right now. If your new pages never appear in that log, they are not part of the picture yet, and no amount of additional publishing will change that on its own. This is the visibility problem I built a WordPress plugin to solve, because nothing in a normal analytics stack reports it.
One warning on the second check. A user agent header is a claim, not proof, and plenty of what calls itself GPTBot is a vulnerability scanner wearing the name. The vendors make this checkable: OpenAI publishes the IP ranges its crawlers operate from, and so does Perplexity. Comparing a request against those ranges is the difference between measurement and fiction.
Two names are always fake when they appear as a user agent, and they are worth knowing because they show up constantly. Google’s crawler documentation states that “Google-Extended doesn’t have a separate HTTP request user agent string.” Apple’s is blunter: “Applebot-Extended does not crawl webpages.” Both are robots.txt tokens for opting out of model training, nothing more. If your analytics shows traffic from either, something is lying to you, and whatever else that source claimed is worth doubting too.
What actually fixed it?
Four changes, none of which involved writing new content.
I deleted the pages that were telling the old story. The fictional team page, an outdated pricing page still describing staff augmentation, two service pages from the previous positioning, and a duplicate blog URL. Every one of them was a retrievable source contradicting the current version of the company. Deleting is not usually the first instinct when you want AI to say better things about you, but the fastest way to stop being described inaccurately is to stop publishing the inaccuracy.
I linked the new work from the homepage. The two case studies existed but were reachable only from the blog archive, several clicks deep. I put them on the homepage with real excerpts. That single change moves a page from the periphery of your site to the center, and internal links are how authority transfers from the pages that have it to the pages that need it.
I wrote the excerpts by hand. An auto-generated excerpt is the first two lines of the post with an ellipsis. A written one is a sentence that states what the project was and what it produced. When an assistant is deciding which of your pages answers a question, the summary you wrote is doing the arguing.
I made sure the new pages were prominent in the machine-readable files. Your sitemap and llms.txt are the closest thing to telling an assistant which pages you consider load-bearing. If your best evidence is buried on page four of a sitemap next to a 2019 tag archive, you have expressed a preference, just not the one you meant.
Then I asked the same question again.
The assessment reversed. AI-specific implementation went from “needs verification” to “demonstrated.” The recommendation went from “worth taking a meeting” to “worth interviewing: definitely.” Same company, same claims, same actual track record. The only thing that changed was which pages were reachable and which were gone.
Total elapsed time: about a day, most of it deleting things.
What does this mean for how you publish?
That publishing a page is the beginning of the work rather than the end of it.
A new URL arrives with no authority. It earns that from links, from being crawled repeatedly, and from time. Which means the important question after you publish something is not whether it is live. It is whether anything points at it, whether it appears in the files that tell machines what matters, and whether the pages contradicting it still exist.
It also means your old content is not neutral. Every page you leave up is a claim about who you are, and the oldest claims are the loudest. If you have repositioned in the last two years, there is a version of your company still being described to buyers, assembled from pages you have not looked at since you published them.
Go read them the way an AI would: as evidence, with no idea which parts you meant to retire.
Freestyle Web Solutions is an AI consulting and automation practice in the San Francisco Bay Area. We build workflow automation, AI agents, data dashboards, and websites for organizations that have outgrown their manual processes.
FAQs
Because it assembles answers from an index built over months or years, and your older pages have accumulated more authority than the ones you published recently. Every page you have ever left online is a candidate source, and the oldest ones usually win. Publishing new content adds to the competition rather than settling it.
There is no fixed schedule, and it depends more on what you change than on waiting. Deleting contradictory pages and linking new pages prominently from your homepage produced a measured reversal on my own site in about a day. Simply publishing and waiting can take months, because a page nothing links to may not be recrawled for a long time.
Delete the ones that make claims you no longer want repeated: retired services, outdated pricing, placeholder content, duplicates. Keep the ones that are still accurate, even if they are old, because their accumulated authority is useful when they point at your current work. The test is not age, it is whether the page tells a story you still want told.
AI crawlers identify themselves in your server logs, split between training crawlers and retrieval crawlers that fetch pages during live conversations. Standard analytics does not report them. Verify any crawler you find against the IP ranges the vendors publish, because scanners routinely impersonate AI crawlers to get past allowlists.

