A lot of AI visibility work assumes that the brands a model already “knows” have a durable advantage. If you are Hoka for running shoes, HubSpot for CRM, or Trello for project management, that makes intuitive sense. Those associations are already embedded in the model’s parametric knowledge.
The more interesting question for marketers is what happens between model updates.
If a competitor is weak or absent in that built-in layer, are they effectively locked out until the next training cycle? Or can live web retrieval change which brands make it into the answer now?
This is what we wanted to test.
We ran the same commercial recommendation prompts (BOFU and buying prompts) with live web search turned off and then turned on. The result was obvious: parametric advantage is not retrieval-proof. Astra commonly dropped brands that were dominant without search (baked in the model), while other brands that barely appeared in the baseline condition broke into the answer once live retrieval was switched on.
We saw the same thing happen with Claude Sonnet 5. Retrieval would reshuffle brand visibility, although the provider-level behavior was different. Astra would narrow the field of brands, while Claude surfaced a broader mix of brands across repeated runs.
Key findings
- Astra narrowed individual answers in 18 of 20 test groups once live search was enabled.
- Across repeated runs, Astra’s overall brand pool shrank in 17 of 20 groups.
- Strong baseline brands lost substantial visibility: Hoka fell from 17 of 20 responses to 1 of 20, Trello from 18 of 20 to 0, and HubSpot from 14 of 20 to 0.
- Other brands broke in once retrieval was available: Elation Health rose from 2 of 20 responses to 15 of 20, and Travelers from 0 of 20 to 8 of 20.
- Claude showed the same basic reshuffling effect, but not the same direction: its broader brand pool expanded in 15 of 20 groups.
- Both models became less brand-dense with search enabled, even though Claude’s answers became much longer.
The practical takeaway is this: a model’s built-in brand associations are a starting position, not necessarily a permanent moat, contrary to popular belief.
What we tested
We built a controlled set of prompts across five categories:
- CRM software
- project management software
- auto insurance
- running shoes
- healthcare software
Within each category, we tested four levels of prompt specificity, from broad category questions to more specific use-case and bottom-of-funnel prompts.
There were 80 prompt phrasings in total. Each was run five times under two conditions:
- Search off: the model answered without live web retrieval.
- Search on: the same model could search the live web before answering.
The primary study included 800 Astra responses and 800 Claude Sonnet 5 responses, for 1,600 current-model responses in total.
The question was specific:
Can live web retrieval materially change which brands make the cut, even when some competitors already have a strong parametric advantage?
For Astra, the answer was yes.
Astra did not preserve the parametric pecking order
In 18 of 20 category-by-prompt-depth groups, Astra mentioned fewer brands per individual answer once search was enabled.
That is significant because retrieval made Astra more selective, not less.
The stronger result appeared when we looked across repeated runs.
For each test group, we also calculated the full pool of unique brands that appeared at least once across all repeated responses. With search enabled, Astra’s brand pool contracted in 17 of 20 groups, expanded in two, and stayed unchanged in one.
So Astra was not just reordering the same set of brands. In most groups, it was drawing from a smaller overall pool once live retrieval was enabled.
That is important for anyone thinking about parametric advantage. A strong position in the model’s built-in knowledge clearly matters, but retrieval can still change who ends up in the final answer.
New competitors can break in
The brand-level results make that clearer.
Several brands that were highly visible without search fell sharply under Astra once retrieval was enabled.
In the relevant test groups:
- Hoka: 17 of 20 responses without search → 1 of 20 with search
- Trello: 18 of 20 → 0 of 20
- HubSpot: 14 of 20 → 0 of 20
At the same time, other brands moved in the opposite direction:
- Elation Health: 2 of 20 → 15 of 20
- Travelers: 0 of 20 → 8 of 20
This means the incumbent advantage is not locked.
A brand can dominate the model’s baseline associations and still lose ground once current web evidence is a part of the equation. A brand that is weak or absent without search can become much more competitive once retrieval is available.
For marketers, that is a much more useful conclusion than simply saying that models have “memory” or parametric knowledge. It means the live web can create a path into the answer between training cycles. So, work you do on your brand does matter, and can impact the answer engine results.
Claude reshuffled the field too, but differently
Claude Sonnet 5 showed the same broad phenomenon: turning on live retrieval changed which brands appeared.
However, the direction was different.
At the individual-answer level, 12 of 20 groups showed fewer brands with search enabled. But across repeated runs, Claude’s broader brand pool expanded in 15 of 20 groups.
That sounds contradictory until you separate brands per answer from unique brands across multiple answers.
Claude could mention fewer brands in any one response while rotating through a wider set of brands from run to run. Astra, by contrast, tended to concentrate on a smaller overall set.
That difference is important if you are measuring AI visibility. A single prompt run can hide a lot of the underlying behavior. The same brand can look stable in one answer and much less stable once you repeat the test across providers and runs.
More words did not mean more brands
Search also changed the shape of the answers themselves.
Astra’s responses became only slightly longer, moving from an average of 237 words to 263. Claude’s nearly doubled, from 276 words to 531.
Despite that difference, both models became less brand-dense.
Astra fell from 15.4 to 10.0 brands per 1,000 words. Claude fell from 20.5 to 10.1.
So the extra text did not translate into proportionally more brand mentions. Retrieval produced more explanation, comparison, evidence, and context around a smaller number of brands per amount of text.
For Astra, that fits the broader selectivity pattern. For Claude, it creates an interesting split: each answer became less brand-dense, while the total pool of unique brands across repeated runs became broader.
The pattern held after OpenAI changed models
We originally ran the OpenAI side of this experiment on GPT-5.6 Sol. Then Astra launched.
Rather than publish a study tied to the previous model generation, we reran the full 800-response OpenAI portion on Astra.
The individual-answer result stayed the same:
- GPT-5.6 Sol: 18 of 20 groups narrowed
- Astra: 18 of 20 groups narrowed
But across repeated runs, Astra became more selective than Sol. The overall pool of unique brands contracted in 17 of 20 groups under Astra, compared with 10 of 20 under Sol.
That gives us more confidence that the main OpenAI pattern was not just a quirk of one model version. At the same time, the exact brand winners and losers did not always replicate. Some did; others changed.
Keep this in mind: the broader behavior can persist even when the individual brands benefiting from it change.
What this means for marketers
The biggest implication is that parametric visibility and retrieval visibility are different layers of the same problem.
If a model already strongly associates your brand with a category, that is valuable. But our results suggest it is not an untouchable position. Live retrieval can materially reshape the final answer.
For brands that are not already dominant in the parametric layer, that is also encouraging. You may not have to wait for the next training cycle to become competitive.
The practical questions become:
- Do we appear when search is off?
- Do we survive when retrieval is enabled?
- Which competitors enter or disappear?
- Does the result hold across repeated runs?
- Does it hold across providers?
- Which third-party pages, reviews, comparisons, and authoritative sources keep showing up around those answers?
This is also why we think off-site brand evidence matters so much in AI search. Reviews, publisher coverage, comparison pages, authoritative mentions, product pages, and category associations can all become part of the retrieval environment a model uses when constructing an answer.
That does not mean a specific citation caused a specific brand movement. We cannot make that claim from this experiment. But it does mean marketers have something they can work on now, rather than treating parametric visibility as a fixed outcome until the next model refresh.
Methodology
The primary study included 1,600 responses:
- 800 from OpenAI Astra
- 800 from Claude Sonnet 5
We tested 80 prompt phrasings across five commercial categories and four levels of prompt specificity. Each prompt was run five times with live web search disabled and five times with search enabled.
That produced:
80 prompts × 5 repetitions × 2 conditions = 800 responses per provider
Astra executed web search in all 400 of 400 search-enabled responses. Claude searched in 361 of 400.
We measured:
- brands appearing per answer
- brand inclusion rates
- the total unique brand pool across repeated runs
- answer length
- brands per 1,000 words
- run-to-run similarity
Brand detection was handled through a normalized detection pipeline with additional checks for ambiguous common-word brands. We also independently recomputed the headline results from the raw response files before publication. The primary findings reproduced with no material discrepancies.
Limitations
This is a controlled marketing experiment, not a claim about the internal architecture of either model.
We can observe what changed when live retrieval was available, but we cannot see every internal decision that led to a brand being included or excluded. We also cannot say that a specific cited source caused a specific brand movement.
Models and retrieval systems change quickly, which is one reason we reran the OpenAI portion after Astra launched. The exact brand-level winners and losers should therefore be treated as time-sensitive.
The broader result is more useful:
A strong parametric position can help, but it does not lock the field. Live retrieval can reshuffle which brands make it into the answer now.
