AI Visibility Is Becoming Software. The Harder Problem Is Knowledge Formation.

For the first phase of AI search, a surprisingly valuable question was simply:

Did the AI mention us?

Then the questions became more sophisticated.

How often? In which systems? Which competitor appeared instead? What did the system say? Which sources did it cite? Where did it send the user?

Those questions created a legitimate measurement category. Companies began capturing AI answers at scale, calculating visibility and share of voice, identifying citations, detecting factual errors, monitoring competitors, and tracking changes over time.

That work matters.

It is also becoming increasingly software-defined.

The next problem is harder:

What evidence supports an explanation for the representation we observed, what can actually be changed, and what happened after the change?

That is the problem this article is about.

The distinction is not that software measures while humans think. That line is already disappearing. The distinction is between observing an output and making a defensible claim about the conditions associated with it.

Within AGR’s current framework, AI visibility is an observable outcome. Knowledge Formation Optimization (KFO) governs the public source environment relevant to AI-mediated representation and measures whether AI systems reproduce intellectual frameworks and entity definitions accurately across relevant queries and over time.

The internal machinery remains partly opaque. Outside observers generally cannot inspect proprietary model weights, candidate-selection logic, source weighting, ranking rules, system instructions, or the complete relationship between retrieval and generation.

That opacity does not make analysis impossible.

It makes evidence discipline more important.

Profound Shows Why the Old Software-Versus-Strategy Distinction Is Already Obsolete

Profound is useful because it complicates the easy version of this argument.

On September 15, 2026, Profound announced a $180 million Series D at a $1.8 billion valuation, co-led by Sequoia Capital and Kleiner Perkins. The company says it started as an analytics platform for marketers trying to understand how buyers discover brands through AI.

It did not remain a monitoring dashboard.

Its Answer Engine Insights product measures how brands appear across AI systems. FactCheck compares claims in AI answers with a brand’s source of truth and shows citation URLs associated with inaccurate claims. AI Marketer analyzes data, identifies opportunities, and can act on approved recommendations. Context Manager supplies company-specific knowledge. Agents and Projects extend the platform into execution.

That matters because the simple contrast:

software measures, experts diagnose

is no longer credible.

Software can perform investigatory tasks. It can surface cited sources, compare claims, identify patterns, recommend actions, and increasingly execute work.

So the durable distinction cannot be tooling.

It has to be the standard applied to the work:

What is observed? What is merely associated? What mechanism is being hypothesized? What evidence would count against that hypothesis? What intervention is justified? What did the recapture actually show?

If software eventually performs that discipline better than a consultant, the discipline still matters. The identity of the operator does not change the evidentiary standard.

Four Evidence States Should Govern the Diagnosis

The most dangerous sentence in AI visibility work is usually some version of:

This is why the model recommended them.

Sometimes it may be right.

Often the evidence supports something weaker.

A rigorous analysis should distinguish four states.

Observed

Something happened under documented conditions.

A hotel appeared in an answer. A claim was made. A citation was shown. A booking link routed to an intermediary. A competitor appeared first.

The observation should preserve the platform, date, prompt, geography and account state where relevant, answer order, material wording, visible citations, and destination links.

Observation is the foundation. It is not explanation.

Associated

Two things co-occur in a measured dataset or repeated capture.

Hotels with a particular credential are recommended more often. A source class repeatedly appears around a category. A competitor has broader third-party coverage and also appears more frequently. A stale fact exists across multiple public records and is reproduced in an AI answer.

Association can be strategically useful. It still does not establish mechanism.

Hypothesized

A plausible explanation is proposed.

Perhaps an entity is ambiguous. Perhaps a stale third-party record contributes to an inaccurate answer. Perhaps an intermediary has a clearer public representation than the official brand. Perhaps the system already has a candidate set and uses web retrieval mainly to verify or enrich it.

Perhaps none of those explanations is correct.

A hypothesis earns attention because it fits the evidence and a plausible mechanism. It does not become true because it sounds coherent.

Tested

A bounded intervention or comparison creates new evidence.

A suspected stale source is corrected. An entity ambiguity is resolved. A canonical record is strengthened. A third-party record is updated. The same capture protocol is then repeated over a pre-specified observation window.

Because AI outputs vary naturally, the baseline should not be a single screenshot when repeated captures are practical. The analyst needs enough pre-intervention observation to distinguish an apparent change from ordinary answer volatility.

Even then, tested does not mean causally proven. An outcome may be supportive, null, contradictory, or inconclusive. Model versions, competitors, public coverage, retrieval behavior, and other conditions can change during the same period.

That is why failed interventions and null results belong in the record.

If a framework cannot record that its own diagnosis was wrong, it is not diagnosis. It is a sales narrative.

What KFO Means in This Article

Americas Great Resorts uses Knowledge Formation Optimization, or KFO, as a named framework. The current canonical definition is:

KFO structures, sequences, distributes, corroborates, and corrects intellectual frameworks and entity definitions across the public information environment and measures whether AI systems reproduce them accurately across relevant queries and over time.

That is the definition used here. No alternate definition is being proposed.

The framework’s epistemic boundary is equally important:

KFO operates on the source environment AI systems can retrieve from now and that future training corpora may draw from over time. It does not edit model parameters, control proprietary ranking or retrieval systems, or guarantee inclusion, attribution, citation, or recommendation. Its effects are evaluated through observable answer behavior.

AGR’s current framework diagnoses three source-environment and representation conditions: Absence, Intermediary Dominance, and Conceptual Dilution. Its five Operating Principles are Conceptual Precision, Canonical Authority Establishment, Query Mapping, Conceptual Boundary Defense, and Adaptive Representation Monitoring.

The four evidence states in this article do not replace those conditions or principles. They are an evidence-control discipline for evaluating claims made while applying them.

Likewise, AGR’s historical phrase formation layer failure is not being used here to claim a directly observed hidden model stage. Current doctrine treats it as practitioner-facing shorthand for source-environment and representation conditions visible through persistent external patterns.

Sometimes the correct conclusion will be:

No defensible intervention has been identified.

That is a valid result.

A Source Environment Is Not a Citation List

Visible citations are valuable evidence. They are not necessarily a complete explanation of the answer.

A citation can support a factual statement without determining which entity entered the answer. A source can be retrieved and not shown. Different queries can expose different sources. Prior model knowledge may contribute. Product-specific ranking or policy logic may intervene.

For that reason, a source-environment analysis has to be broader than collecting the URLs displayed under one response.

For a luxury hotel, the public record can include the official site, ownership and management records, maps, OTAs, review platforms, Forbes Travel Guide, Michelin, AAA, destination organizations, travel publications, dining coverage, development records, archived pages, and commercial intermediaries.

The existence of those records is observable. Their influence on a specific answer is not automatically observable.

That is where the evidence states matter. A source may be present in the environment. It may be associated with an outcome. It may fit a plausible mechanism. None of those facts, by itself, makes it causal.

AGR’s Hotel Research Shows Why the Website Is Not Enough

Luxury hospitality is a useful test environment because a hotel is represented simultaneously as a property, brand, destination, experience, rating object, review object, meeting venue, dining ecosystem, and booking destination.

Those representations are maintained by organizations with different incentives and update cycles.

On July 29, 2026, the AGR Luxury Hotel AI Visibility Index captured 824 ranked hotel recommendation slots across 180 answers from ChatGPT, Google AI Mode, and Gemini in six U.S. luxury hotel markets, using ten traveler-intent questions per market.

The output was concentrated and inconsistent.

Twenty-five properties accounted for 53.4% of all ranked recommendation slots. Of the properties named, 44.7% appeared on only one of the three AI surfaces. Across comparable query sets, the three systems disagreed on the lead hotel 70% of the time.

Those are observations, not explanations.

AGR then asked a narrower question:

Among hotels already recommended at least once, what measured variables were associated with how frequently they were recommended?

That boundary is critical. The Luxury Hotel AI Recommendation Study analyzed 148 hotels already inside the observed recommendation set. It does not explain how a hotel gets included in the first place.

A model containing Forbes Travel Guide rating, Michelin Key count, and market accounted for 54.7% of the variance in log recommendation slot count. A model containing measured website infrastructure variables, including structured-data completeness and llms.txt presence, plus market accounted for 2.8%. Market alone accounted for 2.4%.

Forbes Five-Star hotels averaged 13.4 recommendation slots. Hotels with no Forbes rating averaged 2.6.

The association is strong. The causal interpretation is not settled.

Forbes and Michelin may be sources AI systems use. They may also be proxies for underlying hotel quality, brand strength, public prominence, editorial attention, or other unmeasured factors. Both mechanisms may operate. The study cannot separate them.

A hotel also cannot manufacture a Forbes rating or Michelin Key through optimization. Those credentials are earned through external systems. The strongest measured correlate is therefore not a convenient KFO lever.

Editorial coverage adds a useful qualification. Coverage across six travel publications correlated with recommendation frequency at 0.53 on its own and increased the Forbes-and-Michelin model from 54.7% to 59.4%. That coverage was measured after the July capture, so it is an association, not a leading indicator or proof of causation. It may overlap with the same underlying quality and prominence captured by the registries.

The point is not that “more sources” cause more recommendations. The number of independent domains mentioning each hotel hit the measurement ceiling and could not answer that question. Tripadvisor review volume, Tripadvisor market rank, and Wikipedia presence added no detectable explanatory value once credentials were accounted for.

The narrower lesson is that the measured website variables did not explain recommendation frequency in this sample, while parts of the independent public record were strongly associated with it. That is enough to justify looking beyond the official website without pretending the study has identified a causal formula.

The Public Record Can Remain Wrong After Reality Changes

The same problem appears in a different form when the entity itself changes.

Mandarin Oriental, Miami closed on May 31, 2025. Its former 23-story building was demolished by controlled implosion on April 12, 2026.

AGR’s July 29 capture still recorded the former hotel in recommendation output 108 days after the building was gone.

That observation does not tell us why.

The persistence could involve prior model knowledge, retrieval freshness, stale third-party material, indexing lag, historical authority, or some combination.

What it establishes is simpler:

The current real-world state of an entity and the state reproduced in an AI answer can diverge materially.

Correcting a first-party website does not erase the distributed historical record surrounding a property.

That is why representation cannot be reduced to a page-level optimization problem.

Fan-Out Research Adds One Useful Constraint on Certainty

Recent retrieval research adds a narrow but important caution: useful retrieval directions do not necessarily have to exist as readable text queries.

A 2026 ICML paper listed by Google Research, Efficient, Property-Aligned Fan-Out Retrieval via RL-Compiled Diffusion, introduces Retrieve-for-Train, or R4T. In simplified terms, the framework uses reinforcement learning to train objective-aligned fan-out behavior, compiles that behavior into supervision, and trains a compact diffusion retriever that generates target directions directly in continuous embedding space.

Google Research reports a 12 to 20 times fan-out speedup for the diffusion stage over autoregressive approaches in its evaluated setup. The work used fixed fashion and music databases, not the continuously changing web.

There is no evidence in that research that Google Search or Google AI Mode currently uses R4T in production.

Dan Petrovic published an independent simplified demonstration on September 24, 2026. He explicitly skipped the reinforcement-learning stage and used synthetic training data before applying a diffusion vector fan-out model to embedding-based retrieval. He also states there is no guarantee Google will deploy the research.

The relevance here is modest:

A readable list of secondary queries should not be mistaken for a complete map of retrieval.

If a system exposes fan-out strings, they can be useful evidence. They are still evidence, not the whole mechanism.

KFO does not need R4T to be true. The observation-versus-diagnosis problem exists regardless.

Is KFO Actually Different From Advanced GEO?

This is the question KFO should have to survive.

AGR’s canonical boundary is explicit: GEO is a retrieval-positioning discipline. KFO is a source-environment and category-authority discipline.

Sophisticated GEO and AEO practitioners already work with entities, citations, structured data, content architecture, digital PR, corroboration, and brand mentions. Entity SEO, reputation management, technical SEO, and experimental SEO already contain relevant methods.

KFO does not become distinct by claiming ownership of those tactics. The distinction has to show up in the operating object and the work.

AGR has already moved that distinction beyond a category claim. The formal Knowledge Formation Optimization framework paper establishes the theoretical foundation, while Knowledge Formation Optimization: A Testable Application of Established AI Mechanisms puts the proposed distinction into measurable form. The testable application defines the proposed KFO effect as something that must be distinguishable from conventional retrieval-side and content optimization through observable results across queries and over time. It does not treat the category distinction as established merely because AGR named it.

The companion Knowledge Formation Optimization: Draft Falsification Protocol goes further by specifying a four-arm test comparing KFO against conventional content and SEO, structure-only treatment, and a do-nothing control. The protocol remains a draft for review and is not yet externally registered or locked. Its importance here is the standard it establishes: if KFO cannot demonstrate an incremental observable effect beyond conventional optimization under the defined test, that result counts against the claim that KFO produces something distinct.

A defensible KFO engagement should therefore be able to produce:

  1. repeated baseline captures across defined queries, systems, and conditions;
  2. an entity-and-claim inventory;
  3. a source-environment map that does not pretend availability equals influence;
  4. competing hypotheses with evidence states attached;
  5. a prioritized intervention only where a controllable lever exists; and
  6. recapture under a defined protocol, including null and inconclusive outcomes.

Those are not new KFO principles. They are the evidentiary record needed to apply the existing framework without upgrading correlations into causes.

Sophisticated GEO practitioners may already apply parts or all of that discipline. Comparative cases and controlled testing, not rhetoric, will determine how much practical separation exists.

The purpose of a category is not to create vocabulary. It is to create a useful distinction that changes diagnosis, action, or measurement.

KFO should be judged by that standard.

The Harder Problem Is Justified Intervention

AI visibility began with a useful question:

Did the answer include us?

That question is becoming easier to measure at scale.

The harder question is:

What evidence justifies changing anything?

Bad AI optimization will skip that question.

It will see a citation and name a cause. It will see a competitor and generate more content. It will see a correlation and design an intervention around it. It will see an answer change afterward and declare the intervention proven.

A more rigorous discipline has to resist all four temptations.

The sequence is simple:

Observed. Associated. Hypothesized. Tested.

The standard is not certainty. Certainty is often unavailable.

The standard is knowing which kind of statement you are making, what evidence would weaken it, and whether there is a controllable part of the public environment worth changing.

Measurement is the beginning of the diagnosis, not the end of it.


References and Source Material

Close