A hotel’s appearance in an AI answer is measurable. Its persistence, explanation, and commercial value require different evidence. This companion analysis of AGR’s July 29 capture shows why those distinctions matter.
A luxury hotel appears in an AI answer. Its name is prominent, the description is accurate, and the recommendation matches the traveler’s request. That is a useful observation. For a hotel owner paying to improve AI visibility, the next questions are harder: how consistently does the hotel appear, what distinguishes it from its competitors, and what evidence connects a change in visibility to the work being purchased?
Those questions require different measurements. A screenshot records an appearance. Repeated tests estimate how often it happens under the tested conditions. A comparison of hotels can identify associations with public information. Establishing that a marketing intervention caused an improvement requires a stronger design.
When a visibility provider uses a score to explain why one hotel appears more often than another, both the explanation and the measurement need scrutiny. Our hotel analysis illustrates the first challenge; repeated testing addresses the second.
We examined one part of that problem by returning to the recommendation capture behind the AGR Luxury Hotel AI Recommendation Study. This companion analysis uses revised property matching and source records, with different credential and website measures from the earlier publication. It reuses the July answers; it does not independently replicate the study or test current AI behavior.
Publisher URLs track appearances in this capture
The analysis covers the 148 hotels recommended at least once in the July 29, 2026 capture: 816 recommendation slots across six US markets and three AI platforms. A slot is one appearance in an answer. Hotels occupied between 1 and 26 slots, averaging 5.5.
For each hotel, we measured the number of distinct, property-matched URLs returned by retained searches across six publishing sites: Condé Nast Traveler, Travel + Leisure, AFAR, Robb Report, Fodor’s, and U.S. News Travel. Searches were capped, and the resulting count measures what this collection retrieved. It is not a census of everything written about a hotel. Discussion-forum pages were excluded, and ambiguous names required more specific matching.
Within the July 29 answers, hotels with more property-matched URLs returned by this six-publisher search collection generally occupied more recommendation slots. The rank correlation was approximately 0.55.
We then accounted for market and two documented credential measures: Key counts in Michelin’s 2025 selection, and membership in AAA’s published 2026 Five Diamond list. These dated award rosters support coding membership and non-membership across all 148 hotels, without reconstructing every later status change. Our retained Forbes records establish property-level categories for only 89 hotels. We examine those separately below; the full-sample table is not adjusted for Forbes.
| Measures included | In-sample fit to log slot counts (R²) |
|---|---|
| Market | 2.4% |
| Market, Michelin Keys, and AAA annual-list membership | 37.0% |
| Those measures plus matched publisher URLs | 50.7% |
In this captured set, higher retrieved-URL counts were associated with higher slot counts after the listed market, Michelin, and AAA measures were included. Adding the URL measure improved model fit by 13.7 percentage points. R² describes how closely the model fits differences in log slot counts; it does not allocate causes, explain a percentage of AI decision-making, or predict bookings. The 13.7-point increment does not account for all three credential measures and must be read alongside the Forbes analysis below.
The association with the retrieved-URL measure also remained positive when the model was refitted after removing each market in turn. Restricting the URL measure to matches in page titles or URLs, rather than allowing matches in search snippets, did not remove the relationship.
What changes when Forbes is included
We compared four models on exactly the same 89 hotels to distinguish adding Forbes from changing the sample.
| Measures included, all on the same 89 hotels | In-sample R² |
|---|---|
| Market, Michelin, and AAA | 48.1% |
| Those measures plus publisher URLs | 52.2% |
| Market, Michelin, AAA, and Forbes | 56.4% |
| Those measures plus publisher URLs | 58.8% |
On this fixed sample, the publisher-URL measure adds about 4.0 percentage points before Forbes is included and 2.4 afterward, calculated from unrounded results. These increments depend on the sample and model specification; they do not identify why the measures overlap.
The 89 hotels average 7.46 slots, compared with 2.58 for the other 59. The subset therefore represents a different mix of recommendation frequency. The 13.7-point full-sample increment and 4.0-point subset increment describe different groups; neither can be substituted for the other.
Within the 89 hotels, the publisher-URL coefficient’s 95% bootstrap interval is 0.041–0.556 before Forbes adjustment and −0.033–0.486 afterward. The first stays above zero; the second includes zero. This supports caution about a separately identifiable publisher-URL association after accounting for all three credentials. The interval comparison itself does not prove that Forbes explains the relationship.
The 2026 Forbes Star Award winners were announced on February 11, before the July capture. We read issuer records in September; the collection date is not the award date. However, these retained observations do not independently reconstruct every property’s July status. The other 59 hotels remain unresolved for this variable, rather than being coded as confirmed absence from Forbes.
Coverage, credentials, reputation, and quality may overlap. These models do not establish which pages an AI used. The earlier article separately reports three September ChatGPT sessions dominated by Forbes and Michelin retrieval. That observation concerns different answers and is not evidence that the six publishers generated the July recommendations. AGR study and retrieval observations
Website measurements need the same care
The website comparisons concern how often already-recommended hotels appeared. They cannot establish whether infrastructure helps a hotel enter the recommendation set, supports accurate descriptions, or contributes to bookings.
Of 148 hotels, 123 had usable direct schema observations: we excluded 21 with inferred or template-based values and four without usable direct observations. For llms.txt, 136 were usable; 12 had blocked, failed, or otherwise unusable responses. Those missing observations were not treated as absent features.
On identical samples before and after adjustment, adding lodging-schema presence or a raw schema field count to the market, Michelin, AAA, and editorial model contributed less than 0.1 percentage point of fit. Adding llms.txt presence contributed about 0.1 point. These September observations added little information about July slot counts under this specification. The field count is a different measure from the earlier study’s 0–40 rubric.
An appearance is an observation. How often it repeats matters.
An AI answer can change when the same question is asked again. A hotel’s position in one answer therefore provides limited evidence about persistent visibility.
SparkToro’s company-published research on repeated brand and product prompts found substantial changes in recommendation lists, alongside brands that appeared consistently. It supports measuring appearance frequency over repeated tests. Its volunteer-based design differs from AGR’s capture, and it does not establish a universal repeat count for hotel audits. SparkToro disclosed collaboration with AI tracking vendor Gumshoe. SparkToro research
A 2026 preprint found substantial variability in cited-domain visibility across three generative-search platforms and consumer-product topics, with many differences falling within measurement noise. Its outcome was citation visibility, distinct from hotel recommendations. Author Ronald Sielinski lists an affiliation with IQRush, which offers AI visibility measurement. The paper is a preprint, not established peer-reviewed consensus. Paper and affiliation, IQRush
Commercial tools collect observations on a schedule. Semrush, for example, documents daily tracking of visibility, brand mentions, and domain positions. Those metrics require interpretation: a citation or brand mention is not necessarily a hotel recommendation. A daily schedule alone also does not establish whether a change exceeds ordinary answer variation. Buyers need the provider’s method for that judgment and separate evidence attributing a change to marketing work. Semrush documentation
This standard applies to AGR’s research. The capture underlying this analysis was conducted once. Our bootstrap checks repeatedly resample the observed hotels to assess sensitivity to sample composition. They do not measure how the AI platforms would vary if the questions were run again. Repeated captures are needed to establish the persistence of these findings.
What a hotel owner should ask
Before paying for an AI visibility audit, ask:
- What exactly was tested? Which traveler questions, platforms, markets, languages, dates, and account or personalization settings?
- What does the score count? Hotel recommendations, brand mentions, citations, or rank—and out of how many eligible tests?
- How often was each question repeated? What variation appeared, and what evidence separates an improvement from normal fluctuation?
- How were identities and missing records handled? Can the audit distinguish similarly named properties, brands from hotels, and failed retrievals from verified absence?
- What supports the commercial claim? What connects the measured change to the work performed, and what separate evidence connects it to inquiries or bookings?
For an owner, the association findings justify examining how the property’s identity, credentials, and descriptions appear across the public record. They do not establish that purchasing more coverage will increase recommendations. Any proposed improvement should identify the information problem, preserve a baseline, and be evaluated through repeated observations.
That standard applies equally to AGR’s Knowledge Formation Optimization work. This analysis does not test the effectiveness of KFO. A useful visibility score gives a hotel a documented pattern and enough context to decide what action the evidence supports.
Methodology and scope
The outcome is AGR’s July 29, 2026 capture of 180 answers across ten questions, six markets, and three platforms. Four excluded entities account for eight of the 824 raw entries, leaving 816 slots across 148 recommended hotels. The 67 supplied comparison hotels are retained in the research package but excluded from the regression analysis because their selection rule is unavailable and they were drawn from credential lists.
Publisher coverage was collected in September through retained Brave API responses, with up to 20 results per publisher query. The new analysis uses conservative property-name matching across six specified sites and excludes discussion forums. A returned page may mention the hotel without being primarily about it. No returned match is a zero for this retrieval measure, not proof of no editorial coverage. Publication-date metadata does not prove what a page contained at the time of the July capture.
Models use ordinary least squares on the natural logarithm of slot count, market fixed effects, and the natural logarithm of one plus the publisher-URL count. Michelin Keys enter as a numeric count; AAA annual-list membership enters as a binary variable. The Forbes sensitivity analysis uses category indicators on the 89 hotels with established records.
The report includes 2,000 hotel bootstrap samples stratified by market and alternative coverage definitions. The primary editorial increment’s bootstrap interval is 6.25–21.92 percentage points, conditional on the captured answers and measured features. Held-out-hotel checks also favored the editorial model, but those checks reuse the same capture. Neither procedure measures variation across new AI answers or identifies causation.
Relationship to the earlier publication: the original reported a 54.7% Forbes/Michelin model and 59.4% after editorial coverage. Those results use different credential coding and coverage records from this companion analysis. They are not interchangeable with the new tables. We have not reproduced the original 0–40 website score or reconciled its additional One Key hotel; the original property-level coding behind those figures was not available. This article reports the revised measures explicitly rather than certifying or silently replacing the earlier exhibits.
AGR publishes rankings and offers AI visibility and KFO services. As disclosed in the original study, AGR material was cited in two of the six July markets and one of four September ChatGPT sessions. AGR is therefore both researcher and a publisher whose material appeared in the observed answers. Original study disclosures

