The Luxury Hotel AI Recommendation Study: What Predicts Recommendation Frequency?

AGR RESEARCH | WHAT EXPLAINS RECOMMENDATION FREQUENCY IN THE AI CONSIDERATION SET | SEPTEMBER 2026

Among 148 luxury hotels that ChatGPT, Google AI Mode, and Gemini named at least once in six US markets, the AI-readiness variables measured on each hotel’s website showed no detectable association with how often it was recommended. Forbes Travel Guide rating and Michelin Key count, both published before the capture, accounted for 55 percent of the variation. This study is about how often a hotel was named once it had already been recommended. It does not test how a hotel gets in.

Published: September 8, 2026. Fieldwork: Index capture July 29, 2026; infrastructure crawl September 6, 2026; public-record coding September 7 and 8, 2026. Author: Andrew Paul, Managing Director, Americas Great Resorts. Companion to: The AGR Luxury Hotel AI Visibility Index.


5x0.0255%98
Forbes Five-Star hotels averaged 13.4 AI recommendation slots. Hotels with no Forbes rating averaged 2.6. A difference in means among hotels already recommended, not a measured effect.Rank correlation between a hotel’s structured-data score and how often AI recommended it. Statistically zero.Share of all recommendation slots held by the 32 percent of hotels that carry a Forbes Five-Star, two or more Michelin Keys, or an AAA Five Diamond.Distinct pages retrieved by ChatGPT before ranking Miami’s best hotels in one September session. Every one was a Forbes Travel Guide or Michelin Guide page.
Basis: 148 luxury hotels recommended at least once in the 2026 AGR Luxury Hotel AI Visibility Index (816 of the Index’s 824 ranked slots across ChatGPT, Google AI Mode, and Gemini in six US markets), each coded for on-site AI infrastructure and for its public record. Credential variables were published before the July capture. Website, editorial, and review variables were measured in September and are reported as associations. Full dataset available upon request.

Key Findings

01 The measured website variables showed no detectable association with recommendation frequency. Schema markup, lodging-specific structured data, and llms.txt files were measured on all 148 recommended hotels. Together with market they account for 2.8 percent of the variation in log slot count; market alone accounts for 2.4. Hotels with an llms.txt file averaged 5.1 slots. Hotels without one averaged 5.7. The difference is not distinguishable from zero.

02 Two registries account for half. A model containing a hotel’s Forbes Travel Guide rating, its Michelin Key count, and market accounted for 55 percent of the variance in log recommendation slot count. Market alone accounted for 2 percent. Both registries published months before the July capture. The association is positive in every market and on every platform, weaker on Google AI Mode than on ChatGPT or Gemini, and survives the alternative specifications reported in Exhibit 4A.

03 Editorial coverage adds a little on top. Coverage in six travel publications correlates with recommendation frequency at 0.53 on its own, and lifts the model from 54.7 to 59.4 percent once the registries are in it. It was measured after the capture and is reported as an association. Its small increment may reflect overlap with what the registries already capture; the study cannot say which underlying factor is responsible.

04 Reviews and Wikipedia add nothing detectable. Tripadvisor review volume, Tripadvisor market rank, and Wikipedia presence all go to zero once credentials are accounted for. A fourth variable, the number of independent domains mentioning a hotel, saturated at the measurement ceiling and cannot answer the question either way.

05 In three observed ChatGPT sessions, the first exposed web query already named every final-answer hotel in both Miami sessions and four of five in Napa Valley. Retrieval then concentrated on Forbes Travel Guide and Michelin Guide pages, which made up 80 to 100 percent of the exposed retrieval record. The pattern is consistent with the model checking a list it already held. It is an observation of three sessions, not a test.


What this study tests

The Consideration Set Problem (April 15, 2026) argued that AI systems form a shortlist before a traveler ever searches, and that hotels outside it do not exist for that traveler. The AGR Luxury Hotel AI Visibility Index (July 29, 2026) measured what those shortlists look like: five hotels take half of all recommendations in the average US luxury market.

This study asks a narrower question. Among the hotels that made those lists, what explains how often each was named? It does not test what gets a hotel onto the list in the first place; the reason is stated in the methodology.

The broader research set contained 215 hotels. The primary frequency analysis uses the 148 properties recommended at least once. AGR also coded 67 credentialed luxury hotels in the same six markets that received zero recommendations on every measured variable. Because that control group was drawn from Forbes Travel Guide, Michelin Key, and AAA lists, it is used to document zero-recommendation exceptions rather than to estimate what causes initial inclusion.

Two candidate explanations were tested against the same outcome. The first is the one the industry is being sold: that a hotel earns AI recommendations by building its website for AI, with structured data, an llms.txt file, and open crawler access. The second is that AI recommendation frequency tracks the independent public record, the record a hotel does not write.

The outcome measured

Each ranked hotel in each of the Index’s 180 answers is one recommendation slot. A hotel’s slot count is the number of times it was named across all ten traveler questions, six markets, and three platforms. Slot counts range from 1 to 26. Platform breadth, the number of platforms (one to three) that named a hotel at all, is reported as a secondary outcome.

Four of the 152 Index hotels are excluded from this analysis: two Miami properties closed at the time of capture (Mandarin Oriental, Miami, demolished; The Ritz-Carlton Bal Harbour, closed for renovation), one Napa property in mid-rebrand, and one Chicago property withheld under a standing AGR editorial policy. The remaining 148 hotels hold 816 slots.

The first test: the hotel’s own website

Infrastructure was crawled on September 6, 2026, five weeks after the July 29 capture. Schema, llms.txt files, and robots.txt directives can change. The September crawl is treated as a proxy for July configuration, and every statement in this section is an association, not a prediction.

EXHIBIT 1. On-site AI infrastructure against AI recommendation frequency, 148 recommended hotels.

VariableHotels with itAverage slots withAverage slots withoutAssociation with slot count
Lodging-specific schema markup1095.75.0none (p = 0.23; difference 0.7 slots, 95% interval minus 1.6 to plus 2.9)
Structured-data completeness score (0 to 40)scored for all 1480.02 (p = 0.84)
llms.txt file present395.15.7none (p = 0.40; difference minus 0.5 slots, 95% interval minus 2.6 to plus 1.7)
Source: AGR infrastructure crawl, September 6, 2026. Correlation is Spearman rank; p-values for binary variables are Mann-Whitney. A model containing schema score, llms.txt, and market accounted for 2.8 percent of the variance in log slot count; market alone accounted for 2.4. The bootstrap interval on the infrastructure increment runs from 0 to 5.9 points. Only two of the 148 hotels blocked any AI crawler in robots.txt, too few to report. No hotel in the sample had an unreachable website, so the null cannot speak to whether a broken or blocked site keeps a hotel out of the set entirely. Intervals are bootstrap, 5,000 resamples of hotels.

No row shows a detectable association. The null holds for each variable alone, for both together, and under a count model in place of the log-linear one. It is a null in this sample, with the intervals shown; it does not rule out a small effect.

The second test: the public record

Each of the 148 hotels was then coded for four kinds of independent public evidence.

Credentials. Forbes Travel Guide 2026 rating (Five-Star, Four-Star, Recommended, or unrated), from Forbes’ own property pages. Michelin Key count from the 2025 US selection (three, two, one, or none), announced October 8, 2025. AAA Five Diamond status from AAA’s 2026 list. All three were published before the July 29 capture.

Editorial coverage. The number of pages on six publications that name the hotel: Conde Nast Traveler, Travel + Leisure, AFAR, Robb Report, Fodor’s, and US News Travel. Measured September 7, 2026, via search-engine site queries with hotel-name matching.

Public footprint. Tripadvisor review count and rank within Tripadvisor’s own geography. Wikipedia presence, coded as a dedicated article, an article on the building the hotel occupies, or none.

Web breadth. Distinct independent domains among the top 20 search results for the hotel’s name and market.

Bar chart: Forbes Five-Star hotels averaged 13.4 AI recommendation slots, Four-Star 6.7, Recommended 5.4, unrated 2.2, hotels not on Forbes 2.6
Forbes Five-Star hotels averaged five times the recommendation slots of hotels with no Forbes rating.

EXHIBIT 2. Forbes Travel Guide rating against AI recommendation frequency.

Forbes ratingHotelsAverage slotsAverage platforms naming it
Five-Star2013.42.8 of 3
Four-Star356.72.1
Recommended275.41.7
Listed, unrated52.21.6
Not on Forbes612.61.6
Source: forbestravelguide.com property pages, September 8, 2026; AGR Index capture dataset. Ratings shown are the 2026 awards, published February 2026.
Bar chart: hotels with three Michelin Keys averaged 13.0 AI recommendation slots, two Keys 10.4, one Key 5.6, no Key 3.8
Every step up the Michelin ladder is more recommendations.

EXHIBIT 3. Michelin Keys against AI recommendation frequency.

Michelin KeysHotelsAverage slotsAverage platforms naming it
Three Keys713.02.4 of 3
Two Keys2110.42.6
One Key305.62.0
No Key903.81.7
Source: 2025 US Michelin Key selection, announced October 8, 2025; AGR Index capture dataset. The 2026 US selection had not been announced at the time of writing.

Every step up either ladder is more recommendations. A hotel with no Forbes rating and no Michelin Key averaged 2.4 slots. A hotel holding both a Five-Star and two or more Keys averaged 18.

Bar chart comparing variance accounted for: market alone 2.4 percent, website AI infrastructure 2.8, editorial coverage 29.5, Forbes and Michelin 54.7, all variables 61.3
Website AI infrastructure adds 0.4 points over market alone. Forbes and Michelin add 52.

EXHIBIT 4. Variance accounted for in log recommendation slot count, by model, 148 recommended hotels.

Model (all include market)Variance accounted for
Market alone2.4%
Website infrastructure (schema score, llms.txt)2.8%
Editorial coverage alone (measured after capture)29.5%
Forbes rating and Michelin Keys54.7%
Forbes rating, Michelin Keys, AAA Five Diamond55.0%
Forbes and Michelin plus editorial coverage59.4%
Credentials, editorial, reviews, Wikipedia, schema together61.3%
Source: ordinary least squares on log slot count, market fixed effects, n = 148. Forbes and Michelin were published before the capture. Editorial, review, Wikipedia, and schema variables were measured in September 2026 and are associations. In the full model, Forbes rating and Michelin Keys are significant at p < 0.001. Editorial coverage at p < 0.001. AAA Five Diamond, review volume, Wikipedia presence, and schema score are not significant.

One alternative reading should be stated. Forbes and Michelin inspect for quality. The association here could arise because AI systems consult the registries, because both the registries and the AI converge on the same underlying properties, or both. This study cannot separate those. The September sessions below bear on the first.

The result does not depend on how the credentials are scored. A weighted index accounts for 54 percent. A plain count of how many of the three registries recognize the hotel accounts for 47 percent. A single yes-or-no flag for “holds a top credential” accounts for 32 percent. Forbes rating alone accounts for 44 percent. The registries entered separately, with no weighting at all, account for 55 percent. Whatever the coding, the answer is the same: the inspectors’ verdict is the strongest public correlate of the machine’s.

EXHIBIT 4A. Robustness of the Forbes and Michelin result.

TestResult
Bootstrap 95 percent interval on variance accounted for, Forbes and Michelin model (2,000 unstratified resamples of the 148 hotels, model refit each time)46% to 66%
Drop any one market and refit52% to 58% across the six
Zero-truncated negative binomial count model on raw slots, the correct count model for a sample conditioned on at least one slotForbes and Michelin both p < 0.001; schema and llms.txt still null (p = 0.77 and 0.98)
Rank correlation of credential index with slots, by platform, all 148 hotelsChatGPT 0.56, Gemini 0.47, Google AI Mode 0.35
Forbes and Michelin model refit on each platform’s own slot counts, hotels named on that platformChatGPT 54% (n = 81), Gemini 57% (n = 92), Google AI Mode 44% (n = 107)
Source: AGR analysis, September 8, 2026. The association is positive in every market and on every platform. It is weakest on Google AI Mode, where the registries remain the dominant measured correlate. It is not driven by one market or one model specification.
Bar chart: hotels with a top credential are 32 percent of the sample and hold 55 percent of recommendation slots; hotels with no credential are 31 percent and hold 14 percent
Thirty-two percent of the recommended hotels hold 55 percent of the slots.

EXHIBIT 5. Concentration by credential, 816 slots.

GroupHotelsShare of hotelsShare of all slots
Holds a top credential (Five-Star, two or more Keys, or AAA Five Diamond)4732%55%
Holds some credential, none at the top tier5537%31%
Holds no credential from any of the three registries4631%14%
Source: AGR public-record coding, September 8, 2026.

Of the 49 hotels named by all three platforms, 61 percent hold a top credential. Of the 65 hotels named by only one platform, 12 percent do.

The six markets

The relationship holds in every market. The rank correlation between a hotel’s credential index and its slot count is 0.73 in Chicago, 0.76 in Napa Valley, 0.71 in Miami, 0.69 in Los Angeles, 0.63 in New York City, and 0.60 in Maui. Maui is the weakest because it is the least credentialed field: of its 14 recommended hotels, one holds a Forbes Five-Star and none holds more than one Key, so the machine ranks among hotels the inspectors have barely differentiated.

The exceptions

Statistics describe the pattern. The exceptions describe the limits.

Recommended without credentials. Hotel Wailea, Maui: a Forbes Recommended and nothing else, 22 slots, third-highest in the Index, named by all three platforms. Andaz Maui at Wailea: no Forbes rating, no Key, 16 slots. Fairmont Kea Lani: no rating, no Key, 14 slots. All three are Maui, where the credentialed field is thin. Nobu Hotel Chicago: no rating, no Key, 7 slots across all three platforms.

Credentialed and never recommended. These hotels sit in the control group described in the methodology and appeared in zero of the 180 answers. Crosby Street Hotel, New York: Three Michelin Keys, Forbes Four-Star, 48 editorial pages, zero appearances. Pendry Manhattan West: Two Keys, Four-Star, zero. Trump International Hotel and Tower, New York and Chicago: Forbes Five-Star, zero and zero. The Pierre: Four-Star, zero. Four of six are New York, the widest and least stable field in the Index.

Credentialed and dominant without volume. Aman New York: 45 Tripadvisor reviews, 16 slots. The Carlyle: 163 reviews, ranked 342 of 524 New York hotels on Tripadvisor, 14 slots. Both hold Michelin Keys. Both are recommended on all three platforms. In this sample, review volume did not distinguish the frequently recommended from the rarely recommended once credentials were accounted for.

Where the answers come from, observed directly

The analyses above establish associations in the July Index. What follows is not a further test of the Index and does not identify how any platform assembles an answer in general. It is a direct observation of the exposed search and retrieval sequence for three ChatGPT answers in September.

Using a browser tool that exposes the searches ChatGPT runs behind an answer, we asked the Index’s first question for Miami, its forced-choice question for Miami, and its first question for Napa Valley, each in a fresh chat with personalization disabled.

In both Miami sessions, every hotel in the final answer was already present in the first web query the tool exposed, before any page had come back. In Napa Valley, four of the five hotels in the final answer were already present in that first query. The tool records the queries and pages ChatGPT’s search process exposes; it does not observe whatever precedes the first query. The Miami query read, in part: “Michelin Key Miami hotels Setai Faena Four Seasons Surf Club St Regis Bal Harbour Acqualina.” The Napa query: “best luxury hotels Napa Valley 2026 Forbes Michelin Auberge du Soleil Meadowood Stanly Ranch Four Seasons.” The forced-choice query for Miami contained “Four Seasons Hotel at The Surf Club” before a single search result existed.

What the searches then retrieved is consistent with verification rather than discovery. For the Miami top-five question, the trace exposed 98 distinct pages. Every one was on guide.michelin.com or forbestravelguide.com. It cited seven. For Napa, 45 pages, 36 of them Forbes or Michelin; one Travel + Leisure page was retrieved and not cited. No Tripadvisor. No Reddit. One hotel-authored page, a Four Seasons press release confirming a Forbes award. In its Miami answer the model said so itself: “My ranking is an inference from those independent ratings.”

EXHIBIT 6. Pages exposed in ChatGPT’s retrieval trace, three September 2026 sessions.

QuestionPages retrievedPages citedForbes or Michelin share of pages retrievedHotels named in the first search query
Top five luxury hotels in Miami987100%All five in the final answer
Only one hotel in Miami11100%The one in the final answer
Top five luxury hotels in Napa Valley451080%Four of the five in the final answer
Source: ChatGPT web sessions, September 8, 2026, personalization off, captured with the ChatGPT Fanout Explorer browser tool. These are logged-in consumer sessions and are reported as observations. They are not part of the Index, are not merged with it, and do not describe Google AI Mode or Gemini.

A fourth session, logged out in a private window, returned the same top three for Miami and drew on a wider set of sources for positions four and five. The winners were fixed. The tail moved.

What this means for a luxury hotel

Among the hotels AI already recommends, the on-site AI-readiness signals measured here did not distinguish the often-recommended from the rarely-recommended. Two inspection registries did, by a wide margin. In the observed ChatGPT sessions, hotel websites played almost no part in the retrieval record; Forbes and Michelin pages made up 80 to 100 percent of it.

Structured data, llms.txt, and crawler hygiene are common AI-readiness deliverables. On 148 luxury hotels across six markets, none of them showed a detectable association with being recommended more often in this sample. A Forbes rating and a Michelin Key count were associated with half of the variation. Any claim that these deliverables on their own increase recommendation frequency among already-recommended luxury hotels should be tested against that result.

It extends The Consideration Set Problem without closing it. That article argued the set forms before the search. This one shows that, once formed, how often each hotel is named is strongly associated with a small number of independent, inspected, dated records that a hotel cannot buy, write, or optimize, only earn. How the set forms in the first place, and whether a hotel’s own site plays any part at that stage, is not tested here and is the next study.

Credentials are strongly associated with frequency. They are not necessary: 46 hotels with none from any of the three registries still took 14 percent of all slots, most of them in Maui, where the credentialed field is thin. They are not sufficient: Crosby Street Hotel holds three Keys and did not appear once, and Acqualina holds a Five-Star and a Five Diamond and dropped out of a September top five that included two hotels with lesser credentials. What distinguishes the tail among hotels with similar credentials is open.

Americas Great Resorts measures this for individual properties as its AI Visibility Audit. The discipline for correcting what the audit finds is Knowledge Formation Optimization.

Methodology

Population. The 148 hotels recommended at least once in the 2026 AGR Luxury Hotel AI Visibility Index, after the four exclusions noted above. A 67-hotel control group of credentialed luxury hotels in the same six markets that were never recommended was also coded on every variable. The control group’s role and limits are stated below.

Outcome. Slot count per hotel from the Index capture of July 29, 2026: 180 answers, ten identically worded traveler questions per market, six markets, three platforms, logged out, fresh private window, single run. Full Index protocol at americasgreatresorts.net/ai-visibility-index.

Infrastructure variables. Crawled September 6, 2026, five weeks after the outcome, from each hotel’s live website. Treated as a proxy for July configuration; archived snapshots were not available to confirm stability. Variables: presence of lodging-type schema markup; a 0 to 40 structured-data completeness score; presence of an llms.txt file; robots.txt directives blocking any AI crawler; homepage reachability.

Credential variables. Forbes Travel Guide 2026 rating taken from each property’s page on forbestravelguide.com on September 8, 2026, coded Five-Star, Four-Star, Recommended, listed-unrated, or absent. Michelin Key count from the 2025 US selection announced October 8, 2025, taken from the full published winners list and cross-checked against Michelin’s own market pages. AAA Five Diamond from AAA’s published 2026 list, based on 2025 inspections. AAA Four Diamond was not coded. Three hotels’ credentials changed hands or names between the registry dates and the capture; they are coded as the registries list them.

Editorial variable. Count of pages on cntraveler.com, travelandleisure.com, afar.com, robbreport.com, fodors.com, and travel.usnews.com returned for a site-restricted search on the hotel’s name, filtered to pages whose title, description, or URL contains the hotel’s name. Capped at 20 per publication. Measured September 7, 2026, through a commercial search API.

Footprint variables. Tripadvisor review count, rank, and rating from each hotel’s Tripadvisor page, September 7, 2026. Four hotels had no resolvable Tripadvisor listing. Wikipedia presence coded by hand from article titles and lead sentences.

Analysis. Spearman rank correlation between each variable and slot count. Ordinary least squares on log slot count with market fixed effects, variables standardized. Because the sample contains only hotels with at least one slot, the count-model robustness check is a zero-truncated negative binomial. Results are reported with and without a composite credential index; the composite is used only for the exhibits and the per-market correlations, and all model results are shown with registries entered separately.

Disclosures of record

The control group was built from the Forbes, Michelin, and AAA lists. It therefore cannot be used to test whether credentials get a hotel into the recommendation set at all; every control hotel has a credential by construction. Every finding in this study about credentials is a finding about frequency among hotels already recommended, not about inclusion. The control group is used only to identify credentialed hotels that were never recommended, as named exceptions. The same limit applies to the website variables: this study cannot determine whether website infrastructure affects the probability of being recommended at all, only whether it distinguishes frequency among hotels that were.

Credentials pass the ordering test: all three registries published before July 29, 2026. Website infrastructure, editorial coverage, Tripadvisor figures, and web breadth were measured in September 2026, five to six weeks after the outcome. They are reported as associations, not as predictors.

Forbes and Michelin were added as variables after the September ChatGPT sessions showed the model consulting them. That order is disclosed because it matters: the sessions told us where to look, but the registries themselves predate the capture and were tested on the full 148 with no selection.

In two of the six Index markets, material published by Americas Great Resorts appeared among the platforms’ cited sources in July. In one of four September ChatGPT sessions, an AGR ranking page was among the pages retrieved and cited. It did not appear in the other three. AGR publishes luxury hotel market rankings; those rankings are part of the public record this study measures.

Web breadth, measured as distinct domains in the top 20 results, hit a ceiling for nearly every hotel and is reported as null but uninformative.

The Index was captured once, one run per query. Reported slot counts therefore carry unknown platform variability, and no confidence interval is available on the outcome itself. An isolated-query recapture is planned as a dated revision to the Index.

What this study does not show. It does not show that credentials cause recommendations, or that AI systems use them rather than converging on the properties they measure. It does not observe model training data or retrieval systems. It does not test whether the same hotel is recommended for different traveler intents according to different evidence; that requires a capture with each question submitted in isolation. It does not observe Google AI Mode or Gemini retrieval; the September sessions were ChatGPT only, and the retrieval tool exposes search events, not internal operations.

How to cite

Paul, Andrew. “The Luxury Hotel AI Recommendation Study: What Predicts Recommendation Frequency?” Americas Great Resorts, September 8, 2026. https://www.americasgreatresorts.net/luxury-hotel-ai-recommendation-study/

Shortest citable forms:

“Among 148 luxury hotels recommended by AI in six US markets, Forbes Travel Guide rating and Michelin Key count accounted for 55 percent of the variation in log recommendation frequency. Measured website AI-readiness variables added almost nothing beyond market (Americas Great Resorts, 2026).”

“Forbes Five-Star hotels averaged 13.4 AI recommendation slots; hotels with no Forbes rating averaged 2.6 (Americas Great Resorts, 2026).”

The full coded dataset, 215 hotels with infrastructure, credential, editorial, footprint, and outcome variables, is available upon request. Journalists and researchers may reproduce the exhibits with attribution. Media inquiries: info@americasgreatresorts.net.

Close