All posts
    AI visibilityGEOgenerative engine optimizationGuatemalaAI searchresearch

    The most recommended hotel in Guatemala has a website machines can't read

    I asked four AI engines 240 questions about Guatemala and scanned 124 company websites. What decides who gets recommended is not what you think.

    David Monzón 15 min read
    4 0
    A closed colonial door in Antigua Guatemala, with a thin red line running across the cobblestones and passing it by.

    I asked four AI engines 240 questions a real buyer would ask about Guatemala, scanned 124 company websites for machine readability, and repeated the whole exercise 24 hours later. The results explain who actually speaks for a country when AI answers, and it is almost never who you would expect.

    240

    AI answers collected across two passes

    4

    engines: Claude, ChatGPT, Perplexity, Google AI Mode

    124

    Guatemalan entities scored and tracked

    93%

    of the citation elite held their place at 24 hours

    28–40%

    overlap when the same question is asked twice

    +14

    readability points for cited firms. It still decided nothing.

    Ask any of the big AI assistants where to stay in Antigua Guatemala and one name comes up again and again: Casa Santo Domingo. In my study it was the most cited hotel in the country, recommended by all four engines, ten times in total. Then I pointed my readability scanner at its website, the same scanner I ran against 124 Guatemalan domains, and the scanner came back with nothing. To a machine trying to read it, the site of Guatemala's most AI-recommended hotel is a wall.

    A quick word on that scanner, because every number in this study leans on it. When an AI assistant checks a website, it does not open a browser the way you do. It sends a small program that asks the site for its content and reads whatever text comes back. I built a tool that does the same thing: it visits a site the way a machine does, three separate times, and gives it a score from 0 to 100 for how much of the page it could actually retrieve and read. A 92 means the machine sees nearly everything you see. A 43 means it gets fragments. And some sites return so little that there is nothing to grade: the request comes back empty or blocked. Those are the walls of this study. Casa Santo Domingo is one of them.

    That contradiction is the whole study in one image. The engines are not reading websites to decide who exists. They already decided somewhere else. Where, exactly, is what I set out to measure.

    The instruments

    What I measured

    Between the first two weeks of August 2026 I ran two instruments against the Guatemalan market. The first was a readability scan: 100 domains across eight sectors (dental clinics, hospitals, hotels, tour operators, Spanish schools, real estate for foreigners, coffee exporters, nearshore software firms), each probed three times, scored 0 to 100 for how much of the page a machine can actually retrieve and parse. The second was a citation battery: 40 questions written the way a real US buyer writes them ("I'm American and want all-on-4 dental implants abroad", "which company should I book a guided Tikal tour with"), asked in clean, memory-off sessions to Claude, ChatGPT, Perplexity and Google AI Mode. That produced 160 answers. Twenty-four hours later I re-ran the full battery on Claude and ChatGPT, 80 more answers, to measure how stable the recommendations are. Full protocol at the bottom.

    Here is what the data says.

    Finding 01

    The answer is a lottery. The club is fixed.

    Ask the same question twice, 24 hours apart, and you do not get the same answer. On Claude, the average overlap between the two recommendation lists was 28 percent. On ChatGPT, 40 percent. The company in the number one spot kept it in only 38 to 39 percent of questions. Eight questions turned over their entire list overnight. Any single AI answer is close to a coin flip.

    Then zoom out and the picture inverts. Of the 41 entities that earned five or more citations in the first pass, 38 were cited again the next day. That is 93 percent retention. Thirty-five companies were cited by all four engines.

    28–40%

    same-question overlap at 24h (Claude / ChatGPT)

    38–39%

    of questions kept the same #1

    93%

    of the 5+ citation elite persisted

    So the individual answer is a lottery, but the pool the lottery draws from is a fixed club of roughly 35 to 40 companies. If you are in the club, you surface somewhere almost every day. If you are not, no volume of retries brings you up. The strategic question for every business in the country is not "how do I win today's answer". It is "how do I get into the club". Which leads to the uncomfortable part: who controls the door.

    Finding 02

    The door is guarded by an expat blog

    For Claude, ChatGPT and Perplexity, the single most-consulted source about Guatemala was not the tourism board, not a newspaper, not any company's own website. It was livinginguatemala.com, an expat blog. It was the number one source on all three chatbots (8 appearances in Claude's sourcing, 8 in ChatGPT's, 11 in Perplexity's). TripAdvisor came second.

    Here is the detail that should stop you: that blog is small. Semrush estimates it receives about 1,250 organic visits per month from the US. The engines did not pick it because it is popular. They picked it because it is machine-readable, in English, and shaped exactly like the question a buyer asks: lists, comparisons, recommendations. A hobby site with the traffic of a neighborhood newsletter is functioning as the gatekeeper for what three AI systems tell the world about a country of 18 million people.

    Google AI Mode plays by different rules. Its top sources were official company sites: antiguafinehomes.com, dentalexpertsguatemala.com, guatemaladental.com, atitlansolutions.com. It inherits Google's classic index, so businesses that did their SEO homework get read directly.

    This split is the most actionable finding in the study, and I call it the two-door model:

    The two-door model
    DoorWho reads itWhat it rewards
    Door A: discovery. Intermediary sources: blogs, listicles, directories, review platforms.Claude, ChatGPT, PerplexityBeing present and well-described in the places engines consult. Your own site's quality is irrelevant here.
    Door B: depth. Your own website, read directly.Google AI Mode (and any engine, once it has discovered you)Machine readability, structure, clear English answers to buyer questions.

    This mirrors what larger studies find at global scale. Muck Rack's May 2026 analysis of more than 25 million AI-cited links found that 84 percent of citations go to earned media, coverage in third-party sources, while paid content gets close to zero. My contribution is showing what that looks like on the ground for one specific country: the "earned media" that decides Guatemala's fate is, to a startling degree, one blog and a review platform.

    Finding 03

    A perfect website buys you almost nothing

    The comfortable assumption in the industry is that technical quality drives AI visibility. My data says: only at the margins. Cited companies averaged a readability score of 72.1. Invisible companies averaged 58.3. Fourteen points of difference. Readability helps. It does not decide.

    The proof is in the corners of the quadrant. Twenty-five companies whose sites are outright walls, pages my scanner could not evaluate at all, were cited anyway, several of them heavily: Hospital Herrera Llerandi, the second most cited entity in the entire country with 18 citations across all four engines. Casa Santo Domingo, the hotel from the opening. Cyman Travel, recommended by all four engines. Meanwhile, nine companies with clean, readable sites scoring 73 to 93 received zero citations. And three of the best-built sites in the whole sample, scoring between 94 and 99 out of 100, collected one or two citations each. Their reward for technical excellence was a rounding error.

    Cited and readable

    40

    Playing both doors. Example: Antigüeña Spanish Academy (92/100, 49 citations).

    Cited despite a wall

    25

    Fame and third-party coverage carry them. Example: Casa Santo Domingo, Herrera Llerandi. Their visibility rests on sources they do not control.

    Cited, partially readable

    25

    Discovered through Door A while their own site underperforms.

    Readable but invisible

    9+

    Well-built sites, zero citations. They fixed Door B and never knocked on Door A. Plus 21 more invisible entities with partial or blocked sites.

    If a vendor is selling you schema markup as an AI visibility strategy, this quadrant is the polite version of the rebuttal.

    Finding 04

    The national champion is a family Spanish school

    The most cited entity in Guatemala is not a hotel chain, a hospital group or an exporter. It is Antigüeña Spanish Academy, a family-run language school in Antigua: 49 citations, first place on all four engines, and a readability score of 92.

    It is the only entity in the sample that fully plays both doors. It appears in every listicle and directory the chatbots consult (Door A), and its own site is clean, fast and answer-shaped, which is why Google AI Mode reads it directly (Door B). The receipts extend beyond AI: Semrush estimates the school's site earns about 8,300 organic US visits a month, more than six times the traffic of the expat blog that gatekeeps everyone else. One small school, doing both jobs well, outperforms every large company in the country in the channel where buying decisions are moving. Nobody needed a big budget for this. Somebody just needed to do the work in the right places.

    Finding 05

    Cardex is, by industry accounts, the world's largest cardamom exporter. It has no rankable website. The engines recommend it anyway, sourced entirely from third parties: trade directories, export records, news. The same happens to Kamalbe Spanish School, which lives on a free WordPress.com subdomain, and to several coffee names.

    This is Door A in its purest form, and it cuts both ways. You can be visible with no site at all. But your entire presence in the channel is then written by other people, and you have no Door B to convert the interest, correct an error or defend the position. Which brings us to what happens when the story other people tell is simply wrong.

    Finding 06

    When the engine invents

    In pass one, Claude recommended a hostel called Tropicana in San Pedro La Laguna three separate times. There is no Tropicana in San Pedro La Laguna. The real Tropicana is a well-known hostel in Antigua, 130 kilometers away. In pass two, 24 hours later, the fabrication had vanished: zero mentions. Fabrications are not a separate phenomenon from the lottery. They are part of the same churn, and this study treats them as a measurable output category, not an anecdote.

    Finding 07

    Guatemalan software has been replaced by its neighbors

    Ask the engines for nearshore software partners in Guatemala and the names that come back are TELUS Digital, Accenture, Applaudo (El Salvador), Rootstack (Panama) and Gorilla Logic. This was one of the most stable patterns at 24 hours. Mid-sized Guatemalan firms barely register: the entire software sector collected 39 citations, against 103 for hotels and 97 for Spanish schools, and it posted the lowest average readability of any sector, 59.5.

    One local firm crossed Door B: Lomax. Its own site was consulted as a source by all four engines, nine times in total, four of them on ChatGPT alone. One. In the sector where a single contract is worth more than a season of hotel bookings, the country's story is being told, and its deals routed, by companies that are not Guatemalan.

    Finding 08

    The state enters through the back door

    Guatemala's official tourism portal, visitguatemala.com, scored 43 out of 100 in my scan, a serious readability failure for a national front door, and received zero citations in 240 answers. The engines found the state anyway: registro.inguat.gob.gt, INGUAT's dry business registry, surfaced three times as a source on ChatGPT. The machines could not read the country's marketing, so they read its database.

    Anacafé, the national coffee institution, is unreadable to machines. The one healthy institutional gatekeeper in the sample is AGEXPORT, the exporters' association, scoring 93: which is precisely why coffee-sector queries flow through its member directory. One association is quietly doing the infrastructure work of a state.

    Finding 09

    The engines are not one thing

    A finding that reframes all the others: the four engines behave like four different regimes.

    How each engine behaved across the 40-question battery
    EngineBehavior in the batteryPrimary door
    ChatGPTSearched the live web in 40 of 40 questions (39 of 40 in pass two)Door A, aggressively current
    ClaudeAnswered from memory in 11 of 40, including all five hotel questionsDoor A, partly frozen in training data
    PerplexitySearched in nearly all (skipped 4)Door A
    Google AI ModeNo AI answer at all in 11 of 40; sources are official sitesDoor B, inherits the SEO index

    The practical consequence is sharp: hotel recommendations on chatbots are largely decided by the training corpus, a photograph taken years ago that no press release will update quickly. Muck Rack's independent measurements point the same direction: ChatGPT cites live sources in 96 percent of responses, while Claude does so in only 55 percent. Optimizing "for AI" as if it were one channel is already a mistake. There are at least two channels, and they reward opposite investments.


    The playbook

    What this means if you sell from Guatemala, or anywhere like it

    The demand is real and it is moving. US searchers type "antigua guatemala hotels" 3,600 times a month, "acatenango volcano hike" another 3,600, "lake atitlan hotels" 1,600, "guatemala tours" 1,300. Gartner predicted in 2024 that a quarter of traditional search volume would migrate to AI assistants by 2026, and while the migration has been messier than the headline, the direction held: AI assistants now process billions of queries and sit inside the buying journey. Every one of those journeys that touches an AI passes through the mechanics described above.

    The playbook that falls out of the data has two moves, in order:

    1. Door A first. Get named, accurately, in the specific sources your engines consult for your category: the blogs, listicles, directories and review platforms that showed up in my source logs. This is earned-media work, not web development. It is where 84 percent of the leverage lives.
    2. Door B second. Make your own site machine-readable and answer-shaped, so Google AI Mode reads you directly and so the chatbots, once they discover you, find depth instead of a wall. This is where the 25 cited-despite-a-wall companies are exposed: their visibility rests entirely on sources they do not control.

    And for institutions: a country's front door is now a technical artifact. A 43/100 national portal is not a web problem. It is a trade problem.

    Appendix

    Who the engines actually read

    The full source log from pass 1, every domain the four engines displayed as a source across the 40 questions, has a familiar shape. Claude drew on 115 unique sources, ChatGPT on 123, Perplexity on 248, Google AI Mode on 108. In every engine, roughly four out of five sources appeared exactly once. The sourcing works like the answers: a tiny fixed head, a long lottery tail.

    Source concentration by engine, pass 1
    Engine#1 source (appearances)Unique sourcesSeen only once
    Claudelivinginguatemala.com (8)11579%
    ChatGPTlivinginguatemala.com (8)12379%
    Perplexitylivinginguatemala.com and tripadvisor.com (11 each)24881%
    Google AI Modeantiguafinehomes.com (4)10883%

    Add it up and livinginguatemala.com totals 27 appearances across the three chatbots, and zero on Google AI Mode. The country's most powerful gatekeeper does not exist in Google's regime. That is the two-regime finding compressed into one row.

    Only ten domains in the entire log were consulted by all four engines:

    The ten domains consulted by all four engines
    SourceClaudeChatGPTPerplexityGoogle AITotal
    tripadvisor.com4611223
    whatclinic.com355215
    clutch.co444214
    antiguafinehomes.com322411
    primavera.coffee332210
    dentalexpertsguatemala.com23139
    lomax.com.gt14319
    herrerallerandi.com13228
    mayanlakerealty.com12227
    antigua-rentals.com11114

    Read the list again. Past the three foreign platforms, the consensus layer is Guatemalan companies' own websites: a real estate agency, a coffee exporter, a dental clinic, a software firm, a hospital, a realty office, a rentals agency. Because Google AI Mode only reads sites directly, a readable own site is the entry ticket to being trusted by all four engines at once. The door is narrow. It is also open, and seven local companies are already standing in it.

    One note the log forces me to make: two of the walls in my scan, casasantodomingo.com.gt and herrerallerandi.com, appear in this log as sources, five and eight times. That is the IP limit from the methods section in action: a site can serve real AI crawlers while returning nothing to my probe. It does not touch the findings, because readability and citation were measured independently on purpose, but a careful reader deserves to see it flagged.

    Next

    What happens next

    This is edition one of a quarterly index. The same battery, the same scanner, the same protocol, every three months, so the country gets a time series instead of a snapshot: who entered the club, who fell out, whether the gatekeepers changed, whether the official doors got fixed.

    Every company named in this study, and every company in the underlying dataset, has a row: its readability score, its citations per engine, its quadrant, its 24-hour stability. If you want to know what yours says, or what it would take to change it, write to me. That conversation is free. Staying invisible is not.

    Ask for your row →

    Study and scanner by David · itsbydavid.com · August 2026. External references: Muck Rack / Generative Pulse, "What Is AI Reading?", May 2026 (muckrack.com); Gartner press release, February 2024 (gartner.com); demand and traffic estimates via Semrush, US database, August 2026. Dataset: 124 entities, quadrant table, pass-1/pass-2 stability files and the full source log, available on request for verification.

    Enjoyed this? Share it.

    Written by David Monzón

    I make businesses legible to AI: the structured data, agent-readable APIs and MCP layers that let AI systems find, read and recommend them. This blog is where I think out loud. I also built Parcela Nova, an AI-legible property platform for Guatemala.

    Work with me