Method and results, August 2026

We benchmarked a chatbot against our trip engine

We gave a leading general-purpose model and our own engine the same request: five days, packed, for a couple who like culture, food, views and parks. Nine cities. One run each, no retries, nothing cherry-picked afterwards. Then we checked every venue.

The finding

Ten of the venues it recommended are permanently closed.

Not out of date by a season. Closed for years, in some cases before the model was trained. A traveller following those itineraries arrives at a locked door, on a day they cannot get back.

What it got wrong

  • The Signature Room at the 95th Chicago

    Closed permanently in September 2023. The chatbot listed 360 Chicago, the company that bought the space, as the stop immediately before dinner there.

  • Mona Lisa Ristorante San Francisco

    Closed after roughly 46 years.

  • Park Chow San Francisco

    Closed January 2018; the chain went bankrupt in 2019.

  • Nopalito Inner Sunset San Francisco

    Closed June 2020. The NoPa branch survives, in the wrong neighbourhood for that day.

  • Scala's Bistro San Francisco

    Never reopened after the Sir Francis Drake became the Beacon Grand.

  • The Commissary San Francisco

    Closed February 2021.

  • Pizzuti Collection Columbus

    Closed to the public since March 2020.

  • Louies Modern Sarasota

    Permanently closed 10 March 2019.

  • Lafayette Cemetery No. 1 New Orleans

    Closed to the public by the City since September 2019.

  • Morning Call New Orleans

    Left City Park in 2019. The itinerary placed it "under the park’s oaks".

It also invented the geography

In San Francisco it routed Alcatraz to the mainland as a 16 minute metro ride. Alcatraz is an island. There is no train, and the distance it quoted is a straight line drawn across open water.

In the same city it never mentioned Alcatraz or the Golden Gate Bridge as places to visit at all, and marked every single stop as needing no reservation, including the one attraction that genuinely sells out weeks ahead.

Across all nine cities it published no coordinates for any stop, so none of it can be drawn on a map.

Where we lost

On our second set of five cities, picked at random from our registry and never used to tune anything, the scores came out close to level. It beat us on three things:

  • Day composition

    Its days were evenly balanced, one neighbourhood at a time. Ours were lumpier, and on one city we handed a whole day to a 90 minute waterfall. That bug is fixed; the criticism was fair.

  • Dollar amounts

    It says "$32 per adult". We said "Premium". Its number is more useful, and we are changing ours.

  • Prose

    It writes better sentences than we do. Ours are shorter and plainer.

What we do differently

  • Every stop is a real place with stored coordinates, so the trip draws on a map.
  • Trip structure is deterministic. A model writes the prose; it never picks the places.
  • Travel between stops is measured, and the mode is honest. Where no rail serves a route, we say bus.
  • Attractions that sell out are flagged as needing a booking.
  • We do not have better sentences than a large model, and we are not pretending to.

Method

Fourteen cities in total across two rounds: four chosen deliberately, then five picked at random from our registry to check we had not tuned to the first set, plus the cities named above. Identical prompt for every city, one generation each, no retries. The comparison model was given a plain request and none of our data or instructions. Closures were confirmed against news reports and operator statements at the time of writing. Our own output was scored by the same rubric, and its failures are recorded in the section above rather than left out.

Models change, and a rerun tomorrow will not match this one exactly. The specific closures are checkable regardless: every venue listed here can be looked up.

Free to read, no account. Every stop on them is a real, open place.