We benchmarked a chatbot against our trip engine
We gave a leading general-purpose model and our own engine the same request: five days, packed, for a couple who like culture, food, views and parks. Nine cities. One run each, no retries, nothing cherry-picked afterwards. Then we checked every venue.
Ten of the venues it recommended are permanently closed.
Not out of date by a season. Closed for years, in some cases before the model was trained. A traveller following those itineraries arrives at a locked door, on a day they cannot get back.
What it got wrong
The Signature Room at the 95th
Closed permanently in September 2023. The chatbot listed 360 Chicago, the company that bought the space, as the stop immediately before dinner there.
Mona Lisa Ristorante
Closed after roughly 46 years.
Park Chow
Closed January 2018; the chain went bankrupt in 2019.
Nopalito Inner Sunset
Closed June 2020. The NoPa branch survives, in the wrong neighbourhood for that day.
Scala's Bistro
Never reopened after the Sir Francis Drake became the Beacon Grand.
The Commissary
Closed February 2021.
Pizzuti Collection
Closed to the public since March 2020.
Louies Modern
Permanently closed 10 March 2019.
Lafayette Cemetery No. 1
Closed to the public by the City since September 2019.
Morning Call
Left City Park in 2019. The itinerary placed it "under the park’s oaks".
It also invented the geography
In San Francisco it routed Alcatraz to the mainland as a 16 minute metro ride. Alcatraz is an island. There is no train, and the distance it quoted is a straight line drawn across open water.
In the same city it never mentioned Alcatraz or the Golden Gate Bridge as places to visit at all, and marked every single stop as needing no reservation, including the one attraction that genuinely sells out weeks ahead.
Across all nine cities it published no coordinates for any stop, so none of it can be drawn on a map.
Where we lost
On our second set of five cities, picked at random from our registry and never used to tune anything, the scores came out close to level. It beat us on three things:
Day composition
Its days were evenly balanced, one neighbourhood at a time. Ours were lumpier, and on one city we handed a whole day to a 90 minute waterfall. That bug is fixed; the criticism was fair.
Dollar amounts
It says "$32 per adult". We said "Premium". Its number is more useful, and we are changing ours.
Prose
It writes better sentences than we do. Ours are shorter and plainer.
What we do differently
- Every stop is a real place with stored coordinates, so the trip draws on a map.
- Trip structure is deterministic. A model writes the prose; it never picks the places.
- Travel between stops is measured, and the mode is honest. Where no rail serves a route, we say bus.
- Attractions that sell out are flagged as needing a booking.
- We do not have better sentences than a large model, and we are not pretending to.
Method
Fourteen cities in total across two rounds: four chosen deliberately, then five picked at random from our registry to check we had not tuned to the first set, plus the cities named above. Identical prompt for every city, one generation each, no retries. The comparison model was given a plain request and none of our data or instructions. Closures were confirmed against news reports and operator statements at the time of writing. Our own output was scored by the same rubric, and its failures are recorded in the section above rather than left out.
Models change, and a rerun tomorrow will not match this one exactly. The specific closures are checkable regardless: every venue listed here can be looked up.
Free to read, no account. Every stop on them is a real, open place.