Walk the exhibit floor at HITEC and run the experiment we ran: stop at every booth with the word “voice” on the banner and ask for the demo. Inside an hour you will hear the same phone call eleven times. A warm, unhurried voice answers on the first ring, greets the caller by name, confirms two nights for the Hendersons, mentions the balcony upgrade at exactly the right moment, and wishes everyone a wonderful stay. The accents vary. The call does not.
That is the hotel voice AI market in mid-2026: a crowded field of vendors whose demos have converged so completely that operators are separating them on price, on the confidence of the salesperson, or on a coin flip. And the coin flip is not low-stakes. Skift’s analysis this spring argued that agentic AI is the biggest threat legacy hotel software vendors have ever faced, which is another way of saying that the systems hotels buy this year will help decide who owns the guest relationship for the next decade. Meanwhile the sharpest recap of this HITEC cycle named the real bottleneck without flinching: the industry’s problem has shifted from innovation to implementation, and data is simultaneously the biggest opportunity and the biggest obstacle. Translation for a buyer: the demo tells you almost nothing, because the demo is the one part everybody has solved.
We think the decision can be made rigorous anyway. Every vendor in this market, ourselves included, can be located on four architectural axes, and those axes decide nearly everything about how the system behaves in month six, long after the demo is forgotten. We call the map the Evaluation Grid, and the four questions that place a vendor on it, the Four Questions, are the referee’s rubric this post exists to hand you:
- Who changes the agent when the hotel changes: you, or the vendor’s staff?
- What does the guest see while the agent talks?
- What is the vendor paid for: minutes, or outcomes?
- Can you watch the system fail, or only read its report card?
None of these can be answered by listening to a demo. All of them can be answered in a single direct conversation with any vendor, if you know to ask.
The demos converged because the demo is the easy part
Frontier models made the first five minutes of a voice product effectively free. A competent team can stand up a pleasant, fluent demo agent in a weekend on public infrastructure, and many teams have. What cannot be stood up in a weekend is everything the demo hides: the operational model, the interface surface, the billing incentive, the accountability layer. Those are architecture, and architecture is a set of decisions each vendor committed to months or years ago that no salesperson can restyle on a call. The grid exists because the differences that matter moved to the layers you cannot hear.
The first question is who changes the agent when the hotel changes
A hotel is not a static business. Rates move daily. Packages appear and expire. The restaurant closes for a wedding buyout, the pool deck goes down for resurfacing, the cancellation policy tightens for festival weekend. Every one of those changes has to reach the agent answering your phones, and the first axis is how it gets there.
At one end sits the managed service. Travel Outlook’s Annette is the clearest and most candid expression of the model: the product is explicitly maintained by human experts who continually update the agent. For an operator who wants the phone problem to disappear and has no appetite for another console, that is a legitimate choice. It is closer to retaining an agency than buying software, and it should be evaluated like one, with the vendor’s team sitting inside the loop of every change you make.
At the other end sits the software product, and the industry’s center of gravity is visibly moving there. The clearest evidence is that Canary, one of the largest guest-technology vendors in hospitality, shipped a self-serve AI Agent Studio that puts configuration of the agent directly in the hotel’s hands. When an incumbent at that scale bets on self-serve, it is telling you where the market is going.
The question to ask is concrete. “Our cancellation policy changes tomorrow at 4pm. Walk me through, click by click, how that reaches the agent, and who performs the clicks.” A good answer is a console and a timestamp. A bad answer contains the word “ticket.”
The second question is what the guest sees while the agent talks
A voice with no screen is half an interface.
Booking a hotel room is a visual act. Guests choose rooms from photographs, compare rates side by side, read the fine print on a calendar, and check out on a form. A voice-only agent asks the guest to hold three room categories, two rate plans, and a cancellation policy in working memory, over the phone, while deciding how to spend a thousand dollars. The best human reservation agents compensate for that with skill. Software should compensate for it with a screen.
So the second axis is audio-only versus audio plus synchronized visual, and the word doing the work is synchronized. The guest hears “the corner suite has the wraparound balcony” and sees the balcony at the same moment. The agent quotes a total and the guest watches it itemize. The agent confirms and the confirmation is already on screen. What the guest hears and what the guest sees, in lockstep, for the whole call.
Watch for the escape hatch answer: “we can text a link at the end of the call.” A link at the end is not a synchronized interface. It is the moment the voice gives up and hands the guest to a web form, which is the exact experience the guest called to avoid. This axis is also where two different products quietly separate: a phone-tree replacement that deflects calls, and a revenue interface that closes them. They demo identically. They are not the same machine.
The third question is what the vendor is paid for
The dominant billing unit in voice AI is the minute. The major infrastructure platforms that most hotel-facing vendors build on, Vapi, Retell, and Bland, all publish per-minute rate cards, and vendors built on top of them mostly inherit the unit and pass it upstream to you. There is nothing sinister in the origin. Minutes mirror the underlying telecom and compute costs.
But the billing unit encodes the incentive.
The billing unit is the incentive structure, printed on the invoice.
Under per-minute billing, every call is a cost center, and the number everyone in the system quietly manages is average handle time. The vendor tunes toward shorter calls because shorter calls keep your invoice small and the renewal safe, and no line item anywhere rewards the booking. The perverse case is precise: minute nine of a call, where the guest is deciding between the suite and two standard rooms, is the single most valuable minute of the entire conversation, and per-minute pricing books it as waste. An agent tuned to wrap up gracefully at minute seven leaves that booking on the table, and the invoice looks great.
Per-outcome billing inverts the incentive: the vendor earns when the hotel earns, and the long call that closes becomes the point of the system rather than a cost overrun. The broader industry conversation is already moving this way; the analysis of conversational AI pricing models tracks the shift from per-minute and per-interaction rates toward outcome and resolution-based pricing across the category.
The question to ask: “Which line item on my invoice gets bigger when my direct bookings go up?” A good answer names an outcome. A bad answer names a meter.
The fourth question is whether you can watch the system fail
Every vendor will show you a dashboard. The fourth axis is who decides what is on it.
Vendors are advertising automation rates as high as 97%, and the hospitality-AI trends analysis that flagged that number paired it with exactly the right question: whether a claim of 97% automation counts the messages a bot deflected or the guest problems it actually resolved. That skepticism of black-box numbers is now the industry’s mood, and the pressure is toward explainable AI, a shift HotelTechReport’s Q2 2026 innovation report named a defining trend, with vendors now racing to show their work. An automation percentage without call-level evidence is not a metric. It is a press release.
A glass-box system gives you every call, every transcript, every escalation, and every fumble, searchable by your own team, including the calls that make the vendor look bad. A black-box system gives you a monthly summary of aggregates chosen by the party being graded. The difference matters beyond honesty. Legally, the operator owns what the agent says, a precedent we walked through in our reliability piece, and you cannot own words you are not allowed to read. Operationally, the fumbled calls are the improvement loop; a vendor that hides them from you is also hiding them from the process that is supposed to fix them.
The question to ask: “Pull up the worst call your system handled last week, at any property, and show me how I would have found it myself.” If they can, you are looking through glass. If the answer drifts toward privacy, aggregation, or “our team reviews those internally,” you are looking at paint.
No single axis settles it, but the corner does
Four axes, two sides each: sixteen cells on the grid, and real vendors land in mixed positions for defensible reasons. A managed service is the honest choice for an operator who wants an agency. Per-minute billing is tolerable when the call logs are open enough for you to audit what the minutes bought. No one question is the verdict.
The corners are the verdict. A vendor sitting on the wrong side of all four axes is selling you changes you cannot make, a voice with no screen, an invoice denominated in minutes, and a report card you cannot check. That is not a system. That is a demo with a contract attached.
The opposite corner is what a system looks like: you hold the controls, the guest sees what she hears, the vendor has no reason to rush you off the phone, and every call is on the record. We built FlowStay to sit in that corner, and that is the only sentence of this piece we will spend on ourselves. Put us on the grid the same way you would put anyone on it. That is what the grid is for.
Each of the Four Questions deserves a full piece, and each will get one: the operational model, the synchronized visual layer, the economics of outcome pricing, and glass-box evaluation. We have already published the companion arguments on why reliability is an architecture and why the harness matters more than the model. The grid is the frame that holds all of it.
Back to the exhibit floor.
Eleven booths. One call. The Hendersons book their balcony eleven times, in eleven accents.
Every demo sounded like the future. At most a few of them were.
You cannot hear an architecture in a demo. You can find one with four questions.
Ask who holds the controls. Ask what the guest sees. Ask what the meter charges for. Ask to see the worst call.
You’re not buying a voice. You’re buying an architecture.