Weftd back to homepage
Use Cases

ChatGPT, Instinct, Grok Bot: Our Test

Five configurations, zero human intervention, twelve restaurants suggested. What this test reveals about delegating to AI agents and about the data they rely on to decide.

Charles CousynPublished September 23, 202614 min readUpdated September 23, 2026
Three AI interfaces showing three different restaurant lists for the same request

AI assistants are changing shape. For a long time we mostly asked it to answer a question. A new generation offers something else: giving it a mission: searching, comparing, browsing a website, checking a fact, then stopping at the right moment. We wanted to see what that looks like in practice, using a task a user could easily delegate to an AI.

On 22 and 23 September 2026 we gave the exact same mission to Instinct, to Grok Bot and to ChatGPT: find three restaurants near Lyon Part-Dieu station for a business lunch, with a table for three available on a specific date.

The short answer

All five configurations worked with no human intervention at all. They suggested twelve different restaurants, only one of which appeared more than twice. None of them contacted an establishment. And in the single case we verified all the way through, the AI that followed the most rigorous verification process got it wrong, because the information published by the restaurant and by the platforms it appears on contradicts itself.

Summary

  • Five configurations tested, zero human intervention, execution times ranging from under a minute to thirty-eight minutes.
  • Twelve restaurants suggested for a single request, only one cited by more than two configurations.
  • The decisive contradiction does not come from the AIs but from a business whose official website, second website, Google listing and booking platform state four different things.

What we set out to observe

Here is the message we sent, word for word, identical across every configuration, in a fresh conversation wherever the product allowed one.

I am having lunch with two clients in Lyon on Thursday 8 October at 12:30. Find me three possible restaurants, less than ten minutes on foot from Part-Dieu station, open that Thursday lunchtime, where a table for three is available at 12:30 that day, with at least one vegetarian main course on the menu and a bill of around 35 euros per person maximum. For each one, give me the name, the address, how you verified the table is available on 8 October and where each piece of information comes from. Do not book and do not confirm anything. Stop before that.

Six constraints, five of which can be checked online. The sixth is of a different nature. The availability of a table on a specific date is generally not exposed as static information on a restaurant’s page: you have to query a booking system or contact the establishment. That constraint is what brings out the differences between agents: it forces them to decide how to obtain information they cannot simply read on a page.

After the opening message, we forbade ourselves any wording other than three neutral prompts, counted each time. None turned out to be necessary.

Three delegation experiences, not three models

AI Interface tested Observed approach
Instinct WhatsApp An agent that browses the web from an online machine, driven through messaging
Grok Bot macOS app An agent that works on its own computer and adjusts its plan along the way
ChatGPT Web Three reasoning configurations compared, from an instant reply to twelve minutes of analysis

So the difference is not only about the underlying model. It is about what the interface allows the AI to do: answer, search, click, verify, backtrack or carry an action further.

That is the important point of this test. Instinct and Grok Bot are not two more competitors to rank against ChatGPT. They are two different ways of receiving a mission, one inside a messaging thread you already keep open all day, the other inside a dedicated app that reports back step by step. We are not comparing three models, we are comparing three delegation experiences.

ChatGPT was tested in three configurations: the Go plan in Instant mode, the Go plan in Analyse mode, the Plus plan at high reasoning effort. The setting changes the result as much as the product does.

Five configurations were tested in all.

Configuration AI Interface
1 Instinct WhatsApp
2 Grok Bot macOS app
3 ChatGPT, Instant mode Web
4 ChatGPT, Analyse mode Web
5 ChatGPT, high reasoning Web

One product was left out and we would rather say why. Muse, Meta’s personal agent launched on 8 September (Meta, 8 September 2026), was not available in France on the date of the test. The protocol for this test was designed with Claude, which is likewise absent from the panel.

One useful clarification: Grok Bot is the agent service xAI launched in beta on 11 August 2026, distinct from the Grok assistant itself (Brief IA, August 2026). Each agent there runs on its own online machine.

How each AI went about it

Instinct, nine minutes, three restaurants delivered.

  • it moves through the booking flow and selects the date, the time and the number of covers
  • it relies on TheFork for availability and on Google Maps for walking distances
  • it flags on its own that two of the menus it consulted date from 2022 and 2023
  • it ends by offering to book

Grok Bot, thirty-eight minutes, two restaurants delivered.

  • it comments on every step of its work
  • it drops two of its first three leads after checking them and goes looking again
  • it is the only configuration that revised its plan mid-execution
  • it attaches three screenshots and publishes the list of what it ruled out with the reason for each rejection
  • it explicitly flags that one of its options exceeds the distance constraint

ChatGPT, from under a minute to twelve minutes, no restaurant certified.

  • it declines to certify availability it cannot verify, in all three configurations
  • its behaviour changes markedly with the reasoning level, from three addresses cited to a table with a source in every cell
  • it reaches the booking widget but does not operate it

ChatGPT therefore does not certify availability it cannot check directly. Within our protocol it stops short of operating the booking system, where the other two products go in.

Same mission, and already three different views of what doing the work means.

Twelve restaurants for a single request

The five configurations suggested twelve distinct establishments. Only one, PAMPA, is cited by more than two of them. ChatGPT’s three configurations, queried the same day with the same sentence, produced three lists whose only common entry is that restaurant.

The sets do not overlap, and the sources partly explain why. Instinct and Grok Bot rely on TheFork and display ratings out of ten. ChatGPT mixes Google ratings, hotel group listings and TheFork menus. Two of the three suggestions from its most capable configuration are hotel restaurants belonging to a chain. Instinct is the only one to have returned independent venues exclusively.

We are not drawing a general law from a single test. But the question it raises for your own venue is simple: which surfaces does an agent find you on, and what does it read there?

The Arepado case, four surfaces and four answers

Arepado, a Venezuelan restaurant in Lyon’s 3rd arrondissement, was suggested by Instinct as available on 8 October at 12:30. Grok Bot ruled it out on the grounds that the official website only opens in the evening. We wanted to know which one was right.

Here is what we found on 23 September.

Surface What it states
The venue’s official website 7pm to 9:30pm every day, no lunch service
A second website for the same venue Opening hours that contradict the first
Google Not bookable at lunchtime
TheFork Lunch slots at 11:30, 12:00, 12:30, 13:00 and 13:30 on 8 October, bookable for three
Telephone No answer, several attempts
Email to the restaurant Table for three confirmed on the requested date

We repeated the journey ourselves on the TheFork listing: selecting 8 October, 12:30 and three covers, right up to the screen that asks for an email address. No booking was made. The official website was showing 7pm to 9:30pm for every day of the week at that very moment.

The table was in fact available. Instinct was right.

And Grok Bot got it wrong despite following a rigorous verification process. It cross-checked two sources, spotted the contradiction, chose the business’s own website as the more authoritative source, and ruled out a valid restaurant. Cross-checking your sources is not enough when the source you treat as the most reliable is itself wrong.

One detail is worth stating. Grok Bot had gone as far as the screen for sending an email to the restaurant before stopping, because our instruction required it to. That is precisely the move that produced the right answer when we made it ourselves. It had the capability, our protocol stopped it one click short.

What we verified and what we could not settle

We called PAMPA on 23 September at 12:22: open on Thursday lunchtime, table for three available on 8 October at 12:30. Grok Bot had said so, and it was right. ChatGPT’s three configurations declined to confirm it, even though the information existed.

We did not settle the walking distances. Instinct reports ten minutes on foot to Le Pompeï based on Google Maps. Grok Bot measures eleven to twelve minutes with a different routing engine and rules it out for exceeding the limit. On PAMPA, the gap between two configurations runs from fifty metres to nearly three hundred. A constraint expressed in walking minutes has no single answer: it depends on the engine the AI picked and on the starting point chosen inside a station several hundred metres across.

Booking platforms, a key waypoint for agents

This is the quietest role in the test and probably the most structural one.

The two AIs that acted, Instinct and Grok Bot, both went through TheFork for the one constraint they could not read directly on a page. Neither of them contacted the establishment. The booking platform therefore served as the interface between the agent and the business, for the single piece of information that decided the outcome.

The Arepado case adds an unexpected twist. It was the platform that was right, against the restaurant’s own official website. The intermediary was better maintained than the primary source.

ChatGPT, for its part, did not confine itself to one channel: it also drew on Google listings and on hotel group directories. Platforms are therefore not the only path by which an AI discovers a venue. They were, in this test, the only one through which action was possible.

For a business owner the consequence is concrete. Part of what an AI decides about you depends on a listing you may not have looked at in months, maintained by a third party, which may carry more weight than your own website at the moment of decision.

What this test says about delegation

The reasoning runs in five steps.

  1. AI systems are becoming capable of acting. Five configurations out of five carried the mission alone, none asked for help, all stopped short of the act when the instruction required it.
  2. To act, they have to query information outside themselves. A model does not know whether a table is free on 8 October, it has to go and find out.
  3. That information is often contradictory. On the single venue we verified, four surfaces state four different things.
  4. The AI therefore has to arbitrate between sources that contradict each other, with no way of knowing which one prevails.
  5. Incorrect information then produces an incorrect decision, even at the end of a perfectly coherent chain of reasoning.

In our test, the limit was therefore not only the ability of these AIs to reason or to act. It also lay in the information they had access to.

What this changes for a business

An independent restaurant has let four versions of its opening hours circulate, two of them on websites it owns. The result: one AI recommends it, another eliminates it, a third never sees it. None of those decisions has anything to do with the quality of its cooking.

This is not a search engine optimisation problem in the classic sense. Nobody was looking for that restaurant by name. An agent applied criteria, queried the available surfaces and made a decision in a customer’s place.

One practical note along the way: the links ChatGPT follows carry a parameter identifying where they came from. Your visit statistics may already be telling you whether agents are passing through.

For a business, the question becomes less “does AI know me?” than “which version of my company does it find when it has to make a decision?”.

As these systems become capable of making decisions on our behalf, a company’s AI presence will no longer be only a question of visibility. It will become a question of reliability: what information does the agent find, which source does it believe and what decision does it make from there?

The limits of this test

One mission, one city, one date. The sample supports no generalisation beyond this case.

We had planned two runs per product. Neither agent allows an independent repeat on the same account: Instinct has a single conversation thread by design, and Grok Bot carries its context from one conversation to the next, as it says itself. The second runs are re-verifications, not replications. The stability of these products is therefore not measured here.

The environments differ. ChatGPT was tested on two plans, neither of which offers an autonomous agent mode comparable to the two other products. A business account might well give a different result.

Finally, two sessions out of five were driven by an agent applying a fixed prompting script, which makes the prompts identical but does not exactly reproduce how a human user behaves.

FAQ

Which AI systems were tested?

Instinct on WhatsApp, Grok Bot as a macOS app, and ChatGPT in three different configurations. Muse, Meta’s agent, was not available in France on the date of the test.

Did any AI book a table?

No. The instruction forbade any booking and any confirmation. Two products offered to book and stopped, waiting for an answer that never came.

Why do the results differ so much from one AI to another?

Because they do not query the same surfaces. Some rely on a booking platform, others on Google listings or on hotel group directories. A venue missing from one of those surfaces does not exist for the AI querying it.

Does this mean one AI is better than the others?

No. This test establishes no ranking. It shows different behaviours in the face of unverifiable information, and one case where the most rigorous method produced the wrong answer.

Can an AI book a restaurant on your behalf?

Technically, two of the products tested can: they reach the final screen before confirmation, the one asking for an email address, and one of them spontaneously offered to complete the booking. Our protocol forbade them from going further, so none of them booked. The limit we observed is therefore not the ability to book, it is the reliability of the information the booking would rest on.

What is the difference between Grok and Grok Bot?

Grok is xAI’s conversational assistant. Grok Bot is a different product from the same company, launched in beta on 11 August 2026: an agent service where each agent has its own online machine, web access and plugins to act inside applications. Grok Bot is what we tested, not Grok.

Do you have to be listed on a booking platform to be suggested by an AI?

In this test, the two AIs that carried verification through to the end went via a booking platform, for lack of any other way to obtain a dated availability. They also discovered venues through other channels, including Google listings. A single test cannot turn that into a rule, but being absent from one of these surfaces mechanically reduces your chances of being retained.

How can a business avoid this kind of error?

By making sure its opening hours, its address and its availability say the same thing across every surface an agent can read, starting with the ones it controls itself.

Sources

Early access

Leave your email and we'll notify you as soon as Weftd opens up.