Guides

Instinct AI Review

Independent tests show Instinct can complete useful real-world tasks, but they also document costly and destructive mistakes. There is no reliable overall success-rate benchmark yet.

By InstinctAI Wiki Editorial TeamLast verified 2026-10-026 min read

Instinct is interesting because independent testers have documented both genuinely useful task completion and serious execution mistakes. Any review that presents only one side misses the product's central tradeoff.

We do not give Instinct a numeric score. There is no representative public task-success rate, error rate, or customer-satisfaction benchmark that would justify a precise rating.

Instead, this review asks three questions:

  1. What has Instinct completed successfully in hands-on tests?
  2. What has gone wrong?
  3. What kind of supervision do those results imply?

What Instinct did well in independent tests

WIRED's September 24 hands-on test reported several meaningful successes.

The tester used Instinct for restaurant reservations, received a warning about a phishing email, and had the agent handle an airline schedule-change situation that resulted in a full refund worth about $550 and a rebooked return trip.

Those are not trivial chatbot demonstrations. They are real workflows involving external services, changing travel conditions, and financial consequences.

Business Insider's multi-staffer test also reported useful results, including:

  • comparing products across price tiers;
  • creating an itinerary;
  • booking a commuter ticket;
  • obtaining an $80 refund for a forgotten subscription after a free trial;
  • building a detailed merchandise buying plan.

Taken together, these tests support the idea that Instinct can be more useful than a chat-only assistant when a task involves research plus execution.

What went wrong

The failures matter just as much because an action-taking assistant can create real costs.

In WIRED's test, the user instructed Instinct to cancel a delayed DoorDash order only if a refund was available. The agent canceled the order anyway and the tester lost $64.

Business Insider documented other mistakes across its tests. These included:

  • paying the wrong Oxford driving-zone charge;
  • clearing personal scheduling data during a Notion overhaul;
  • returning inaccurate movie-showing information.

These are different kinds of failure: financial, destructive, and informational.

That range is important. A reliability problem in an agent is not limited to hallucinated text; it can affect real accounts and real transactions.

Instinct's own Terms warn about this

The observed failures are consistent with the limitations in Instinct's Terms of Service.

The Terms warn that outputs and actions may be incorrect or incomplete and tell users to verify results before relying on them. They also acknowledge that actions can have consequences and may not always be reversible.

That does not make every mistake acceptable. It does mean users should not interpret an autonomous action as guaranteed correct merely because Instinct was able to execute it.

Is Instinct easy to use?

WIRED's tester described the messaging-oriented form factor as easier to direct than a more open-ended chatbot experience because Instinct can suggest and execute concrete tasks.

That is a subjective usability observation, not a universal fact.

Business Insider's testing adds a useful counterpoint: testers emphasized the importance of clear, precise instructions and checking the result. In other words, a convenient interface does not eliminate the need for supervision.

Where Instinct looks strongest

Based on the evidence available now, the most convincing use cases are tasks where:

  • the desired outcome is concrete;
  • the user can inspect the final result;
  • a mistake can be caught before it becomes irreversible;
  • there is a clear external confirmation, such as a reservation or refund notice.

This is an editorial synthesis of the documented behavior, not a measured product benchmark.

The Task Database is more useful than a broad review if you want to evaluate a specific workflow. It shows task-level evidence, risk, required permissions, limitations, and last-verification dates.

Where extra caution is warranted

High-impact actions deserve more scrutiny.

Examples include:

  • purchases;
  • cancellations;
  • refunds;
  • changes to personal data;
  • messages sent on your behalf;
  • actions involving sensitive accounts.

The reason is simple: an agent can be useful precisely because it has permission to do things. The same permission increases the cost of a misunderstanding.

Our planned safety cluster goes deeper on permissions and privacy. For now, the Instinct Terms are an important first-party source for understanding the product's own warnings.

Is Instinct reliable?

There is not enough public evidence to assign a credible overall reliability percentage.

The current evidence is mostly a collection of product claims, media reporting, and individual hands-on tests.

Those tests are valuable because they show what actually happened, but they are not a statistically representative benchmark across thousands of tasks.

A statement like "Instinct succeeds 90% of the time" would be invented unless supported by a transparent dataset or company metric.

Is Instinct worth using?

That depends on the task and on how comfortable you are supervising an agent with access to real services.

The evidence supports two conclusions at the same time:

  • Instinct can save meaningful effort and complete tasks that a chat-only assistant would leave to the user.
  • Instinct can also misunderstand conditions, return bad information, or take an unwanted action.

For low-impact, easily checked tasks, that tradeoff may feel very different than it does for payments, cancellations, or destructive account changes.

The useful decision is therefore task-specific, not a single site-wide verdict.

How to use this review

Treat this page as the broad evidence summary, then move to the page closest to your intent:

As more independent tests appear, the review should be updated by adding dated evidence rather than by silently changing a score.

Frequently asked questions

Is Instinct AI good?

Independent tests show both useful successes and material failures. There is not enough evidence for a universal "good" or "bad" verdict across all tasks.

Has Instinct successfully completed real tasks?

Yes. WIRED and Business Insider reported successful reservations, refunds, product research, itinerary work, ticket booking, and other tasks.

Has Instinct made costly mistakes?

Yes. WIRED reported a $64 loss after an incorrect DoorDash cancellation, and Business Insider documented other erroneous or destructive actions.

What is Instinct's success rate?

No statistically representative public success-rate benchmark is established by our current evidence set.

Should I verify Instinct's work?

Yes. Instinct's own Terms say outputs and actions can be incorrect or incomplete and tell users to verify results before relying on them.

Related guides

Use cases

Instinct AI Tasks

Which Instinct tasks are documented by official sources or independent reports, and what permissions and risks do they involve?

Read guide

Safety

Instinct AI Safety

What safety risks and safeguards are documented for actions, purchases, connected services, and autonomous behavior?

Read guide

Comparisons

Instinct AI Alternatives

Which currently available assistants overlap with Instinct on messaging, persistent context, computer use, or real-world actions?

Read guide

Safety

Instinct AI Privacy

What information can Instinct collect or access, how is model training described, and what deletion controls are documented?

Read guide