There is no evidence-based way to answer "Is Instinct safe?" with a simple yes or no.
Instinct's own Terms warn that outputs and actions can be wrong, incomplete, or irreversible. Its Privacy Policy explicitly identifies autonomous-agent risks such as unintended payments, unintended communications, incorrect outputs, and misleading third-party instructions designed to influence the agent.
At the same time, reporting describes safeguards under development, and independent tests include positive examples such as phishing detection.
The accurate picture is risk management, not risk elimination.
What does Instinct officially warn about?
Instinct's Terms contain the clearest first-party safety boundary.
They warn that:
- outputs and actions may be incorrect or incomplete;
- actions may not always be reversible;
- users should verify results before relying on them;
- high-impact use requires meaningful human review.
The Terms also say Instinct may implement safeguards or confirmation requirements but does not warrant that those safeguards will prevent every erroneous action.
That is an important distinction. A confirmation system can reduce risk without turning autonomous execution into a guarantee.
What autonomous-agent risks does the Privacy Policy identify?
Instinct's Privacy Policy explicitly discusses risks that are specific to an agent that can act.
Examples include:
- unintended payments;
- unintended communications;
- incorrect or incomplete outputs;
- misleading third-party instructions intended to influence the agent.
That last category is closely related to the broader problem often called prompt injection: an agent encounters malicious or misleading content while browsing or interacting with an external service and is pushed toward an action the user did not intend.
The policy also states that no security measures are impenetrable.
When does human review matter most?
Human review matters more as the consequence of an action rises.
Examples include:
- sending an external message;
- making a purchase;
- canceling an order or subscription;
- changing a booking;
- deleting or overwriting account data;
- changing an external account;
- sharing sensitive information.
A low-risk research task can often be checked after the fact. A destructive or costly action may be difficult to reverse.
That is why the Task Database assigns risk levels instead of treating all tasks as equivalent.
What safeguards have been reported?
Reuters reported that Instinct was implementing enhanced privacy and security measures including:
- isolated sandboxes;
- short-lived credentials;
- a system intended to detect subtle inaccuracies.
Those are reported measures being implemented. They should not be rewritten as proof that every account, task, or connected service is fully protected from all failure modes.
The useful question is not whether safeguards exist. It is whether a particular high-impact action is still independently verifiable before or after execution.
Has Instinct detected a real security threat?
WIRED reported a positive example: Instinct warned the tester not to open an email that turned out to be a phishing attempt.
That is evidence of one successful security-related judgment.
It is not evidence of universal phishing protection.
A user should still apply normal security practices and should not assume every malicious message, website, attachment, or instruction will be detected correctly.
Has Instinct made harmful mistakes?
Yes.
WIRED documented a DoorDash case where the user said the order should be canceled only if a refund was available. Instinct canceled it anyway, and the tester reported a $64 loss.
Business Insider reported other harmful mistakes, including:
- paying the wrong Oxford driving-zone charge;
- clearing personal scheduling data during a Notion overhaul.
These examples matter because they show different failure modes:
- misunderstanding a condition;
- performing the wrong financial action;
- making a destructive data change.
They reinforce the value of scoped permissions and explicit checkpoints.
What is a safer way to delegate a high-risk task?
Break the task into stages.
Instead of:
Handle this subscription.
Use a request like:
Check the cancellation and refund rules. Show me what will happen, including any lost benefits or fees. Do not make the final change until I approve.
Instead of:
Book the best flight.
Use:
Research options for these exact dates and airports. Show total price, baggage rules, restrictions, and the final itinerary. Do not purchase until I approve.
The principle is to separate research, decision, authorization, and verification.
Should you minimize permissions?
Yes, where the product and connected service allow it.
If a task only needs calendar read access, broader write permissions create unnecessary exposure. If a task only requires product research, payment access is unnecessary until you actually want to buy.
A good permission strategy is:
- connect only the service needed for the task;
- grant only the capability the task requires;
- verify high-impact actions;
- disconnect services you no longer need;
- remember that disconnecting is not necessarily the same as deleting previously indexed data.
For the last point, see Instinct AI Privacy.
Is Instinct risk-free if it asks for confirmation?
No.
Instinct's Terms explicitly avoid promising that safeguards or confirmation requirements eliminate erroneous actions.
Confirmations can still be useful. They give the user a chance to notice a wrong amount, wrong recipient, wrong date, or wrong action.
But a confirmation is only as useful as the information shown and the attention given to it.
What tasks deserve the most caution?
Based on consequence rather than product marketing, the highest-scrutiny categories include:
- payments and purchases;
- refunds and cancellations;
- travel bookings;
- messages or calls sent on your behalf;
- account changes;
- deletion or editing of personal data.
For documented examples, see Instinct AI Shopping, Instinct AI Refunds, and Instinct AI Review.
Frequently asked questions
Is Instinct AI safe?
Instinct documents safeguards and independent testing includes positive security behavior, but official policies and hands-on tests also document meaningful risk. No verified source supports calling it risk-free.
Can Instinct make irreversible mistakes?
Its Terms say actions may not always be reversible, and independent tests have documented costly or destructive mistakes.
Does Instinct protect against phishing?
WIRED reported one successful phishing-warning example. That does not establish comprehensive phishing detection.
Does Instinct use security sandboxes?
Reuters reported that Instinct was implementing isolated sandboxes and short-lived credentials. We keep that claim attributed because it comes from reporting, not a detailed public security specification.
Should I review actions before Instinct completes them?
For high-impact tasks, meaningful human review and explicit checkpoints are the safest interpretation of Instinct's own Terms and the observed failure evidence.