Is Meta Muse safe? Secure VM, Sentinel and approvals explained
Meta Muse runs in an isolated VM where a Sentinel approves every action and the agent never sees real credentials. What that covers, and what it misses.
Meta Muse is built on the assumption that its own agent will be attacked and will sometimes be fooled, and most of its safety story is about limiting what a fooled agent can reach rather than about preventing the fooling. That is the right design, and it is unusually well documented: Meta's engineering post describes a separate component, Sentinel, that is the only thing allowed to approve a connector action or a network request, and a credential store the agent never reads from. The two decisions that are still yours are the ones to spend your attention on: the model-training toggle is on when you start, and a connector Muse writes for you is one nobody at Meta has reviewed.
#What is Meta Muse actually protecting you against?
Meta's security post, "How We Built Safety Into Muse", published 8 September 2026 by Tarek Sheasha, states the threat model in the first few paragraphs. It says that no matter how strong the model is, "any agent like this will still make mistakes, and it will sometimes be attacked via the data it reads", and that the system is therefore designed to "assume the agent may be under attack and limit the potential damage": "the harness runs in its own isolated cell, it doesn't see real credentials, and every interaction with the outside world runs through a Sentinel which the agent can't override".
That is a containment argument rather than a prevention argument, and it is worth reading it as one. Meta is not claiming Muse cannot be manipulated. It is claiming that manipulating the agent does not by itself get an attacker your credentials or your data, because the agent is not the thing that holds either.
The post also names the framing it designed against. It quotes Simon Willison's "lethal trifecta", the combination of access to private data, exposure to untrusted content, and a way to communicate externally, and treats the presence of all three in a personal agent as the condition to engineer around. If you have read our write-up of the Clawjacked vulnerability, this is the same shape of problem, approached from the architecture side rather than the patch side.
#What is the Muse Secure VM?
Every Muse user gets a dedicated cloud machine. The Muse Privacy Policy, effective 17 September 2026, says Muse "runs on an environment in the cloud called a virtual machine", that this is "the core environment where Muse keeps information you share and things you create together", and that "Every Muse user's VM is isolated so no one else's agent can access it".
Inside that VM the split that matters is not between you and other users but between the agent and the security services. Meta describes it plainly: "The right mental model is two isolated security domains on one box, not an LLM powered agent with root."
The agent, your files and every tool it runs sit in a systemd-nspawn runtime container. Meta says "Root inside the runtime cell is mapped to an unprivileged host user so runtime cell root is not host root", that the cell has its own root filesystem separate from the host, a virtual network interface, filtered system calls and limited kernel capabilities.
Four services sit deliberately outside that cell, which is the part that does the work:
hatch-safety, which "operates an independent set of models and classifiers which inspect requests and responses to/from core model inference". Meta is explicit about why it lives outside: "Running these outside of the runtime cell means that attackers cannot disable these protections."privsepworkers, which "execute built-in connector code with tightly scoped privileges, keeping the connected credentials out of scope of the agent". Each worker "is identified by its cgroup and has explicit credential allowlists", and Meta gives the consequence as an example: "A calendar worker cannot ask authd for an email credential simply by changing a request parameter."hatch-authd, the credential store. Meta says OAuth tokens for services you connect "are stored in your VM, not in centralized Meta infrastructure".- Sentinel, covered next.
Want step-by-step guides for this and more?
ClawDocx Pro includes 500+ curated prompts, setup guides, SKILL.md files, and templates — everything to make your AI agent unstoppable.
See plans & pricing#What does Sentinel do, and why does it matter more than the container?
Sentinel is the single most load-bearing claim in Meta's security story. The post calls it "a separate host-side agent from your Muse" and "the sole permission authority for approval to perform actions with connectors to third-party services and for all egress over the network", and states the division of labour in one sentence: "Muse proposes actions, but only Sentinel can grant permission to perform action."
For connectors, the runtime cell submits a request describing the connector, the method, the class of action, its scope and the context of what you asked for, from which Sentinel "generates a user-visible purpose for the request" and then allows, denies or asks you.
For network traffic, Meta says "Every concrete network request is governed by Sentinel at egress", and that Sentinel can evaluate "the hostname, the resolved and final destination IP address, the port, protocol, HTTP method, path, and the actual decoded request". It also says SSRF restrictions stop "an apparently public hostname resolving private infrastructure after DNS lookup".
One mechanism here is worth knowing by name because it explains why Muse does not ask you about everything. Meta calls it "tainted egress": kernel-level data flow tracking in which "Each tool execution process starts in a clean state and becomes tainted if it reads user data", implemented with eBPF. A clean request to an allowlisted destination can pass without bothering you; a tainted one loses auto-allow and goes back to the approval flow. That is a real exfiltration control rather than a convenience feature, and it is the reason an approval prompt from Muse carries more information than a generic permission dialog.
#How does Muse keep credentials away from the agent?
This is the part of the design that most directly answers "what happens when the agent is fooled". Meta says Sentinel "performs just-in-time credential insertion for any requests which need a secret or auth token", and that code in the runtime cell "only ever sees a surrogate token, minted by authd". The real credential is substituted at the network boundary, after the request has been authorised. Meta draws the conclusion itself: "The agent never sees real tokens, which means any attempt to coerce the agent to reveal the actual secrets via prompt-injection or otherwise is futile."
Browser sign-ins work the same way. Meta says a custom client UI captures your username and password, routes them "directly to authd", and stores them outside the runtime cell; what you type "is not visible to your main agent, but gets injected into the browser window at the point of need". The browser sub-agent "sees an accessibility tree snapshot of the page" rather than the raw DOM, cannot run JavaScript in the page context, and is paused entirely while you take over or while a credential is being filled.
Two more specifics are worth having, because they target the exact escalation paths an attacker would use:
- Email. Meta says Muse's email connector "filters out one-time tokens, password reset links, and login magic links via both deterministic filters and a classifier model". Your inbox is the reset path for most of your other accounts, so this is the single most useful filter in the product.
- Payments. Meta says that on checkout pages it detects the checkout and prompts for approval "with the exact details of the purchase every time", and that for new merchants "a single-use card number is issued", tied to that merchant, that amount and a limited period. The launch partner is Stripe Link, with Shop Pay described as coming soon.
#What are the approval choices, and what does each one actually grant?
Meta documents two layers here, and conflating them is easy. The engineering post describes the grant types Sentinel can mint: "Muse has support for obtaining one-time, session-scoped, task-scoped, time-bounded, or perpetual permission", and insists these are real capability grants: "Approvals granted via the human in the loop system are strict capabilities, not conversational suggestions. They're bound to the particular connector/destination and use case."
What you see in the app is the Help Center's list. The five choices, in Meta's own wording:
| Choice | What Meta says it does | How wide it is |
|---|---|---|
| Allow once | "Muse proceeds this one time" | One action |
| Allow for this task | "Muse can take this type of action for the entire task" | One task, one action type |
| Allow for this site | "Muse can take this type of action for this website in the future without asking again" | One website, indefinitely |
| Always allow | "Muse can take this type of action for this Connector in the future without asking again" | One connector, indefinitely |
| Deny | "Muse won't proceed this one time" | One action |
The two indefinite rows are the ones to be deliberate about. "Allow for this site" and "Always allow" both remove the prompt for that action type in future, and nothing in Meta's wording time-bounds either.
How often you are asked depends on a setting. Meta's Help Center says that "If you select Ask for some actions, your Muse will ask for permission before every write action and important read actions", and the Privacy Policy sets the default conservatively but not absolutely: "By default, Muse will not take many important actions, like sending an email, without your approval." The word doing the work there is "many". Meta is also open that silence is intended: "The point is not to ask the user about everything. Read-only, previously allowed, or demonstrably low-risk actions can proceed without interruption."
Meta says you can audit what happened afterwards by visiting "your Activity log by tapping your assistant icon to see a chronological record of actions your Muse has taken and permissions you've given it".
#What is still your decision?
Three things, and none of them is covered by the architecture above.
The model-training toggle is on when you start. The Privacy Policy says "You decide whether we can use your interactions with Muse to train and improve our AI models", then: "This setting is on when you first use Muse. You can change it at any time using the toggle in your Settings under 'Data controls.'" Two details matter. Turning it off is retroactive: "Changes to this setting also apply to previous interactions." And it covers more than your typing, because "Your interactions with Muse can also include information from services Muse uses to perform tasks for you, including Connectors", governed by the same toggle. Meta says that if you leave it on it removes "certain categories of personally identifiable information like names, email addresses, phone numbers and Social Security Numbers" and disassociates interactions from your account.
Nobody at Meta reviews a custom connector. If a service is not in the connector list, Meta's Help Center says you "can ask Muse to create a Custom Connector", and then warns in its own words: "Meta doesn't review custom connectors or how they use your information, so grant access with caution and review the provider's privacy policies." Every other connector was built with the service provider; a custom one was built by an agent, for you, unreviewed. Treat authorising one the way you would treat installing an unaudited extension, and apply the same questions we set out in our skill security audit.
Deletion is not forgetting. The Privacy Policy says you can delete messages, side chats and files, and reset Muse entirely, but also that "After you delete something, Muse may still 'remember' information it learned from what you deleted". Meta's data-management page says the same and points at the Forget skill, and the Privacy Policy notes that disconnecting a connector stops the data flow while data it already used "may remain in Muse's memories and your conversation history". You can read what it holds: Meta says you can ask directly, or look at files like MEMORY.md.
One account-shape decision sits alongside these. The Privacy Policy says you can run Muse on a separate account or in the same Accounts Center as your other Meta accounts, and that if you keep them separate "Muse won't share your Muse interactions or data on your VM with other Meta Products". Meta's connectors page says Facebook, Instagram and Threads "are connected automatically if you have your accounts in the same Accounts Center", so the Accounts Center choice silently decides three of your connectors.
#What does Meta say it has not solved?
More than most vendors, which is a reason to take the rest of the document seriously.
On prompt injection, the post's conclusion says "Prompt injection remains an open problem in the industry" and that "Muse will sometimes make mistakes", and frames the whole design as bounding the impact rather than removing the cause. The defences it does claim are stacked: model training it describes as "close to SOTA" on its own injection evals, untrusted-input labelling in the harness, and "An ensemble of multiple prompt injection detection classifiers" running outside the cell.
On Meta's own access, the post is specific and limited: today's architecture "restricts access to your data by Meta personnel through operational policies", and "It does not prevent Meta from accessing data when necessary to support, secure or operate the service." The fix Meta describes is future tense. Muse Confidential VM "is intended to cryptographically and verifiably prevent Meta from accessing data in your VM", is planned "later this year", is in use "with a small group of trusted testers", and its design and source code have been made available to external auditors. Until that ships, the honest summary is that Muse separates your data from other users and from Meta's ad systems, and does not separate it from Meta.
On advertising, the Privacy Policy and the post agree: "Muse doesn't share your conversations or the data in your virtual machine with Meta ad systems", and the Privacy Policy adds that this holds "even if your Accounts Center includes other Meta Products". The post then names the leak that remains, which is behavioural rather than technical: "When Muse browses the internet, it will appear as your activity", so a site Muse visits for you may later show you an ad.
Meta also put money behind external scrutiny. The post says the bug bounty programme is open to anyone and "awards up to $300,000 for valid reports, including up to $130,000 for successful prompt injection attempts that affect one user". A vendor paying six figures for single-user injections is a vendor that expects them to exist.
#So should you connect your accounts?
A reasonable sequence, given what is documented:
- Decide the account shape first, before connecting anything. A separate account keeps Muse data out of your other Meta products and stops Facebook, Instagram and Threads connecting themselves.
- Turn the training toggle off if you want it off, in Settings under Data controls. It is on until you do, and switching it off reaches back over what you have already said.
- Start read-only. Meta says that where the underlying service supports it, Muse separates read and write access, and offers finer-grained control than the OAuth scopes the service exposes. Read access to a calendar is a much smaller bet than write access to an inbox.
- Spend your approvals carefully. Prefer "Allow once" and "Allow for this task" while you are learning what it does. "Always allow" and "Allow for this site" are indefinite.
- Avoid custom connectors for anything sensitive until you have read the provider's terms yourself, because Meta has not.
- Read the Activity log after the first few real tasks. It is the only place that shows you what was actually done rather than what was proposed.
For what Muse is and what it costs, see our Meta Muse platform page and the wider comparison in personal AI agents compared. If you run agents on your own machines rather than in a vendor VM, the equivalent controls are in our security hardening guide.
One last framing, because it is the thing most reviews of Muse get backwards. The question is not whether Muse can be attacked; Meta says it can and pays for proof. The question is what an attacker gets when it works. On Meta's own description the answer is: not your tokens, not your passwords, not an unapproved purchase, and not a request to a destination Sentinel refused. What it does get is whatever the agent was already allowed to do, which is exactly why the approval choices and the read-only default are where your attention belongs.
Frequently asked questions
- Is Meta Muse safe to connect to your email?
- Safer than a CLI agent holding the same token, because Muse's email connector runs outside the agent's container and the agent never sees the credential. Meta also says the connector filters out one-time tokens, password reset links and login magic links, which is the specific path an attacker would use to take over your other accounts. The residual risk is that the agent can still be talked into sending or reading the wrong thing, so keep write access off until you have watched it work.
- Does Meta train its models on Muse conversations?
- Yes, unless you turn it off. The Muse Privacy Policy says the setting is on when you first use Muse, and that you change it with a toggle in Settings under Data controls. Turning it off also applies to your previous interactions. Information Muse pulled in from connectors is covered by the same toggle.
- Can Meta Muse be hijacked by prompt injection?
- Meta says it can. Its security post calls prompt injection an open problem in the industry and says Muse will sometimes make mistakes. The design assumes the agent will be fooled and limits what a fooled agent can do: it holds no real credentials, and a separate component called Sentinel has to approve every connector action and every outbound network request.
- Does Meta review the custom connectors Muse builds?
- No. Meta's Help Center says Meta does not review custom connectors or how they use your information, and tells you to grant access with caution and read the provider's privacy policy. A custom connector is the one part of the Muse surface where no Meta engineer has looked at what you are about to authorise.
- Can Meta see the data in your Muse VM?
- Today, yes, in defined circumstances. Meta's security post says the architecture restricts access by Meta personnel through operational policies, and that it does not prevent Meta from accessing data when necessary to support, secure or operate the service. Meta says a Muse Confidential VM intended to prevent that cryptographically is planned for later in 2026 and is with trusted testers and external auditors.