Muse: Meta's personal agent for long-running tasks
muse
Meta's personal AI agent, released in the US on 2026-09-08 across iOS, Android and web: runs on a cloud VM, takes a goal in chat, and lets you watch it drive a browser.
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- Launch platforms (Meta)
- Vendor Claim · 2026-09
- MATURITY
- Product
- research → demo → product → production
Editor's takeMuse belongs at the top of the documents & office ladder not because "Meta shipped another model", but because it moves the two real agent problems into the product layer: permissions and observability. Most office tasks have side effects — sending mail, placing orders, changing files. An agent that cannot bound its read/write scope and cannot let you watch it work will not be trusted with real work no matter how good the demo is. Sentinel's read/write split and the watchable browser are aimed squarely at that, which is why it is closer to usable than "another chat box".
Confidence: C (vendor claim). Checkable: the 2026-09-08 release, US-first on iOS/Android/web, execution on a cloud VM running Meta's own models, the 23 September additions (macOS control, an agent email address) and the commerce partner list (Meta's own announcements plus public entries). Not checkable: task success rate and the reliability of cross-application actions — there is no third-party task-level evaluation, so the card's reading states the launch platforms and no completion rate at all.
Where it draws the line between an agent and a chatbot
Meta defines Muse as a personal agent that "carries out long-running tasks on a user's behalf, rather than answering queries in a single exchange." That is not marketing phrasing — it decides three engineering properties: tasks must survive across time, they must be authorised, and they must be observable. The design maps onto exactly those three.
Three design decisions that matter
One: it runs on a cloud VM, not your device. Muse executes on a virtual machine inside Meta's cloud, driven through a chat interface. The upside is that a task does not occupy your laptop and you do not have to stay online. The cost is that data leaves your machine — the opposite trade-off from tools that run locally.
Two: a browser you can watch. When the agent drives the web, the browser is visible. For office work this is arguably the whole ballgame: being able to see what it is doing is the precondition for handing over tasks with side effects — filling forms, placing orders, looking things up. Opaque execution demos well and gets trusted with nothing.
Three: Sentinel separates read from write. A permissions system governs which services the agent may reach, and it distinguishes read access from write access. This is the part that looks like putting a lock on the thing doing the work — letting it read your mail and letting it send mail on your behalf differ by an order of magnitude in risk.
You can also set the agent's name, avatar and communication style. That reads as decoration, but it is part of giving something that acts on your behalf a stable identity.
Release and what followed
Muse was released on 8 September 2026, initially in the United States, on iOS, Android and the web. It runs on the latest generation of Meta's models — the first of which, Muse Spark, arrived in April 2026 — with the line led by Meta's first chief AI officer, Alexandr Wang.
On 23 September Meta announced more: a Realtime Avatar model (video conversations with a visual representation of the agent), the ability to control applications on macOS, an email address of the agent's own so users can forward messages to it or copy it into threads, and wake-word access through Meta's smart glasses. Commerce partners announced the same day included Stripe, Shopify, Best Buy, Gap, Sephora, Walmart, Wayfair, Expedia, Instacart, GitHub and Notion, alongside a palm-sized device called Muse Charm.
Boundaries
US-first, mobile and web at launch (desktop control arrived only in September). The vendor-claimed parts — how well long tasks actually complete, how reliable cross-application actions are — have no third-party task-success rate behind them. What can be checked is the release itself, the platforms, the permission model and the partner list.