OPEN SOURCE DEEP DIVE
mobile-mcp: the mobile automation MCP server driven by the native accessibility tree
mobile-next/mobile-mcp (8.2k stars, Apache-2.0) connects iOS simulators and real devices plus Android emulators and real devices to any MCP client, including Claude Code, Codex, Gemini and Copilot. Its first principle is accessibility-first: read the native accessibility tree instead of screenshots, no vision model and no image tokens, falling back to screenshots with coordinates only when necessary. Roughly 30 tools in seven groups cover device management, remote devices through Mobile Next Cloud, app management, screen interaction with recording, input and navigation, logs and crashes, and batch commands that compress N rounds into one. The README examples are all cross-app long-horizon tasks, the hardest class for mobile agents.
What it is
mobile-mcp (mobile-next/mobile-mcp, 8,216 stars, 712 forks, TypeScript, Apache-2.0, homepage mobilenext.ai) is an MCP server that lets AI agents and LLMs drive native iOS and Android applications and devices through a single platform-agnostic interface - simulators, emulators and physical devices alike. The README subtitle is blunt about the goal: no need for distinct iOS or Android knowledge. It names the clients it works with: Claude Code, Codex, Gemini, GitHub Copilot, Antigravity, or any MCP-compatible client.
Which section of this site does it belong to? AI Coding (code) as the primary domain, with agents as secondary. The reason is not that it touches source code, it is that its consumers are coding harnesses. It trains no model and runs no agent loop; it turns a phone into one more tool group on the tool surface of a coding agent. Add one mcp add line in Claude Code or Codex and that coding agent can immediately boot a device, tap the UI and read crash reports. At the same time it executes long-horizon multi-step cross-app tasks, which is evidence for the agents domain (multi-step tool use, OS tasks, long-horizon autonomy), so both domains apply.
Fact sheet
| Field | Value |
|---|---|
| Repository | mobile-next/mobile-mcp |
| Stars / forks | 8,216 / 712 |
| Language / licence | TypeScript / Apache-2.0 |
| Shape | MCP server (stdio) plus a Streamable HTTP server (--listen) |
| Targets | iOS Simulator, iOS real device (USB and trusted), Android emulator, Android real device (adb with USB debugging authorised) |
| Prerequisites | Xcode command line tools, Android Platform Tools, Node.js v20+ |
| Install | npx -y @mobilenext/mobile-mcp@latest; one-liners per client: claude mcp add, codex mcp add, gemini mcp add, amp mcp add, a Cursor deeplink, plus install badges for VS Code and Goose |
| Topics | agent, android, emulator, ios, mcp, mobile, physical, real, simulator |
The core design: accessibility first, not a vision model reading screenshots
This is the decision that defines the project, and the README puts it first in the feature list: it drives apps from the native accessibility tree - no vision model, no image tokens - and falls back to screenshots plus coordinates only when needed. The companion claim is structured, deterministic output: it reads real UI elements with their coordinates and properties, which removes the ambiguity of screenshot-only approaches.
Why that is decisive in cost terms:
| Route | Cost per step | Element identity | Failure mode |
|---|---|---|---|
| Screenshot-only GUI agent | An image per step, so image tokens accumulate linearly with step count | Inferred visually; the same button may resolve differently between steps | Coordinate drift, mis-taps, token cost runaway on long tasks |
| mobile-mcp (accessibility first) | Structured text, cheap; screenshots only when necessary | The real element and its properties and coordinates, handed over by the platform | Depends on how well the target app is annotated for accessibility; poor annotation is exactly when the visual fallback is needed |
Read it against entries already on this site. zai-org/Open-AutoGLM (26,321 stars) is the phone agent model route - train a model that can look at a screen. zai-org/CogAgent is the earlier VLM-based GUI agent in the same family. joshhhhhan/VISTA is more extreme still, playing ARC-AGI-3 from raw pixels. mobile-mcp goes the opposite way: it trains no model at all and instead makes the operating system hand over the interface structure. The three routes are not mutually exclusive, and which one wins in practice depends on how completely a target app is annotated for accessibility. On the ruler of cost per step, however, the structured route is overwhelmingly cheaper, and that is precisely what lets it live inside a daily coding harness rather than only in a lab demo.
Tool surface: about 30 MCP tools in seven groups
| Group | Tools |
|---|---|
| Device management | mobile_list_available_devices (simulators, emulators, real devices), mobile_get_screen_size, mobile_get_orientation / mobile_set_orientation, mobile_set_location (override or clear GPS), mobile_clipboard (read or replace) |
| Remote devices (Mobile Next Cloud) | mobile_login_to_cloud_provider (browser-based device-code login), mobile_list_remote_devices (models available to reserve), mobile_allocate_remote_device (reserve a physical cloud device exclusively), mobile_release_remote_device |
| App management | mobile_list_apps, mobile_get_foreground_app, mobile_launch_app (by package name), mobile_terminate_app, mobile_install_app (.apk, .ipa, .app, .zip), mobile_uninstall_app |
| Screen interaction | mobile_take_screenshot, mobile_save_screenshot, mobile_list_elements_on_screen (elements with coordinates and properties), mobile_click_on_screen_at_coordinates, mobile_double_tap_on_screen, mobile_long_press_on_screen_at_coordinates, mobile_swipe_on_screen (four directions), mobile_start_screen_recording / mobile_stop_screen_recording (saved to a video file) |
| Input and navigation | mobile_type_keys (into the focused element, optional submit), mobile_press_button (HOME, BACK, VOLUME_UP, VOLUME_DOWN, ENTER and friends), mobile_open_url (open URLs in the device browser, which is also the deep link entry point) |
| Logs and crashes | mobile_get_device_logs (logcat on Android, unified log on iOS, optionally saved to a file), mobile_list_crashes, mobile_get_crash (full report by id) |
| Batching | mobile_batch_commands (run several tools in sequence in one call, for example click then type then click, optionally listing screen elements at the end) |
Two tools are easy to overlook and matter most. mobile_batch_commands collapses N agent round trips into one, so a fixed sequence like tap, type, tap stops paying a model inference per step. mobile_get_crash and mobile_get_device_logs are why this belongs to development rather than only to automation: the agent edits mobile code, installs the build, runs it, hits a crash, reads the crash report back and keeps fixing - the whole loop closes inside one tool surface.
What people actually prompt it to do: long-horizon cross-app journeys
The example prompts in the README are not "tap this button". They are multi-app, multi-step journeys with verification loops. A few, paraphrased from the originals:
- Find the video called "Beginner Recipe for Tonkotsu Ramen", like it, write the comment "this was delicious, will make it next Friday", then share it with the first contact in the WhatsApp list.
- Find and download a free Pomodoro app with more than 1k stars, register with my email, work out how to start the timer, and once the timer is running go back to the app store, rate it 5 stars and leave a comment.
- Open Substack, search for "Latest trends in AI automation 2025", open the first article, highlight the section titled "Emerging AI trends", save it to the reading list and comment with a summary of a random paragraph.
- Open ClassPass, search for yoga classes tomorrow morning within 2 miles, book the highest-rated class at 7 AM, confirm the reservation, then set a timer on the phone for that slot.
- Open Eventbrite, search for AI startup meetups this weekend in Austin TX, pick the most popular one, register and RSVP yes, then create a calendar reminder.
- Check tomorrow weather forecast for Berlin and send the summary to a named contact over WhatsApp, Telegram or Slack, then thumbs-up their reply.
- Schedule a Zoom meeting titled "AI Hackathon" for tomorrow at 10AM with a one hour duration, copy the invitation link and send it by Gmail to the team address.
These all share one shape: several apps, confirmation that each step landed before the next one, and search or ranking judgement in the middle. That is exactly the shape a screenshot-per-step agent pays the most for, and exactly the shape structured element output makes affordable. The use case distribution is not an accident; it is what the accessibility-first decision selects for.
Operational and security knobs
| Variable | Effect | Why it matters |
|---|---|---|
MOBILEMCP_AUTH | Requires a Bearer token on the Streamable HTTP server (--listen); every request must then send Authorization: Bearer <token> | Exposing this over HTTP means putting "operate a real device" on the network. It must not run unauthenticated |
MOBILEMCP_DISABLE_TELEMETRY | Turns off anonymous usage telemetry (PostHog and Scarf) | A hard requirement inside corporate networks and regulated environments |
MOBILEMCP_ALLOW_UNSAFE_URLS | Lets mobile_open_url open non-standard URL schemes; blocked by default | On an agent-driven device, deep links are a real attack surface for launching arbitrary apps and private schemes. Default-deny is the correct choice |
MOBILEMCP_LEGACY_ROBOT | Uses the legacy platform-specific robots for Android devices and physical iOS devices; iOS simulators keep using mobilecli | A compatibility escape hatch, and evidence that the underlying layer is now unified on mobilecli |
The privacy statement is clean: mobile-mcp runs locally and communicates only with the devices you connect. With no physical device attached you can run headless against a simulator or emulator - on Android start one with avdmanager or the emulator command, on iOS run xcrun simctl list then xcrun simctl boot "iPhone 16".
Where it sits in the Mobile Next toolkit
mobilecli: the universal device CLI that mobile-mcp is built on. It controls devices, simulators and emulators from the command line or a JSON-RPC API, which means MCP is only one front end - the underlying capability does not depend on MCP existing.mobilewright: described as "Playwright for mobile". The stated path is deliberate - explore with the agent first, then graduate to mobilewright when you need repeatable, deterministic iOS and Android tests. Exploration on mobile-mcp, consolidation on mobilewright.- Mobile Next Cloud: the same stack rented, real devices on demand. The onboarding is the interesting part: you tell the agent "log in to mobile next cloud and then show me which remote devices are available to me" and the agent drives those four remote tools itself. The tool surface becomes the product entry point.
Our reading
- It is a tool surface, not a model, so it belongs in the code and agents domains rather than in llm. Apply the criterion this site uses for llm - "replace it and the intelligence of the model or the cost per token changes". Replace mobile-mcp and the model is unchanged; what changes is which devices the model can reach.
- Its leverage comes from making the platform hand over structure, which is the most underrated route in mobile agents today. Progress on the visual route will keep eroding that advantage, but as long as accessibility trees exist the structured route stays ahead on cost and determinism - and cost is what decides whether something reaches daily CI rather than staying a demo.
- It is the same class of asset as
GLips/Figma-Context-MCP, already in the code domain, pointed at a different reality. One hands design context to a coding agent, the other hands a mobile device to a coding agent. Neither is an agent; both add a new interface to the real world for an agent. This layer of MCP tool surfaces is the fastest-growing part of AI coding and deserves to be tracked as its own subsection.