Agent-driven mobile QA infrastructure: how Limrun works in 2026
Computer use came to the desktop first. Give a model a screen, a mouse, and a keyboard, and it clicks through software the way a person does. Mobile took longer, because the screen in question usually lives on a physical phone in someone's hand or a simulator on someone's laptop, and neither is reachable from wherever the model runs.
Agent-driven mobile QA infrastructure is what fixes that. It is a cloud-hosted iOS simulator or Android emulator, created through an API, that an AI agent can open your app on and drive: tapping, typing, scrolling, reading what came back, and deciding whether the flow worked. No local install, no device plugged into anything, no human moving a thumb.
This page is about how that works in practice, what it replaces, and where to start if you are testing your own app rather than selling testing as a service.
What an AI agent can do inside a mobile app
The mechanics are less mysterious than the phrase "computer use" suggests. An agent driving a mobile app runs a tight loop: read the screen, take an action, read the screen again.
The reading step is where most of the quality comes from. A screenshot alone gives a model pixels, and it has to guess what is tappable. The accessibility tree gives it structure: every element on screen with its label, accessibility identifier, type, and frame. An agent working from the tree knows there is a button labeled "Continue" and where it sits, instead of inferring that from a rendering.
Limrun returns both together, in one call, through ios-screenshot-and-element-tree. The guidance in the docs is blunt about it: call this before every action.
From there the agent has a full interaction surface.
Taps by selector, not coordinates. Coordinates break the moment a layout shifts or a device size changes. Selectors survive both.
lim ios tap-element --ax-label "Sign in"
lim ios tap-element --ax-unique-id startButton
lim ios tap-element --type Button --ax-label "Done"
Selectors accept --ax-unique-id, --ax-label, --ax-label-contains, --type, --title, --title-contains, and --ax-value, and they combine to narrow a match. Elements the tree can see get scrolled into view automatically.
Typing with real key events.
lim ios type "test@example.com" --enter
lim ios set-text "P@ssw0rd!" --focused
lim ios press-key enter --modifier shift
type sends real keystrokes into the focused field, which exercises the same input path a person would. It needs a focused field and errors when there is none, so tap the field first. set-text writes the exact value through accessibility and skips keyboard behavior, which is faster when you are setting up state rather than testing the keyboard.
Scrolling, swiping, and gestures.
lim ios scroll down --amount 300
lim ios swipe --from 200,300 --to 200,600
Deep links, app lifecycle, and hardware buttons. The agent can open myapp://orders/42 directly instead of navigating six screens to reach it, launch and terminate apps by bundle ID, and press hardware buttons (home, lock, side, applePay, softwareKeyboard) through batched buttonDown and buttonUp actions.
Batched actions. A known sequence goes in one round trip instead of one per action:
lim ios perform \
--action type=tap,x=100,y=200 \
--action "type=typeText,text=Hello World"
Verification beyond the screen. UI that renders correctly can still be failing underneath, so the agent pulls app logs:
lim ios app-log com.example.MyApp --tail 200
Put together, that is enough for the flows teams actually care about: login, signup, form fill and validation, navigation, search, deep link handling, and regression passes over screens that broke before. In-app purchase flows work too, through StoreKit's local test environment, with no Apple ID or sandbox tester involved. See In-app purchases.
Android works the same way through android-screenshot-and-element-tree, android-use, and the CLI equivalents, with the emulator reachable over an ADB tunnel so existing tooling attaches unchanged.
Computer use for mobile, and getting a human to look
An agent that finishes a run and reports "all flows passed" is asking for trust. Someone should be able to check without reproducing the whole setup.
Live preview links. Every Limrun instance exposes a Signed Stream URL with its token in the URL fragment. Open it in any browser and you are looking at the running device, and you can drive it yourself. No sign-in, no install, no Mac. This is the fastest way to check on an agent mid-run, and it works for people who do not have a development environment at all.
Demo videos. For asynchronous review, the agent records what it did:
lim ios record start
# agent drives the flow
lim ios record stop -o /tmp/login-regression.mp4
Quality accepts integers 5 through 10, and 5 is the default. Pass --presigned-url instead of -o to send the file straight to your own bucket, which is the usual choice when the recording is going to be attached to a ticket or a pull request comment.
Preview links on pull requests. Build, upload the artifact, and post a link:
lim xcode build . --scheme MyApp --upload my-app-pr-42.zip
The reviewer opens https://console.limrun.com/preview?asset=<name>&platform=ios and taps through the branch. Product and design get to review the actual build instead of a screenshot in a comment thread. Build products uploaded this way expire 14 days after the last upload by default, and each new upload pushes the expiry out again.
Embedded devices. If review should happen inside your own tool, <RemoteControl /> from @limrun/ui renders a live iOS or Android device in a web app. Your backend keeps the API key and the browser only ever sees a per-instance URL and token.
What this replaces
Most teams testing their own mobile app are running some version of the same setup, and it has predictable failure points.
The device drawer. A shelf of phones, half of them dead, one of them the only device that reproduces the bug. Someone owns charging them. Onboarding a new engineer means finding a spare. Remote simulators remove the hardware entirely for the build-and-verify loop, though real devices still matter for release checks. For those, the same remote build can upload straight to TestFlight or publish to a Google Play testing track, so testers install it on their own phones.
One simulator per laptop. A local simulator is fine until you need three at once, or until QA needs to see what the engineer sees, or until the engineer is on Windows. Cloud instances are created through an API call and shared with a link, so a flake that only one person can reproduce becomes a URL everyone can open. Anyone with the link can control the instance, so share it only with people who should have access, and let one person drive at a time.
Manual regression passes. The pre-release run through twenty screens that takes an afternoon and gets skipped when the release is urgent. An agent runs it on every branch, and it costs whatever a few minutes of instance time costs.
Test scripts that rot. Selector-based automation written against a moving UI needs constant maintenance, which is why so much of it gets abandoned. An agent reading the accessibility tree at runtime adapts to layout changes that would break a hardcoded coordinate or a brittle XPath. It is not maintenance-free, but the failure mode is softer.
Waiting on a Mac for CI. macOS runner availability and pricing depend on the provider, and a CI job gives reviewers nothing to tap. With the Mac on Limrun's side, a standard ubuntu-latest job builds and tests an iOS app with no macOS runner in the loop.
None of this argues against the frameworks you already use. Maestro runs against a remote simulator from a Linux runner with your flows and the Maestro CLI unmodified, aside from a few commands that do not apply remotely. Appium works through @limrun/appium-xcuitest-driver, a fork of the upstream driver. XCTest runs through lim xcode test, which streams one line per test case as it finishes and exits non-zero when anything fails. Agent-driven interaction covers the exploratory and hard-to-script work; your existing suites keep doing what they are good at.
How Limrun fits
Limrun runs the parts that need a Mac and exposes them through one CLI, SDKs, and MCP endpoints: real-time remote Xcode, cloud iOS simulators, cloud Android emulators, asset storage for builds, and browser previews with recording and WebUSB install to a physical device.
For an agent, wiring it up is two decisions. Give the agent the CLI for builds and lifecycle:
npm install --global lim
lim skills install
And give it MCP for interaction, either the organization endpoint or a per-instance server:
claude mcp add --transport http limrun https://mcp.limrun.com/mcp
Scope each run with labels so instances are reusable and cleanup is precise, and set an inactivity timeout so an interrupted agent does not leave a device running.
FAQ
Can an AI agent really tap through my mobile app?
Yes. The agent reads the accessibility tree to see what is on screen, sends taps, typing, and scrolls by selector, then reads the tree again to confirm what changed. It works on flows a person would run by hand: login, signup, search, checkout, deep links. Whether it does that well depends mostly on how well your app is labeled for accessibility, which is worth improving regardless.
What is computer use for mobile apps?
The same idea as desktop computer use, applied to a phone screen. A model is given a device it can see and control, and it operates the app the way a person would rather than calling your code directly. On mobile the device is usually a cloud simulator or emulator, since that is the most practical way to give a model a screen it can reach over the network.
How long does setup take?
Installing the CLI and authenticating is a few minutes. Getting an agent driving a real flow depends more on your app than on the infrastructure: a well-labeled app works almost immediately, and an app with unlabeled custom controls needs accessibility work first.
Does this replace our QA engineers?
It changes what they spend time on. Manual regression passes and device wrangling shrink. Deciding what to test, reviewing agent output, and chasing the failures that matter do not. It works best when QA writes the flows and reviews the recordings.
What should we automate first?
Whatever you re-run most and dread most. For most teams that is the login and onboarding path, since it gates everything else and breaks often. Get one flow reliable end to end, put its recording in the pull request, and expand from there.
Can non-engineers use any of this?
Yes, through preview links. A live device in a browser tab needs no install, no account, and no development environment. That is usually how product and design get pulled into review.
Related: Run an iOS simulator · MCP server · Maestro · Appium · Test with XCTest · PR previews