One AI supervising another vendor's agent team

Claude → xAI Grok Bots 18 MCP tools Live since 1 October 2026

xAI's Grok Bots are always-on agents that share a cloud computer: a browser, a terminal and a file system they work on around the clock. They are steered only by messaging them in the Grok app. I built a bridge that lets Claude, from any chat on the web or a phone, watch what each Bot is doing, hand them tasks, unblock them when they get stuck on a click, generate images and video with Grok Imagine, and run commands on the computer they share.

The problem: a team of agents you cannot see into

I run eleven Grok Bots with distinct jobs, from operations and career research to a studio Bot and one that looks after my servers. Each has its own conversation in the Grok app, and together they had logged around eleven hundred actions on their shared computer. The only way to know what any of them was doing was to open the app and scroll through one conversation at a time.

Three things were missing. There was no overview: which Bot is active, what it did last and when. There was no way for another system to give a Bot work, because xAI documents no API for messaging a Grok Bot, only the desktop and mobile apps. And when a Bot stalled on something a person would fix in two seconds, such as a cookie banner or a confirmation dialog in its browser, it simply waited.

Meanwhile most of my day already runs through Claude. The question became whether Claude could act as a supervisor for a team of agents built by a different company, on infrastructure neither vendor controls.

What is new about it

Cross-vendor supervision. Agent platforms usually orchestrate their own agents. Here an Anthropic model observes, tasks and assists agents built by xAI, through the one surface both can touch: the computer the agents work on. Neither product was changed to make this possible.

Tasks without an API. Because a Grok Bot cannot be messaged programmatically, tasks travel through an inbox and outbox convention on the shared file system. One standing instruction, pasted once to each Bot, turns any always-on agent into something another AI can delegate to. The same pattern works for any agent that can read files.

Observability built from what the platform already writes. The Bots keep an audit log of every shell command, browser page and tool call, plus full transcripts. The bridge reads these in place, from the end of the file backwards, so Claude can follow an agent's last steps in under a second without copying gigabytes of history.

A human-in-the-loop for clicks. When a Bot is stuck in its browser, Claude takes a screenshot of the agents' desktop, clicks or types, and checks the result. Captchas, passwords and two-factor codes are explicitly left to me.

How it works

The connector is a remote MCP server written in Python with FastMCP. It runs on my own server behind Nginx and Google sign-in, and it reaches the agents' computer only over SSH, through a reverse tunnel that the agents' computer opens itself. Nothing on that computer listens on the internet. Every tool turns into one shell command sent over that tunnel, which keeps the server small and the behaviour easy to test locally against the same commands.

Claude connects as a custom connector in claude.ai, so the whole system is usable from a phone. Every call is written to an audit log with the account, the tool, its arguments and how long it took.

Claude web or phone MCP server 18 tools, Google sign-in Audit log every call Grok Bots' shared cloud computer Shell and files commands, read, write Grok CLI + Imagine answers, images, video Agent logs activity, transcripts Task mail inbox, outbox, done Desktop and browser where the Bots work; screenshot, click, type Bots run here, steered from the Grok app HTTPS SSH tunnel

Use cases

A morning overview of the whole team. "Which Bots worked overnight and what did they do?" Claude lists every Bot with its last action and time, then reads the latest steps of the busy ones and summarises them in a few lines.

Delegating from a phone. During a meeting I ask Claude to give the operations Bot a task. Claude drops it in that Bot's inbox, and later reads the result from its outbox and tells me whether it is done, still pending or answered with a question.

Unsticking a Bot. A Bot's browser is waiting on a consent dialog. Claude takes a screenshot, clicks the right button, takes another screenshot to confirm, and tells the Bot to carry on.

Creative work on the same login. Grok Imagine runs on the Bots' computer with the same grok.com account, so Claude can generate or edit an image and show it in the chat, or start a video clip and report where the file landed.

Operations on the agents' machine. Disk space, a stuck process or a file a Bot produced: Claude runs the command or reads the file directly, without opening a terminal.

Health checks after a restart. If the computer restarts, the status tool reports whether the tunnel, the SSH server and the Grok CLI are back, and gives the exact command to fix whichever is not.

The eighteen tools

  • run_command, read_file, write_file, list_dir. The computer itself, with time limits and output caps.
  • ask_grok, grok_image, grok_video. The Grok CLI and Grok Imagine, signed in with the owner's account.
  • list_agents, agent_activity, agent_transcript. Who the Bots are, what they did and what they said.
  • send_task, agent_replies, agent_mail_setup. Delegation through the inbox and outbox.
  • desktop_screenshot, desktop_click, desktop_type, desktop_key. Hands on the Bots' desktop.
  • status. What is up, what is down and how to fix it.

Every tool tells claude.ai whether it only reads or can change something, so viewing tools run without a confirmation and anything that acts asks first.

Design decisions

The agents' computer opens the tunnel, not the other way round. No port is exposed on the machine the Bots use. The key it uses to reach my server is restricted so that it can hold that one tunnel and do nothing else.

Every command proves where it is running. During setup the tunnel was once started on the wrong machine, so the connector briefly talked to my own server while believing it was the agents' computer. Now every command begins with a check of the remote machine's identity, and refuses to run if it cannot read it or if it is the wrong one. A random token marks the end of each command, so the output of a command can never be mistaken for the connector's own signals.

Survive a platform that resets itself. The agents' computer wipes its system folders on restart and keeps the home folder. The SSH host key therefore lives in the home folder, and a keeper script there restores the SSH server and the tunnel automatically. Without this, every restart would have looked like an impostor.

Never probe in a way the platform punishes. Modern OpenSSH penalises repeated connections that never sign in, and every tunnelled connection arrives from the same local address. A naive "is the port open" check would have locked the connector out of the agents' computer, so the check reads the local socket table instead of connecting.

Read logs from the end. One Bot's transcript alone holds about seventeen hundred entries. The helper that runs on the agents' computer reads files backwards in blocks, so the last twenty entries cost the same whatever the history.

How it was built

The build was itself a multi-agent exercise across three vendors. Claude wrote the design and a task-by-task plan; the Grok CLI implemented each task with tests first; the Codex CLI reviewed every task and the whole branch, including a security scan. Codex's reviews caught real defects before anything went live, among them a directory listing that a crafted folder name could have turned into a delete, and the lockout risk described above. The agent and desktop tools were then built directly, and everything ships with 126 automated tests, several of which run against a private SSH server. The full story is on the blog.

Common questions

Can Claude control xAI Grok Bots?

Not directly, because xAI documents no API for messaging a Grok Bot. This bridge works through the cloud computer the Bots share: Claude reads their activity logs and transcripts there, leaves tasks in a per-agent inbox that each Bot has been told to check, and can see and click the desktop where the Bots run their browser.

How does a task reach a Grok Bot without an API?

Through files. Each Bot gets an inbox, an outbox and a done folder on its computer. The owner pastes one standing instruction to each Bot in the Grok app, telling it to check its inbox every few minutes, do the task, write the result to the outbox and move the task to done. Claude writes tasks atomically and reads results later. Delivery is not instant, and it depends on the Bot following the instruction.

Is it safe to give an AI a shell on another machine?

It is a deliberate full-control design, so the safety comes from the edges: only one Google account can sign in, the machine is reached only through a tunnel it opens itself, the connector trusts one pinned host key, every command first proves it is running on the right machine, and every call is written to an audit log. Claude is instructed never to solve captchas or type passwords and two-factor codes.

What happens when the agents' computer restarts?

The computer resets its system folders on restart and keeps only the home folder, so a small keeper script stored there starts the SSH server with a persistent host key and reopens the tunnel, retrying every ten seconds. The connector's status tool tells Claude exactly what is down and which command fixes it.

Limitations and what comes next

Task delivery is only as prompt as each Bot's inbox check, and it relies on the Bot honouring its standing instruction, so it suits delegation rather than conversation. Approvals that a Bot requests inside the Grok app happen in xAI's cloud and cannot be clicked from the agents' computer. Transcripts reach the computer by synchronisation, so the very latest message can lag.

The next steps follow from that: a live progress view for long Grok jobs using the CLI's streaming output, notifications when a Bot replies, and a health view across all the AI tools I run, so one question to Claude answers what every agent, local or cloud, is doing right now.