Blog · Multi-agent engineering
Claude, Grok and Codex on one build: one plans, one writes, one reviews
Splitting one build across three vendors worked better than asking which single tool is best. Claude wrote the design and a task-by-task plan, the Grok CLI implemented each task with tests first, and the Codex CLI reviewed every task and the whole branch. Codex caught real defects before anything shipped, including an option injection that could turn a directory listing into a delete. The price is supervision: an implementer running with auto-approve once edited files outside its task.
Why use three tools instead of picking one?
“Claude Code vs Codex” is one of the most searched AI tooling questions this year, and the honest answer is that comparing them misses the point. Models from different vendors make different mistakes. A reviewer from another vendor does not share the writer's blind spots, which is the whole value of review. So for a recent production build, an MCP connector that lets Claude supervise a team of Grok Bots, I gave each tool one job.
Who did what?
| Role | Tool | Output |
|---|---|---|
| Design, plan, controller | Claude | A spec, a plan of small tasks with exact tests, a ledger of every decision |
| Implementer | Grok CLI, headless | Code and tests per task, test-first, one commit each |
| Reviewer | Codex CLI, read-only sandbox | Spec compliance and quality verdicts per task, a final branch review and security scan |
Each task went through a loop: implement, review, fix, scoped re-review, with at most five fix rounds and a fresh implementer from round four. Disagreements with the plan went to me, not to the agents.
What did the reviewer actually catch?
- A delete hiding in a list. The directory listing tool passed a path straight to
find. A folder named-deletewould have been read as an action. Grok checked that GNU find on the server still parses expressions after--, so the fix prefixes such paths with./. - A safety check with a race. The guard against talking to the wrong machine ran once per connection, and a reconnect could skip it. It now runs on every command, and fails closed if it cannot read the machine's identity.
- Output that could lie. End-of-command markers were matched anywhere in the output, so a command could print one and fake success. Markers now carry a random token and must be the last line.
- Silent truncation. Large listings and error output were cut without saying so; now they say so, and the important tail survives.
- Deploy scripts that trusted the happy path. Seven issues, from disabled host-key checks to an Nginx edit that could fail half-written. All now validate, swap atomically and roll back.
One task, the SSH layer, needed four fix rounds. Every round fixed something real, and the final version is far better than my own plan for it was.
What went wrong?
While proving that a shell snippet worked on temporary copies of two shell configuration files, the implementer's path substitution missed, and it edited the real files on the server. Its clean-up then removed a line that had been there before. I caught it by checking the files myself, and it was a one-line restore. The rule I took from it: an implementer running with auto-approve gets an explicit list of directories it may write to, and scratch work goes in a temporary directory, every time.
Was it worth it?
Yes, for anything that will run unattended. The connector shipped with 126 automated tests, several against a private SSH server, and every review finding is either fixed or recorded with a reason. For a quick feature I now build inline and review afterwards; for infrastructure, the three-vendor loop pays for itself the first time it catches a delete in a list.
Questions people ask
Is Claude Code or Codex better?
For production work the more useful question is how to combine them. Models from different vendors catch each other's mistakes, so using one to write and another to review finds defects a single tool misses.
Can the Grok CLI run headless?
Yes. grok -p with --output-format json runs a prompt non-interactively and returns the answer, session id and cost, and --resume continues a session, which makes it usable as an implementer in a pipeline.
How do you stop an AI coding agent editing the wrong files?
Give it an explicit write scope, make it use a temporary directory for experiments, run reviewers read-only, and check the files yourself after any auto-approved run.
What does multi-agent code review find that tests miss?
Design-level problems: unsafe input reaching a command, checks that can be bypassed by timing, and outputs that can be spoofed. Tests prove the cases you thought of; a reviewer from another vendor brings different cases.
About the author. Muhammad Tayyab Ilyas is an Applied AI & Solutions Engineer in Barcelona who builds and operates MCP servers, multi-agent systems and the infrastructure under them. He runs multi-agent delivery in production on LoopCodeLab.