# Codex vs Claude Code: methods across three rounds

Recorded on October 4, 2026. These are 36 task runs across a few small, fictional tasks, using Astra, Sol, Fable 5 and Opus 5. Malak supplied daily-use preferences and ranked two sets of anonymous writing. Codex executed model commands and independent checks.

| Round | Task runs | Detailed record |
|---|---:|---|
| Client email, dashboard and filter update | 12 | Original pilot below |
| Subscription manager and pause/resume update | 8 | [Reliability checks](reliability-results.md) |
| Unfamiliar repair, explanation, save correction and fresh-chat use | 16 | [Repair and saved-note checks](round3-results.md) |

Opus ranked first and Sol second in both blind writing tasks. The second order was Opus, Sol, Fable, Astra. Read [the explanation samples and reveal](explanation-samples.md). Automatic background memory, each model's own browser operation and voice remain untested. The saved-correction test deliberately used explicit project notes. Test cases are not independent projects or statistical trials.

The following record describes only the original 12-run pilot. Its timings and ordering remain unchanged.

---

# Codex vs Claude Code: firsthand pilot

Date: October 4, 2026. Status: 12 scored runs complete; blind preference ranking recorded.

This pilot tests a fictional client email, a small sales dashboard, and a fresh-session edit to that dashboard. It does not establish broad model superiority. Malak’s preferences for Codex’s conversational writing, computer use and voice remain separate firsthand opinions.

## Setup

Approved pairings: Astra vs Fable, Sol vs Opus. Exact models: gpt-6-astra, claude-fable-5, gpt-6.1-sol, claude-opus-5. These results must not be described as a test of newer Claude versions. Codex CLI 0.160.0; Claude CLI 2.1.246. Both used existing subscription authentication. Medium reasoning was selected; matching labels do not mean equal compute.

Models ran sequentially in fresh CLI sessions and separate folders with identical synthetic inputs. No production/client data was submitted. Writing had tools disabled/read-only. Coding permitted local file operations and Node checks. Codex had broader shell access; Claude had Read/Write/Edit and Node-prefixed Bash authorization. Consequently these are observed application-run completion times, not controlled inference-speed measurements. Subscription capacity, service load and built-in prompts may differ.

Codex desktop control was blocked by the computer-use tool. We did not bypass it. Malak delegated the route choice; we selected CLI execution and independent browser verification. The evaluator used Chrome through computer/browser-use tools to check the finished dashboards. Participants themselves did not operate the browser.

## Writing

Each model got the same 90–140-word client-email brief. All four retained the required facts, conditional handover and administrator-introduction request. All met the word limit and basic mechanical constraints.

| Model | Completion time | Body words |
|---|---:|---:|
| Astra | 9.58 s | 97 |
| Fable 5 | 5.86 s | 117 |
| Sol | 9.96 s | 102 |
| Opus 5 | 5.34 s | 120 |

Body counts include greeting/signoff and exclude subject. Codex outputs were shorter; Claude outputs completed sooner on this one task. Malak submitted the blind preference order C, B, D, A before the key was revealed:

| Preference | Sample | Model |
|---|---|---|
| 1 | C | Claude Opus 5 |
| 2 | B | Codex Sol (gpt-6.1-sol) |
| 3 | D | Codex Astra (gpt-6-astra) |
| 4 | A | Claude Fable 5 |

Within the approved pairings, Astra was preferred to Fable and Opus was preferred to Sol. This is a personal ranking of one client-email output per model, not a general writing-quality score. Reasons for the ranking and editing time were not supplied. It qualifies the initial preference for Codex writing: Opus won this particular blind test.

## Initial dashboard

All four scored outputs passed independent browser checks: full total $1,070 / 8 orders; client totals North $470, Cedar $280, Harbor $320; September 2–5 inclusive $679.50 / 5 orders; an October range with zero orders and a clear empty state; reset restores the full data. All showed no horizontal overflow at a 390-pixel viewport, and initial mobile layouts were visually inspected.

| Model | Completion time | Browser functionality |
|---|---:|---|
| Astra | 179.51 s | Pass |
| Fable 5 | 57.78 s | Pass |
| Sol | 179.86 s | Pass |
| Opus 5 | 112.84 s | Pass |

Functional passes do not imply perfect instruction-following. Opus reported writing `/tmp/dash_test.js` outside its assigned folder, contrary to the prompt, and that file was observed on disk. It also encountered a denied shell command while probing browser tooling. Keep this exception visible when using the results.

## Follow-up edit

Each model received its own saved dashboard in a new folder and fresh session, then the same request to add a client filter, combine it with date filters, and reset both. This tests editing an existing artifact, not retaining a long conversation. All four passed independent Chrome checks: default totals, Cedar Works plus September 2–5 ($80 / 2 orders), October empty range, reset clearing both filters, and no horizontal overflow at 390 pixels.

| Model | Completion time | Browser functionality |
|---|---:|---|
| Astra | 121.22 s | Pass |
| Fable 5 | 90.26 s | Pass |
| Sol | 113.73 s | Pass |
| Opus 5 | 156.39 s | Pass |

Fable finished sooner than Astra; Sol finished sooner than Opus. These are single runs with different testing choices, not reliable rankings.

Fable’s follow-up reported temporary test scaffolding in /tmp, contrary to the folder restriction, and encountered a denied non-Node shell command. Opus also reported writing its follow-up test harness to `/tmp/dash-test/test.js`, outside its assigned folder. These are separate instruction-following issues from whether the resulting dashboards work. No outside-folder writes were observed in the Codex command logs reviewed. Claude logs do not expose all intermediate operations, so this is not a complete filesystem audit.

## Retries and evidence limits

- Initial Codex CLI 0.150.1 rejected the requested models as requiring an update. Updated through the built-in official updater and retried; failed setup attempts remain saved and are excluded from task times.
- Initial Claude coding runs lacked a shell tool. Fable (42.41 s) and Opus (66.84 s) were rerun from scratch with Node access. Those first runs remain saved but are excluded from the table.
- The runner had a local text-mode setup error before any successful task submission; fixed before scored runs.
- One scored run per model/task. No statistical conclusion, no randomized ordering, no repeated-task variance estimate, and no inference about large-repository coding.
- Timing uses local monotonic subprocess start-to-exit, including CLI overhead and testing decisions, excluding evaluator browser checks. Model response latency alone is not measured.
- Claude JSON logs contain final results/usage/denials rather than a full intermediate tool transcript. Model-reported checks are distinguished from evaluator-observed UI outcomes. No actual billing-cost comparison is made from list-price telemetry.
- B1 public browser research was not run. OS computer use, voice, live interruption handling and conversation memory were not tested. Screenshots were inspected in tool output; no local screenshot files were retained.

## Article claims supported so far

All four models built a working small dashboard from the same brief. Claude completed these particular writing and initial coding runs sooner. Codex produced shorter emails. In the blind email ranking, Malak preferred Opus, then Sol, Astra and Fable. This supports a task-specific preference, alongside her broader real-use preference for Codex. These results do not substantiate “Codex is faster” or “Codex is better at computer use.” Preserve conflicting evidence in the article.

## Evidence location

Raw prompts, model responses, timings, outputs and checks: `.system/benchmarks/codex-claude-2026-10-04/`.
Writing: `writing/astra-attempt2`, `writing/fable`, `writing/sol-attempt2`, `writing/opus`.
Scored coding: `coding/astra`, `coding/fable-attempt2`, `coding/sol`, `coding/opus-attempt2`.
Follow-up: `followup/astra`, `followup/fable`, `followup/sol`, `followup/opus`.
The runner is `run.py`; browser observations are summarized in `verification.json`. Blind samples are also saved alongside this report. The private evaluation key is `writing-key.json` in the scratch evidence folder.
