# Four agents tested on a repair, saved corrections and a plain-language explanation

Test date: 4 October 2026. Original research for the Codex vs Claude Code article. This report and the applications are local drafts.

All four agents fixed the three seeded bugs and passed every independent repair check. All four also saved the project corrections and followed them in a fresh chat without a reminder. Claude finished the repair faster in both pairings in this round. Malak ranked the explanations Opus 5 first, Sol second, Fable 5 third and Astra fourth.

The correction test used explicitly saved project notes. Automatic background memory and delayed memory extraction were not tested.

## The repair and saved-note checks

| Agent | Independent repair checks | Browser checks | Saved-note and CSV checks |
|---|---:|---|---:|
| Codex Astra | 61/61 | pass | 8/8 |
| Claude Fable 5 | 61/61 | pass | 8/8 |
| Codex Sol | 61/61 | pass | 8/8 |
| Claude Opus 5 | 61/61 | pass | 8/8 |

These are cases for one small application and one export function. They are not separate projects or repeated statistical trials. Each agent received one repair request, one explanation request, and a two-session correction test: 16 model calls in total.

Exact models: `gpt-6-astra`, `claude-fable-5`, `gpt-6.1-sol`, `claude-opus-5`. The Claude versions match our earlier pilot; this round does not test newer Claude releases. All runs used medium effort through Codex CLI 0.160.0 or Claude Code 2.1.246. Calls ran sequentially using existing subscription logins.

## Each agent fixed a project it hadn't built

The evaluator created the same small Node application for every agent. The original report counted pending payments, counted repeat event IDs again, and kept only the last partial refund. The business requirements were available in the README; the evaluation cases were outside each agent's project and were not supplied in the prompt.

Before any repair run, the broken version passed only 12 of 61 checks. An independently written reference fix passed all 61. The suite contains 21 named cases and 40 deterministic generated cases. It checks repeated events, conflicting retries, partial refunds, status and month boundaries, exact cents, empty months, negative revenue, CSV output, and input mutation. The cases test the supplied business contract, including its explicit first-occurrence rule for repeated IDs.

Every completed repair changed only the original calculation module, `src/ledger.cjs`, and added its own tests. The agents preserved the original interface, endpoint code, CSV formatter, fixture and smoke test. Independent evaluation ran after completion against the delivered files. The evaluator did not repair the outputs or send corrective follow-ups before scoring.

The browser checks exercised September and October, displayed each CSV and reloaded the page. September showed $150 in sales, $30 in refunds and $120 net revenue from four eligible events. October showed a $5 refund and negative $5 net revenue from one event. These are fictional fixtures.

This task checks a small repair with clear requirements. Large repositories, unfamiliar dependencies and architectural refactors remain outside its scope. The agents generated code through their CLIs; the evaluator performed the browser checks. This round does not compare the models' own computer-use abilities.

## Each agent used its own saved corrections in a fresh chat

The teaching request contained three project rules:

1. Display pending contacts as “Needs a person.”
2. Leave email addresses out of contact CSV exports.
3. Use the filename `contacts_review.csv`.

Each agent saved the rules in `PROJECT.md`. The evaluator then launched a completely new process in the same project and requested a contact-export function. That prompt told the agent to read the project notes but did not repeat the three answers.

At the session boundary, `PROJECT.md` was the only prior-context file in each workspace. There was no existing exporter, sample output, teaching prompt or previous answer to copy. Logs and prompts stayed outside the workspace. Snapshots and file hashes preserve what was available.

The eight checks covered the three remembered rules plus CSV structure, approved status, empty input, escaping, final newline and unchanged input. No reminder was sent. Both tools used the same explicit note-reading workflow, with the earlier controlled configuration disabling automatic project instruction injection and session continuity.

This establishes continuity through saved project notes. It does not establish spontaneous recall or a native memory winner. The installed Codex CLI reported its memories feature disabled, and the documented background process has an idle eligibility period. An immediate restart would confound configuration with memory quality. We did not change the owner's global memory settings. See the [official memory guide](https://learn.chatgpt.com/docs/customization/memories), [Codex configuration reference](https://learn.chatgpt.com/docs/config-file/config-reference), and [Claude memory guide](https://code.claude.com/docs/en/memory).

## Malak preferred Opus, then Sol, in the blind explanation test

Every agent received the same factual packet and a request to explain the bug, fix and next check to a business owner in at most 180 words. All four wrote an answer. No tools were used during these explanation calls. The supplied facts described an example fix; the answers are a controlled writing exercise, not claims that those sessions ran tests themselves.

The evaluator randomized the answers into a new A-D order and preserved the original text. [Read the four anonymous explanations](explanation-samples.md).

Malak submitted D, B, A, C before the model names were revealed on 4 October 2026. That maps to:

| Preference | Sample | Model | Words | Within 180 words |
|---|---|---|---:|---|
| 1 | D | Claude Opus 5 | 187 | No |
| 2 | B | Codex Sol | 156 | Yes |
| 3 | A | Claude Fable 5 | 167 | Yes |
| 4 | C | Codex Astra | 164 | Yes |

This records Malak's preference on one explanation prompt. She supplied no reasons, so none are inferred. Factual coverage and instruction compliance remain separate: Opus was her favorite despite exceeding the requested length. The count includes Markdown list numerals; it still exceeds 180 without them. The earlier C, B, D, A ranking belongs to a different writing task.

## Completion times varied by task

| Agent | Repair | Save correction | Fresh-chat implementation | Command denials: repair/save/fresh chat |
|---|---:|---:|---:|---:|
| Codex Astra | 86.50s | 24.71s | 46.66s | 0/0/0 |
| Claude Fable 5 | 62.81s | 8.09s | 31.66s | 0/0/0 |
| Codex Sol | 88.45s | 19.50s | 41.08s | 0/0/0 |
| Claude Opus 5 | 52.14s | 10.18s | 37.97s | 1/0/0 |

These times run from CLI process start to exit, including each agent's own checks and explanation. They exclude evaluator preparation and browser verification. The two CLIs used different permission controls, and equal effort labels do not establish equal compute. Command denials are recorded in the table. One run per task cannot establish a general speed ranking.

## What this adds to the article

The repair result is a tie on the checks we ran. Saved project notes also worked across fresh chats in the cases checked. Malak's second blind writing ranking puts Opus first and Sol second on this explanation task.

The earlier integration tests remain separate evidence: the two Codex outputs rejected unexpected redirects and inconsistent success counts, while the tested Claude outputs accepted them. This round adds a case where the correctness results matched and Claude completed the task sooner. The article should report both outcomes.

## Prompts and evidence

- [Predeclared protocol](round3-results-protocol.md)
- [Unfamiliar-app request](round3-results-repair.txt)
- [Shared explanation facts and request](round3-results-explain.txt)
- [Save-correction request](round3-results-memory-teach.txt)
- [Fresh-chat request](round3-results-memory-probe.txt)
- [Independent repair checks](round3-results-repair-tests.cjs)
- [Saved-note checks](round3-results-memory-tests.cjs)
- [Final evidence audit](round3-results-final-audit.json)
- [Browser observations](round3-results-browser-evidence.jsonl)

The evidence folder preserves the seed app, all four repaired applications, per-run CLI output and metadata, saved-note snapshots, evaluator results and screenshots.

Example browser proof from the Opus repair; all four retained the same original interface:

![Repaired monthly sales report showing $120 net revenue](round3-results-opus-repair.jpg)
