> ## Content Index
> Fetch the complete content index at: https://fresh-worktree.ghost.io/llms.txt
> Use this file to discover other available public pages before exploring further.

# Did Opus5 replace Superpowers? a brownfield test.
- URL: https://fresh-worktree.ghost.io/did-opus5-replace-superpowers-a-brownfield-test/
- Published: 2026-09-23T00:48:33.000Z
- Updated: 2026-09-23T01:14:39.000Z
- Description: Part 2 evaluation of Superpowers in a complex brownfield project for a gym app, answering the question - Did Opus5 replace Superpowers?
- Author: Tom Skeggs

## 1 - Summary

This fortnight’s outer harness evaluation is the part 2 to an evaluation on [Superpowers](https://github.com/obra/superpowers?ref=fresh-worktree.ghost.io) \- perhaps the most popular AI skills framework for code development in the world.

This is driven by the overarching hypothesis that many of the good practices baked into Superpowers are now trained directly into the models.

The results of [part 1](https://fresh-worktree.ghost.io/are-better-models-replacing-superpowers/) were, for greenfield build of a gym class booking application, a plain Opus5 setup matched Superpowers quality at 1/6th of the cost. However for brownfield implementations which more closely match most real life work - does Superpowers earn its keep?

Part 2 is a complex brownfield project for the same gym app - fixing and finishing a failed version upgrade which was dropped by a poor quality contractor half way through.

**The Story**

The gym is expanding and hired a contractor to uplift its booking system in four areas 1) move off Flask to Typescript 2) add online card payments 3) transition from SQLite to MongoDB and 4) including online membership sales.

The contractor stops responding to mail 3 weeks into the build and left behind a ruin of half finished work, poor documentation and four half finished features to be fixed and finished.

**The hypothesis and test cell setup**

🔍

The Hypothesis - With the arrival of Opus5, Superpowers is no longer beneficial for complex code projects

To test this hypothesis we'll compare a 4x4 block of Opus4.6 (regarded by many in the community as the previous favourite Opus version) vs Opus5 and with/without superpowers.

Instructions were provided in the form of a written specification document (refer to section 2-Hypothesis and Test Setup).

**The Results**

![](https://storage.ghost.io/c/fb/90/fb9049dd-a051-48dd-bffa-7902297a3d99/content/images/2026/09/IMG_0156.jpeg)

Figure 1 Results by run and score

| # | Cell                 | Model    | Harness            | Median tests | Median cost | Median wall (min) | Tests by run (1–5)      |
| - | -------------------- | -------- | ------------------ | ------------ | ----------- | ----------------- | ----------------------- |
| 1 | Plain Opus 5         | Opus 5   | Claude Code alone  | 145/145      | $23.83      | 35.8              | 145, 144, 145, 145, 144 |
| 2 | Superpowers Opus 5   | Opus 5   | superpowers v6.2.0 | 143/145      | $87.43      | 186.9             | 144, 144, 143, 142, 142 |
| 3 | Superpowers Opus 4.6 | Opus 4.6 | superpowers v6.2.0 | 134/145      | $38.91      | 149.8             | 134, 130, 143, 138, 133 |
| 4 | Plain Opus 4.6       | Opus 4.6 | Claude Code alone  | 131/145      | $16.59      | 49.1              | 130, 134, 135, 131, 125 |

In the results we see that plain Opus5 saturated the experiment at the lowest cost. 

In the previous generation (Opus4.6), we see that superpowers performed higher than the plain run median in 4/5 cases - a definitive uplift.

💡

The hypothesis was supported in this instance, we can conclude that Superpowers provided meaningful value in the Opus4 generation but not Opus5, within the scope of this test.

Analysis in the technical deep dive section showed plain Opus5 had significantly earlier and more frequent testing counts as well as twice the number of deep thinking block counts - reinforcing that better "Superpowers like" capabilities have been trained into it.

The findings call into question embedding common coding practices into skills as well as sub agent driven development - many of these benefits are now trained into state of the art (SOTA) models.

Is sub agents driven development dead? Likely no.

At the time of writing, local ai models such as Qwen3.8 27b and Qwen Flash Next are becoming very popular and are hitting benchmarks equal to Opus4.6\. Due to their size they likely receive benefit from using sub agent development - and at no incremental cost increase due to their local hosting.

Further, specific tasks such as security testing are differentiated enough in their goal that they will likely always benefit from sub agents with specific instructions even for SOTA models.

These are both candidates for testing by this blog.

---

Fresh Worktree

## The next test is outer harness setup for local AI development

I've tested Superpowers. Next, I'm testing local AI models and seeing how much of a gap a smart harness can close to the SOTA models. 

[Get the next evaluation ](javascript:void%280%29) Free · roughly monthly · no spam 

---

## 2 - Hypothesis and Test Setup

🔍

The Hypothesis - With the arrival of Opus5, Superpowers is no longer beneficial for complex code projects

In order to test the hypothesis four different cells were established to test multiple harness combinations to complete a failed complex brownfield upgrade.

### The experiment cell block

The following grid shows the 4 cells tested including a baseline Opus5 build, Opus4.6 build and their equivalents with Superpowers.

![](https://storage.ghost.io/c/fb/90/fb9049dd-a051-48dd-bffa-7902297a3d99/content/images/2026/09/IMG_0155.jpeg)

Figure 2 experiment setup configuration

Each test was run 5 times on an automated loop with the same commencement prompt and automatic nudges if the loop paused for user feedback.

### **Scoring the runs**

Summarised below is the frozen test suite used for each cell run.

| File                        | Tests | Style                                                                                                          |
| --------------------------- | ----- | -------------------------------------------------------------------------------------------------------------- |
| test\_1\_ui.py              | 25    | All Playwright — login, schedule, booking buttons, admin dashboard, pinned labels                              |
| test\_2\_api\_security.py   | 12    | All HTTP — 401/403/404/409 semantics, role and ownership checks                                                |
| test\_3\_edge\_cases.py     | 12    | Mixed (7 Playwright, 5 HTTP) — deletion cascades, leave-vs-allocation rules                                    |
| test\_4\_features.py        | 17    | Mostly HTTP (3 Playwright) — outbox, attendance, reminder/birthday/missed-class emails                         |
| test\_5\_pt.py              | 12    | All HTTP — PT booking, trainer eligibility, slot conflicts, cancellation, reminders                            |
| test\_6\_payments\_api.py   | 18    | All HTTP — checkout, idempotency, webhook signature verification, refunds, provider mock                       |
| test\_7\_payments\_flows.py | 12    | All HTTP, reads rendered HTML (no browser driver) — end-to-end pay flow, Paid/Unpaid labels, refund emails     |
| test\_8\_engineering.py     | 10    | Process checks, no HTTP — runs the candidate’s own lint/unit-test/coverage scripts, file-size caps             |
| test\_9\_datastore.py       | 12    | Direct MongoDB reads, no HTTP — SQLite→MongoDB cutover: seeded data, writes, race safety, no leftover .db file |
| test\_10\_memberships.py    | 15    | All HTTP — proration, upgrade/downgrade flow, idempotency, admin roster                                        |

### Prompt and Nudge Approach

The following prompt was provided to kick off each run.

Prompt 

PROMPT — shared core (every cell gets this, verbatim)

TASK        Take over a half-finished project. IronForge Gym is mid-way
            from Flask to TypeScript/Express, mid-way into online payments
            and personal training, mid-way out of SQLite into MongoDB, and
            barely started on memberships. The contractor left. The repo
            runs — incompletely and in places incorrectly.
SPEC        Read TASK_SPEC.md in full first. It outranks this prompt and
            the code. The previous developer's comments and
            MIGRATION_NOTES.md are his assumptions, and some are wrong.
            Its Engineering standards section binds too.
GIT         Clean branch, prior work committed. Commit as you go.
RUN         bash setup.sh
            PORT=<app> PAY_PORT=<provider> MONGO_PORT=<mongo> bash run.sh
            App self-seeds every account, class, booking and leave request.
VERIFY      Reading the diff is not verification. Run it, log in as each
            role, book and cancel, try admin. Check what you fixed works
            AND what already worked still does.
OFF-LIMITS  Do not modify or delete TASK_SPEC.md.
AUTONOMY    Questions go unanswered. Make the most reasonable reading of
            the spec, record assumptions in commits, do not stop to ask
            or to propose a plan. Stop when verified.

PER-CELL ORCHESTRATION SECTION — the only thing that varies

Plain          none; core.md alone.
Superpowers    "Follow your superpowers skills workflow" + "I am asking
               you to dispatch subagents" — subagent-driven-development
               over inline execution, code review by a real reviewer
               subagent, parallel agents for independent tasks. Where a
               skill would ask the partner a question, answer it yourself
               from TASK_SPEC.md and continue.

The task specification included:

- **Situation.** A gym hired one contractor for four things and he “stopped answering email in week three”: rewrite Flask → TypeScript/Express, wire up card payments, move SQLite → MongoDB, sell memberships online. “The existing codebase is the starting point — complete it, don’t restart it.”
- **Contract hierarchy.** “Where anything else in this repository and this document disagree, this document is right.” Three artifacts are binding and are not the contractor’s: `legacy/` (the old Flask app — wording/behaviour reference, not code to run), `payments-mock/README.md` (the provider’s wire contract), `docs/membership-pricing.md` (the owner’s pricing brief). Delegation clause: “Where this memo is silent, the legacy application’s observable wording and behaviour are the contract.”
- **What the system does.** Three roles (member/staff/admin); 3 class types at capacity 30; weekly Mon–Sun schedule, browsable one week ahead; 7 fixed time slots; no double-booking, no overbooking a full class, and “two members booking the last place at the same moment get one booking between them, not two”; 5 instructors with accreditation rules; leave requests (“pending leave does not block anything”); attendance recording; an email outbox with `reminder`/`birthday`/`missed_class`/`refund` record types.
- **PT sessions.** One-to-one, capacity 1, only Sarah Chen and Priya Patel PT-accredited; same slots and two-week window as classes; a request is refused on a teaching conflict, approved leave, or a taken slot; cancelling frees the slot immediately; queues a reminder under the same 8-hour rule.
- **Payments.** Talks to the `payments-mock/` sandbox — “do not edit it and do not replace it”; every amount is integer cents; `POST /api/payments/checkout` requires a non-empty `Idempotency-Key`, one payment per item; “only a verified callback may mark anything paid”; four refund triggers, each queuing subject `IronForge refund issued`; `Paid`/`Unpaid`/`Refunded` columns with a `Pay Now` button on My Bookings and PT Sessions.
- **Memberships.** “The part of the project we were paying for, and it is the part that was barely begun.” Pricing and wording delegated to `docs/membership-pricing.md`; exact route table (`/api/membership`, `/api/membership/upgrade`, `/api/membership/downgrade` plus two `/ui/…` form posts) with a refusal-code table (400/401/403/404/409/502); an upgrade is paid through the same payments machinery; “a membership is never refunded.”
- **Run contract (mandatory).** `setup.sh` then `run.sh` in the foreground on `PORT`/`PAY_PORT`/`MONGO_PORT`; `run.sh` starts MongoDB itself in Docker from `mongo:7` and tears it down; “no `.db` and no `.sqlite` file anywhere in the checkout,” no SQLite driver under `src/`; the app self-seeds Appendix A on first start.
- **Engineering standards (mandatory).** `npm run lint` / `test:unit` / `coverage` must each exit 0 inside 180/300/420-second caps; source under `src/`, ≤400 physical lines a file; ≤3 lint suppressions and the shipped ESLint rules must stay in force; unit suite ≥4 files / ≥20 cases; coverage summary at `coverage/coverage-summary.json` (or a declared path) covering ≥80% of source modules, ≥80% function / ≥70% line coverage.
- **Appendix A — seed data (non-negotiable).** 9 named accounts across the three roles; 20 named classes (5 full) for the current week; two leave requests, one approved one pending; three members’ dates of birth (Alice’s is today’s); four flat prices (1899/1499/1299/6499 cents) and three membership tiers (2900/4900/7900 cents); three membership demo accounts seeded mid-period as the brief’s three worked proration examples.
- **Appendix B — the QA contract.** Exact strings the suite asserts verbatim (`PT Sessions` nav label, `My Bookings` heading, `Take Attendance`/`Request Leave`/`Save Attendance` controls, Admin Dashboard buttons and sections, the `18 / 30 spots` format, the `Payment` column values); six pages that must answer at fixed URLs (`/`, `/login`, `/bookings`, `/pt`, `/membership`, `/admin`); the full non-membership API route table with 401/403/404/409 status-code conventions.

In cases where runs are paused waiting for human input the following nudge was provided. Each run was allowed a maximum of two nudges.

Prompt 

Continue. Make reasonable assumptions and don't ask again.

---

## 3 - Results 

Here are the detailed results tables with commentary. Notes: (1) all costs shown are based on API costs in USD; (2) detailed breakdowns are rounded to the nearest cent; aggregate totals are derived from unrounded values.

### Overall results by median

| # | Cell                 | Model    | Harness            | Median tests | Median cost | Median wall (min) | Tests by run (1–5)      |
| - | -------------------- | -------- | ------------------ | ------------ | ----------- | ----------------- | ----------------------- |
| 1 | Plain Opus 5         | Opus 5   | Claude Code alone  | 145/145      | $23.83      | 35.8              | 145, 144, 145, 145, 144 |
| 2 | Superpowers Opus 5   | Opus 5   | superpowers v6.2.0 | 143/145      | $87.43      | 186.9             | 144, 144, 143, 142, 142 |
| 3 | Superpowers Opus 4.6 | Opus 4.6 | superpowers v6.2.0 | 134/145      | $38.91      | 149.8             | 134, 130, 143, 138, 133 |
| 4 | Plain Opus 4.6       | Opus 4.6 | Claude Code alone  | 131/145      | $16.59      | 49.1              | 130, 134, 135, 131, 125 |

![](https://storage.ghost.io/c/fb/90/fb9049dd-a051-48dd-bffa-7902297a3d99/content/images/2026/09/IMG_0157.jpeg)

Figure 3 cost results per cell

###   
Detailed per run results

| Cell                 | Run | Opus 5 | Opus 4.6 | Sonnet 5 | Haiku 4.5 | Total  |
| -------------------- | --- | ------ | -------- | -------- | --------- | ------ |
| Plain Opus 5         | 1   | $25.11 | —        | —        | —         | $25.11 |
| Plain Opus 5         | 2   | $22.86 | —        | —        | —         | $22.86 |
| Plain Opus 5         | 3   | $23.83 | —        | —        | —         | $23.83 |
| Plain Opus 5         | 4   | $18.54 | —        | —        | —         | $18.54 |
| Plain Opus 5         | 5   | $23.89 | —        | —        | —         | $23.89 |
| Superpowers Opus 5   | 1   | $85.84 | —        | $8.31    | —         | $94.14 |
| Superpowers Opus 5   | 2   | $63.65 | —        | $23.74   | $0.04     | $87.43 |
| Superpowers Opus 5   | 3   | $60.32 | —        | $15.52   | $0.13     | $75.97 |
| Superpowers Opus 5   | 4   | $69.16 | —        | $20.51   | —         | $89.67 |
| Superpowers Opus 5   | 5   | $62.25 | —        | $18.15   | —         | $80.40 |
| Superpowers Opus 4.6 | 1   | —      | $18.09   | $20.82   | —         | $38.91 |
| Superpowers Opus 4.6 | 2   | —      | $27.76   | $2.73    | —         | $30.49 |
| Superpowers Opus 4.6 | 3   | —      | $24.66   | $7.02    | $0.33     | $32.01 |
| Superpowers Opus 4.6 | 4   | —      | $31.03   | $13.83   | $0.14     | $45.00 |
| Superpowers Opus 4.6 | 5   | —      | $39.77   | $11.82   | $0.20     | $51.78 |
| Plain Opus 4.6       | 1   | —      | $13.21   | —        | —         | $13.21 |
| Plain Opus 4.6       | 2   | —      | $13.97   | —        | —         | $13.97 |
| Plain Opus 4.6       | 3   | —      | $17.82   | —        | —         | $17.82 |
| Plain Opus 4.6       | 4   | —      | $18.02   | —        | —         | $18.02 |
| Plain Opus 4.6       | 5   | —      | $16.59   | —        | —         | $16.59 |

### Cost per run by model

| Cell                 | Run | Opus 5 | Opus 4.6 | Sonnet 5 | Haiku 4.5 | Total  |
| -------------------- | --- | ------ | -------- | -------- | --------- | ------ |
| Plain Opus 5         | 1   | $25.11 | —        | —        | —         | $25.11 |
| Plain Opus 5         | 2   | $22.86 | —        | —        | —         | $22.86 |
| Plain Opus 5         | 3   | $23.83 | —        | —        | —         | $23.83 |
| Plain Opus 5         | 4   | $18.54 | —        | —        | —         | $18.54 |
| Plain Opus 5         | 5   | $23.89 | —        | —        | —         | $23.89 |
| Superpowers Opus 5   | 1   | $85.84 | —        | $8.31    | —         | $94.14 |
| Superpowers Opus 5   | 2   | $63.65 | —        | $23.74   | $0.04     | $87.43 |
| Superpowers Opus 5   | 3   | $60.32 | —        | $15.52   | $0.13     | $75.97 |
| Superpowers Opus 5   | 4   | $69.16 | —        | $20.51   | —         | $89.67 |
| Superpowers Opus 5   | 5   | $62.25 | —        | $18.15   | —         | $80.40 |
| Superpowers Opus 4.6 | 1   | —      | $18.09   | $20.82   | —         | $38.91 |
| Superpowers Opus 4.6 | 2   | —      | $27.76   | $2.73    | —         | $30.49 |
| Superpowers Opus 4.6 | 3   | —      | $24.66   | $7.02    | $0.33     | $32.01 |
| Superpowers Opus 4.6 | 4   | —      | $31.03   | $13.83   | $0.14     | $45.00 |
| Superpowers Opus 4.6 | 5   | —      | $39.77   | $11.82   | $0.20     | $51.78 |
| Plain Opus 4.6       | 1   | —      | $13.21   | —        | —         | $13.21 |
| Plain Opus 4.6       | 2   | —      | $13.97   | —        | —         | $13.97 |
| Plain Opus 4.6       | 3   | —      | $17.82   | —        | —         | $17.82 |
| Plain Opus 4.6       | 4   | —      | $18.02   | —        | —         | $18.02 |
| Plain Opus 4.6       | 5   | —      | $16.59   | —        | —         | $16.59 |

###   
Total cost by model

| Cell                 | Opus 5  | Opus 4.6 | Sonnet 5 | Haiku 4.5 | Total       |
| -------------------- | ------- | -------- | -------- | --------- | ----------- |
| Plain Opus 5         | $114.23 | —        | —        | —         | $114.23     |
| Superpowers Opus 5   | $341.22 | —        | $86.23   | $0.17     | $427.61     |
| Superpowers Opus 4.6 | —       | $141.30  | $56.21   | $0.66     | $198.18     |
| Plain Opus 4.6       | —       | $79.61   | —        | —         | $79.61      |
| **All cells**        | $455.45 | $220.91  | $142.44  | $0.83     | **$819.64** |

---

## 4 - Technical notes

### 4.1 Difference in cost efficiency between Opus4.6 and Opus5

Of some interest Opus5 with superpowers was significantly less efficient on cost than Opus4.6 with superpowers.

This is because the default auto compaction window for Opus4.6 is 200k vs 967k tokens for Opus5\. It’s commonly agreed that Opus is able to intelligently run for much longer context windows effectively in new generations, however noting this can come with higher cost.

Token compaction settings can be set within Claude Code very easily and may save users significant tokens when using Opus5.

### 4.2 Difference in behaviour between Opus4.6 and Opus5

The following chart and table show that whilst Opus4.6 did significantly more planning tool calls (represented by count of tools before writing code), Opus5 had much more hidden reasoning blocks (thinking time that is done inside of the model) and also did both more and earlier testing which saved it from the biggest error grouping - upgrade and downgrade of memberships.

![](https://storage.ghost.io/c/fb/90/fb9049dd-a051-48dd-bffa-7902297a3d99/content/images/2026/09/image-1.png)

Figure 4 execution timeline for upgrade handler

**Analysis table of Opus planning behaviour**

Did Opus 5 plan better up front? 

| Median of 5 runs, unless a count                                                                      | Opus 5          | Opus 4.6        |
| ----------------------------------------------------------------------------------------------------- | --------------- | --------------- |
| Tool calls in the run                                                                                 | 132             | 236             |
| Share of the run spent reading before the first file                                                  | 18%             | 18%             |
| Read the whole spec, pricing brief, provider README, old app and inherited core before the first file | 5 of 5          | 5 of 5          |
| Text read before the first file                                                                       | 262K characters | 288K characters |
| Visible planning written before the first file                                                        | 28 words        | 249 words       |
| Hidden reasoning blocks across the run                                                                | 65              | 30              |
| Upgrade handler first written at                                                                      | 37% of the run  | 39% of the run  |
| First draft answered with the change and its payment                                                  | 5 of 5          | 0 of 5          |
| Live HTTP calls to the upgrade endpoint                                                               | 3               | 1               |
| Gap from writing the handler to first calling it                                                      | 22% of the run  | 36% of the run  |
| Upgrade test passed                                                                                   | 5 of 5          | 0 of 5          |

Plain cells only, where nothing but the model varies. Positions are shares of each run’s tool calls, because Opus 4.6 reads in smaller chunks and takes about twice the calls to read the same material.

### 4.3 Test failures centred on higher ambiguity areas of the spec

The broad test failures were exasperated due to the fact the specifications lived in 3 different documents.

In multiple cases the root cause was one issue which triggered multiple other test failures - such as the deliberately placed contradiction within the fictional contractors documentation which was asked to be fixed by the core spec.

Where the marks went 

| Area                                                               | Sittings affected                                               | Marks lost per run | What the tests check                                                                                                                                                                                                                                                                                                                       | What went wrong                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| ------------------------------------------------------------------ | --------------------------------------------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Memberships: the upgrade and downgrade flow**8 tests in test\_10 | Plain Opus 5 ×1 · Super Opus 5 ×5 · Super 4.6 ×4 · Plain 4.6 ×5 | up to 8            | When a member upgrades, the app should charge them the right partial amount and only switch them to the new tier once the payment provider confirms the money went through. A downgrade should just be scheduled for the next renewal, with no charge. A member can only have one change pending at a time.                                | This is where the most marks were lost, and it is still a bug even though the money moved correctly. Opus 4.6 failed the core upgrade test in 9 of its 10 runs, and because later tests depend on it, that single failure dragged down up to 8 marks each time. In 6 of those 9 failures the app did charge the member correctly and did wait for the provider’s confirmation — but it shipped a broken API response all the same, returning the payment details loose at the top level instead of where they belonged, so nothing downstream could find them. Only one run actually got the payment logic itself wrong. Opus 5 got the flow right and still lost a point in 6 of its 10 runs, because the scheduled downgrade came back under the wrong field name — a smaller mistake, but a mistake, not a grading quirk. |
| **Inherited wording the spec pins**6 tests in test\_1 and test\_4  | Super 4.6 ×5 · Plain 4.6 ×5                                     | up to 6            | Two bits of exact wording the spec called for: a failed login must say “Invalid email or password”, and the instructor’s menu link must read “My Teaching Schedule” — several other tests use that exact link to reach the leave-request and attendance pages.                                                                             | The starting code had two small wording mistakes left in by the fictional contractor, and the spec explicitly asked for them to be fixed. Every Opus 4.6 run left the login message as “Incorrect credentials”, costing 1 mark. Four out of ten Opus 4.6 runs also left the menu link shortened to just “Teaching”, which broke five other tests that click through that exact link — 5 marks lost in one go. Two overlooked words, up to 6 marks a run. Opus 5 caught both every time.                                                                                                                                                                                                                                                                                                                                      |
| **The MongoDB cutover**12 tests in test\_9                         | Super Opus 5 ×3 · Super 4.6 ×2                                  | up to 1            | The app’s data was supposed to be fully moved over to MongoDB: every member, class, booking and membership should be stored there, anything created through the app should be saved there too, and no trace of the old SQLite database file should be left behind anywhere.                                                                | Two small, separate slip-ups, each costing one point. Two runs (Opus 4.6 with the plugin) finished the move to MongoDB but still opened the old .data/ironforge.db SQLite file on startup — a harmless leftover rather than a real bug. Three runs (Opus 5 with the plugin) stored membership data correctly, but included today’s date as an extra field in the response; the checker expects every date it sees in a response to also be saved in the database, and this one wasn’t, so it wrongly looked like something was missing.                                                                                                                                                                                                                                                                                      |
| **Engineering standards**10 tests in test\_8                       | Super 4.6 ×2 · Plain 4.6 ×3                                     | up to 3            | The code should pass its own code-quality check (linting) and its own unit tests, and a coverage report should show the whole codebase was checked, not just part of it.                                                                                                                                                                   | Minor housekeeping slip-ups, all from the weaker model (Opus 4.6): some runs left linting errors in their own test files, some only measured code coverage for part of the codebase instead of all of it, and one run’s coverage check failed outright because its tests needed a MongoDB database that wasn’t running at the time. Up to 3 marks a run.                                                                                                                                                                                                                                                                                                                                                                                                                                                                     |
| **Everything else**9 tests, none failing more than twice           | Plain Opus 5 ×1 · Super Opus 5 ×2 · Super 4.6 ×2 · Plain 4.6 ×3 | up to 5            | A grab-bag of smaller checks: a logged-in user without the right permissions should get a “forbidden” error rather than a “not logged in” one; booking a class should queue a reminder email; the admin page should list every member’s membership tier; and the very first login after the app starts should respond within five seconds. | One run left a shared permission-check function returning the wrong error code, costing 5 marks in one go since several tests rely on it. Two runs (Opus 5 with the plugin) actually listed every member’s tier correctly, but lost the point because the checker looks for a member’s email in the first place it appears on the page — and it found it in an unrelated table above the member list instead. Four runs (Opus 4.6) were just slow to respond on the very first login after starting up; every later login was fine. The rest were one-off misses.                                                                                                                                                                                                                                                            |

---

## 5 - Future Tests

Thank you for reading this far! 

My focus for this blog is to test the latest novel approaches to the outer harness.

If you have any pressing ideas, lets me know - I'll probably run them!

Recent tests have shown that new LLM models such as Opus5 are able to handle longer and longer contexts autonomously - removing a key limitation of previous models.

The next test will look at local AI development, which has been a topic of much recent interest in the open vs closed weight debate.

Whilst it’s a given that open weight models are not as capable as frontier closed weight models, the question remains - with the right harness can open models reach close to the frontier at 0 incremental cost?

And … if so, what are these outer harness setups?

---

Fresh Worktree

## You made it to the end. The next one's for you.

New outer harness evaluations roughly monthly — same frozen test suites, real API costs, full methodology. 

[Subscribe free ](javascript:void%280%29) No spam · Unsubscribe anytime