Measuring solver readiness before production: a v0 proposal

CoW’s consistency metric measures reliability from production data: wins, settlements, bid quality. That means it’s zero for a solver that isn’t live yet, and on most chains there’s no shadow environment to generate wins in. So at the moment you’d most want to know whether a solver is reliable, there’s nothing to measure.

This is an attempt at the missing piece: replay real auctions against a solver’s endpoint, with the same time budget the driver actually gives it, and see what comes back. Six metrics, one verdict, a tool anyone can run on their own endpoint.

I ran it on kaisersolver. All three chains came out REVIEW. The report says why.


0. Disclosure and scope

I run kaisersolver: production on Arbitrum One and Base, staging on BNB, and working on the Polygon and Gnosis integrations. I wanted something like this in hand for those conversations, and that’s the position it’s written from.

It proposes no new requirement. Nothing here is a bar a solver must clear. It is a self-assessment a solver can run against its own endpoint, a shared vocabulary for the staging and production conversation with the core team, and a set of definitions anyone can reproduce on their own endpoint. Every number below is about kaisersolver only.

An earlier readiness report I published (docs/readiness/kaisersolver-base-2026-08-22.md in the tool’s repo) measured against an assumed 20-second budget and a 25-auction sample. It is superseded by the method here; the differences are exactly the ones this document fixes.


1. Goal

CIP-85 introduced consistency rewards “intended to incentivize consistent, reliable solver behavior, broad token coverage, and other aspects the core team considers important for maintaining healthy competition.” [ Snapshot ] Consistency metric v2 (felixhenneke, June 2026; in effect since June 30) measures that in production: settlement success rate times bid quality, over the accounting week, from data that exists only once a solver is live and winning. [ Consistency metric v2 · Solver rewards | CoW Protocol Documentation ]

This document proposes the upstream complement: a definition of reliable behaviour that can be computed for a solver before it is live, from a replay of real auctions against its endpoint, using the time budget the driver actually gives it.


2. What exists today

  • Onboarding ( Joining The CoW Protocol Solver Competition | CoW Protocol Documentation ). There are no prerequisites for moving from shadow to staging beyond KYC, though it is recommended that a solver manages to submit and win on shadow first. Production follows staging with the next Tuesday release. Solvers joining under the CoW DAO bonding pool are asked to start on BNB Chain first; other L2s follow “relatively soon”; mainnet “will require further evaluation after some time of solving on L2’s.” The criteria for that evaluation are not published.
  • Shadow competition currently runs on Arbitrum and mainnet only (same page, FAQ). On every other supported chain a solver has no production-orderflow environment before staging.
  • Competition rules ( Solver competition rules | CoW Protocol Documentation ). A solution is valid if it proposes at least one order eligible to contribute to score and respects Uniform Directional Clearing Prices, meaning all orders trading the same tokens in the same direction receive the same price, with an exception for orders with hooks. Winner selection is a fair combinatorial auction (CIP-67): the best single-directed-pair bids form a reference outcome, batched bids that score below the reference on any directed pair are filtered out, then winners are chosen among the survivors. A settlement is valid if it executes the winning solution as bid (solver, score, amounts), honours hooks, and lands before or at the auction deadline; violations can mean immediate denylisting until manual inspection, enforced by the circuit-breaker validator (GitHub - cowprotocol/circuit-breaker-validator: Library that contains the validation logic for circuit breaker · GitHub solver identity, 1-to-1 trade mapping, exact amounts, hooks).
  • Consistency metric v2 (docs, rewards page). Success rate = auction-order pairs won and settled in time ÷ auction-order pairs won. Bid quality: each executed order distributes a weight of one among the solvers that bid on it, in proportion to proposed surplus, counting only bids that pass fairness filtering (equivalent to the forum post’s surplus-relative-to-winner form). Metric = product. A solver that does not win any order in the period has a success rate of zero.
  • Instruments the core team already runs ( GitHub - cowprotocol/services: Off-chain services for CoW Protocol · GitHub at e911013, 2026-09-11). Driver, per solver: solutions{solver,result} with results Success / SolutionNotFound / DeadlineExceeded / SolverHttpError / SolverDeserializeError / SolverDtoError / SolverCustomError; dropped_solutions{solver,reason}; remaining_solve_time and used_solve_time histograms (crates/driver/src/infra/observe/). Autopilot, per driver: solve{driver,result} with success / timeout / no_solutions / error / deny_listed; settle; settled (crates/autopilot/src/run_loop.rs). Shadow autopilot, per driver: results, wins, performance_rewards (crates/autopilot/src/shadow.rs).
  • The /solve request carries a deadline. By default the autopilot’s solve-deadline is 10 s (or block-synced); the driver subtracts a 500 ms HTTP buffer and forwards 80 % of what remains as the solver’s deadline (crates/configs/src/autopilot/run_loop.rs, crates/driver/src/domain/time.rs, crates/driver/src/infra/config/file/mod.rs).

3. The gap

The protocol’s reliability metric is undefined at the moment the readiness decision is made (it is zero for a solver with no wins), and on every supported chain except Arbitrum and mainnet there is no shadow environment to generate wins in. The instruments above exist, but no threshold on them is documented publicly. This document proposes something that could be.


4. How readiness is measured

cow-backtester --readiness (v0.11.0, MIT, eth_abi as its only runtime dependency; GitHub - KaiserSolver/cow-backtester: Offline backtester, A/B harness, and counterfactual scorecard for CoW Protocol solvers — replay recent auctions against your own /solve endpoint · GitHub ) replays real auctions against an endpoint:

  1. Enumerate settlements on chain over a block window. For each, fetch the auction body published to the instance bucket for that auction, and its competition record from the public API.
  2. POST that body to the endpoint under test with deadline set to now + the observed driver budget for the chain and lane (§6). Measure the response with the tool’s own clock.
  3. Validate and score every returned solution against the auction’s orders, reference prices and the historical field.
  4. Aggregate over the window into six metrics and a verdict.

Timing. The endpoint routes on its current view of liquidity, not the auction block’s. The tool therefore runs in watch mode, replaying each auction within minutes of settlement; the report prints the median and p95 replay lag. Answer rate, transport, deadline misses, latency and validity do not depend on that lag. Capture ratio does, and is read accordingly (§5).

Environment. The auctions are always production auctions. The endpoint under test may be a staging or not-yet-activated deployment. That is deliberate: it measures readiness on production orderflow before production.

Load. Replay runs at 2 concurrent requests. On its first full day (2026-09-14) it was 0.27 % of the Arbitrum instance’s /solve requests and 0.24 % of Base’s (650 of 236,563 and 1,429 of 591,023); in its first full hour, 0.62 % and 1.09 %. The BNB instance under test receives no driver traffic yet, so the replay was its only load.

Provenance, labelled on every metric:

  • public: computed from the competition API or chain; verifiable by anyone.
  • replay: measured by the tool posting to the endpoint. Verifiable by anyone on their own endpoint; for the numbers in §9, by the core team, whose driver already talks to this endpoint. Cross-checkable against the driver’s own used_solve_time / solutions{result} series for the same window.
  • protocol-reused: the definition is the protocol’s, cited, not restated.

Reproducibility. Every run archives the bodies it used with their SHA-256 and prints a complete reproduce: line (chain, env, block range, bodies directory, resolved budget, evidence floor, tool version, engine build). The input side of every number in §9 is archived. The output side requires access to the endpoint; if anyone on the core team wants to name a block range, I will run it and publish the manifest.

Run it on your own endpoint.

pip install cow-backtester==0.11.0
cow-backtester --chain base --env prod --readiness \
  --rpc-url <your-rpc> \
  --solver-url https://<your-endpoint-base> --solver-name mine \
  --watch 60 --workers 2 --archive-bodies ./bodies --archive-dir ./records

--env prod selects production auction bodies; point --solver-url at whichever deployment you want to test (the tool appends /solve). --archive-dir keeps the competition records the fairness check needs. On a chain without an observed budget (§6) the tool refuses to produce a verdict; pass --solve-timeout <seconds> with what your endpoint sees as deadline − receipt, and send me the distribution.


5. Proposed definitions

Terms. Attempted: auctions in the window with a settlement on chain, a fetched body, an order set matching the settlement, and a reconstructable winner baseline. Returned: attempted auctions where the endpoint answered with at least one solution. Replayed: attempted auctions where the endpoint answered with a parseable response, empty or not. Auctions with no settlement are not observed by this method; every exclusion is counted and printed with its reason.

Window. A contiguous block range per chain, reported with its wall-clock span, attempted count and attempted-per-hour. Minimum 500 attempted auctions for a verdict; a capped or under-floor sample reports REVIEW at best.

metric:       answer rate                      provenance: replay
numerator:    replayed
denominator:  attempted
mirrors:      driver solutions{result=Success|SolutionNotFound} vs the rest;
              autopilot solve{result}
v0 thresholds: PASS ≥ 90 %  ·  WARN ≥ 50 %  ·  FAIL < 50 %
metric:       transport errors                 provenance: replay
numerator:    HTTP error status, unreachable, oversized body,
              unparseable JSON, wrong schema
denominator:  attempted
mirrors:      driver SolverHttpError / SolverDeserializeError / SolverDtoError;
              autopilot solve{result=error}
v0 thresholds: PASS = 0  ·  WARN ≤ 1 %  ·  FAIL > 1 %
metric:       deadline misses (solve-side)     provenance: replay
numerator:    socket timeout, or an answer that arrived after the stamped deadline
denominator:  attempted
mirrors:      driver DeadlineExceeded; autopilot solve{result=timeout}
note:         settlement-side misses (won but not settled by the auction
              deadline) are consistency v2's success rate and are not
              redefined here
v0 thresholds: PASS = 0  ·  WARN ≤ 1 %  ·  FAIL > 1 %
metric:       solve latency                    provenance: replay
statistic:    p50 / p95 / max of round-trip time over answered auctions
              (a dead endpoint's fast failures do not enter the sample)
budget:       the observed driver budget for the chain and lane (§6),
              stamped into the request as `deadline`
mirrors:      driver used_solve_time{solver,kind=auction}
v0 thresholds: PASS p95 ≤ 50 % of budget  ·  WARN ≤ 100 %  ·  FAIL > 100 %
metric:       solution validity                provenance: replay + protocol-reused
numerator:    auctions with ≥ 1 solution that
              (i)  passes every feasibility check (limit, fill, fee, numeric,
                   duplicate) and is scorable, i.e. has a reference price for
                   its surplus token; this approximates "eligible to
                   contribute to score", and zero-surplus fills are counted
                   separately and printed;
              (ii) has one price per token for every trade it touches: UDCP,
                   satisfied by construction on a /solve response, checked
                   structurally (the protocol's hooks exception does not arise
                   within a single solution);
              (iii) survives the CIP-67 fairness filter against the historical
                   field, as the rules page states it and as implemented in
                   the autopilot's winner-selection crate: the reference
                   outcome comes from single-directed-pair bids, and a batched
                   bid is filtered when any directed pair scores below the
                   reference
denominator:  returned
mirrors:      driver dropped_solutions{reason}; API solutions[].filteredOut
note:         where no competition record exists for an auction, fairness is
              reported "not evaluated", never as passed
v0 thresholds: PASS ≥ 90 %  ·  WARN ≥ 50 %  ·  FAIL < 50 %
metric:       capture ratio                    provenance: replay + public
numerator:    Σ over attempted auctions of the endpoint's best disjoint
              combination of fairness-surviving valid solutions, native wei
denominator:  Σ over the same auctions of the on-chain winner's before-fee
              surplus at uniform prices, decoded from settlement calldata and
              valued at the auction's reference prices (direct settlements
              exact; wrapper-routed a delivered lower bound, counted and
              disclosed)
reading:      a near-time comparison of the endpoint's routing against what
              won, for the same order set; the aggregate analogue of
              consistency v2's bid quality, computable for a solver with zero
              wins. Not a re-run of the auction.
disclosure:   auctions the endpoint's operator won are included; their count
              is printed. The decoded winner surplus can disagree with the
              protocol's referenceScore on auctions with unreliable reference
              prices; such auctions are flagged (see guards) and v0.1 adds a
              sanity check against referenceScore and reports capture on both
              bases.
v0 thresholds: PASS ≥ 50 %  ·  WARN > 0 %  ·  FAIL ≤ 0 %

Verdict. READY iff every check passes, at least 500 auctions were attempted, and the sample was not capped. NOT READY iff any check fails. REVIEW otherwise.

Guards and auxiliary lines. A printed report has more than six lines. Bid coverage (share of answered auctions carrying ≥ 1 solution; warns below 50 %, never fails) describes how broadly the endpoint bids. Price plausibility warns when the endpoint’s claimed surplus exceeds ten times the decoded winner surplus on any auction. That is a guard on the measurement, not a judgement of the solver. Scan coverage and field coverage warn when the RPC missed blocks or more than 5 % of settlements could not be matched to a body. Any of these holds a verdict at REVIEW, by design: a report that cannot vouch for its own inputs should not say READY.

On the thresholds. They are v0 defaults in a per-chain table, shipped with a single default profile. At the 500-auction floor, “WARN ≤ 1 %” means at most five misses; “PASS p95 ≤ 50 % of budget” is a headroom convention for production load and driver-side variance. Both are starting points for calibration, not claims.


6. Per-chain profiles

chain settlement deadline (blocks) observed settle-lane budget p50 / p95 (s) observation status
Arbitrum One 11 4.84 / 4.90 engine logs, 94,575 requests, 2026-09-08 → 09-14 calibrated (own data)
Base 4 4.62 / 4.70 engine logs, 60,191 requests, 2026-09-10 → 09-14 calibrated (own data)
BNB 5 2.35 / 2.36 engine logs, 138 requests, 2026-09-07 → 09-13 (thin) calibrated (own data, thin)
Mainnet 3 — — no own-data observations yet
Gnosis 3 — — no own-data observations yet
Polygon 8 — — no own-data observations yet
Avalanche 8 — — no own-data observations yet
Linea 4 — — no own-data observations yet
Ink 5 — — no own-data observations yet
Plasma 5 — — no own-data observations yet

Settlement deadlines are from Solver competition rules | CoW Protocol Documentation as read on 2026-09-14 (a third-party mirror of the docs shows older values; the docs.cow.fi page is the source). The onboarding page also lists Optimism and Sepolia as endpoint networks, and the local-testing page lists Lens; none has a published settlement deadline, so they are not in the table until one exists.

Observed budgets are reconstructed from engine logs as deadline − start of solve, a slight under-estimate of the wire budget; the wire value replaces them in v0.1 once the engine logs it per request. An independent upper bound corroborates them: the archived bodies’ deadline minus the auction-start block timestamp is p50 5.89 s on Arbitrum, 5.36 s on Base and 3.53 s on BNB over the §9 windows, about a second above the engine-side figure on each chain. The difference is the driver’s send offset. Quote-lane budgets are a separate regime (≈ 2.82 s on all observed chains) and are out of scope for v0.

“No own-data observations yet” is a request: a solver on any of those chains who runs the tool with --solve-timeout set to its observed budget, or who sends me the deadline − receipt distribution, calibrates that row.


7. What a verdict is, and is not

A verdict is a statement about an endpoint over a window: it answered, it answered in time, what it answered was valid under the protocol’s own rules, and how its surplus compared to what actually won. It is not an admission decision, not an EBBO judgement, not a ranking of solvers, and not a prediction of rewards.

8. Non-goals

  • No verdicts on other solvers. This document covers kaisersolver only.
  • No settlement-side rules. Settlement validity is the circuit-breaker validator’s; settlement success is consistency v2’s.
  • No EBBO. Capture ratio compares to what won, not to a best-execution oracle.
  • No attribution the data cannot support. Where a measurement cannot separate two causes, the report says so.

9. Worked example: kaisersolver, 2026-09-14 → 15

Tool 0.11.0 (PyPI, 2026-09-14). Engine builds 3a36e59e5 (Arbitrum, BNB) and e777ab394 (Base), boot-line git_sha, private repo, cited for provenance; the endpoint does not expose it over /solve. Full reports with every check line, exclusion list and archive manifest: docs/readiness/ in the tool repo, commit e7b3b73.

Windows (block range · attempted-auction span · attempted · rate): Arbitrum blocks 505098473 to 505365109 · 2026-09-14 14:05Z → 09-15 08:50Z (18.75 h) · 1,030 · 55/h. Base blocks 51300926 to 51336483 · 14 13:06Z → 15 08:51Z (19.75 h) · 2,098 · 106/h; the window starts after a submission-behaviour change on Base at 10:59Z on 09-14, so it does not straddle it. BNB blocks 121850336 to 122000731 · 14 14:02Z → 15 08:51Z (18.8 h) · 5,970 · 317/h. Median replay lag 0.6 / 0.7 / 1.1 min (Base p95 45.7 min from the first window’s backlog; steady state under a minute).

chain endpoint attempted answer bid coverage transport deadline misses latency p50 / p95 (s) vs budget validity (fairness evaluated / filtered) capture (own wins) per-bid median vs winner verdict
Arbitrum One prod 1,030 100.0 % 47.7 % 0 0 0.79 / 1.98 vs 4.84 (0.41×) 99.8 % (491/491 · 13) 36.4 % (24) 0.979 REVIEW
Base prod 2,098 100.0 % 49.9 % 0 0 0.76 / 2.36 vs 4.62 (0.51×) 100.0 % (1,037/1,047 · 95) 63.1 % (72) 0.955 REVIEW
BNB not yet activated 5,970 100.0 % 23.9 % 0 0 0.92 / 1.53 vs 2.35 (0.65×) 99.9 % (1,426/1,427 · 1) 0.0 %; 8.7 % excluding one valuation artefact (0) 0.952 REVIEW

What passes everywhere. Over 9,098 attempted auctions the endpoints answered every request, missed no deadline, produced no transport error, and every bid auction but two carried at least one valid, fairness-surviving solution (1,443 + 3,485 + 3,142 UDCP checks, 0 violations). The reliability checks are clean on all three chains.

Why REVIEW on all three. No check fails; every REVIEW is a warn. Bid coverage: kaisersolver returns a solution on 47.7 % of Arbitrum auctions, 49.9 % of Base auctions and 23.9 % of BNB auctions; the guard warns below 50 %, and Base is two auctions short of it. Latency headroom: Base p95 is 0.51× budget and BNB 0.65×, both over the 0.5× line; Arbitrum at 0.41× passes. Capture: 36.4 % on Arbitrum, under the 50 % line; the per-bid column shows why. The median of our surplus over the winner’s on unflagged bids is 0.979 (Arbitrum), 0.955 (Base), 0.952 (BNB), so the gap is coverage of large auctions we do not enter, not price quality. Price plausibility: 8, 110 and 385 auctions where our claimed surplus exceeded ten times the decoded winner surplus. Those are reference-price problems in the tool’s valuation; they are disclosed, and they hold the verdict at REVIEW, as the guards are meant to.

BNB, specifically. This is the weakest row and it stays in. It was measured on a keyless public RPC: 4,246 of 150,396 blocks were not scanned and 10.3 % of settlements could not be matched to a body (tx_uid_mismatch 396, body_uid_mismatch 317), so the field is under-counted. One auction (25459284, a 2 USDC order) carries a reference price that values the received tokens at ≈ 154,000 BNB; decoded winner surplus on that single auction moves capture from 8.7 % to 0.0 %. Both are v0.1 items (keyed RPC; winner-surplus sanity check against referenceScore).

Exclusions. About 3 % of the field on Arbitrum (uid mismatches between transaction and body, wrapper settlements the decoder could not attribute), 1 % on Base (wrapper settlements), 10 % on BNB as above. Bodies archived: 1,030 / 2,098 / 5,970 (one per attempted auction) with a SHA-256 manifest.

Reproduce (one line per chain; --bodies-dir replays from the archive, so the input side is fixed):

cow-backtester --chain arbitrum-one --env prod --from-block 505098473 --to-block 505365109 \
  --bodies-dir <archive> --solve-timeout 4.84 --min-evidence 500 \
  --rpc-url <rpc> --solver-url <endpoint> --solver-name kaisersolver --readiness
  # cow-backtester 0.11.0 · engine build 3a36e59e5
cow-backtester --chain base --env prod --from-block 51300926 --to-block 51336483 \
  --bodies-dir <archive> --solve-timeout 4.62 --min-evidence 500 \
  --rpc-url <rpc> --solver-url <endpoint> --solver-name kaisersolver --readiness
  # cow-backtester 0.11.0 · engine build e777ab394
cow-backtester --chain bnb --env prod --from-block 121850336 --to-block 122000731 \
  --bodies-dir <archive> --solve-timeout 2.35 --min-evidence 500 \
  --rpc-url <rpc> --solver-url <endpoint> --solver-name kaisersolver --readiness
  # cow-backtester 0.11.0 · engine build 3a36e59e5

Manifest SHA-256: 0c809a600e676f998a0a7bf550ef810b87889e72b6ccc1845851dcb00f2cfa89 (9,650 lines, 2026-09-15 09:36:16Z)


10. Open questions for the solver team

  1. When a solver moves from staging to production, or is considered for mainnet, are any thresholds applied to the shadow autopilot’s per-driver wins / performance_rewards, or to the production solve{driver,result} series? If so, which? Then this proposal can reference them rather than duplicate them.
  2. For a solver under the CoW DAO pool, would it be acceptable to share aggregated used_solve_time and solutions{result} for that solver’s own endpoint over a stated window? That would let replay-measured latency and answer rate be cross-checked against the driver’s view, which is the cleanest test of the method.
  3. Are the observed settle-lane budgets in §6 what the autopilot configuration predicts on those chains? If per-chain solve-deadline values can be stated, the profile table cites them instead of inferring them.

11. Next

v0.1 after comments: threshold calibration per chain from whatever observations come in; wire-deadline budgets replacing the log-reconstructed ones; capture on the protocol’s referenceScore basis alongside the decoded one, with a sanity check that retires the valuation guard; capture excluding own wins; quote-lane profile; the driver-side cross-check if question 2 is answered. The tool already computes the bid-quality term of consistency v2 counterfactually from competition records (--reward-ev); tying that to the readiness window is the natural bridge between this and the production metric. Starting in October I will publish a readiness report per chain, in this format, every month in docs/readiness/, so whatever the definitions become after this thread, there is a running record measured against them. Tool: GitHub - KaiserSolver/cow-backtester: Offline backtester, A/B harness, and counterfactual scorecard for CoW Protocol solvers — replay recent auctions against your own /solve endpoint · GitHub . Reports: docs/readiness/ in the same repo.

1 Like

Update, eleven days in.

First outside contribution. @ZeroNine ran the tool on Plasma and hit a case the 0.11.1 winner-surplus check couldn’t see: three auctions valued at one bogus reference price masked each other, so capture read ≈ 0 % for every solver. The fix is theirs (the new tests failed on main and pass with it) and shipped in 0.11.2. Thank you.

§6 now has a Plasma row, the first from someone other than me. It’s the tool’s upper bound (archived deadline − auction-start block timestamp): p5 / p50 / p95 5.253 / 5.452 / [5.725] s, min 5.237 s, over 520 auctions (2026-08-17 → 09-16). Two findings came with it:

  • Plasma’s settlement deadline is 10 to 11 blocks in the data (312 and 207 auctions), not the 5 the competition-rules page lists. The winning settlement landed at start + 7 blocks on 513 of 520.
  • The upper bound has a hard floor on three of the four chains measured so far. Min / p5 is 5.237 / 5.253 s on Plasma and 5.295 / 5.322 s on Base, with a softer 5.314 / 5.440 s on Arbitrum. BNB shows no floor. So it isn’t Plasma-specific, but I can’t yet explain BNB.

That sharpens open question 3 for the solver team: is there a configured minimum solve time, and does it apply on BNB? And is Plasma’s settlement deadline 5 blocks or 10 to 11?

Tool. 0.11.3 prints p5 / p99 / min / max of the upper bound so a floor shows up in the report (on GitHub now, PyPI shortly). Reports have moved to GitHub - KaiserSolver/kaisersolver-readiness: Readiness record for kaisersolver on CoW Protocol: every cow-backtester --readiness run, checksummed · GitHub.

Still open: Mainnet, Gnosis, Polygon, Avalanche, Linea and Ink have no rows. Run cow-backtester --readiness --compete on your own endpoint, or send your deadline − receipt distribution with sub-second receipt timestamps (whole-second stamps can’t resolve this), and I’ll add the row.

1 Like

Thanks for the tool, the proposal and the update above. Here is the full Plasma run behind that row, with the numbers §10 asked for and a few things the run surfaced.

Everything below is local or read-only. ZeroNine ran on loopback on a workstation. Chain access was public JSON-RPC reads plus two public HTTPS GETs (the S3 instance bucket and api.cow.fi). Nothing was signed or broadcast.

The run

cow-backtester 0.11.1 (f09dc7d), --readiness --compete --archive-bodies, solve timeout 5.2 s. Window: blocks 30040878 to 32632878, 30.0 days, auctions from 2026-08-17T17:15Z to 2026-09-16T14:42Z.

523 settlements found, 523 auctions formed, 520 attempted, 520 replayed, 338 returned a solution. Six auctions were excluded (3 tx_uid_mismatch, 3 body_uid_mismatch). Every block in the window was scanned; zero failed getLogs ranges.

# Metric Measured Flag
1 Answer rate 100% (520/520) PASS
2 Transport errors 0 PASS
3 Deadline misses 0 PASS
4 Solve latency p50 6 ms, p95 15 ms, max 48 ms of a 5,200 ms budget PASS
5 Solution validity 100% of bid auctions (338 checked, 0 violations) PASS
6 Capture ratio 0% as reported by 0.11.1 (see below) WARN
Bid coverage 65% PASS

Tool verdict: REVIEW. The 500-auction floor was cleared with 520.

Capture: 0% was the guard, not the solver

Three auctions (8939411, 8952964, 9040629) carried 100.0% of the window’s winner surplus. All three price the same buy token at a referencePrice of about 5.0e37, about eleven orders of magnitude above anything else in the window. The 0.11.1 artefact guard tests one auction against the rest of the window combined, so a cluster of three defeats it: the largest is about 0.5x the other two. That is now fixed upstream (PR #1, merged 2026-09-23, thanks for the fast review).

With the three removed: capture 44.18% over the 517 real auctions, 99.26% over the 338 auctions ZeroNine bid, per-bid median ratio 1.000x.

Two caveats on those numbers. They are a replay of one routing path, the one that reads public liquidity, under the tool’s uniform-price scoring. And every auction in the set was settled by someone else first. They are not a live result.

Plasma budgets (the §6 numbers)

Upper bound from public data. original_deadline minus the auctionStartBlock timestamp, the tool’s own measure, over the 520 attempted auctions:

p5 p50 p95 max
5.253 s 5.452 s 5.728 s 10.564 s (one auction)

The floor is sharp at 5.237 s and p99 is 5.763 s. One correction to the row above: 5.725 s is p95 by nearest rank, which is what I sent first; the tool’s own percentile convention gives 5.728 s on the same rows, and the table should carry that one. Like for like on this measure, Plasma’s p50 to p95 spread (0.28 s) sits between Base (0.15 s) and Arbitrum One (0.49 s) in the 2026-09-14 reports. So Plasma is tight like the others, not special.

On open question 3, a configured minimum solve time: the code has one, and I think it is the floor. In services, the autopilot sets its deadline as now plus min_solve_time, optionally moved forward to a block boundary (crates/autopilot/src/run_loop.rs, pick_solve_deadline_impl), and the driver gives solvers now plus 0.8 x (that deadline minus a 500 ms HTTP buffer) (crates/driver/src/domain/time.rs). With both lags at zero the floor is 0.8 x (min_solve_time minus 0.5 s). A floor of 5.237 s implies a min_solve_time of about 7.05 s. The Base and Arbitrum floors in the update above (5.295 s and 5.314 s) imply about 7.1 s by the same formula. The spread above the floor is 0.2 x the driver’s own lag, plus any block alignment. BNB having no floor would then mean a different min_solve_time or a different alignment setting there, which the code cannot tell me. I cannot see the deployed configuration, so this is a reading of the code, not a measurement of the config.

My receipt-side measure, from 3,842 auctions captured on ZeroNine: the receipt stamp in that capture is whole seconds, so it cannot give sub-second percentiles. What it can say is that the deadline’s integer second was the receipt second plus 5 on 3,841 of 3,842. That is consistent with the public upper bound sitting 0.2 to 0.5 s above a receipt-side figure, which is where an auction-start-side bound should sit. That whole-second stamp is why the row is the upper bound and not a receipt-side figure. A receipt-side figure will follow once the stamp records milliseconds.

Supply. 18.2 settlements per day over 70 days (1,271 settlement transactions), every one a single trade. The S3 auction-body archive keeps bodies for about 30 days, and the edge is sharp: 3 of 3 present at 29.1 days, 2 of 2 at 30.0 days, 0 of 3 at 30.9 days and beyond. So the reachable corpus on Plasma is about 523 auctions against the 500 floor, a 4.6% margin. The competition API does not evict; it answered out to 69 days.

Four things the run surfaced

1. Capture is conditioned on settlement by construction. The corpus is built from Trade events, so an auction nobody won cannot be in it. “Capture over all auctions” means over all auctions someone settled. On Plasma that is about 0.1% of the auction stream (18 settle per day out of roughly 18,500). One sentence in the spec would cover it. Two small additions would help a specialist solver without weakening the metric: print the conditional capture (over auctions the solver bid) with the same prominence as the coverage-adjusted one, and let a solver declare a scope (a token or pair set) so the denominator can optionally be restricted to auctions it was eligible for.

2. Rank is not resolvable on Plasma at the tool’s stated tolerance. rank_vs_field inserts the replay’s uniform-price surplus into a field of CIP-38 scores, and documents that the two bases agree to within 0.2%. On the auctions ZeroNine bid, the gap between its surplus and the winner’s score was 0.01 to 0.43 bps. 0.2% is 20 bps. When the gap to the winner is inside the basis tolerance, the rank should be suppressed or flagged as unresolved rather than reported.

3. enumerate_settlements could take a step size. step=800 is hard-coded, so a 30-day Plasma window is 3,240 serial eth_getLogs calls. The same endpoint returned a 200,000-block span for the same filter in 0.2 s, so 13 calls would do. The run pushed the endpoint into 429s, which the tool handled correctly (backoff, rotation, range-splitting, and failed ranges reported rather than zeroed). A --getlogs-step flag, default unchanged, would avoid it.

4. A replay does not drive the solver’s data plane. The harness controls the auction and the clock. It does not feed whatever live data a solver’s pricing depends on. A path that needs a fresh feed can refuse every auction under replay while a second path answers, and every metric above reads healthy. That is what happened here: all 338 bids came from one routing path, and the tool cannot see which path answered. A line in the spec could say that a readiness verdict covers what the replay exercised, and a solver’s report could name which path produced its bids.

1 Like