CoW’s consistency metric measures reliability from production data: wins, settlements, bid quality. That means it’s zero for a solver that isn’t live yet, and on most chains there’s no shadow environment to generate wins in. So at the moment you’d most want to know whether a solver is reliable, there’s nothing to measure.
This is an attempt at the missing piece: replay real auctions against a solver’s endpoint, with the same time budget the driver actually gives it, and see what comes back. Six metrics, one verdict, a tool anyone can run on their own endpoint.
I ran it on kaisersolver. All three chains came out REVIEW. The report says why.
0. Disclosure and scope
I run kaisersolver: production on Arbitrum One and Base, staging on BNB, and working on the Polygon and Gnosis integrations. I wanted something like this in hand for those conversations, and that’s the position it’s written from.
It proposes no new requirement. Nothing here is a bar a solver must clear. It is a self-assessment a solver can run against its own endpoint, a shared vocabulary for the staging and production conversation with the core team, and a set of definitions anyone can reproduce on their own endpoint. Every number below is about kaisersolver only.
An earlier readiness report I published (docs/readiness/kaisersolver-base-2026-08-22.md in the tool’s repo) measured against an assumed 20-second budget and a 25-auction sample. It is superseded by the method here; the differences are exactly the ones this document fixes.
1. Goal
CIP-85 introduced consistency rewards “intended to incentivize consistent, reliable solver behavior, broad token coverage, and other aspects the core team considers important for maintaining healthy competition.” [ Snapshot ] Consistency metric v2 (felixhenneke, June 2026; in effect since June 30) measures that in production: settlement success rate times bid quality, over the accounting week, from data that exists only once a solver is live and winning. [ Consistency metric v2 · Solver rewards | CoW Protocol Documentation ]
This document proposes the upstream complement: a definition of reliable behaviour that can be computed for a solver before it is live, from a replay of real auctions against its endpoint, using the time budget the driver actually gives it.
2. What exists today
- Onboarding ( Joining The CoW Protocol Solver Competition | CoW Protocol Documentation ). There are no prerequisites for moving from shadow to staging beyond KYC, though it is recommended that a solver manages to submit and win on shadow first. Production follows staging with the next Tuesday release. Solvers joining under the CoW DAO bonding pool are asked to start on BNB Chain first; other L2s follow “relatively soon”; mainnet “will require further evaluation after some time of solving on L2’s.” The criteria for that evaluation are not published.
- Shadow competition currently runs on Arbitrum and mainnet only (same page, FAQ). On every other supported chain a solver has no production-orderflow environment before staging.
- Competition rules ( Solver competition rules | CoW Protocol Documentation ). A solution is valid if it proposes at least one order eligible to contribute to score and respects Uniform Directional Clearing Prices, meaning all orders trading the same tokens in the same direction receive the same price, with an exception for orders with hooks. Winner selection is a fair combinatorial auction (CIP-67): the best single-directed-pair bids form a reference outcome, batched bids that score below the reference on any directed pair are filtered out, then winners are chosen among the survivors. A settlement is valid if it executes the winning solution as bid (solver, score, amounts), honours hooks, and lands before or at the auction deadline; violations can mean immediate denylisting until manual inspection, enforced by the circuit-breaker validator (GitHub - cowprotocol/circuit-breaker-validator: Library that contains the validation logic for circuit breaker · GitHub solver identity, 1-to-1 trade mapping, exact amounts, hooks).
- Consistency metric v2 (docs, rewards page). Success rate = auction-order pairs won and settled in time ÷ auction-order pairs won. Bid quality: each executed order distributes a weight of one among the solvers that bid on it, in proportion to proposed surplus, counting only bids that pass fairness filtering (equivalent to the forum post’s surplus-relative-to-winner form). Metric = product. A solver that does not win any order in the period has a success rate of zero.
- Instruments the core team already runs ( GitHub - cowprotocol/services: Off-chain services for CoW Protocol · GitHub at
e911013, 2026-09-11). Driver, per solver:solutions{solver,result}with results Success / SolutionNotFound / DeadlineExceeded / SolverHttpError / SolverDeserializeError / SolverDtoError / SolverCustomError;dropped_solutions{solver,reason};remaining_solve_timeandused_solve_timehistograms (crates/driver/src/infra/observe/). Autopilot, per driver:solve{driver,result}with success / timeout / no_solutions / error / deny_listed;settle;settled(crates/autopilot/src/run_loop.rs). Shadow autopilot, per driver:results,wins,performance_rewards(crates/autopilot/src/shadow.rs). - The
/solverequest carries adeadline. By default the autopilot’ssolve-deadlineis 10 s (or block-synced); the driver subtracts a 500 ms HTTP buffer and forwards 80 % of what remains as the solver’s deadline (crates/configs/src/autopilot/run_loop.rs,crates/driver/src/domain/time.rs,crates/driver/src/infra/config/file/mod.rs).
3. The gap
The protocol’s reliability metric is undefined at the moment the readiness decision is made (it is zero for a solver with no wins), and on every supported chain except Arbitrum and mainnet there is no shadow environment to generate wins in. The instruments above exist, but no threshold on them is documented publicly. This document proposes something that could be.
4. How readiness is measured
cow-backtester --readiness (v0.11.0, MIT, eth_abi as its only runtime dependency; GitHub - KaiserSolver/cow-backtester: Offline backtester, A/B harness, and counterfactual scorecard for CoW Protocol solvers — replay recent auctions against your own /solve endpoint · GitHub ) replays real auctions against an endpoint:
- Enumerate settlements on chain over a block window. For each, fetch the auction body published to the instance bucket for that auction, and its competition record from the public API.
- POST that body to the endpoint under test with
deadlineset to now + the observed driver budget for the chain and lane (§6). Measure the response with the tool’s own clock. - Validate and score every returned solution against the auction’s orders, reference prices and the historical field.
- Aggregate over the window into six metrics and a verdict.
Timing. The endpoint routes on its current view of liquidity, not the auction block’s. The tool therefore runs in watch mode, replaying each auction within minutes of settlement; the report prints the median and p95 replay lag. Answer rate, transport, deadline misses, latency and validity do not depend on that lag. Capture ratio does, and is read accordingly (§5).
Environment. The auctions are always production auctions. The endpoint under test may be a staging or not-yet-activated deployment. That is deliberate: it measures readiness on production orderflow before production.
Load. Replay runs at 2 concurrent requests. On its first full day (2026-09-14) it was 0.27 % of the Arbitrum instance’s /solve requests and 0.24 % of Base’s (650 of 236,563 and 1,429 of 591,023); in its first full hour, 0.62 % and 1.09 %. The BNB instance under test receives no driver traffic yet, so the replay was its only load.
Provenance, labelled on every metric:
- public: computed from the competition API or chain; verifiable by anyone.
- replay: measured by the tool posting to the endpoint. Verifiable by anyone on their own endpoint; for the numbers in §9, by the core team, whose driver already talks to this endpoint. Cross-checkable against the driver’s own
used_solve_time/solutions{result}series for the same window. - protocol-reused: the definition is the protocol’s, cited, not restated.
Reproducibility. Every run archives the bodies it used with their SHA-256 and prints a complete reproduce: line (chain, env, block range, bodies directory, resolved budget, evidence floor, tool version, engine build). The input side of every number in §9 is archived. The output side requires access to the endpoint; if anyone on the core team wants to name a block range, I will run it and publish the manifest.
Run it on your own endpoint.
pip install cow-backtester==0.11.0
cow-backtester --chain base --env prod --readiness \
--rpc-url <your-rpc> \
--solver-url https://<your-endpoint-base> --solver-name mine \
--watch 60 --workers 2 --archive-bodies ./bodies --archive-dir ./records
--env prod selects production auction bodies; point --solver-url at whichever deployment you want to test (the tool appends /solve). --archive-dir keeps the competition records the fairness check needs. On a chain without an observed budget (§6) the tool refuses to produce a verdict; pass --solve-timeout <seconds> with what your endpoint sees as deadline − receipt, and send me the distribution.
5. Proposed definitions
Terms. Attempted: auctions in the window with a settlement on chain, a fetched body, an order set matching the settlement, and a reconstructable winner baseline. Returned: attempted auctions where the endpoint answered with at least one solution. Replayed: attempted auctions where the endpoint answered with a parseable response, empty or not. Auctions with no settlement are not observed by this method; every exclusion is counted and printed with its reason.
Window. A contiguous block range per chain, reported with its wall-clock span, attempted count and attempted-per-hour. Minimum 500 attempted auctions for a verdict; a capped or under-floor sample reports REVIEW at best.
metric: answer rate provenance: replay
numerator: replayed
denominator: attempted
mirrors: driver solutions{result=Success|SolutionNotFound} vs the rest;
autopilot solve{result}
v0 thresholds: PASS ≥ 90 % · WARN ≥ 50 % · FAIL < 50 %
metric: transport errors provenance: replay
numerator: HTTP error status, unreachable, oversized body,
unparseable JSON, wrong schema
denominator: attempted
mirrors: driver SolverHttpError / SolverDeserializeError / SolverDtoError;
autopilot solve{result=error}
v0 thresholds: PASS = 0 · WARN ≤ 1 % · FAIL > 1 %
metric: deadline misses (solve-side) provenance: replay
numerator: socket timeout, or an answer that arrived after the stamped deadline
denominator: attempted
mirrors: driver DeadlineExceeded; autopilot solve{result=timeout}
note: settlement-side misses (won but not settled by the auction
deadline) are consistency v2's success rate and are not
redefined here
v0 thresholds: PASS = 0 · WARN ≤ 1 % · FAIL > 1 %
metric: solve latency provenance: replay
statistic: p50 / p95 / max of round-trip time over answered auctions
(a dead endpoint's fast failures do not enter the sample)
budget: the observed driver budget for the chain and lane (§6),
stamped into the request as `deadline`
mirrors: driver used_solve_time{solver,kind=auction}
v0 thresholds: PASS p95 ≤ 50 % of budget · WARN ≤ 100 % · FAIL > 100 %
metric: solution validity provenance: replay + protocol-reused
numerator: auctions with ≥ 1 solution that
(i) passes every feasibility check (limit, fill, fee, numeric,
duplicate) and is scorable, i.e. has a reference price for
its surplus token; this approximates "eligible to
contribute to score", and zero-surplus fills are counted
separately and printed;
(ii) has one price per token for every trade it touches: UDCP,
satisfied by construction on a /solve response, checked
structurally (the protocol's hooks exception does not arise
within a single solution);
(iii) survives the CIP-67 fairness filter against the historical
field, as the rules page states it and as implemented in
the autopilot's winner-selection crate: the reference
outcome comes from single-directed-pair bids, and a batched
bid is filtered when any directed pair scores below the
reference
denominator: returned
mirrors: driver dropped_solutions{reason}; API solutions[].filteredOut
note: where no competition record exists for an auction, fairness is
reported "not evaluated", never as passed
v0 thresholds: PASS ≥ 90 % · WARN ≥ 50 % · FAIL < 50 %
metric: capture ratio provenance: replay + public
numerator: Σ over attempted auctions of the endpoint's best disjoint
combination of fairness-surviving valid solutions, native wei
denominator: Σ over the same auctions of the on-chain winner's before-fee
surplus at uniform prices, decoded from settlement calldata and
valued at the auction's reference prices (direct settlements
exact; wrapper-routed a delivered lower bound, counted and
disclosed)
reading: a near-time comparison of the endpoint's routing against what
won, for the same order set; the aggregate analogue of
consistency v2's bid quality, computable for a solver with zero
wins. Not a re-run of the auction.
disclosure: auctions the endpoint's operator won are included; their count
is printed. The decoded winner surplus can disagree with the
protocol's referenceScore on auctions with unreliable reference
prices; such auctions are flagged (see guards) and v0.1 adds a
sanity check against referenceScore and reports capture on both
bases.
v0 thresholds: PASS ≥ 50 % · WARN > 0 % · FAIL ≤ 0 %
Verdict. READY iff every check passes, at least 500 auctions were attempted, and the sample was not capped. NOT READY iff any check fails. REVIEW otherwise.
Guards and auxiliary lines. A printed report has more than six lines. Bid coverage (share of answered auctions carrying ≥ 1 solution; warns below 50 %, never fails) describes how broadly the endpoint bids. Price plausibility warns when the endpoint’s claimed surplus exceeds ten times the decoded winner surplus on any auction. That is a guard on the measurement, not a judgement of the solver. Scan coverage and field coverage warn when the RPC missed blocks or more than 5 % of settlements could not be matched to a body. Any of these holds a verdict at REVIEW, by design: a report that cannot vouch for its own inputs should not say READY.
On the thresholds. They are v0 defaults in a per-chain table, shipped with a single default profile. At the 500-auction floor, “WARN ≤ 1 %” means at most five misses; “PASS p95 ≤ 50 % of budget” is a headroom convention for production load and driver-side variance. Both are starting points for calibration, not claims.
6. Per-chain profiles
| chain | settlement deadline (blocks) | observed settle-lane budget p50 / p95 (s) | observation | status |
|---|---|---|---|---|
| Arbitrum One | 11 | 4.84 / 4.90 | engine logs, 94,575 requests, 2026-09-08 → 09-14 | calibrated (own data) |
| Base | 4 | 4.62 / 4.70 | engine logs, 60,191 requests, 2026-09-10 → 09-14 | calibrated (own data) |
| BNB | 5 | 2.35 / 2.36 | engine logs, 138 requests, 2026-09-07 → 09-13 (thin) | calibrated (own data, thin) |
| Mainnet | 3 | — | — | no own-data observations yet |
| Gnosis | 3 | — | — | no own-data observations yet |
| Polygon | 8 | — | — | no own-data observations yet |
| Avalanche | 8 | — | — | no own-data observations yet |
| Linea | 4 | — | — | no own-data observations yet |
| Ink | 5 | — | — | no own-data observations yet |
| Plasma | 5 | — | — | no own-data observations yet |
Settlement deadlines are from Solver competition rules | CoW Protocol Documentation as read on 2026-09-14 (a third-party mirror of the docs shows older values; the docs.cow.fi page is the source). The onboarding page also lists Optimism and Sepolia as endpoint networks, and the local-testing page lists Lens; none has a published settlement deadline, so they are not in the table until one exists.
Observed budgets are reconstructed from engine logs as deadline − start of solve, a slight under-estimate of the wire budget; the wire value replaces them in v0.1 once the engine logs it per request. An independent upper bound corroborates them: the archived bodies’ deadline minus the auction-start block timestamp is p50 5.89 s on Arbitrum, 5.36 s on Base and 3.53 s on BNB over the §9 windows, about a second above the engine-side figure on each chain. The difference is the driver’s send offset. Quote-lane budgets are a separate regime (≈ 2.82 s on all observed chains) and are out of scope for v0.
“No own-data observations yet” is a request: a solver on any of those chains who runs the tool with --solve-timeout set to its observed budget, or who sends me the deadline − receipt distribution, calibrates that row.
7. What a verdict is, and is not
A verdict is a statement about an endpoint over a window: it answered, it answered in time, what it answered was valid under the protocol’s own rules, and how its surplus compared to what actually won. It is not an admission decision, not an EBBO judgement, not a ranking of solvers, and not a prediction of rewards.
8. Non-goals
- No verdicts on other solvers. This document covers kaisersolver only.
- No settlement-side rules. Settlement validity is the circuit-breaker validator’s; settlement success is consistency v2’s.
- No EBBO. Capture ratio compares to what won, not to a best-execution oracle.
- No attribution the data cannot support. Where a measurement cannot separate two causes, the report says so.
9. Worked example: kaisersolver, 2026-09-14 → 15
Tool 0.11.0 (PyPI, 2026-09-14). Engine builds 3a36e59e5 (Arbitrum, BNB) and e777ab394 (Base), boot-line git_sha, private repo, cited for provenance; the endpoint does not expose it over /solve. Full reports with every check line, exclusion list and archive manifest: docs/readiness/ in the tool repo, commit e7b3b73.
Windows (block range · attempted-auction span · attempted · rate): Arbitrum blocks 505098473 to 505365109 · 2026-09-14 14:05Z → 09-15 08:50Z (18.75 h) · 1,030 · 55/h. Base blocks 51300926 to 51336483 · 14 13:06Z → 15 08:51Z (19.75 h) · 2,098 · 106/h; the window starts after a submission-behaviour change on Base at 10:59Z on 09-14, so it does not straddle it. BNB blocks 121850336 to 122000731 · 14 14:02Z → 15 08:51Z (18.8 h) · 5,970 · 317/h. Median replay lag 0.6 / 0.7 / 1.1 min (Base p95 45.7 min from the first window’s backlog; steady state under a minute).
| chain | endpoint | attempted | answer | bid coverage | transport | deadline misses | latency p50 / p95 (s) vs budget | validity (fairness evaluated / filtered) | capture (own wins) | per-bid median vs winner | verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Arbitrum One | prod | 1,030 | 100.0 % | 47.7 % | 0 | 0 | 0.79 / 1.98 vs 4.84 (0.41×) | 99.8 % (491/491 · 13) | 36.4 % (24) | 0.979 | REVIEW |
| Base | prod | 2,098 | 100.0 % | 49.9 % | 0 | 0 | 0.76 / 2.36 vs 4.62 (0.51×) | 100.0 % (1,037/1,047 · 95) | 63.1 % (72) | 0.955 | REVIEW |
| BNB | not yet activated | 5,970 | 100.0 % | 23.9 % | 0 | 0 | 0.92 / 1.53 vs 2.35 (0.65×) | 99.9 % (1,426/1,427 · 1) | 0.0 %; 8.7 % excluding one valuation artefact (0) | 0.952 | REVIEW |
What passes everywhere. Over 9,098 attempted auctions the endpoints answered every request, missed no deadline, produced no transport error, and every bid auction but two carried at least one valid, fairness-surviving solution (1,443 + 3,485 + 3,142 UDCP checks, 0 violations). The reliability checks are clean on all three chains.
Why REVIEW on all three. No check fails; every REVIEW is a warn. Bid coverage: kaisersolver returns a solution on 47.7 % of Arbitrum auctions, 49.9 % of Base auctions and 23.9 % of BNB auctions; the guard warns below 50 %, and Base is two auctions short of it. Latency headroom: Base p95 is 0.51× budget and BNB 0.65×, both over the 0.5× line; Arbitrum at 0.41× passes. Capture: 36.4 % on Arbitrum, under the 50 % line; the per-bid column shows why. The median of our surplus over the winner’s on unflagged bids is 0.979 (Arbitrum), 0.955 (Base), 0.952 (BNB), so the gap is coverage of large auctions we do not enter, not price quality. Price plausibility: 8, 110 and 385 auctions where our claimed surplus exceeded ten times the decoded winner surplus. Those are reference-price problems in the tool’s valuation; they are disclosed, and they hold the verdict at REVIEW, as the guards are meant to.
BNB, specifically. This is the weakest row and it stays in. It was measured on a keyless public RPC: 4,246 of 150,396 blocks were not scanned and 10.3 % of settlements could not be matched to a body (tx_uid_mismatch 396, body_uid_mismatch 317), so the field is under-counted. One auction (25459284, a 2 USDC order) carries a reference price that values the received tokens at ≈ 154,000 BNB; decoded winner surplus on that single auction moves capture from 8.7 % to 0.0 %. Both are v0.1 items (keyed RPC; winner-surplus sanity check against referenceScore).
Exclusions. About 3 % of the field on Arbitrum (uid mismatches between transaction and body, wrapper settlements the decoder could not attribute), 1 % on Base (wrapper settlements), 10 % on BNB as above. Bodies archived: 1,030 / 2,098 / 5,970 (one per attempted auction) with a SHA-256 manifest.
Reproduce (one line per chain; --bodies-dir replays from the archive, so the input side is fixed):
cow-backtester --chain arbitrum-one --env prod --from-block 505098473 --to-block 505365109 \
--bodies-dir <archive> --solve-timeout 4.84 --min-evidence 500 \
--rpc-url <rpc> --solver-url <endpoint> --solver-name kaisersolver --readiness
# cow-backtester 0.11.0 · engine build 3a36e59e5
cow-backtester --chain base --env prod --from-block 51300926 --to-block 51336483 \
--bodies-dir <archive> --solve-timeout 4.62 --min-evidence 500 \
--rpc-url <rpc> --solver-url <endpoint> --solver-name kaisersolver --readiness
# cow-backtester 0.11.0 · engine build e777ab394
cow-backtester --chain bnb --env prod --from-block 121850336 --to-block 122000731 \
--bodies-dir <archive> --solve-timeout 2.35 --min-evidence 500 \
--rpc-url <rpc> --solver-url <endpoint> --solver-name kaisersolver --readiness
# cow-backtester 0.11.0 · engine build 3a36e59e5
Manifest SHA-256: 0c809a600e676f998a0a7bf550ef810b87889e72b6ccc1845851dcb00f2cfa89 (9,650 lines, 2026-09-15 09:36:16Z)
10. Open questions for the solver team
- When a solver moves from staging to production, or is considered for mainnet, are any thresholds applied to the shadow autopilot’s per-driver
wins/performance_rewards, or to the productionsolve{driver,result}series? If so, which? Then this proposal can reference them rather than duplicate them. - For a solver under the CoW DAO pool, would it be acceptable to share aggregated
used_solve_timeandsolutions{result}for that solver’s own endpoint over a stated window? That would let replay-measured latency and answer rate be cross-checked against the driver’s view, which is the cleanest test of the method. - Are the observed settle-lane budgets in §6 what the autopilot configuration predicts on those chains? If per-chain
solve-deadlinevalues can be stated, the profile table cites them instead of inferring them.
11. Next
v0.1 after comments: threshold calibration per chain from whatever observations come in; wire-deadline budgets replacing the log-reconstructed ones; capture on the protocol’s referenceScore basis alongside the decoded one, with a sanity check that retires the valuation guard; capture excluding own wins; quote-lane profile; the driver-side cross-check if question 2 is answered. The tool already computes the bid-quality term of consistency v2 counterfactually from competition records (--reward-ev); tying that to the readiness window is the natural bridge between this and the production metric. Starting in October I will publish a readiness report per chain, in this format, every month in docs/readiness/, so whatever the definitions become after this thread, there is a running record measured against them. Tool: GitHub - KaiserSolver/cow-backtester: Offline backtester, A/B harness, and counterfactual scorecard for CoW Protocol solvers — replay recent auctions against your own /solve endpoint · GitHub . Reports: docs/readiness/ in the same repo.