This is the canonical spec route99 runs on: the code implements these rules, step by step. Every number on the rankings page is produced by the six rules below, in this order. Read it top to bottom and you can reconstruct any figure by hand.
Two kinds of input go in: measured facts, produced by Enso Shield and never altered by us, and your setup (the controls you set). Everything else on the page is derived from these six rules. DFC Research re-grades and re-plots; it does not change a measured value.
Integrations quote several aggregators per swap and take the best result, so every grade treats an aggregator as one member of that set. That is why no-quote weighs less as your set grows (why multi-aggregator ↗).
Every rule and sub-rule is directly linkable. Use the beside any heading to copy its address, or cite it by number: § 2.3, § 3.6.
Measured facts come from Enso Shield unchanged; your setup is everything else; nothing here re-writes a measured value.
Six figures per aggregator, from Enso Shield's 7-day window retrieved 18 Jul 2026.
InputsCollected by quoting every aggregator on the identical swap and re-simulating against live chain state:
The gap used throughout this page is Shield's immediate, same-block re-simulation — the first of the five points Shield measures (t=0, then +12s, +24s, +36s, +48s).
It is therefore drift at zero elapsed time, the most conservative of the five, and the reason § 2.8 models a separate survival term for real execution delay rather than reusing this number as-is. Pinned-block figures are not used.
Underquoting is carried in the dataset and never costs an aggregator points.
Underquote rate is carried in the dataset and deliberately not scored: a fill better than quoted costs the integration nothing, so grading it would punish a provider for quoting conservatively.
Coverage and pricing come from provider docs; the controls you set are the rest.
Coverage and API pricing come from each provider's own docs, linked in every row.
Your setup is the rest: chain, Simulate on/off, latency budget, monthly volume, the Standard/Negotiated pricing lens, and the route you select.
Measured facts are never changed, only re-graded for your setup.
One resolver answers “how many chains does this provider support” for every surface on the site, in a fixed order of authority.
A chain count is resolved once, in this order, and every surface reads the same answer: live registry, then declared, then enumerated.
A declared count is a claim, not a measurement. Where a provider claims more chains than this dataset can name, both figures are shown — the claim, and how many of them are named here. Nordstern declares 100+; this dataset names 57.
Chain counts are informational only and never scored. Needing 60 chains and needing 2 are different jobs, so no criterion in Rule 3 reads this number, and the coverage tie-break in § 2.11 uses measured chains rather than claimed ones.
The rankings column, its tooltip, the radar's coverage axis, the coverage bars, the chain finder and the data appendix all call the same resolver, and a live refresh repaints all six together. Previously the refresh updated the coverage chart alone, so the other surfaces could show a stale figure until a control was touched — the page disagreed with itself. route99SelfTest() now asserts the precedence, the fallback and that a live reading leaves the declared snapshot intact.
A multi-chain selection ranks on the unweighted mean of the selected chains' measured rows; the gauge floors on the weakest selected chain.
When more than one chain is selected, each aggregator's inputs are the unweighted mean of its measured rows on the selected chains where it is measured. An aggregator measured on any selected chain is included; the reliability gauge's floor (Rule 2) is then taken over the selected chains only, so a member absent from one of your picks still binds the gauge there.
This is the same pooled-convenience construction — and carries the same caveat — as the All chains view: Enso Shield's fairness guarantee holds within one chain, so a pooled row summarises different route populations and is an overview, not a like-for-like benchmark. Per-chain views remain the comparable read. All chains itself stays Enso's own pooled aggregate, not our average.
The headline % is the chance your route returns at least k usable quotes on a request, computed from each member's own measured inputs.
A quote counts only if it arrives, arrives in time, survives simulation, and survives its own overquote drift.
The headline % is the chance your route (the set of aggregators you selected) returns at least k usable quotes on a request. An aggregator is usable when it returns a quote, arrives inside your latency budget, when Simulate is on survives simulation, and survives its own overquote drift.
The target defaults to 99.9% and is switchable; an aggregator's green band is the per-aggregator failure rate at which N independent aggregators still clear it.
The target defaults to 99.9% (0.1% failure) and is switchable — 99%, 99.9%, or 99.99% — from the Setup bar's Target control; every number below illustrates the 99.9% default.
An aggregator's “green” band for redundancy-covered failures is the per-aggregator rate at which N independent aggregators still clear the active target — which depends on both N and k, capped at 10%, beyond which correlated failures dominate.
At k=1 that is the closed form target^(1/N); at k≥2 there is no closed form, so the band is solved numerically against the same Poisson-binomial in § 2.4.
The band must use the same k as the selected goal; at the price goal's defaults a k=1 band is 5.4× too lenient.
The reliability threshold must use the same k as the selected ranking goal.
Route size N | 3 |
|---|---|
Usable quotes required k | 2 |
k = 1 band | 10% |
correct k = 2 band | 1.84% |
| difference | 5.4× more lenient |
A single number for both goals was wrong by more than a rounding step. Under Optimize for price the route needs two usable quotes, not one, because price cannot be compared with only one result.
At its default N=3 the k=1 band is 5.4× more lenient than the k=2 truth — it painted green on aggregators this profile's own arithmetic disqualifies.
k.Aggregators are treated as independent and the route's reliability is the exact probability of ≥ k usable quotes.
With Simulate on, aggregators are treated as independent and the route's reliability is the exact probability of ≥ k usable quotes across them (a Poisson-binomial sum).
k = 2 for “Optimize for price” (you need two quotes to compare); k = 1 otherwise.
With simulation enabled, adding another aggregator cannot reduce route reliability: a weak aggregator simply contributes fewer usable quotes, and the P(≥k) sum is monotone.
The route inherits the average revert and overquote-drift risk of the set, so adding an aggregator can lower the expectation.
With Simulate off, adding an aggregator can lower reliability: one high-revert or heavy-overquote aggregator drags the expectation down for every swap it wins on price. The gauge names the offending aggregator when this happens.
| Property | Simulation On | Simulation Off |
|---|---|---|
| Reliability input | Simulated usable quotes | Historical average failure assumptions |
| Effect of adding an aggregator | Cannot reduce reliability | Can reduce expected reliability |
| Weak aggregator impact | Contributes fewer usable quotes | Can increase average revert or drift risk |
| Revert treatment | Quote becomes unusable | Average revert rate affects expectation |
| Overquote treatment | Simulation detects unusable execution | Average overquote drift affects expectation |
| Mathematical behavior | P(≥k) is monotone | Set averages may worsen |
| UI behavior | Shows route reliability | Identifies the aggregator causing degradation |
A two-parameter log-normal is fitted to the two published quantiles, with two guards and one acknowledged limit.
σ floor of 0.05 stops a degenerate p95 ≈ med from collapsing the distribution to a point mass. On current data the smallest real σ is 0.139, so it never binds — it is a guard against future rows, not an active adjustment.p95 = max(p95, med × 1.02) guard handles the arithmetically impossible case of a reported p95 at or below the median (rounding, or a thin sample), which would otherwise give a negative or zero σ and an inTime of 0 or 1 with no middle.Both are stated here because a reader reconstructing inTime by hand from the two published quantiles would otherwise not reproduce our number.
A two-point fit has no slack in it. A log-normal has two free parameters, and this page has exactly two published quantiles — median and p95 — to fit them from. That pins the curve with zero degrees of freedom.
There is no third point to check the shape against, so a real latency distribution with a heavier tail than log-normal (a slow RPC, a congested block) will read as more “in time” at your budget than it actually is, especially past p95 where the fit is pure extrapolation, not data.
This is a property of fitting any two-parameter curve to two quantiles, not a bug in this implementation, and it is the reason the guard above exists rather than a shape check: with only two points there is nothing left to check the shape against.
Treat inTime as the best estimate the published data supports, most reliable near the median and least reliable in the tail your budget actually depends on.
Shield measures immediate re-simulation failure, which is not the same event as a revert after your own execution delay.
Enso Shield measures failure of an immediate re-simulation against the latest block, with RPC-error simulations excluded.
That is not the same event as a revert after your own execution delay — wallet confirmation, gas movement, settlement windows and mempool exposure all sit between a quote and a landed trade.
With Simulate on the distinction barely matters, because the check happens pre-trade either way.
With Simulate off this page uses that number as a modelled proxy for revert risk, and the oqSurv term (§ 2.8) exists precisely because the proxy is known to understate delayed execution.
Read the sim-off gauge as the best available estimate, not as a measured revert rate.
An inflated quote eats the slippage cushion before the market moves, so overquote drift enters usable as a bounded probability.
A slippage tolerance is normally sized off the quoted output, not the true one. An inflated quote eats into that cushion before any real market movement happens, so the same drift that costs score points also raises the odds the trade reverts instead of executing — this is exactly the mechanism Enso Shield's own quote-decay research describes.
It matters more than a same-block check suggests: Shield's benchmark quotes and simulates essentially instantly (same block), but real routes often don't execute that fast — a wallet-confirmation click or a bridge settlement window can run well past the seconds-scale delay where Shield's own data shows overquote gaps continuing to widen.
So the measured sim rate is closest to revert risk at zero elapsed time, not at realistic execution time, and oqSurv is this app's stand-in for that gap: it starts from the same oqCost priced in § 3.6, gated gentler under Simulate (a pre-commit check catches most drift before it's acted on) and steeper without it (the quote is committed, cushion and all).
Unlike the score weight in § 3.6, this term is not gated by profile: reliability is a physical estimate of what actually happens, not a preference, so it applies the same way under every goal.
A quote that was never inflated cannot fail because it was inflated, so at most oq% of an aggregator's quotes can be lost this way: oqSurv is floored at 1 - oq/100.
That bound is what separates the two uses of oqCost. This one is a probability and has to stay physically plausible; the § 3.6 grade is a preference scale, free to be as steep as it needs to be to rank aggregators apart.
For a while both read off one shared number, which sounded tidy but forced one of them to be wrong: a grade steep enough to separate the field implied more quotes failing than were ever overquoted, and a probability gentle enough to be plausible left the grade unable to punish anything. They now share the input and calibrate separately.
The gauge models N aggregators quoting in parallel on every request, which is only true where all N are live.
Coverage is part of realized reliability, not a separate cosmetic score.
The score computed from all measured rows across chains. On the pooled All chains view the members are often not all live: Barter is measured on Ethereum alone while claiming 8 chains, OogaBooga on HyperEVM alone.
Reliability recomputed from each chain's own measured rows, restricted to the route members present on that chain.
Every chain served by a route must have at least k live members.
When fewer than k route members are live on a chain, price comparison is impossible and P(≥k) = 0 — exactly zero rather than merely low. The route does not satisfy the goal on that chain.
On “All chains” the headline is the weakest chain the route serves, not the pooled average.
Per-chain reliability is recomputed from each chain's own measured rows, restricted to the route members present, and on All chains the headline is that realised floor — the weakest chain the route serves.
The full per-chain breakdown is on hover over the number. The pooled average appears in “How it's computed” whenever the two differ.
Conclusion. A route can show 99.92% pooled and realise 91.2% on a chain where two of its four members do not exist. The pooled score can materially overstate the reliability actually realized on the weakest chain — and with fewer than k members present, price cannot be compared there at all.
Coverage enters assembly in two bounded ways, and never licenses breaking a goal's own filter.
Auto-pick guarantees that every chain the route serves has at least k members live on it, because below that price cannot be compared there at all and P(≥k) is exactly zero rather than merely low.
Short chains are topped up with the highest-scoring eligible aggregator measured there, never one that would break the goal's own filter.
Where two aggregators sit within 0.15 points, the one measured on more chains wins, since on the pooled view an absent member contributes nothing.
Coverage never licenses breaking a goal's own filter. Where a goal's eligible pool genuinely cannot reach a chain within the 5-member cap, the route stays as it is and the headline shows the shortfall honestly — only five aggregators are measured on all six chains, so this binds exactly where the field is thinnest.
Six criteria in three pairs, each yielding a points-kept fraction on a fixed absolute scale, averaged under the active goal's weights.
Every criterion yields a points-kept fraction gc in [0,1]; the score is their weighted average, times ten.
Each aggregator is graded on six criteria in three pairs: reliability (no-quote, sim-failure), price (all-in cost, overquote), latency (median, p95).
No-quote and sim-failure grade on a raw scale, then pass through a difficulty curve so perfection stays at 10 but small imperfections still cost real points.
reliability pair · gc = 1 − √(1 − g)
These zero-points aren't the value that “should” mean total failure — they're calibrated with roughly the same headroom above the worst aggregator Enso Shield currently measures that Cost's own zero-point already carries (38 bps against a 27.06 bps worst case, 1.40×).
Applied here: the worst measured no-quote rate is 59.52% and the worst sim-failure is 20.84%, so 85% and 30% give each about that same margin (1.43× and 1.44×).
Sim-off stays deliberately twice as strict as sim-on, because an uncaught failure reverts on-chain rather than being caught pre-trade.
An earlier calibration used 10% and 10/5%, tight enough that an aggregator with 95% real-world availability graded as if it barely worked at all. A single extreme outlier (an aggregator failing outright, most of the time) still lands near the bottom; an aggregator that's merely imperfect no longer reads as broken.
This is the correction that produced the numbers above, and it is worth stating plainly because the previous ones were justified against the wrong population.
The constants used to be sized on the all-chain view (~42% no-quote, ~13% sim-failure), but grading runs on per-chain rows, where the real worst cases are 59.52% (Eisen on BSC) and 20.84% (Velora on Base). That left 1.008× and 0.96× headroom instead of 1.4×.
0.96× means the scale had run out: Velora on Base graded exactly 0.000 on sim-failure, so the criterion could no longer tell 21% apart from 45%. A zero-point a live row can reach is not a zero-point.
Every constant is sized against the worst value any gradeable row can take, and the check is re-run against per-chain data on each refresh.
A deadline adds a compliance question without removing the speed question, so a budget re-aims the pair instead of replacing half of it.
The latency pair is graded two ways, because setting a budget changes the question. Every goal but Low latency measures against a 1,200 ms budget by default (Low latency uses a stricter 400 ms); Off is the deliberate opt-out, for reading raw speed with no deadline attached.
A deadline adds a compliance question without removing the speed question, because 400 ms can clear a 400 ms budget and still feel slow.
The tail is never dropped: p95 is what sets σ, so a heavy tail directly lowers the in-time share.
In the weights table (§ 3.7), Median and p95 are the primary and secondary latency criteria; with a budget set they read as compliance and speed.
Graded on an anchored scale: below L each halving earns a point, above L each extra L/2 costs one.
“How fast is it?”, no deadline. L is the half-credit point (800 ms typical, 2,200 ms tail).
budget = Off only · log below L, linear above
On this field the anchored scale has an unreachable top, and a budget is the only lever with real separation in it.
A category grade of 8.0 needs roughly a 100 ms median and a 275 ms p95, and 9.0 needs 50 ms/137 ms. Nothing measured here is that fast.
The whole 15-aggregator field lands in 3.4–8.0 with the top two points of the scale dead — and since green starts at 8.0, only the single fastest aggregator ever cleared it, rendering the second-fastest in the field amber. That reads as “this aggregator is mediocre at latency,” which is false.
Re-weighting cannot repair it: the median and p95 grades are collinear on this data, so however the pair is weighted the latency category moves by at most ~0.13 points. A budget is the only lever with real separation in it.
1,200 ms rather than 800 because at 800 the curve cliffs just past the anchor — a 1,025 ms and a 1,133 ms aggregator both collapse to 2.4, overstating the difference between 800 ms and 1,000 ms — where 1,200 grades them 6.6 and 5.1: clearly worse, not annihilated.
And not 400 ms everywhere, because 400 is genuinely selective (a 576 ms median falls to 2.8): right for the goal whose job is finding fast aggregators, wrong as a global default.
A budget is a deadline, not a ratio. The anchored curve grades ratios of milliseconds and is deliberately gentle, which caps how far apart two aggregators can land: 93 ms and 628 ms are 2.75 doublings apart, so no one-point-per-doubling curve can separate them by more than 2.75 points, whatever the anchor.
Against a 400 ms deadline that is the wrong answer, because the first aggregator beats it on 98% of quotes and the second on 7%. Grading compliance by the in-time share puts the steepness exactly at the deadline, where the decision turns, while the speed criterion keeps rewarding genuinely fast aggregators inside it.
The difficulty curve assumes perfection is attainable: 0% reverts is, and several aggregators reach it. A 0 ms quote is not. Grading latency from zero crushed the whole field, so the scale anchors at a realistic half-credit point instead, keeping a reachable top end while staying strict about real slowness.
Because that scale's 10 is unreachable, the latency column bands against the achievable range when the budget is Off (green from 7.5, amber from 5) rather than the 8.0/5.5 cut every other column uses. With a budget set the scale is reachable and the ordinary cut applies.
All-in cost and overquote carry their own reciprocal gc with no square-root curve, so cost is never double-penalised.
The two price criteria carry their own reciprocal gc (no square-root curve), so cost is never double-penalised. Fees come from the cost model (Rule 4); overquote is costed as expected loss from inflated quotes.
points kept vs all-in cost · gc = 1 / (1 + bps/2)
Overquote is weighted only under Optimize for price — every other profile carries it at w=0, so an aggregator that overquotes isn't scored down by a goal that isn't price.
The criterion is still computed and shown in every breakdown; outside price it just costs zero score points (rendered as “–”, not “0”).
The reliability gauge (Rule 2) starts from the same oqCost but calibrates it separately, as a bounded probability rather than a preference scale; that half is also not profile-gated, since reliability estimates what happens rather than what you prefer.
The two look like the same kind of number and are not. All-in cost is a realized bps figure. oqCost is an expected one: a frequency multiplied by a magnitude, which collapses toward zero even when both inputs are alarming.
Grading an expected cost on a realized-cost scale was a real defect, not a matter of taste: with a divisor of 8 the curve emptied out around 152 bps, so the criterion Optimize for price advertises as one of its two leading concerns could not move a ranking — the field's worst overquoter gave up half a point out of ten, and seven of thirteen aggregators lost nothing at all.
The divisor is sized on per-chain rows for the same reason the reliability pair is: the worst measured oqCost is 10.03 bps (Odos on Ethereum, 63.4% of quotes inflated), not the 2.09 bps of the all-chain view.
A divisor of 0.5 empties the curve at 9.5 bps — below that worst case, leaving the field's heaviest overquoter past the end of the scale. At 0.75 it empties near 14.25 bps: 1.42× headroom, matching Cost's 1.40× and the reliability pair's 1.43×.
The active profile sets the weights: the leading pair is emphasised, everything else stays at 1, overquote is 0 outside Optimize for price, and Balanced folds overquote's zeroed weight into cost so no pair is left underweighted.
The active profile sets the weights. The leading pair is emphasised; everything else stays at 1 — except overquote, which is 0 everywhere except Optimize for price. In Balanced, where no pair leads, cost carries overquote's zeroed weight (×2) instead of just dropping it, so the price pair still totals the same as reliability and latency:
| Profile | No-quote | Sim-fail | Cost | Overquote | Median | p95 |
|---|---|---|---|---|---|---|
| Balanced | 1 | 1 | 2 | 0 | 1 | 1 |
| Optimize for price | 1 | 1 | 3 | 2 | 1 | 1 |
| Low revert rate | 1 | 3 | 1 | 0 | 1 | 1 |
| Low latency | 1 | 1 | 1 | 0 | 3 | 2 |
Elsewhere its weight is 0: the criterion is still computed and its bps cost still shows in every score breakdown, but it keeps zero points from every other goal's ranking.
Low revert rate and Low latency already say as much in their own descriptions (§ 5.4); this is that same intent enforced in the weights, not just the copy.
Balanced is the one place this needs a second step: dropping overquote's weight there without compensating would leave the price pair (cost + overquote) totalling half the weight of reliability (no-quote + sim-failure) and latency (median + p95) — quietly contradicting "no single pair dominates" even though every other number in the row still read as 1. Cost absorbs the difference so all three pairs total the same weight.
With more than one aggregator in the route, No-quote weight halves (×0.5) on every profile except Balanced: a missing quote from one aggregator is simply covered by another.
Simple view rolls the six into three categories: Reliability = no-quote + sim-failure, Cost = all-in cost + overquote, Latency = median + p95, each the weighted average × 10 of its members.
Every cap, divisor and anchor above is a fixed constant, never a percentile, a rank, or a curve fitted to the current dataset.
The worst value in the measured field is used only as a sanity check when choosing a constant, never as an input while grading: nothing here is a percentile, a rank, or a curve fitted to the current dataset.
A score means the same thing across every data refresh, so an aggregator improving its own numbers moves its own score and nobody else's. And removing an aggregator, even the worst one on a criterion, leaves every other score bit-identical: drop the field's heaviest overquoter and no remaining grade shifts by so much as a rounding step.
The alternative, grading on a curve against the current field, would mean an aggregator's score changed when a competitor changed, and week-to-week comparisons would be meaningless.
The radar plots raw measured dimensions on linear axes and deliberately does not use this rule's curves.
It plots the raw measured dimensions on plain linear axes, sharing this rule's endpoints (no-quote empties at 85%, sim-failure at 30%) but none of its transfer functions.
That is on purpose: the curves above are each shaped differently on purpose (square-root for reliability, reciprocal for price, anchored log for latency) because each encodes its own judgment about what counts as bad, which is right for collapsing six things into one number and wrong for a shape you read axis-against-axis.
The radar also splits overquote into rate and magnitude across two axes where the score multiplies them into one oqCost, and blends median with p95 where the score keeps them apart.
Use the radar to compare an aggregator's shape, the table to compare its score.
One thing the radar does borrow: its radius is the square root of each value, so the area a polygon encloses is proportional to the value rather than to its square. Plotting radius directly would overstate every comparison, since area is what the eye actually measures.
API cost is expressed as bps at your monthly volume, from each provider's published plan, with three distinct provenance states.
A floor, a ceiling, and where offered a subscription tier, interpolated across volume.
API cost is expressed as bps at your monthly volume, from each provider's published plan: a floor, a ceiling, and (where offered) a subscription or custom/enterprise tier.
Negotiated amortises published paid plans; only a genuinely unpriced custom tier shows as 0 bps.
In Standard you see the published/free-tier fee at your volume.
In Negotiated, a provider that publishes a paid plan is priced at its amortised rate — bpsMin + subMax/V×1e4, the formula already above — because that is what the plan costs per swap at your volume, and it is arithmetic rather than assumption.
Only a provider advertising a custom/enterprise tier with no published rate at all is shown as 0 bps under Negotiated, since there is no number to compute instead. That 0 is contingent, so those scores carry a ! whose tooltip names what it takes. Genuinely-free APIs are already 0 with no marker.
Any score resting below a provider's published ceiling carries the !, in either lens.
A provider with no published low end (no subMax plan to amortise) is still interpolated toward its disclosed range's bottom as volume rises, because that is what the range is for.
But nothing published confirms that low end is reachable without negotiating — it is the same unconfirmed floor Negotiated shows explicitly at 0, just reached gradually instead of in one step.
The ! is not gated to the Negotiated lens: any score resting below that provider's published ceiling (bpsMax, the ordinary Standard-tier number) carries it, in either lens. Only the ceiling itself, and rates from providers with a real amortisable plan, are unflagged.
Pricing every negotiable provider at 0 bps penalised disclosure and was the largest single ranking lever on the page.
Pricing every negotiable provider at 0 bps threw away a number this model already computes, and it made a provider publishing a $799/mo plan score identically to one publishing nothing — the lens penalised disclosure.
It was also the largest single ranking lever here: at $10M/mo it moved one aggregator seven places on a 2.94-point swing, entirely on an assumption.
Worse, picking a volume above $1M used to flip the lens on automatically, so a user who only chose a volume was shown contingent rates they never opted into.
Volume now sets volume. Negotiated is a deliberate click, and it now changes the fewest numbers it honestly can.
Two situations look alike and are graded differently: a concealed rate is scored at the field's worst, a genuinely absent fee is read as 0 and flagged.
If a provider charges a fee whose rate it does not publish, that silence conceals a real number, and scoring it as free would make “tell nobody your prices” the most profitable thing a provider could do here.
Those are scored at the worst published effective rate in the field at your volume, flagged in the breakdown.
If instead no fee exists to find — nothing in the docs, no fee field, fee recipient or fee line in returned quotes — that absence is itself weak evidence, and it points at 0. Those are read as 0 bps and marked “unverified”.
Imputing a punitive rate there would publish a number we have positive reason to believe is false, about a named team; the flag carries the doubt instead of the score.
An affirmative “we charge nothing” is a third thing again — a disclosure — and counts as published.
Inside the score, fees are capped at 60 bps for the curve, and overquote is costed separately (Rule 3).
If a rate here is wrong, telling us is the fix.
A goal sets the weights, a target k, a starting set size and default controls, then walks the ranked list until the route clears your reliability target.
Ranking uses the same profile-weighted score the table sorts on, evaluated at the size the route will actually be.
Choosing a profile sets the weights (Rule 3), a reliability target k, a starting set size, and default controls, then assembles your route.
Ranking uses the same profile-weighted score the table sorts on, evaluated at the size the route will actually be, so the checked aggregators are the top rows of the pool this goal is drawing from (no “why isn't #2 selected?”).
The target only decides how many to add: walk down the ranked list, adding aggregators until the route clears your reliability target, or the next aggregator is too weak to move it.
One thing legitimately moves a checked row away from the top of the visible table, and it is deliberate: the target walk, which stops as soon as your reliability target is cleared and so may not reach a row that scores well but adds nothing.
Drift is priced, never gated: it costs score points in proportion to its size.
No goal applies a hard eligibility gate. Drift is priced, never gated: it costs score points in proportion to its size, and an aggregator that drifts badly ranks last on its own.
Optimize for price used to drop anything with |drift| above 0.1 bps.
The visible symptom was an aggregator showing the field's second-highest score and the word FILTERED in the same row.
Scoring takes the candidate size explicitly and the target walk iterates to a fixed point.
What is not deliberate, and has been fixed, is ranking that depended on what was selected before the pick: the no-quote weight halves at more than one aggregator, so scoring during assembly used to read the previous route's size.
Scoring now takes the candidate size explicitly and the target walk iterates to a fixed point, so assembly and the table are always scored in one context.
Four goals, each with its own leading pair, target k, starting size and default controls.
Growth stops at a 3% usable floor and a 5-aggregator cap — practical limits, not derived ones.
usable probability is under 3% — it cannot meaningfully move P(≥k), so adding it is integration work for nothing.Both are practical limits rather than derived ones; they only bind when the target is out of reach, and the gauge's caption names the binding chain when it happens.
While the weakest chain the route serves is under target, the next pick is aimed at that chain and ranked by what it does there.
Growth is coverage-aimed: while the weakest chain the route serves is under target, the next pick is aimed at that chain — an addition absent from the binding chain cannot move the number the gauge shows.
The aimed pick is ranked by what it does to that chain, not by its overall score.
Those are different questions, and ranking the repair by score picked an aggregator that scored well overall but was weak precisely where the route was short, leaving the chain under target when the target was reachable.
A repair can pull in an aggregator this goal scores poorly. It is there for coverage, not as an endorsement — the route is quoted in parallel and the best result wins, so a costly member only ever wins when it is genuinely the best net quote, and the table still shows its real score.
If no aggregator can help the binding chain while the pooled target is met, growth stops rather than spending slots that change nothing.
The moment you hand-edit a row the set is yours, and no control re-picks over it again.
Auto-pick is a convenience, not a rule: it seeds the smallest set that reaches this goal's target and keeps that set fresh while the set is still auto-managed, so changing any control (volume, latency, simulation, pricing, chain) re-runs the pick under the new settings.
The moment you hand-edit a row, the set is yours and no control re-picks over it again: knobs re-grade your set, they never replace it. The goal stays active the whole time and its weights keep scoring every row, so hand-editing is not leaving the goal, it is evaluating a different set under it.
Restore auto-pick hands the set back to the goal.
What this model assumes, where it stops, who publishes it, and how to falsify it.
Correlated failures miss every aggregator at once and set a floor no redundancy clears.
The gauge treats aggregator failures as independent. Correlated failures (honeypots, transfer-tax tokens, genuinely illiquid routes) miss every aggregator at once and set a floor no redundancy clears.
Reaching sub-0.1% in practice is two jobs: enough fast, honest, independent quotes (modelled here) and pre-trade filtering of intrinsically unexecutable routes (not in this dataset).
The gauge stops at 99.999%+, because past the fourth digit the number is a statement about the independence assumption, not about the aggregators.
The gauge is capped at 99.999%+, and that is not a rounding choice. The arithmetic above will happily return 99.99999% for a set of four or more strong aggregators, and printing that as 100% would be the model asserting something it has no standing to assert.
§ 6.1 says correlated failures set a floor no redundancy clears, and that floor is not in this dataset. Past the fourth digit the number is a statement about the independence assumption, not about the aggregators. So the display stops there.
Read anything above 99.9% as “the modelled quote-supply problem is solved for this set; what remains is the correlated-failure and unexecutable-route problem, which is a different job.”
route99 is published by DFC Research and some measured teams are club members, which is exactly why every figure is Enso Shield's.
route99 is built by DFC Research, the research arm of the DeFi Founders Club; some measured teams are club members.
That is exactly why every figure comes unchanged from Enso Shield, every ranking follows these same published rules for every aggregator, and the source is linked throughout so any figure can be verified directly (Enso methodology ↗).
Enso is scored on the same six criteria and the same published rules as everyone else, with no special handling.
Enso is scored here on the same six criteria and the same published rules as everyone else, with no special handling: 11th of 15 under Balanced, 12th under Optimize for price, 7th under Low revert rate, 13th under Low latency (each at that goal's own default settings).
We do not adjust Enso's inputs, weights, or curve differently from any other row, and the numbers above are reproducible from the same public dataset every other score comes from.
Run route99SelfTest() in the console: it asserts every formula on this page against a hand-computable expected value.
“The code matches this page” is an assertion, and given who publishes route99 it should be cheap to falsify.
Open your browser console on this page and run route99SelfTest(). It asserts every formula above against a hand-computable expected value — Poisson-binomial at k=1 and k=2 with unequal inputs, the green band at both k, the log-normal inversion, latency-curve continuity and clamping, every zero-point and divisor, dollars-to-bps, both pricing lenses, and the display cap — then prints a pass/fail table.
If any row fails, this page and the code have drifted, and one of them is wrong.
route99SelfTest() proves it.Every assumption behind the cost and coverage columns, generated at page load from the same embedded dataset the rankings use.
Both tables are generated at page load from the same embedded dataset the rankings use — they cannot drift from what's shown there.
Every assumption behind the cost and coverage columns, in one place. Both tables are generated at page load from the same embedded dataset the rankings use — they cannot drift from what's shown there. This is about consistency, not freshness: the dataset itself is the dated snapshot named above, re-read the same way every time the page loads. Hover a rate for the full sourcing note.
If any figure is wrong, tell us ↗: corrections ship with the next data refresh, and disagreement with a published number is exactly the review this page exists to invite.
Effective bps per provider at each monthly volume, under both pricing lenses, with provenance flags.
Standard = the published or free-tier rate at that monthly volume, with paid plans amortised into bps where that is cheaper. Negotiated = genuinely unpriced custom/enterprise tiers at their contingent 0 (Rule 4).
LP/pool fees, bridging, gas and consumer-UI fees are excluded throughout. Effective-rate columns are in bps.
A checkmark grid of measured rows per chain, against each provider's own claimed network count.
✓ marks a chain with a measured Enso Shield row in this dataset — the checkmark grid makes gaps visible at a glance, which is exactly what names the shortfall in Optimize for price's Polygon coverage (§ 2.11).
“Claimed” is the provider's own network count linked to its source (live registry where available, otherwise official docs); hover it for source quality and how much sits beyond what's measured here, which we cannot independently verify.
For the interactive version — pick your chains, see who covers them — use the finder in the Network coverage section of the rankings page.
Even the best aggregators miss more than 1 in 1,000 swaps. Serious integrations quote several per swap and take the best result. We grade each one for that setup, against your goals, to your reliability target (99.9% default, switchable). Why multi-aggregator ↗
Reliability and latency: Enso Shield's 18 Jul 2026 weekly snapshot. Pricing: provider docs, same pull. Chain coverage: refreshed daily. Full sourcing →
Enso's measured values, unchanged: only the colour bands and score weights move with your setup. The score weighs six criteria in four categories (reliability, cost, overquote, latency); your profile sets which lead. In the goal profiles, redundancy eases no-quote once you run two or more aggregators. Without simulation, sim-failure is used as a modelled proxy for your on-chain revert rate (Rule 2). On All chains the rows pool every chain an aggregator is measured on, so they summarise different route populations — use it for a field-wide overview and a per-chain tab for a like-for-like comparison. Click rows to build your set.
| Aggregator | Score1–10 · weighted | Reliabilityno-quote · sim-fail | Latencymedian · p95 | Costfees · overquote | No-quoteavailability | Sim failurereverts | Median lat.speed | p95 lat.tail | Overquoteprice holds | Fill gapavg drift | Costbps | Chainssupported |
|---|
Six dimensions at once; further out is better on every axis. Hover an axis label for its exact definition. Selected aggregators in colour, the rest of the field faint behind them. These are the raw measured dimensions, not the graded ones: the table's 1–10 scores put the same inputs through per-criterion curves and profile weights, so use this to compare an aggregator's shape and the table to compare its score. Rings mark 25 / 50 / 75 / 100%, spaced so enclosed area tracks value rather than exaggerating it.
Your selected set, traced gate by gate. Thin lines: each aggregator draining as quotes fail to return, arrive late, or fail simulation. Bold line: the chance your goal's quote requirement is met — at least one usable quote, or at least two under Optimize for price, where you need a second to compare against. Redundancy is the gap between them.
A quote only counts if it arrives in time. One aggregator: its whole latency distribution has to fit your budget, tail included. Several: they cover each other's tails, so a slow p95 stops being fatal. Chips show the share arriving in time — modelled from median and p95 together (Rule 2), not from either threshold alone.
Two views of one failure: a quoted price that does not hold. Horizontal: how often quotes are inflated. Vertical: how much fills drift. Bottom-left is honest and tight.
A little drift is normal: prices move in the milliseconds between quote and fill, so a few hundredths of a bp reflects latency, not padding. It becomes a cost signal past the 0.05 bps noise band.
Quality is one axis; price is the other. This compares only the published API-fee layers - subscription plus platform fee in bps - at your volume; it excludes gas, bridge, LP and other costs outside the API itself. Undisclosed stays undisclosed, not assumed zero.
Quality is half the picture; an aggregator must also support your chains. There is no single registry. Counts come from live provider APIs where available, otherwise official docs, linked per row.
All measurements are produced by Enso Shield, an independent benchmark that quotes every aggregator on the identical swap and re-simulates against live chain state. DFC Research changes none of the numbers; we only re-grade and re-plot them for your setup.
Your route's headline % is the chance that at least k of your selected aggregators return a usable quote: one that arrives in time and survives simulation. Aggregators are treated as independent, and the target defaults to 99.9% — switchable to 99% or 99.99% from the Setup bar's Target control.
Six criteria in three pairs (reliability, price, latency), each graded 0–10 on published curves. Your profile sets the weights; the Standard/Negotiated lens sets how API fees are priced. Same rules for every aggregator.
route99 is built by DFC Research; some measured teams are DeFi Founders Club members, and Enso — the source of every reliability and latency figure — is itself scored here (11th of 15 Balanced, 7th Low revert, 12th price, 13th latency, at each goal's default settings). No number here is ours, every ranking follows the same published rules for every row, and every figure links back to its source.