Six agent-based engines, validated on six axes: cross-implementation replication, calibration against published targets, Monte Carlo dispersion, ensemble convergence, one-at-a-time sensitivity, and the numerical limits imposed by deterministic chaos. Every number on this page is computed by scripts/validate.ts and read from a generated artefact — none of it is written by hand.
Engines validated
6
Calibration checks
11 / 11
Bit-exact field comparisons
97.9%
Independent Monte Carlo runs
520
Each engine exists twice: once in TypeScript and once, independently transcribed, in Python. The two share nothing but a seeded Mulberry32 stream. Agreement across every reported field is strong evidence that both implement the specification rather than the same mistake.
| Engine | Fields compared | Bit-exact | Max relative error | Worst field |
|---|---|---|---|---|
| Offshore logistics | 51 | 100.0% | 0 | — |
| Aviation network | 39 | 100.0% | 0 | — |
| AUV swarm | 39 | 87.2% | 7.08e-3 | swarm_spread |
| Spill response | 36 | 100.0% | 0 | — |
| Licensing rounds | 33 | 100.0% | 2.64e-16 | cumulative_union_take |
| Supplier ecosystem | 39 | 100.0% | 0 | — |
Five of six engines reproduce bit-exactly across three seeds. The swarm engine does not, and section 6 explains why that is a property of the system rather than a defect in the port.
Where the literature or the public record gives a number, the model is held to it. These are acceptance tests, not illustrations: a failure here means the engine is not reproducing the behaviour it claims to.
Published probabilistic modelling puts beaching above 40-50% along Amapá and French Guiana during JFMA.
Target is the reported 12.5% probability of oil entering the Amazon river mouth in the JFMA period.
Reported transit to the coast is 10-20 days. The model reports the median beaching day, not the leading edge.
With the ITCZ north and NBC recirculation active, beaching should fall below 15%.
Reported as a residual ~2% in the JASO period.
Real winning profit-oil bids: Sépia 27.88%, Aram 29.96%, Atapu 31.68%, Libra 41.65%. The model should clear in that neighbourhood at mid-cycle Brent.
Disabling consortium formation removes the averaging of independent geological signals, so the winner's valuation sits further from the truth. Observed: 1.52% with consortia, 3.85% without.
A hotter market pushes every entrant closer to its own ceiling, so the selection on optimism sharpens. Observed: 1.52% at Brent 72 with 8 IOCs, 5.77% at Brent 95 with 11.
The random-geometric-graph connectivity threshold r_c = L·sqrt(ln N/(πN)) should land near a quarter of the field's characteristic scale for a dozen vehicles, matching the reported 25% fragmentation tipping point.
The effective number of stable sub-populations should peak inside the reported 200-500 m band, with degradation below and collapse above.
A floating forward stock point should convert most long base-to-field legs into short shuttles. This is the mechanism the whole hub business case rests on.
Every headline KPI, across independent seeds, with a 95% interval on the mean. The coefficient of variation is the number to read: where it is large, a single run tells you almost nothing and any conclusion needs an ensemble behind it.
How the ensemble mean of each engine's primary KPI moves as seeds are added. Where the mean is still drifting at n = 5, five runs is not a result — the auction engine is the clearest example.
| Engine | Metric | n = 5 | n = 10 | n = 20 | n = 40 | n = 80 | Drift |
|---|---|---|---|---|---|---|---|
| Offshore logistics | cost_per_well_day | 43,983 | 36,363 | 36,548 | 38,587 | 36,772 | 16.4% |
| Aviation network | avg_pax_delay_h | 4.23 | 3.88 | 4.65 | 4.63 | 4.49 | 6.1% |
| AUV swarm | coverage_pct | 23.33 | 30.00 | 24.17 | 25.00 | 25.00 | 7.1% |
| Spill response | oil_beached_pct | 45.80 | 45.18 | 44.58 | 44.50 | 44.50 | 2.8% |
| Licensing rounds | winning_profit_oil | 42.73 | 42.02 | 41.85 | 40.34 | 40.31 | 5.7% |
| Supplier ecosystem | lc_compliance_pct | 58.91 | 59.06 | 59.29 | 59.59 | 59.74 | 1.4% |
Drift is the change in the ensemble mean between the smallest and largest ensemble. Anything above about 10% is a warning that the headline number for that engine should be quoted with an interval, not as a point.
Elasticity of a target metric to a proportional change in one parameter, each measured on its own ensemble. This is what separates the parameters that drive an engine from the ones that decorate it.
| Engine | Parameter | Target metric | Δ | Base | Perturbed | Elasticity |
|---|---|---|---|---|---|---|
| Offshore logistics | PSV fleet size | cost_per_well_day | 20% | 36,377 | 31,583 | -0.66 |
| Offshore logistics | Weather severity | cost_per_well_day | 20% | 36,377 | 125,239 | 12.21 |
| Offshore logistics | Reorder trigger | stockout_hours | 20% | 469.04 | 420.92 | -0.51 |
| Offshore logistics | Deferred production cost | cost_per_well_day | 20% | 36,377 | 41,262 | 0.67 |
| Aviation network | SBMI fuel index | share_sbmi | -10% | 25.19 | 27.28 | -0.83 |
| Aviation network | Disruption rate | flights_cancelled | 50% | 6.33 | 16.33 | 3.16 |
| AUV swarm | Acoustic range | niching_index | 50% | 1.65 | 1.58 | -0.09 |
| AUV swarm | Sensor noise floor | coverage_pct | 50% | 28.33 | 18.33 | -0.71 |
| Spill response | Response delay | oil_recovered_pct | -50% | 0.56 | 0.75 | -0.70 |
| Spill response | Distance offshore | oil_beached_pct | 30% | 44.15 | 38.20 | -0.45 |
| Licensing rounds | Signal dispersion | winners_curse_pct | 50% | -0.91 | 8.56 | base 0 |
| Licensing rounds | Number of bidders | winning_profit_oil | 40% | 40.08 | 44.29 | 0.26 |
| Supplier ecosystem | LC requirement | lc_compliance_pct | 30% | 58.90 | 64.42 | 0.31 |
| Supplier ecosystem | LC requirement → penalties | penalty_usd_m | 30% | 0.00 | 47.23 | base 0 |
| Supplier ecosystem | Supplier base size | avg_lead_time_m | -30% | 1.10 | 6.43 | -16.08 |
| Supplier ecosystem | PEDEFOR investment | project_delay_m | 100% | 1.18 | 0.59 | base 0 |
Three results are worth reading closely. Weather severity has an elasticity above 12 on logistics cost per well-day, because a worse sea state pushes the fleet past the point where it can hold cover and the deferred-production term takes over — the response is strongly non-linear, and a linear planning model will not see it.
Acoustic range shows an elasticity near zero on the swarm's niching index. That is not insensitivity: the response is an inverted U, and the default sits near its peak, so a one-sided perturbation reads flat. The range sweep below is the honest view.
The local-content requirement has no elasticity on penalties because the baseline penalty is exactly zero. A policy that moves an outcome from nothing to something is precisely the interesting case, so it is flagged rather than silently reported as no effect.
The swarm engine's central result, swept rather than perturbed. Three regimes are visible, and the middle one is the only one in which a decentralised multi-target inspection campaign works.
| Acoustic range | Target coverage | Niching index | Regime |
|---|---|---|---|
| 50 m | 22.9% | 1.20 | Fragmented — vehicles hold station |
| 100 m | 29.2% | 1.62 | Fragmented — vehicles hold station |
| 150 m | 27.1% | 1.35 | Fragmented — vehicles hold station |
| 200 m | 31.2% | 1.77 | Niching — stable sub-populations |
| 250 m | 31.2% | 2.02 | Niching — stable sub-populations |
| 300 m | 37.5% | 2.01 | Niching — stable sub-populations |
| 350 m | 31.2% | 1.84 | Niching — stable sub-populations |
| 400 m | 33.3% | 2.05 | Niching — stable sub-populations |
| 450 m | 31.2% | 1.66 | Niching — stable sub-populations |
| 500 m | 31.2% | 1.62 | Niching — stable sub-populations |
| 600 m | 29.2% | 1.57 | Over-coupled — global convergence |
| 700 m | 20.8% | 1.11 | Over-coupled — global convergence |
| 800 m | 18.7% | 1.05 | Over-coupled — global convergence |
The swarm engine is the one place where the TypeScript and Python implementations diverge. Measuring how the divergence grows with horizon settles whether it is a porting defect or a property of the dynamics.
| Mission steps | Max relative divergence | Field |
|---|---|---|
| 5 | 0 | — |
| 10 | 0 | — |
| 20 | 0 | — |
| 40 | 0 | — |
| 80 | 1.83e-14 | swarm_spread |
| 160 | 1.50e-9 | swarm_spread |
| 320 | 1.03e-3 | swarm_spread |
| 400 | 2.34e-3 | swarm_spread |
Finding
The two implementations agree exactly for the first forty steps, then diverge at a rate of roughly 0.0799 per step — an e-folding time of about 13 steps. That is the signature of a positive Lyapunov exponent, not of a transcription error: the PRNG streams are identical, so the seed of the divergence can only be a last-place difference in how V8 and CPython evaluate the logarithm, cosine and exponential inside the Box–Muller draw and the intensity field. Exponential amplification does the rest.
Two consequences worth stating to anyone relying on this engine. First, no cross-language reimplementation of a chaotic swarm can be bit-exact past a few hundred steps, and demanding it would be a category error. Second, and more usefully, individual vehicle trajectories in this regime are not reproducible in any meaningful engineering sense — only the aggregate statistics are. Coverage, targets held, network fragmentation and the niching index all still agree exactly between the two implementations at every horizon tested. Those are the quantities to specify a mission against; a particular vehicle's track is not.
Regenerate with `npx tsx scripts/validate.ts`. Every figure below is computed, not authored.