AlphaIQ foresight fractal markAlphaIQ
v1.0
Data specification
Calibration and deployment
Six engines

Data requirements

What a production deployment inside Petrobras would need in order to calibrate, run and defend each engine — the field, its granularity, the likely system of record, why the model needs it, and whether a public or vendor proxy is good enough to stand in during a pilot.

None of the constants currently in the engines came from Petrobras data. They are industry-typical values, published figures and modelling choices, and the platform should be read as a structural model awaiting calibration. This page is the specification for that calibration. It is deliberately explicit about the items for which no usable proxy exists, because those are the items that determine whether a pilot can produce a defensible number or only a defensible shape.

Ingestion tiers
What is needed to start, what is needed to be believed, and what is needed to run the platform as an operational tool
Tier 0 · Demonstration
No client data — the current state of the platform

Runs entirely on the constants documented in the Model Documentation tab. Suitable for exploring mechanism and for scenario framing workshops. Every absolute number is illustrative. Nothing on this tier should appear in a business case.

Tier 1 · Pilot
Public and vendor data only, roughly 6 to 10 weeks of work

Substitutes the highest-leverage items that have credible external proxies: WAVEWATCH III or ERA5 wave hindcasts for the logistics weather process, HYCOM currents and ASCAT winds for the spill trajectory field, ADIOS weathering on a published analogue assay, METAR archives for base weather in the aviation engine, ANP round results and well data for the auction engine, and the published capital plan for the supplier engine. AIS covers vessel movements well enough to sanity-check cycle times. The output is a model whose behaviour can be argued about seriously, with cost levels still marked as indicative.

  • —Highest value per unit of effort: HYCOM plus ASCAT for the spill engine, and ERA5 for logistics weather. Both are free, both replace the crudest approximation in their engine.
  • —Lowest value at this tier: anything requiring internal cost data, because a partial cost calibration is more misleading than none.
  • —Deliverable: a calibrated structural model plus an explicit register of which constants remain uncalibrated.
Tier 2 · Production
Internal systems of record, governed feeds

Requires read access to the plant historian for FPSO tank telemetry, the chartering and procurement contract databases, the supply base terminal operating system, aviation contractor reporting, the drilling schedule and the supplier registry. Aggregated POB rosters. At this tier the model produces cost per well-day figures that can carry a decision, and the validation tab becomes a backtest against held-out history rather than a consistency check.

  • —Feeds should be scheduled extracts into a governed landing zone, not live system access. Nothing in the model needs sub-daily latency except during an active spill response.
  • —Every internal feed needs a named data owner and a documented refresh cadence before it is wired in.
  • —The spill engine is the exception: to be useful during an incident it needs near-real-time met-ocean forecast ingestion, which is an operational integration rather than a calibration one.
Data inventory by engine
Field, granularity, source of record, model dependency and proxy availability

Data quality gates
Checks that must pass before a feed is allowed to drive a calibrated run

Gate 1 · Completeness and coverage

  • —Time series must cover at least two full seasonal cycles for anything feeding a weather, wave or current process, so that the seasonal term is estimated rather than assumed.
  • —Gaps must be explicit. A missing AIS hour is not a stationary vessel, and a missing tank reading is not a constant level; both must be flagged rather than interpolated silently.
  • —Entity coverage must be stated as a percentage of the fleet, the platform set or the supplier base. A calibration on the six best-instrumented FPSOs is a calibration on the six best-instrumented FPSOs.

Gate 2 · Consistency and reconciliation

  • —Volume reconciliation: diesel loaded at the base, less delivered to platforms, less vessel consumption, must close within a stated tolerance. A mass balance that does not close means the model will be calibrated to a leak in the data.
  • —Cross-source agreement: AIS-derived port calls must reconcile against the terminal operating system's call records before either is trusted.
  • —Unit and currency discipline: every field carries an explicit unit, a currency and a price base year. The engines mix m³, tonnes, nautical miles and kilometres deliberately; an unlabelled column is a defect.
  • —Timezone discipline: all timestamps normalised to UTC on ingestion, with the local offset retained as a separate field, because crew-change waves are a local-time phenomenon.

Gate 3 · Plausibility and outlier handling

  • —Physical bounds per field, checked on ingestion: wave heights within a stated range for the basin, vessel speeds within the design envelope, tank levels within capacity, profit-oil offers within the regulated band.
  • —Outliers are quarantined and reviewed, never dropped automatically. In this domain the tail is usually the signal: the storm that closed the basin for four days is the observation that sizes the fleet.
  • —Step-change detection on every series, to catch instrument replacements, contract changes and system migrations that would otherwise be read as behaviour.

Gate 4 · Provenance and reproducibility

  • —Every calibrated constant carries a citation to the extract that produced it, with the extract date and the query or filter used.
  • —Calibration extracts are immutable and versioned. A model result is quoted as config plus seed plus code commit plus data version; any one of the four changing invalidates the comparison.
  • —A held-out period is reserved before calibration begins, not chosen afterwards, so that the validation tab reports a genuine out-of-sample result.
  • —Derived fields are computed in the pipeline, not in the engine, so that the same derivation is used by the model and by the validation.
Data residency, regulatory and privacy notes
Constraints that shape the architecture, not just the paperwork

Exploration and production data under the ANP regime

Seismic, well and reservoir data acquired under Brazilian exploration and production contracts are subject to confidentiality periods and to reporting obligations to the ANP through the BDEP. Anything derived from a block data package inherits those constraints. For the auction engine this matters directly: block resource estimates and pre-drill volumetrics cannot be moved into a shared analytical environment without checking the applicable confidentiality window and the contractual restrictions on disclosure to partners in other consortia.

LGPD and personnel data

POB rosters, rotation calendars and flight manifests are personal data under the Lei Geral de Proteção de Dados. The aviation engine does not need identities: it needs counts by cluster, by day and by discipline. The correct pattern is to aggregate inside the source system and to transfer only the aggregate, so that no personal data ever enters the modelling environment. Where a discipline breakdown is fine-grained enough to be re-identifying on a small installation, it should be suppressed or banded. This is not merely a compliance posture; it also removes the need for a legal basis assessment on the analytics platform itself.

Residency and hosting

The platform runs entirely in the browser: the TypeScript engines execute client-side and the Python mirror executes under Pyodide, also client-side. No simulation input or output leaves the user's machine by default. The only components with a server dependency are the optional AI assistant and any hosted calibration pipeline; both should be deployed in-region, and the assistant should not be given access to calibration extracts. Where a Petrobras deployment requires processing to remain within Brazil, the calibration pipeline should run on internal infrastructure and only the resulting constants — not the underlying extracts — should be shipped with the application bundle.

Commercial sensitivity

Charter rates, contract award values, supplier delivery performance and hurdle rates are commercially sensitive to counterparties as well as to competitors. Calibrated constants derived from them are still sensitive: a day rate recovered from a published cost-per-well-day figure is a disclosure. Any external publication of results from a calibrated deployment should go through the same review as the underlying contract data, and the demonstration tier should remain clearly separated from the calibrated tier in the build so the two cannot be confused.

Nothing on this page constitutes legal advice. The regulatory references are working summaries intended to shape the data architecture; the applicable obligations for a given block, contract or dataset should be confirmed with the relevant legal and regulatory functions before any feed is established.