Research & evidence

Evidence before assertion.

Polaris is showing promising development performance while its protection controls keep weaker challengers from replacing better methods. The next step is to demonstrate that potential on representative industrial data.

Current statusPromising development signalControlled demonstration completedIndustrial shadow validation is the next gate

Development performance

Positive results—with protection when a model falls short.

The current controlled demonstration evaluates three synthetic monthly demand segments across 18 validation observations. It shows the potential of the selection and uncertainty framework, not production or customer performance.

Demonstration onlySynthetic data · 3 segments · 18 validation points
Weighted forecast accuracy75.7%

Twelve-month forecast-volume-weighted result across the demonstration segments.

Selected forecast value added+7.1 pp

Improvement versus the selected baseline across the controlled snapshot.

Prediction-interval coverage83.3%

Observed coverage against an 80% target in the demonstration.

Champion gate coverage100%

Every evaluated segment passed through the configured selection gate.

93.3%

Strongest substantial segment Alpha / Core achieved 93.25% accuracy with positive forecast value added.

97.5%

Protected segment Gamma / Flex retained its baseline when the challenger would have reduced performance.

2.1

Current provider generation Improved aggregate demonstration accuracy by 25.1 percentage points over the earlier 2.0 snapshot.

Scientific methodology

Built to support decisions—not just produce forecasts.

Our validation process is designed to reduce the business risk of acting on a model that has not earned trust.

01

Define the question

Name the dataset, decision horizon, forecast cadence and operational use before evaluating a model.

02

Freeze the comparison

Fix the data split, candidate version, metrics and baselines before results are examined.

03

Test through time

Use time-ordered holdouts or rolling origins so future observations never leak into training.

04

Measure several failure modes

Evaluate scale-adjusted error, directional bias, forecast value added and interval coverage.

05

Retain the safer method

A challenger that fails the agreed gate does not replace the current champion or simple baseline.

06

State the boundary

Engineering verification, forecast performance and customer value are separate evidence levels.

Public proxy stress test

A harder test revealed the next performance priority.

A reproducible study used 24 monthly tourism series selected from 168 eligible series, with a six-month holdout. The data are public, but they are a proxy—not industrial product-order histories. The result identified cadence handling and calibration as concrete development priorities.

Polaris WMAPE16.7%Native monthly evaluation
Seasonal-naive WMAPE7.1%Lower error; baseline retained
Polaris bias−1.7%Aggregate directional bias
P10–P90 coverage31.2%Below the nominal 80% interval
Aggregate WMAPELower is better
Polaris16.7%
Seasonal naive7.1%

Development conclusion: retain the seasonal baseline for this proxy cohort while improving candidate selection and interval calibration.

What the experiment taught us

Input frequency was a major integration variable. A dated application route produced 76.3% WMAPE; preserving the series at its native monthly cadence reduced Polaris WMAPE to 16.7%. This was a substantial correction, but it did not overturn the baseline result.

Business interpretation

What this approach means for a planning team.

The objective is not to replace a working planning method with a more complicated one. It is to identify where better evidence can improve a decision—and where the current method should remain.

Protect what already works

No change without demonstrated improvement.

A simple baseline or existing process remains in place when a new model cannot show a reliable advantage.

Make uncertainty usable

Plan around ranges, not false precision.

Visible uncertainty helps teams prepare scenarios, review risk and choose where human attention matters most.

Start with limited exposure

Test value before operational commitment.

A no-write shadow pilot can compare decisions and outcomes without changing production systems or automated workflows.

Potential & benefits

Better evidence can improve the quality and timing of planning decisions.

When validated on representative business data, Polaris is designed to help teams focus attention earlier, compare options consistently and make uncertainty part of the planning process.

Evidence boundary

These are intended benefits to be measured in a controlled pilot. They are not yet established customer outcomes.

01

Earlier risk visibility

Identify demand ranges and potential exceptions before they become urgent planning problems.

02

More focused planner attention

Direct expert review toward uncertain or high-impact cases instead of treating every forecast equally.

03

Stronger scenario decisions

Compare plausible outcomes and trade-offs with a consistent evidence trail rather than isolated assumptions.

04

Safer model adoption

Keep current methods in place until a challenger demonstrates a reliable advantage under agreed conditions.

05

Greater decision accountability

Preserve the forecast version, evidence, scenario and human override behind each reviewed decision.

Interpretation boundary

What this evidence does not prove.

01

Industrial performance

Tourism series do not reproduce the cadence, sparsity, promotions, lead times or constraints of industrial demand.

02

Production reliability

A benchmark does not establish uptime, integration quality, safe operational behavior or maintainability in a customer environment.

03

Customer value

Forecast metrics alone do not prove improved inventory, service, planning time or economic outcomes.

04

Vendor comparison

These results cannot be compared with vendor claims using different datasets, horizons, metrics or evaluation protocols.

Next evidence gate

Prove relevance using real planning conditions.

The next step is a focused evaluation using representative industrial demand, agreed success measures and no operational writes.

  1. 01

    Freeze the software version, representative dataset and evaluation protocol.

  2. 02

    Run rolling-origin evaluation on industrial demand at its native cadence.

  3. 03

    Compare against seasonal-naive and demand-appropriate intermittent baselines.

  4. 04

    Report WMAPE or MASE, bias, forecast value added and interval coverage by horizon and demand type.

  5. 05

    Publish a reproducible release note with exclusions, failures and known limitations.

  6. 06

    Claim superiority only after the model consistently clears baseline-protection gates.

  7. 07

    Evaluate operational value separately through an approved, no-write shadow pilot.

Benchmark source

Tourism Monthly Dataset, Zenodo DOI 10.5281/zenodo.4656096, licensed CC BY 4.0. Results shown are aggregate values from the Phaneon public-proxy evaluation and should be read with the limitations above.