Twelve-month forecast-volume-weighted result across the demonstration segments.
Research & evidence
Evidence before assertion.
Polaris is showing promising development performance while its protection controls keep weaker challengers from replacing better methods. The next step is to demonstrate that potential on representative industrial data.
Development performance
Positive results—with protection when a model falls short.
The current controlled demonstration evaluates three synthetic monthly demand segments across 18 validation observations. It shows the potential of the selection and uncertainty framework, not production or customer performance.
Improvement versus the selected baseline across the controlled snapshot.
Observed coverage against an 80% target in the demonstration.
Every evaluated segment passed through the configured selection gate.
Strongest substantial segment Alpha / Core achieved 93.25% accuracy with positive forecast value added.
Protected segment Gamma / Flex retained its baseline when the challenger would have reduced performance.
Current provider generation Improved aggregate demonstration accuracy by 25.1 percentage points over the earlier 2.0 snapshot.
Scientific methodology
Built to support decisions—not just produce forecasts.
Our validation process is designed to reduce the business risk of acting on a model that has not earned trust.
Define the question
Name the dataset, decision horizon, forecast cadence and operational use before evaluating a model.
Freeze the comparison
Fix the data split, candidate version, metrics and baselines before results are examined.
Test through time
Use time-ordered holdouts or rolling origins so future observations never leak into training.
Measure several failure modes
Evaluate scale-adjusted error, directional bias, forecast value added and interval coverage.
Retain the safer method
A challenger that fails the agreed gate does not replace the current champion or simple baseline.
State the boundary
Engineering verification, forecast performance and customer value are separate evidence levels.
Public proxy stress test
A harder test revealed the next performance priority.
A reproducible study used 24 monthly tourism series selected from 168 eligible series, with a six-month holdout. The data are public, but they are a proxy—not industrial product-order histories. The result identified cadence handling and calibration as concrete development priorities.
Development conclusion: retain the seasonal baseline for this proxy cohort while improving candidate selection and interval calibration.
Input frequency was a major integration variable. A dated application route produced 76.3% WMAPE; preserving the series at its native monthly cadence reduced Polaris WMAPE to 16.7%. This was a substantial correction, but it did not overturn the baseline result.
Business interpretation
What this approach means for a planning team.
The objective is not to replace a working planning method with a more complicated one. It is to identify where better evidence can improve a decision—and where the current method should remain.
No change without demonstrated improvement.
A simple baseline or existing process remains in place when a new model cannot show a reliable advantage.
Plan around ranges, not false precision.
Visible uncertainty helps teams prepare scenarios, review risk and choose where human attention matters most.
Test value before operational commitment.
A no-write shadow pilot can compare decisions and outcomes without changing production systems or automated workflows.
Potential & benefits
Better evidence can improve the quality and timing of planning decisions.
When validated on representative business data, Polaris is designed to help teams focus attention earlier, compare options consistently and make uncertainty part of the planning process.
These are intended benefits to be measured in a controlled pilot. They are not yet established customer outcomes.
Earlier risk visibility
Identify demand ranges and potential exceptions before they become urgent planning problems.
More focused planner attention
Direct expert review toward uncertain or high-impact cases instead of treating every forecast equally.
Stronger scenario decisions
Compare plausible outcomes and trade-offs with a consistent evidence trail rather than isolated assumptions.
Safer model adoption
Keep current methods in place until a challenger demonstrates a reliable advantage under agreed conditions.
Greater decision accountability
Preserve the forecast version, evidence, scenario and human override behind each reviewed decision.
Interpretation boundary
What this evidence does not prove.
Industrial performance
Tourism series do not reproduce the cadence, sparsity, promotions, lead times or constraints of industrial demand.
Production reliability
A benchmark does not establish uptime, integration quality, safe operational behavior or maintainability in a customer environment.
Customer value
Forecast metrics alone do not prove improved inventory, service, planning time or economic outcomes.
Vendor comparison
These results cannot be compared with vendor claims using different datasets, horizons, metrics or evaluation protocols.
Tourism Monthly Dataset, Zenodo DOI 10.5281/zenodo.4656096, licensed CC BY 4.0. Results shown are aggregate values from the Phaneon public-proxy evaluation and should be read with the limitations above.