Skip to main content

ForesightEval Protocol

A quality standard for strategic foresight

When AI writes a scenario analysis for your board, how do you know it's any good? ForesightEval is the protocol we built to answer that question — seven measurable dimensions that separate foresight you can stake a decision on from analysis that merely reads well.

ForesightEval · ScorecardAI-Driven Public Sector 2030

Strategic Anticipation Quotient

8.6/ 10

Seven dimensions · every score decomposable to its evidence

The problem

Looking right is not being right

  1. 01

    Fluency masks failure

    Models produce authoritative prose that reads like strategy. But fluency is a surface property — it tells you nothing about whether the causal reasoning holds up.

  2. 02

    Benchmarks test the wrong thing

    Existing benchmarks score isolated predictions. Foresight is a different discipline — its value lies in stress-testing strategy against multiple futures, not calculating the probability of one.

  3. 03

    Alignment kills honesty

    Modern AI models are trained to be helpful. That training teaches them to agree, avoid discomfort, and default to consensus. For risk management, where the entire point is naming uncomfortable truths, this is a structural failure.

Our approach

Three principles, built into every score

  1. 01

    Measure what matters

    It is simple to score whether a model’s probability estimate was correct. It is hard to score whether a scenario is coherent, whether it surfaces the disruption a board hasn’t considered, or whether it translates into action inside ninety days. ForesightEval does the hard version, because the easy version is not what strategy teams actually need.

  2. 02

    Penalize comfort, reward courage

    The most dangerous AI foresight is the kind that quietly agrees with the strategy already on the table. ForesightEval explicitly scores whether a model named the uncomfortable scenario, challenged the assumption, or blinked. Analysis that only confirms what leadership already believes does not pass the bar.

  3. 03

    Every score, fully decomposable

    A quality metric you cannot audit is not a quality metric. Every ForesightEval score breaks down to its seven dimensions, each dimension to its evidence, each piece of evidence to its source. Scenarios inherit the same discipline through Bayesian anchoring (Tetlock, Shell, IPCC) — probabilities move only on triggered signposts or materially new claims, never from a fresh model run.

Differentiators

What makes it different

  1. 01

    Analysis quality, not prediction accuracy

    Existing benchmarks ask "did AI get the answer right?" We ask whether the reasoning is sound, the scenarios are structurally distinct, and the output is actionable at board level.

  2. 02

    Anti-sycophancy by design

    Frontier models produce sycophantic answers roughly a third of the time. ForesightEval explicitly measures whether the model had the courage to name the scenario that threatens your strategy.

  3. 03

    Fully decomposable scoring

    Every score traces back to specific passages, evidence, and reasoning chains. No black box — because a black-box quality metric is a contradiction in terms.

  4. 04

    Two evaluation tracks

    Retrospective backtests against known outcomes, plus live foresight on unresolved topics. Both tracks are designed to eliminate data contamination and hindsight bias.

The framework

Seven dimensions. One Strategic Anticipation Quotient

Each dimension targets a specific way AI-generated foresight fails in practice — failures we have seen repeatedly in client work across regulated industries. The seven dimensions combine into a composite score, the Strategic Anticipation Quotient, with every sub-score fully decomposable so you can see where an analysis is strong and where it quietly breaks.

Tier IThe narrative coreWhat any piece of strategic foresight must do before anything else matters.
  1. 01

    Scenario Quality

    Are the scenarios logically consistent, structurally distinct, and systemically grounded? A set that varies only one dial is a sensitivity analysis wearing a costume.

  2. 02

    Epistemic Grounding

    Can every material claim be traced to evidence and an explicit reasoning chain? Analysis that cannot show its work is indistinguishable from fluent confabulation.

Tier IIThe strategic edgeWhat separates foresight that changes decisions from foresight that comforts them.
  1. 03

    Unpalatable Truths

    Does the analysis name the scenario that threatens the current strategy, or does it retreat to the consensus? This is where today’s models fail hardest.

  2. 04

    Weak Signal Detection

    Does the model surface disruptions not yet in mainstream conversation, or does regression to the mean bury them under signals everyone already tracks?

  3. 05

    Actionability

    Can a C-level executive act on this inside ninety days? Does it interact with real regulatory thresholds, capex cycles, and risk appetite?

Tier IIIThe quality moderatorsWhat keeps the above trustworthy over time.
  1. 06

    Living Foresight

    Does the analysis stay alive — updating as signposts trigger, revising as uncertainties resolve — or is it a static PDF that ages on delivery?

  2. 07

    Behavior Fidelity

    Are projections of customer and stakeholder behavior grounded in validated models, or aesthetic placeholders? Conditional — scored only when synthetic panel data exists.

  3. 08

    Explainability of the Score

    Every SAQ is fully decomposable. Per-dimension scores, evidence, specific passages. No black box — a black-box quality metric is a contradiction in terms.

Methodology

How we measure it

Two evaluation tracks, designed to eliminate data contamination and hindsight bias.

Retrospective track

Frozen Context Snapshot

A model receives a bounded corpus up to a specific past date and nothing after it. It generates scenarios and recommendations. We score against the reality that subsequently unfolded — with evaluators blinded to chronological context to prevent hindsight bias.

First backtests: CEE energy transition (2021–2025), EU banking digitalization (2020–2025), AI capability development (2022–2026).

Prospective track

Live Foresight

Multiple models generate foresight today on unresolved topics. We evaluate at 12, 24, and 36 months. Contamination-proof by construction — the future has not happened yet.

Also measures whether models retreat from contrarian positions under sustained user pressure.

In practice

Every Future Space carries a ForesightEval score

ForesightEval currently runs as the internal quality layer on every Future Space DSGHT.ai publishes. The score is calculated before release, visible on the analysis page, and decomposable to the per-dimension level — so the quality claim can be audited against the evidence.

This is not yet a cross-model benchmark — that track opens with the first retrospective backtests later in 2026. What follows is the standard DSGHT.ai holds its own production work to, published openly rather than kept internal.

AI-Driven Public Sector 2030

CEE · 2030 horizonCompleted April 2026

Strategic Anticipation Quotient

8.6/ 10

DimensionScoreNote
Scenario Quality9.0Structurally distinct 2×2 matrix, probabilities sum to 100 %
Epistemic Grounding10Historical analogies, complete structural consistency
Unpalatable Truths10Sovereign Algocracy scenario directly challenges comfort zone
Weak Signal Detection7.8Relies on well-publicised cases; fringe signals underrepresented
Actionability9.0Tension-linked recommendations tied to regulatory milestones
Living Foresight7.5Static probabilities; no temporal metadata or signpost tracking
Explainability7.0Claims metadata missing from artifact; citations unverifiable
View full analysis

Scored by the DSGHT.ai internal pipeline. Cross-model scoring, human-vs-AI comparison, and retrospective backtests are on the 2026 roadmap.

Roadmap

From internal protocol to industry standard

  1. 01

    Internal protocol

    ForesightEval started as our internal quality layer — the standard DSGHT.ai holds its own production work to.

  2. 02

    Scores on every Future Space

    Every Future Space we publish carries a ForesightEval score with all sub-metrics visible. When the model hedges or blinks, the score says so.

  3. 03

    Open to partners over 2026

    We’re opening the methodology to early partners — buyers of AI-generated strategy who want to evaluate their own pipelines, and producers who want independent scoring. We publish openly because a quality standard only one vendor can audit is not a standard.