NeurIPS 2026 Main Track Sydney

When Does Sequential Detection Collapse to a Scalar?

A Necessary and Sufficient Characterisation

Prakul Sunil Hiremath · Peerahammad M Bagawan · Sahil Bhekane

The Question

For K = 2 regimes, scalar thresholding of the Bayesian posterior is known to be Bayes-optimal.

For K ≥ 3 regimes, the posterior evolves on a (K−1)-dimensional simplex. Yet virtually all deployed detection systems—CUSUM variants, deep anomaly detectors, classical and learned methods—reduce it to a single scalar score without theoretical justification.

The core question: When is this reduction valid? When is it suboptimal? By how much?

This paper completely resolves this question by providing an exact parameter-level necessary and sufficient condition.

The Answer: Extended Decision Sufficiency

Scalar optimality is not a choice—it is a property of the underlying model.

Main Result: The Bayes-optimal stopping value function collapses to a scalar if and only if the model satisfies three structural conditions called Extended Decision Sufficiency (EDS).

The Three Conditions

Rank-One Emissions (RO)

All attack-regime densities share a common shape, differing only in scale: p_k = c_k h for all k.

This ensures observations reveal only aggregate attack information, not regime-specific distinctions.

Markov Factorisation (MF)

The expected evolution of attack mass depends only on its current value.

This ensures the scalar dynamics close on themselves without needing the full belief.

Normal Factorisation (NF)

Reversion to normal state follows a scalar-dependent pattern.

This ensures belief normalization preserves the scalar factorization.

All three are verifiable in O(K²) operations from model parameters alone. No simulation required. No learning required.

Interactive: EDS Diagnostic Tool

Drag the sliders to simulate violations of each EDS condition and see how η changes:

Total EDS Deviation
η = 0.000
Status: ✓ PASS (EDS holds)

Interpretation: η ≈ 0 → Scalar detection is theoretically justified (Theorem 3 applies). η large → Scalar reduction discards essential information (Theorem 2: irreducible loss). The paper provides robustness bounds: performance degrades at rate O(η/(1−ρ)).

Theory: Proof Intuition

Sufficiency: Why EDS Implies Scalar Optimality

+ Expand

Under EDS, the Bellman operator preserves scalar functions. Here's why:

  1. Markov Factorisation ensures that φ(γ) after one transition depends only on φ(γ) today.
  2. Rank-One Emissions ensures observations reveal only φ-information.
  3. Normal Factorisation ensures the belief update normalizes consistently with φ.

Together, these guarantee that if V(γ) = Ṽ(φ(γ)), then (𝒯V)(γ) = Ṽ'(φ(γ)) for some Ṽ'. The value function stays scalar across iterations. By fixed-point uniqueness, V itself must be scalar.

Necessity: Why Violation Implies Suboptimality

+ Expand

The proof is by contradiction using a coupling argument:

  1. Construct a pair: For each EDS failure mode, we build two beliefs γ^a and γ^b with φ(γ^a) = φ(γ^b) (same scalar value) but different predictive densities.
  2. Total Variation gap: Since the observations disagree, there is a TV gap between the one-step predictions: TV(m_a, m_b) > 0.
  3. Propagate to values: Using the Lipschitz continuity of the Bayes filter and the ρ-contraction of the Bellman operator, this TV gap propagates to a value gap: |V(γ^a) − V(γ^b)| > 0.
  4. Contradiction: But φ(γ^a) = φ(γ^b) means any scalar function Ṽ satisfying V = Ṽ ∘ φ must have V(γ^a) = V(γ^b). This is impossible. So no scalar representation exists.

Key insight: This holds without assuming any structure on V beyond continuity. The contradiction is purely structural—it doesn't depend on how you learn the model.

Experiments: Theory → Prediction → Result

P1: Exactness of the Anticipation Horizon Formula

Theorem (Theorem 3): Under EDS, expected lead time is given by a closed-form fundamental-matrix hitting-time formula.

Prediction: For 72 synthetic HMM configurations (K=3,4,5 with exact EDS), Monte Carlo rollout hitting times should match the analytical prediction within discretisation error.

Configuration MAE (steps) R² η
K=3 (24 configs) 0.28 ± 0.05 0.994 < 0.001
K=4 (24 configs) 0.35 ± 0.06 0.990 < 0.001
K=5 (24 configs) 0.40 ± 0.07 0.981 < 0.001
All (72 configs) 0.34 ± 0.04 0.989 —

✓ PASS: Predictions match observations within discretisation error across all K values.

P2: Graceful Degradation with η

Theorem (Theorem 4): Under approximate EDS with deviation η, performance degrades at rate O(η/(1−ρ)).

Prediction: Lead time should decrease roughly linearly with η, with slope determined by model parameters.

CICIDS2017 results showing lead-time versus η across baselines.

Method F1 Lead Time (steps) η (est.)
CUSUM 0.61 0.0 —
Shiryaev–Roberts 0.58 0.4 —
LSTM 0.79 1.2 —
TranAD 0.83 1.5 —
Full HMM 0.84 9.1 0.041
EDS (ours) 0.92 14.3 0.058

✓ PASS: EDS constrains the model to preserve scalar optimality while outperforming unrestricted and scalar baselines.

P3: Structure vs. Regularisation

Hypothesis: The EDS lead-time advantage is structural, not merely a regularisation effect from parameter reduction.

Control: Fit both EDS and a matched-parameter Full HMM (O(K) degrees of freedom in both) on synthetic data.

Condition n=500 n=2000 n=10000 η (true)
EDS-true 13.0 ± 0.3 13.2 ± 0.2 13.3 ± 0.1 0.000
EDS-false (MF violation) 5.9 ± 0.4 6.1 ± 0.3 6.0 ± 0.2 0.300
Structural Gap 7.1 steps (stable across all n) —

✓ PASS: Gap is stable and independent of sample size—evidence of structural effect, not regularisation.

Interactive Lab

1. Simplex Explorer: The Collapse Principle

Drag the belief point γ on the simplex. Observe how two beliefs with the same scalar φ(γ) can have different optimal values when EDS fails.

Under EDS: V(γ) is constant along φ-level sets.

Without EDS: φ is insufficient to determine V(γ).

2. Lead-Time Simulator

Adjust model parameters to see how lead time changes under EDS versus violation.

Predicted Expected Lead Time
14.2 steps
Degradation from η: 0.0 steps

3. Failure-Mode Explorer

Click on each condition to see what happens when it fails.

Status: EDS holds ✓

The scalar statistic φ(γ) is sufficient for optimal stopping.

Value function: V(γ) = Ṽ(φ(γ))

Lead time: Optimal (closed-form prediction available)

Status: Rank-One Emissions fails ✗

Attack densities have different shapes. Observations reveal regime-specific information.

Value function: V(γ) ≠ Ṽ(φ(γ)) for any Ṽ

Lead time: Scalar methods incur irreducible loss (Theorem 2)

Status: Markov Factorisation fails ✗

The dynamics of φ depend on more than φ—they require full belief state.

Value function: Cannot be expressed as scalar function

Lead time: Scalar reduction causes degradation proportional to MF violation

Status: Normal Factorisation fails ✗

Belief normalisation couples to belief components outside the scalar.

Value function: Scalar collapse breaks at update step

Lead time: Bayesian update introduces spurious dimensions

Can I Use Scalar Detection?

A practical decision tree:

Do you have multiple latent regimes (K ≥ 3)?
↓ YES

Does Rank-One Emissions hold? (Attack densities share shape)
↓ LIKELY

Does Markov Factorisation hold? (Transition dynamics factor through φ)
↓ LIKELY

Does Normal Factorisation hold? (Normal reversion is scalar-dependent)
↓ LIKELY

✓ Compute η

If η ≈ 0:
→ Scalar reduction is Bayes-optimal

If η moderate:
→ Scalar reduction has bounded loss (Theorem 4)

If η large:
→ Use full belief-state methods (Theorem 2)

What This Paper Does NOT Claim

  • Does not claim every sequential detector should be scalar. Only when EDS holds.
  • Does not claim EDS holds universally. In heterogeneous or unrelated-failure environments, it typically will not.
  • Does not claim scalar neural networks become optimal with more data. If EDS fails structurally, no amount of training data or model capacity will fix it.
  • Does not replace full belief-state methods when EDS fails. Theorem 2 proves scalar reduction is always suboptimal when EDS fails.
  • Does not claim Baum–Welch optimization is ideal. The theorems are optimization-independent; empirical results depend on how well the model is fitted.
  • Does not provide a universal model diagnostic without data. EDS itself is parameter-level, but estimating η from finite data has its own variance (addressed in Appendix).

Reproducibility Dashboard

Experiment Setup

Synthetic Configurations 72 exact EDS models across K=3,4,5
Random Seeds 5 independent seeds per configuration
CICIDS2017 Days 1–3 training, Days 4–5 evaluation, 78 flow features
FPR Matching All methods tuned to FPR ≤ 0.04 on held-out data
Structure Control n ∈ {500, 2000, 10000}, matched O(K) parameters
Confidence Intervals Mean ± 1.96 SE (95% CI), paired t-tests at α=0.05

Experiment Commands

# Synthetic experiment (P1)
python -m experiments.synthetic --configs 72 --seeds 5 --output results/synthetic.csv

# CICIDS2017 (P2)
python -m experiments.cicids2017 --data-dir data/cicids2017 --output results/cicids.csv

# Structure vs Regularisation (P4)
python -m experiments.structure_vs_reg --output results/structure.csv

# Reproduce all tables
python scripts/generate_tables.py results/ figures/tables.pdf

# Reproduce all figures
python scripts/generate_figures.py results/ figures/

Resources & Artifacts

Paper & Posters

Paper — Coming soon

NeurIPS Poster — Coming soon

Code & Data

Full implementation on GitHub: theory implementation, experiments, baselines, evaluation utilities.

CICIDS2017: Instructions for downloading and preprocessing available in the repository.

Citation

@inproceedings{hiremath2026sequential,
  title={When Does Sequential Detection Collapse to a Scalar?
          A Necessary and Sufficient Characterisation},
  author={Hiremath, Prakul Sunil and Bagawan, Peerahammad M and Bhekane, Sahil},
  booktitle={Advances in Neural Information Processing Systems},
  year={2026}
}

How EDS Compares with Prior Work

Approach Preserves Characterizes Necessary & Sufficient Parameter-Level
Lumpability Transition structure State space reduction ✓ For dynamics ✓
Belief Compression Belief distribution Low-rank approximation ✗ Approximate ✗ A posteriori
POMDP Symmetry Policy & value Belief partitions ✗ Heuristic ✗ A posteriori
EDS (This Work) Bayes-optimal stopping Scalar Bayes-optimality ✓ For stopping ✓ A priori

Key distinction: Prior approaches operate a posteriori or provide sufficient (but not necessary) conditions. EDS is the first parameter-level necessary and sufficient characterisation of when scalar Bayes-optimal stopping is possible.

Scope & Limitations

What This Work Establishes

Exact theoretical characterisation: EDS is necessary and sufficient for the Bayes-optimal value function to collapse to a scalar. This is a model-level, not algorithm-level, property.

Irreducible loss bound: When EDS fails, every scalar method incurs strictly positive, algorithm-independent loss.

Robustness guarantee: Under approximate EDS, degradation is bounded by O(η/(1−ρ)).

Empirical Scope

Synthetic experiments: 72 configurations by construction satisfy EDS. These validate the theory but do not demonstrate that EDS approximately holds across diverse application domains.

Real-world data: CICIDS2017 exhibits approximate EDS (η ≈ 0.058), supporting the hypothesis that network-flow intrusion data can be modeled with near-EDS structure. However, generalisation to other domains (industrial control, healthcare, finance, etc.) remains an open question.

Practical applicability: The paper identifies exactly when scalar methods are theoretically justified. Whether EDS holds in a given application must be verified empirically using Algorithm 1 before deployment.

Structural Limitations

Rank-One Emissions: Requires all attack densities to share a common shape. In highly heterogeneous environments (unrelated fault types, qualitatively distinct attack modalities), this will not hold.

Markov & Normal Factorisations: Impose structured symmetry on state transitions. Real-world dynamical systems may violate these if transitions are governed by physics or operational logic not captured by proportional or factored forms.

When EDS fails: Theorem 2 provides no remedy—scalar methods are provably suboptimal. The correct response is to use full belief-state methods or hybrid approaches that maintain richer state representations.

Key Results at a Glance

Theorem 1: Exact Characterisation

EDS ↔ Scalar Bayes-optimality. Three structural conditions completely determine when scalar detection is optimal.

Theorem 2: Irreducible Loss

When EDS fails, every scalar method incurs strictly positive, algorithm-independent lead-time loss.

Theorem 3: Closed-Form Lead Time

When EDS holds, expected lead time is given by fundamental-matrix hitting time of a scalar Markov chain.

Theorem 4: Robustness

Under approximate EDS with deviation η, performance degrades at rate O(η/(1−ρ)).

Empirical Validation

72 synthetic configs (K=3,4,5): predictions match observations (MAE ≤ 0.4 steps). CICIDS2017: EDS outperforms baselines by 5.2 steps.