The Question
For K = 2 regimes, scalar thresholding of the Bayesian posterior is known to be Bayes-optimal.
For K ≥ 3 regimes, the posterior evolves on a (K−1)-dimensional simplex. Yet virtually all deployed detection systems—CUSUM variants, deep anomaly detectors, classical and learned methods—reduce it to a single scalar score without theoretical justification.
This paper completely resolves this question by providing an exact parameter-level necessary and sufficient condition.
The Answer: Extended Decision Sufficiency
Scalar optimality is not a choice—it is a property of the underlying model.
The Three Conditions
Rank-One Emissions (RO)
All attack-regime densities share a common shape, differing only in scale: p_k = c_k h for all k.
This ensures observations reveal only aggregate attack information, not regime-specific distinctions.
Markov Factorisation (MF)
The expected evolution of attack mass depends only on its current value.
This ensures the scalar dynamics close on themselves without needing the full belief.
Normal Factorisation (NF)
Reversion to normal state follows a scalar-dependent pattern.
This ensures belief normalization preserves the scalar factorization.
All three are verifiable in O(K²) operations from model parameters alone. No simulation required. No learning required.
Interactive: EDS Diagnostic Tool
Drag the sliders to simulate violations of each EDS condition and see how η changes:
Interpretation: η ≈ 0 → Scalar detection is theoretically justified (Theorem 3 applies). η large → Scalar reduction discards essential information (Theorem 2: irreducible loss). The paper provides robustness bounds: performance degrades at rate O(η/(1−ρ)).
Theory: Proof Intuition
Sufficiency: Why EDS Implies Scalar Optimality
Under EDS, the Bellman operator preserves scalar functions. Here's why:
- Markov Factorisation ensures that φ(γ) after one transition depends only on φ(γ) today.
- Rank-One Emissions ensures observations reveal only φ-information.
- Normal Factorisation ensures the belief update normalizes consistently with φ.
Together, these guarantee that if V(γ) = Ṽ(φ(γ)), then (𝒯V)(γ) = Ṽ'(φ(γ)) for some Ṽ'. The value function stays scalar across iterations. By fixed-point uniqueness, V itself must be scalar.
Necessity: Why Violation Implies Suboptimality
The proof is by contradiction using a coupling argument:
- Construct a pair: For each EDS failure mode, we build two beliefs γ^a and γ^b with φ(γ^a) = φ(γ^b) (same scalar value) but different predictive densities.
- Total Variation gap: Since the observations disagree, there is a TV gap between the one-step predictions: TV(m_a, m_b) > 0.
- Propagate to values: Using the Lipschitz continuity of the Bayes filter and the ρ-contraction of the Bellman operator, this TV gap propagates to a value gap: |V(γ^a) − V(γ^b)| > 0.
- Contradiction: But φ(γ^a) = φ(γ^b) means any scalar function Ṽ satisfying V = Ṽ ∘ φ must have V(γ^a) = V(γ^b). This is impossible. So no scalar representation exists.
Key insight: This holds without assuming any structure on V beyond continuity. The contradiction is purely structural—it doesn't depend on how you learn the model.
Experiments: Theory → Prediction → Result
P1: Exactness of the Anticipation Horizon Formula
Theorem (Theorem 3): Under EDS, expected lead time is given by a closed-form fundamental-matrix hitting-time formula.
Prediction: For 72 synthetic HMM configurations (K=3,4,5 with exact EDS), Monte Carlo rollout hitting times should match the analytical prediction within discretisation error.
| Configuration | MAE (steps) | R² | η |
|---|---|---|---|
| K=3 (24 configs) | 0.28 ± 0.05 | 0.994 | < 0.001 |
| K=4 (24 configs) | 0.35 ± 0.06 | 0.990 | < 0.001 |
| K=5 (24 configs) | 0.40 ± 0.07 | 0.981 | < 0.001 |
| All (72 configs) | 0.34 ± 0.04 | 0.989 | — |
✓ PASS: Predictions match observations within discretisation error across all K values.
P2: Graceful Degradation with η
Theorem (Theorem 4): Under approximate EDS with deviation η, performance degrades at rate O(η/(1−ρ)).
Prediction: Lead time should decrease roughly linearly with η, with slope determined by model parameters.
CICIDS2017 results showing lead-time versus η across baselines.
| Method | F1 | Lead Time (steps) | η (est.) |
|---|---|---|---|
| CUSUM | 0.61 | 0.0 | — |
| Shiryaev–Roberts | 0.58 | 0.4 | — |
| LSTM | 0.79 | 1.2 | — |
| TranAD | 0.83 | 1.5 | — |
| Full HMM | 0.84 | 9.1 | 0.041 |
| EDS (ours) | 0.92 | 14.3 | 0.058 |
✓ PASS: EDS constrains the model to preserve scalar optimality while outperforming unrestricted and scalar baselines.
P3: Structure vs. Regularisation
Hypothesis: The EDS lead-time advantage is structural, not merely a regularisation effect from parameter reduction.
Control: Fit both EDS and a matched-parameter Full HMM (O(K) degrees of freedom in both) on synthetic data.
| Condition | n=500 | n=2000 | n=10000 | η (true) |
|---|---|---|---|---|
| EDS-true | 13.0 ± 0.3 | 13.2 ± 0.2 | 13.3 ± 0.1 | 0.000 |
| EDS-false (MF violation) | 5.9 ± 0.4 | 6.1 ± 0.3 | 6.0 ± 0.2 | 0.300 |
| Structural Gap | 7.1 steps (stable across all n) | — | ||
✓ PASS: Gap is stable and independent of sample size—evidence of structural effect, not regularisation.
Interactive Lab
1. Simplex Explorer: The Collapse Principle
Drag the belief point γ on the simplex. Observe how two beliefs with the same scalar φ(γ) can have different optimal values when EDS fails.
Under EDS: V(γ) is constant along φ-level sets.
Without EDS: φ is insufficient to determine V(γ).
2. Lead-Time Simulator
Adjust model parameters to see how lead time changes under EDS versus violation.
3. Failure-Mode Explorer
Click on each condition to see what happens when it fails.
Status: EDS holds ✓
The scalar statistic φ(γ) is sufficient for optimal stopping.
Value function: V(γ) = Ṽ(φ(γ))
Lead time: Optimal (closed-form prediction available)
Status: Rank-One Emissions fails ✗
Attack densities have different shapes. Observations reveal regime-specific information.
Value function: V(γ) ≠ Ṽ(φ(γ)) for any Ṽ
Lead time: Scalar methods incur irreducible loss (Theorem 2)
Status: Markov Factorisation fails ✗
The dynamics of φ depend on more than φ—they require full belief state.
Value function: Cannot be expressed as scalar function
Lead time: Scalar reduction causes degradation proportional to MF violation
Status: Normal Factorisation fails ✗
Belief normalisation couples to belief components outside the scalar.
Value function: Scalar collapse breaks at update step
Lead time: Bayesian update introduces spurious dimensions
Can I Use Scalar Detection?
A practical decision tree:
↓ YES
Does Rank-One Emissions hold? (Attack densities share shape)
↓ LIKELY
Does Markov Factorisation hold? (Transition dynamics factor through φ)
↓ LIKELY
Does Normal Factorisation hold? (Normal reversion is scalar-dependent)
↓ LIKELY
✓ Compute η
If η ≈ 0:
→ Scalar reduction is Bayes-optimal
If η moderate:
→ Scalar reduction has bounded loss (Theorem 4)
If η large:
→ Use full belief-state methods (Theorem 2)
What This Paper Does NOT Claim
- Does not claim every sequential detector should be scalar. Only when EDS holds.
- Does not claim EDS holds universally. In heterogeneous or unrelated-failure environments, it typically will not.
- Does not claim scalar neural networks become optimal with more data. If EDS fails structurally, no amount of training data or model capacity will fix it.
- Does not replace full belief-state methods when EDS fails. Theorem 2 proves scalar reduction is always suboptimal when EDS fails.
- Does not claim Baum–Welch optimization is ideal. The theorems are optimization-independent; empirical results depend on how well the model is fitted.
- Does not provide a universal model diagnostic without data. EDS itself is parameter-level, but estimating η from finite data has its own variance (addressed in Appendix).
Reproducibility Dashboard
Experiment Setup
| Synthetic Configurations | 72 exact EDS models across K=3,4,5 |
| Random Seeds | 5 independent seeds per configuration |
| CICIDS2017 | Days 1–3 training, Days 4–5 evaluation, 78 flow features |
| FPR Matching | All methods tuned to FPR ≤ 0.04 on held-out data |
| Structure Control | n ∈ {500, 2000, 10000}, matched O(K) parameters |
| Confidence Intervals | Mean ± 1.96 SE (95% CI), paired t-tests at α=0.05 |
Experiment Commands
python -m experiments.synthetic --configs 72 --seeds 5 --output results/synthetic.csv
# CICIDS2017 (P2)
python -m experiments.cicids2017 --data-dir data/cicids2017 --output results/cicids.csv
# Structure vs Regularisation (P4)
python -m experiments.structure_vs_reg --output results/structure.csv
# Reproduce all tables
python scripts/generate_tables.py results/ figures/tables.pdf
# Reproduce all figures
python scripts/generate_figures.py results/ figures/
Resources & Artifacts
Paper & Posters
Paper — Coming soon
NeurIPS Poster — Coming soon
Code & Data
Full implementation on GitHub: theory implementation, experiments, baselines, evaluation utilities.
CICIDS2017: Instructions for downloading and preprocessing available in the repository.
Citation
title={When Does Sequential Detection Collapse to a Scalar?
A Necessary and Sufficient Characterisation},
author={Hiremath, Prakul Sunil and Bagawan, Peerahammad M and Bhekane, Sahil},
booktitle={Advances in Neural Information Processing Systems},
year={2026}
}
How EDS Compares with Prior Work
| Approach | Preserves | Characterizes | Necessary & Sufficient | Parameter-Level |
|---|---|---|---|---|
| Lumpability | Transition structure | State space reduction | ✓ For dynamics | ✓ |
| Belief Compression | Belief distribution | Low-rank approximation | ✗ Approximate | ✗ A posteriori |
| POMDP Symmetry | Policy & value | Belief partitions | ✗ Heuristic | ✗ A posteriori |
| EDS (This Work) | Bayes-optimal stopping | Scalar Bayes-optimality | ✓ For stopping | ✓ A priori |
Key distinction: Prior approaches operate a posteriori or provide sufficient (but not necessary) conditions. EDS is the first parameter-level necessary and sufficient characterisation of when scalar Bayes-optimal stopping is possible.
Scope & Limitations
What This Work Establishes
Exact theoretical characterisation: EDS is necessary and sufficient for the Bayes-optimal value function to collapse to a scalar. This is a model-level, not algorithm-level, property.
Irreducible loss bound: When EDS fails, every scalar method incurs strictly positive, algorithm-independent loss.
Robustness guarantee: Under approximate EDS, degradation is bounded by O(η/(1−ρ)).
Empirical Scope
Synthetic experiments: 72 configurations by construction satisfy EDS. These validate the theory but do not demonstrate that EDS approximately holds across diverse application domains.
Real-world data: CICIDS2017 exhibits approximate EDS (η ≈ 0.058), supporting the hypothesis that network-flow intrusion data can be modeled with near-EDS structure. However, generalisation to other domains (industrial control, healthcare, finance, etc.) remains an open question.
Practical applicability: The paper identifies exactly when scalar methods are theoretically justified. Whether EDS holds in a given application must be verified empirically using Algorithm 1 before deployment.
Structural Limitations
Rank-One Emissions: Requires all attack densities to share a common shape. In highly heterogeneous environments (unrelated fault types, qualitatively distinct attack modalities), this will not hold.
Markov & Normal Factorisations: Impose structured symmetry on state transitions. Real-world dynamical systems may violate these if transitions are governed by physics or operational logic not captured by proportional or factored forms.
When EDS fails: Theorem 2 provides no remedy—scalar methods are provably suboptimal. The correct response is to use full belief-state methods or hybrid approaches that maintain richer state representations.
Key Results at a Glance
Theorem 1: Exact Characterisation
EDS ↔ Scalar Bayes-optimality. Three structural conditions completely determine when scalar detection is optimal.
Theorem 2: Irreducible Loss
When EDS fails, every scalar method incurs strictly positive, algorithm-independent lead-time loss.
Theorem 3: Closed-Form Lead Time
When EDS holds, expected lead time is given by fundamental-matrix hitting time of a scalar Markov chain.
Theorem 4: Robustness
Under approximate EDS with deviation η, performance degrades at rate O(η/(1−ρ)).
Empirical Validation
72 synthetic configs (K=3,4,5): predictions match observations (MAE ≤ 0.4 steps). CICIDS2017: EDS outperforms baselines by 5.2 steps.