๐Ÿ•ต๏ธ Detection Tool ๐Ÿ”ฌ Zenodo ๐Ÿ“ฆ PyPI ๐Ÿ”— GitHub โœ๏ธ Medium Article ๐Ÿผ Pandas & Polars โœ“ mypy strict

temporal-leaks

Valgrind for Time-Series ML

Automatically detect look-ahead bias before it costs you

A Python package that catches temporal leakage in your feature pipelines. Automatically test for data leakage, find which columns are compromised, and get actionable hints about where the bugs are โ€” all without modifying your code.

Pure Python, zero dependencies. Supports Pandas and Polars.

"Your backtest shows 40% annual returns. You deploy. You lose money."

Look-ahead bias (data leakage) is the silent killer of quant strategies and forecasting models. Somewhere in your feature pipeline, a rolling average peeked at tomorrow's prices. Your tests passed. Your backtests looked amazing. And then reality hit.

The Problem: Future Data in Your Past Features

In time-series machine learning, look-ahead bias occurs when a feature computed for timestamp t inadvertently uses data from timestamps t+1, t+2, โ€ฆ t+n.

This is devastatingly easy to introduce โ€” and none of these will raise an error:

โŒ Centered Rolling Window

# center=True looks FORWARD
df["roll_mean"] = df["price"].rolling(
    window=5,
    center=True
).mean()

โœ“ Backward-Only Rolling Window

# center=False (default) looks only back
df["roll_mean"] = df["price"].rolling(
    window=5,
    center=False
).mean()

โŒ Forward Shift

# shift(-1) reads the NEXT row
df["next_return"] = df["return"].shift(-1)

โœ“ Backward Shift

# shift(+1) reads the PREVIOUS row
df["prev_return"] = df["return"].shift(1)

โŒ Global Normalization

# Uses ENTIRE dataset (future!)
df["znorm"] = (
    (df["price"] - df["price"].mean())
    / df["price"].std()
)

โœ“ Expanding Normalization

# Only uses PAST data
mean = df["price"].expanding().mean()
std = df["price"].expanding().std()
df["znorm"] = (df["price"] - mean) / std

None of these bugs will raise an error. Your tests will pass. Your backtests will look amazing. And then reality hits when you deploy to live trading.

The Solution: Temporal Perturbation Test

temporal-leaks uses an automated temporal perturbation test that works like a memory debugger (Valgrind) but for time-series data.

How It Works

Timeline:  โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บ
                                    T (checkpoint)
                                    โ”‚
Past โ—€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”คโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บ Future
                                    โ”‚

Step 1:  Run pipeline on ORIGINAL data
         baseline_features = pipeline(df)

Step 2:  MUTATE the FUTURE
         df_perturbed[t > T] = ๐Ÿ”ฅ (noise / sign-flip / NaN)

Step 3:  Re-run pipeline on PERTURBED data
         perturbed_features = pipeline(df_perturbed)

Step 4:  Compare PAST rows only (t โ‰ค T)
         If baseline_features[tโ‰คT] โ‰  perturbed_features[tโ‰คT]
         then PAST depends on FUTURE โ†’ ๐Ÿšจ LEAK DETECTED!

Key insight: If your past features are truly causal, mutating the future should not change them. If they change, future data crept in somewhere.

The test is simple, deterministic, and reproducible. No black-box ML required.

Quick Start

Installation

pip install temporal-leaks

Minimal Example

import pandas as pd
import numpy as np
from temporal_leaks import TemporalAudit

# Create sample data
df = pd.DataFrame({
    "ts": np.arange(500),
    "price": np.random.default_rng(42).normal(100, 5, size=500),
})

# โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
# โœ“ CLEAN PIPELINE
# โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
def causal_features(df: pd.DataFrame) -> pd.DataFrame:
    out = df.copy()
    # Expanding window: only looks at past โœ“
    out["expanding_mean"] = out["price"].expanding(min_count=1).mean()
    # shift(+1): looks at previous row โœ“
    out["lag1"] = out["price"].shift(1)
    return out

auditor = TemporalAudit(mode="nullify", random_seed=42)
report = auditor.check(df, timestamp_col="ts", pipeline_fn=causal_features)
print(report)
# โœ“ CLEAN โ€” leakage_score=0.0000


# โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
# โœ— LEAKING PIPELINE
# โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€
def leaking_features(df: pd.DataFrame) -> pd.DataFrame:
    out = df.copy()
    # center=True: peeks at future rows โœ—
    out["centred_roll"] = out["price"].rolling(
        11, center=True, min_periods=1
    ).mean()
    return out

try:
    auditor.check(df, timestamp_col="ts", pipeline_fn=leaking_features)
except TemporalLeakageError as exc:
    print(exc)
    # TemporalLeakageError: leakage_score=0.4812
    #   Breached columns (1):
    #     โ€ข [HIGH] column='centred_roll' effect_size=0.4812

Decorator API

from temporal_leaks import temporal_audit

@temporal_audit(timestamp_col="ts", mode="noise", random_seed=42)
def build_features(df: pd.DataFrame) -> pd.DataFrame:
    df = df.copy()
    df["expanding_mean"] = df["price"].expanding(min_count=1).mean()
    return df

# Audit runs automatically on every call
result = build_features(df)  # raises TemporalLeakageError if leak detected

HTML Audit Reports

Every audit generates a beautiful, standalone HTML report you can share or archive:

report = auditor.check(df, "ts", leaking_features)

# Write a standalone HTML report
with open("audit_report.html", "w") as f:
    f.write(report.to_html())

The report includes:

HTML reports are great for sharing audit results with your team, archiving for compliance, or including in documentation.

API Reference

TemporalAudit

TemporalAudit(
    mode: Literal["noise", "sign_flip", "nullify"] = "noise",
    random_seed: int = 42,
    delta_threshold: float = 1e-8,
    leakage_threshold: float = 0.0,
    ignore_columns: list[str] | None = None,
)
Parameter Type Description
mode str noise: adds Gaussian noise ยท sign_flip: multiplies by โˆ’1 ยท nullify: sets NaN
random_seed int Integer seed for determinism. Fully reproducible across runs.
delta_threshold float Minimum cell-level change to count as "different". Suppresses float noise.
leakage_threshold float If leakage_score > leakage_threshold, raise error. Set to 1.1 to always return report.
ignore_columns list[str] Output columns to skip during comparison (e.g., timestamps, IDs).

Perturbation Modes

noise

Adds Gaussian noise to future rows: ฮผ=0, ฯƒ=2ร—column_std. Safe for pipelines that handle noise.

sign_flip

Multiplies numeric values by โˆ’1. Good for testing sign-sensitive factors (momentum, etc).

nullify

Replaces future values with NaN. The strictest test โ€” use this first.

Severity Classification

Severity Effect Size Interpretation
๐ŸŸฆ LOW effect_size < 0.15 Minimal leakage. May be noise or minor bug.
๐ŸŸจ MEDIUM 0.15 โ‰ค effect_size < 0.40 Noticeable leakage. Investigate and fix.
๐ŸŸง HIGH 0.40 โ‰ค effect_size < 0.75 Significant leakage. Feature is compromised.
๐ŸŸฅ CRITICAL effect_size โ‰ฅ 0.75 Severe leakage. Feature is unreliable.

Backends & Performance

Supported DataFrames

๐Ÿผ Pandas

Full support for pd.DataFrame. Results returned as Pandas DataFrames.

๐Ÿ“Š Polars

First-class Polars support. Recommended for large frames (10M+ rows). Results returned as Polars DataFrames.

Benchmarks

Measured on Apple M2 Pro, 16 GB RAM:

Dataset Rows Columns Backend Mode Time
Synthetic prices 1,000,000 5 Polars nullify ~1.1 s
Synthetic prices 10,000,000 5 Polars nullify ~3.2 s
Equity features 500,000 20 Pandas noise ~2.8 s

Performance tip: For large datasets (>100k rows), use Polars. It's typically 5-10x faster than Pandas for temporal operations.

Testing & Development

Run the Test Suite

# Install dev extras
pip install -e ".[dev]"

# Run full test suite
pytest tests/ -v

# With coverage
pytest tests/ --cov=temporal_leaks --cov-report=term-missing

Code Quality

The codebase is maintained to strict standards:

Type Checking

mypy

strict mode

Linting

ruff

all rules

CI/CD

GitHub

Actions

Contributing

Pull requests are welcome! Please make sure ruff check . and mypy temporal_leaks/ pass before submitting.

Why temporal-leaks?

๐ŸŽฏ Automatic Detection

No manual inspection. No guessing. The tool finds leaks automatically.

๐Ÿ“‹ Detailed Reports

Identifies exactly which columns leak, where, and why. HTML reports ready to share.

๐Ÿ”ฌ Deterministic

Reproducible across runs. Fixed seeds. No randomness in the results.

โšก No Dependencies

Pure Python. Works with Pandas, Polars, or your own dataframe.

โœ“ Decorator API

One decorator. Audit runs automatically on every call. Ship with confidence.

๐Ÿ› ๏ธ Polars Native

First-class Polars support. Lightning-fast on large datasets.

temporal-leaks โ€” Valgrind for Time-Series ML

MIT License ยท Python 3.9+ ยท Pandas & Polars ยท Open Source

"Catch the bugs before they cost you."