Valgrind for Time-Series ML
Automatically detect look-ahead bias before it costs you
A Python package that catches temporal leakage in your feature pipelines. Automatically test for data leakage, find which columns are compromised, and get actionable hints about where the bugs are โ all without modifying your code.
Pure Python, zero dependencies. Supports Pandas and Polars.
"Your backtest shows 40% annual returns. You deploy. You lose money."
Look-ahead bias (data leakage) is the silent killer of quant strategies and forecasting models. Somewhere in your feature pipeline, a rolling average peeked at tomorrow's prices. Your tests passed. Your backtests looked amazing. And then reality hit.
In time-series machine learning, look-ahead bias occurs when a feature computed for timestamp t inadvertently uses data from timestamps t+1, t+2, โฆ t+n.
This is devastatingly easy to introduce โ and none of these will raise an error:
# center=True looks FORWARD
df["roll_mean"] = df["price"].rolling(
window=5,
center=True
).mean()
# center=False (default) looks only back
df["roll_mean"] = df["price"].rolling(
window=5,
center=False
).mean()
# shift(-1) reads the NEXT row
df["next_return"] = df["return"].shift(-1)
# shift(+1) reads the PREVIOUS row
df["prev_return"] = df["return"].shift(1)
# Uses ENTIRE dataset (future!)
df["znorm"] = (
(df["price"] - df["price"].mean())
/ df["price"].std()
)
# Only uses PAST data
mean = df["price"].expanding().mean()
std = df["price"].expanding().std()
df["znorm"] = (df["price"] - mean) / std
None of these bugs will raise an error. Your tests will pass. Your backtests will look amazing. And then reality hits when you deploy to live trading.
temporal-leaks uses an automated temporal perturbation test that works like a memory debugger (Valgrind) but for time-series data.
Timeline: โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโบ
T (checkpoint)
โ
Past โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโคโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโบ Future
โ
Step 1: Run pipeline on ORIGINAL data
baseline_features = pipeline(df)
Step 2: MUTATE the FUTURE
df_perturbed[t > T] = ๐ฅ (noise / sign-flip / NaN)
Step 3: Re-run pipeline on PERTURBED data
perturbed_features = pipeline(df_perturbed)
Step 4: Compare PAST rows only (t โค T)
If baseline_features[tโคT] โ perturbed_features[tโคT]
then PAST depends on FUTURE โ ๐จ LEAK DETECTED!
Key insight: If your past features are truly causal, mutating the future should not change them. If they change, future data crept in somewhere.
The test is simple, deterministic, and reproducible. No black-box ML required.
pip install temporal-leaksimport pandas as pd
import numpy as np
from temporal_leaks import TemporalAudit
# Create sample data
df = pd.DataFrame({
"ts": np.arange(500),
"price": np.random.default_rng(42).normal(100, 5, size=500),
})
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
# โ CLEAN PIPELINE
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
def causal_features(df: pd.DataFrame) -> pd.DataFrame:
out = df.copy()
# Expanding window: only looks at past โ
out["expanding_mean"] = out["price"].expanding(min_count=1).mean()
# shift(+1): looks at previous row โ
out["lag1"] = out["price"].shift(1)
return out
auditor = TemporalAudit(mode="nullify", random_seed=42)
report = auditor.check(df, timestamp_col="ts", pipeline_fn=causal_features)
print(report)
# โ CLEAN โ leakage_score=0.0000
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
# โ LEAKING PIPELINE
# โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
def leaking_features(df: pd.DataFrame) -> pd.DataFrame:
out = df.copy()
# center=True: peeks at future rows โ
out["centred_roll"] = out["price"].rolling(
11, center=True, min_periods=1
).mean()
return out
try:
auditor.check(df, timestamp_col="ts", pipeline_fn=leaking_features)
except TemporalLeakageError as exc:
print(exc)
# TemporalLeakageError: leakage_score=0.4812
# Breached columns (1):
# โข [HIGH] column='centred_roll' effect_size=0.4812from temporal_leaks import temporal_audit
@temporal_audit(timestamp_col="ts", mode="noise", random_seed=42)
def build_features(df: pd.DataFrame) -> pd.DataFrame:
df = df.copy()
df["expanding_mean"] = df["price"].expanding(min_count=1).mean()
return df
# Audit runs automatically on every call
result = build_features(df) # raises TemporalLeakageError if leak detectedEvery audit generates a beautiful, standalone HTML report you can share or archive:
report = auditor.check(df, "ts", leaking_features)
# Write a standalone HTML report
with open("audit_report.html", "w") as f:
f.write(report.to_html())The report includes:
HTML reports are great for sharing audit results with your team, archiving for compliance, or including in documentation.
TemporalAudit(
mode: Literal["noise", "sign_flip", "nullify"] = "noise",
random_seed: int = 42,
delta_threshold: float = 1e-8,
leakage_threshold: float = 0.0,
ignore_columns: list[str] | None = None,
)| Parameter | Type | Description |
|---|---|---|
mode |
str |
noise: adds Gaussian noise ยท sign_flip: multiplies by โ1 ยท nullify: sets NaN |
random_seed |
int |
Integer seed for determinism. Fully reproducible across runs. |
delta_threshold |
float |
Minimum cell-level change to count as "different". Suppresses float noise. |
leakage_threshold |
float |
If leakage_score > leakage_threshold, raise error. Set to 1.1 to always return report. |
ignore_columns |
list[str] |
Output columns to skip during comparison (e.g., timestamps, IDs). |
Adds Gaussian noise to future rows: ฮผ=0, ฯ=2รcolumn_std. Safe for pipelines that handle noise.
Multiplies numeric values by โ1. Good for testing sign-sensitive factors (momentum, etc).
Replaces future values with NaN. The strictest test โ use this first.
| Severity | Effect Size | Interpretation |
|---|---|---|
| ๐ฆ LOW | effect_size < 0.15 |
Minimal leakage. May be noise or minor bug. |
| ๐จ MEDIUM | 0.15 โค effect_size < 0.40 |
Noticeable leakage. Investigate and fix. |
| ๐ง HIGH | 0.40 โค effect_size < 0.75 |
Significant leakage. Feature is compromised. |
| ๐ฅ CRITICAL | effect_size โฅ 0.75 |
Severe leakage. Feature is unreliable. |
Full support for pd.DataFrame. Results returned as Pandas DataFrames.
First-class Polars support. Recommended for large frames (10M+ rows). Results returned as Polars DataFrames.
Measured on Apple M2 Pro, 16 GB RAM:
| Dataset | Rows | Columns | Backend | Mode | Time |
|---|---|---|---|---|---|
| Synthetic prices | 1,000,000 | 5 | Polars | nullify | ~1.1 s |
| Synthetic prices | 10,000,000 | 5 | Polars | nullify | ~3.2 s |
| Equity features | 500,000 | 20 | Pandas | noise | ~2.8 s |
Performance tip: For large datasets (>100k rows), use Polars. It's typically 5-10x faster than Pandas for temporal operations.
# Install dev extras
pip install -e ".[dev]"
# Run full test suite
pytest tests/ -v
# With coverage
pytest tests/ --cov=temporal_leaks --cov-report=term-missingThe codebase is maintained to strict standards:
Type Checking
mypy
strict mode
Linting
ruff
all rules
CI/CD
GitHub
Actions
Pull requests are welcome! Please make sure ruff check . and mypy temporal_leaks/ pass before submitting.
No manual inspection. No guessing. The tool finds leaks automatically.
Identifies exactly which columns leak, where, and why. HTML reports ready to share.
Reproducible across runs. Fixed seeds. No randomness in the results.
Pure Python. Works with Pandas, Polars, or your own dataframe.
One decorator. Audit runs automatically on every call. Ship with confidence.
First-class Polars support. Lightning-fast on large datasets.
temporal-leaks โ Valgrind for Time-Series ML
MIT License ยท Python 3.9+ ยท Pandas & Polars ยท Open Source
"Catch the bugs before they cost you."