Documentation • Source Code • PyPI
Distance Metrics • Clustering • Forecasting • Imaging • Changepoints • Preprocessing
polars-ts is a batteries-included time series toolkit built on Polars. It gives you Rust-accelerated distance metrics, 10+ clustering algorithms, a full forecasting stack (classical → deep → foundation-model → agentic), Bayesian and causal inference, streaming/online learning, and multi-agent domain applications — all from a single pip install, no heavyweight frameworks required.
| Pain point | How polars-ts helps |
|---|---|
| "I need DTW but scipy is slow" | 12 distance metrics compiled to native code via Rust + Rayon, orders of magnitude faster on large panels |
| "I want to cluster time series but tslearn/sktime have too many deps" | K-Medoids, K-Shape, HDBSCAN, Spectral, Hierarchical, K-Means DBA, KASBA, CLARA/CLARANS, U-Shapelets, Contrastive, DEC/IDEC — all built-in |
| "Setting up a forecast pipeline takes too long" | ForecastPipeline wires up lags, rolling stats, calendar features, target transforms, and any sklearn model in 5 lines |
| "I want to use foundation models / LLMs for forecasting" | Chronos, TimesFM, Moirai adapters for zero-shot; Time-LLM, LLM-PS for reprogrammed LLM forecasting; N-BEATS, PatchTST native DL |
| "I need Bayesian methods" | Kalman filters, BSTS, Bayesian ETS, Bayesian VAR, GP regression, MCMC wrapper, particle filters — all built-in |
| "I want multi-agent forecasting for my domain" | Hierarchical agent systems for energy, supply chain, healthcare, and industrial IoT, plus portfolio MARL and RL anomaly detection |
| "I don't know which clustering method to pick" | auto_cluster sweeps methods × distances × k values and returns the best result with evaluation scores |
| "Polars doesn't have time series functions" | Mann-Kendall, Sen's slope, CUSUM, PELT, decomposition, ACF/PACF — all group-aware and Polars-native |
import numpy as np
import polars as pl
import polars_ts as pts
from datetime import datetime, timedelta
# Build a small panel of 20 hourly series
dates = [datetime(2023, 1, 1) + timedelta(hours=i) for i in range(100)]
df = pl.concat([
pl.DataFrame({
"unique_id": [f"s_{i}"] * 100,
"ds": dates,
"y": (np.sin(np.linspace(0, 4 * np.pi, 100)) + np.random.normal(0, 0.1, 100)).tolist(),
})
for i in range(20)
])
# Cluster by shape similarity
result = pts.auto_cluster(df, methods=["kmedoids", "spectral"], distances=["sbd", "dtw"])
print(result.best_labels) # DataFrame[unique_id, cluster]
# Forecast with a full ML pipeline
from sklearn.linear_model import Ridge
pipe = pts.ForecastPipeline(Ridge(), lags=[1, 7, 14], rolling_windows=[7], calendar=["day_of_week"])
pipe.fit(df)
forecasts = pipe.predict(df, h=7)
# Detect changepoints
breaks = pts.pelt(df, cost="meanvar", penalty=10)Column defaults are
unique_id,ds,ythroughout. Passid_col=,time_col=,target_col=to override.
pip install polars-timeseriesExtras for optional features:
pip install "polars-timeseries[clustering]" # HDBSCAN, DBSCAN, spectral (sklearn + scipy)
pip install "polars-timeseries[forecast]" # SCUM, auto_arima (statsforecast)
pip install "polars-timeseries[decomposition]" # Fourier decomposition (polars-ds)
pip install "polars-timeseries[all]" # EverythingRequires Python 3.12+ and Polars 1.30+.
import polars as pl
import polars_ts as pts
df = pl.DataFrame({
"unique_id": ["A"] * 5 + ["B"] * 5,
"y": [1.0, 2.0, 3.0, 2.0, 1.0,
1.0, 3.0, 5.0, 3.0, 1.0],
})
result = pts.compute_pairwise_dtw(df, df)result = pts.auto_cluster(
df,
methods=["kmedoids", "spectral", "kshape"],
distances=["sbd", "dtw"],
k_range=range(2, 6),
)
print(result.best_method, result.best_k, result.best_score)
print(result.best_labels) # DataFrame[unique_id, cluster]from sklearn.ensemble import GradientBoostingRegressor
import polars_ts as pts
pipe = pts.ForecastPipeline(
GradientBoostingRegressor(),
lags=[1, 2, 7],
rolling_windows=[7],
calendar=["day_of_week", "month"],
target_transform="log",
)
pipe.fit(train_df)
forecasts = pipe.predict(train_df, h=7)
# forecasts: DataFrame[unique_id, ds, y_hat]pipe = pts.ForecastPipeline(
Ridge(),
lags=[1, 2, 7],
past_covariates=["temperature"], # lagged automatically
future_covariates=["is_holiday"], # looked up from future_df
)
pipe.fit(train_df) # train_df has temperature + is_holiday columns
forecasts = pipe.predict(train_df, h=7, future_df=future_df)# Fit ARIMA(1,1,1) and forecast 12 steps ahead
fitted = pts.arima_fit(df, order=(1, 1, 1))
forecast = pts.arima_forecast(fitted, h=12)
# Or use automatic order selection
forecast = pts.auto_arima(df, h=12, season_length=12)# PELT — multiple changepoints with mean/variance cost
breaks = pts.pelt(df, cost="meanvar", penalty=10)
# breaks: DataFrame[unique_id, changepoint_idx, ds]
# Bayesian Online Changepoint Detection
probs = pts.bocpd(df)# Holt-Winters seasonal forecast
result = pts.holt_winters_forecast(df, h=12, season_length=12, seasonal="additive")import polars as pl
import polars_ts as pts
df = pl.DataFrame({
"group": ["A"] * 10 + ["B"] * 10,
"y": list(range(10)) + [10 - x for x in range(10)],
})
result = df.group_by("group").agg(
pts.mann_kendall(pl.col("y")).alias("trend"),
pts.sens_slope(pl.col("y")).alias("slope"),
)df = pl.DataFrame({
"unique_id": ["A"] * 48,
"ds": list(range(48)),
"y": [10 + 5 * (i % 12 > 5) + 0.5 * i for i in range(48)],
})
result = pts.seasonal_decomposition(df, freq=12, method="additive")All distance functions return a tidy DataFrame with columns [id_1, id_2, <metric>]. A unified compute_pairwise_distance(method=...) API lets you swap metrics with a single string.
| Metric | Function | Key Parameters |
|---|---|---|
| Dynamic Time Warping | compute_pairwise_dtw |
method: standard, sakoe_chiba, itakura, fast |
| Derivative DTW | compute_pairwise_ddtw |
Shape-sensitive comparison |
| Weighted DTW | compute_pairwise_wdtw |
g: weight sharpness |
| Move-Split-Merge | compute_pairwise_msm |
c: move cost |
| Edit Distance (Real Penalty) | compute_pairwise_erp |
g: gap value |
| Longest Common Subsequence | compute_pairwise_lcss |
epsilon: matching threshold |
| Time Warp Edit Distance | compute_pairwise_twe |
nu: stiffness, lambda_: deletion cost |
| Shape-Based Distance | compute_pairwise_sbd |
Cross-correlation based |
| Frechet Distance | compute_pairwise_frechet |
Geometric coupling distance |
| Edit Distance on Real Sequences | compute_pairwise_edr |
Edit-operation cost |
| Multivariate DTW | compute_pairwise_dtw_multi |
metric: manhattan, euclidean |
| Multivariate MSM | compute_pairwise_msm_multi |
c: move cost |
| Method | Function | When to use |
|---|---|---|
| K-Medoids (PAM) | kmedoids |
Known k, any distance metric, interpretable medoids |
| K-Shape | KShape |
Shape-based grouping via cross-correlation centroids |
| Spectral (KSC) | spectral_cluster |
Non-convex clusters, graph Laplacian structure |
| Contrastive | ContrastiveClusterer |
Self-supervised representation learning + clustering |
| DEC / IDEC | DECClusterer |
Autoencoder-based deep clustering |
| HDBSCAN | hdbscan_cluster |
Unknown k, varying density, noise detection |
| DBSCAN | dbscan_cluster |
Fixed-radius neighbourhood, noise detection |
| Hierarchical | agglomerative_cluster |
Dendrogram visualization, flexible linkage |
| K-Means DBA | kmeans_dba |
DTW Barycentric Averaging centroids |
| KASBA | kasba, KASBAClusterer |
k-means-accelerated Rust clustering for large panels |
| CLARA | clara |
Scalable k-medoids via sampling |
| CLARANS | clarans |
Randomized k-medoids neighbourhood search |
| U-Shapelets | shapelet_cluster |
Interpretable sub-sequence patterns |
| ROCKET / MiniRocket | rocket_features, minirocket_features |
Random convolutional kernel feature extraction |
| Auto-cluster | auto_cluster |
Sweep methods × distances × k, pick the best |
Evaluation: silhouette_score, davies_bouldin_score, calinski_harabasz_score
Classification: knn_classify (distance-based k-NN), TimeSeriesKNNClassifier (OOP), KShapeClassifier (centroid-based), RocketClassifier, InceptionTimeClassifier, ResNetClassifier (deep)
- Mann-Kendall test — non-parametric trend detection (Rust)
- Sen's slope — robust trend magnitude estimation (Rust)
- CUSUM — cumulative sum changepoint detection (Rust)
- PELT — multiple changepoints with mean/variance/meanvar cost functions
- BOCPD — Bayesian Online Changepoint Detection
- Regime detection — Hidden Markov Model state inference
- Seasonal decomposition — additive or multiplicative (classical)
- Fourier decomposition — harmonic decomposition with configurable frequencies
- Decomposition features — trend/seasonal strength extraction (simple or MSTL)
- Anomaly flagging — residual-based anomaly detection from any decomposition
- Lag features — create lagged versions of a target column per group
- Rolling features — rolling window aggregations (mean, std, min, max, sum, median, var)
- Calendar features — extract day_of_week, month, quarter, is_weekend, etc.
- Fourier features — sin/cos pairs for seasonal modelling
- Target encoding — smoothed categorical encoding by target mean
- Holiday features — binary holidays + distance-to-holiday (requires
holidayspackage) - Interaction features — cross-term column generation
- Time embeddings — cyclical sin/cos encoding for time components
- Log transform — log1p / expm1 with automatic validation and lossless inversion
- Box-Cox transform — parametric power transform with configurable lambda
- Differencing — configurable order and seasonal period with metadata for lossless inversion
All transforms are group-aware, invertible, and accessible via the df.pts namespace.
- Missing value imputation — forward/backward fill, linear interpolation, mean, median, seasonal
- Outlier detection — z-score, IQR, Hampel filter, rolling z-score
- Outlier treatment — clip (winsorize), median replacement, interpolation, null
- Temporal resampling — downsample/upsample with configurable aggregation
- Expanding window CV — growing training window cross-validation
- Sliding window CV — fixed-size training window cross-validation
- Rolling origin CV — general rolling-origin with configurable initial/fixed train size
- SCUM — ensemble model combining AutoARIMA, AutoETS, AutoCES, and DynamicOptimizedTheta
- ARIMA/SARIMA — explicit
(p,d,q)order viastatsmodels(arima_fit/arima_forecast) or automatic selection viastatsforecast(auto_arima) - Baseline models — naive, seasonal naive, moving average, and FFT-based forecasts
- Exponential smoothing — SES, Holt's linear, Holt-Winters (additive/multiplicative, Rust-accelerated)
- Multi-step strategies —
RecursiveForecasterandDirectForecaster - ForecastPipeline — end-to-end ML pipeline with feature engineering + transforms
- GlobalForecaster — cross-series panel model with optional ID encoding
- N-BEATS / PatchTST — native deep learning forecasters (
NBEATSForecaster,PatchTSTForecaster, requirestorch) - Native multivariate deep —
MultivariatePatchTST,iTransformerForecasterfor jointly forecasting correlated series - Time-LLM / LLM-PS — LLM-reprogrammed forecasting adapters (requires
torch) - Foundation models — Chronos, TimesFM, Moirai zero-shot forecasting
- Agentic forecasting —
TimeSeriesScientistmulti-agent pipeline (Curator → Planner → Forecaster → Reporter)
- QuantileRegressor — one model per quantile level with CRPS-compatible output
- Conformal prediction — distribution-free intervals with coverage guarantees
- EnbPI — Ensemble Batch Prediction Intervals with adaptive online updates
- WeightedEnsemble — equal, manual, or inverse-error-optimized weights
- StackingForecaster — meta-learner trained on out-of-fold predictions
- Metrics — MAE, RMSE, MAPE, sMAPE, MASE, CRPS
- Kaboudan metric — model robustness evaluation via block-shuffle backtesting
- Bias detection & correction — mean, regression, quantile mapping
- Calibration diagnostics — calibration table, PIT histogram, reliability diagram
- Residual diagnostics — ACF, PACF, Ljung-Box test
- Permutation importance — model-agnostic feature importance
- VAR — Vector Autoregression with OLS fitting and multi-step forecasts
- Granger causality — F-test for causal relationships between series
- GARCH — volatility modelling and conditional variance forecasting
- Forecast reconciliation — bottom-up, top-down, middle-out, MinTrace (OLS/shrink/covariance/CV), and PERMBU
- Grouped / cross-sectional hierarchies — reconcile across multiple non-nested grouping dimensions
- Kalman Filter / RTS Smoother — linear state-space models
- BSTS — Bayesian Structural Time Series (local level/trend + seasonality + regressors)
- Bayesian ETS — exponential smoothing with posterior predictive
- Bayesian VAR — vector autoregression with Minnesota/Normal-Wishart priors
- Gaussian Process regression — non-parametric Bayesian forecasting
- MCMC wrapper — adapter for NumPyro/PyMC backends
- Unscented / Ensemble Kalman Filter — nonlinear state-space models
- Particle Filter / SMC — fully nonlinear/non-Gaussian
- Bayesian anomaly scoring — posterior predictive p-values, Bayes factors
- CausalImpact — Bayesian structural time series counterfactual analysis
- Synthetic control — donor-pool weighted counterfactual estimation
- Placebo tests — significance testing via counterfactual permutation
- Unified backtest framework —
backtestandcompare_modelsfor fit → predict → score across rolling windows with multiple models
- StreamingETS — online exponential smoothing (SES, Holt, Holt-Winters) with single-observation updates
- StreamingKalmanFilter — incremental Kalman filtering for state-space models
- StreamingGlobalForecaster — incremental panel model for SGD-compatible estimators
- SlidingWindowManager — memory-efficient windowed state management
- Decomposition-based — residual threshold anomaly flagging
- Isolation Forest — unsupervised anomaly detection on engineered features
- RL anomaly agents — multi-agent consensus (
AnomalyOrchestratorover Z-score, rolling-std, and MAD agents) for adaptive thresholding
- Portfolio MARL — cooperating risk, return, and allocation agents (
MARLOrchestrator,PortfolioEnv) for decision optimization - ForecastEnv / PortfolioEnv / AnomalyEnv — Gymnasium-compatible environments for RL over time series
Hierarchical multi-agent systems for specific verticals, built on the core forecasting stack:
- Energy / demand —
EnergyGridOrchestrator: regional → grid → household hierarchy with weather, renewable, and demand-response agents - Supply chain —
SupplyChainOrchestrator: demand sensing, promotion effects, inventory, and multi-echelon coordination - Healthcare —
ClinicalOrchestrator: EHR/vital monitoring, sepsis warning, escalation, and treatment agents (with federated averaging) - Industrial IoT —
MaintenanceOrchestrator: spectral features, health index, RUL estimation, and predictive maintenance scheduling
- ModelRegistry — version, persist, and load models with metadata
- Experiment / Run — track parameters, metrics, and artifacts across runs
- Imaging transforms — GASF/GADF (
to_gasf,to_gadf), Markov Transition Field (to_mtf), recurrence plots (to_recurrence_plot), spectrogram/scalogram, and path-signature images - RQA & signature features — recurrence quantification (
rqa_features) and signature features for downstream models - Vision embeddings —
extract_vision_embeddingsfor pretrained-CNN features over imaged series
- NeuralForecast — convert to/from N-BEATS, PatchTST, N-HiTS format
- PyTorch Forecasting — convert to/from TFT, DeepAR format
- HuggingFace — convert to Dataset for Chronos, TimesFM, Lag-Llama
- Chronos / MOMENT embeddings — foundation model feature extraction for clustering
- Foundation forecast — ChronosForecaster, TimesFMForecaster, MoiraiForecaster (zero-shot)
- LLM forecast — TimeLLMForecaster, LLMPSForecaster (reprogrammed LLM)
- ForecastEnv — Gymnasium-compatible RL environment for decision making
The notebooks/ directory contains 14 end-to-end tutorials:
| # | Topic | Notebook |
|---|---|---|
| 01 | Data wrangling & exploration | Open |
| 02 | Feature engineering & transforms | Open |
| 03 | Forecasting fundamentals | Open |
| 04 | ML forecasting pipelines | Open |
| 05 | Uncertainty & calibration | Open |
| 06 | Changepoint & anomaly detection | Open |
| 07 | Time series similarity & clustering | Open |
| 08 | Multivariate & volatility | Open |
| 09 | Ensembles & reconciliation | Open |
| 10 | Ecosystem adapters | Open |
| 11 | Time series imaging | Open |
| 12 | Advanced feature extraction | Open |
| 13 | Agentic forecasting | Open |
| 14 | KASBA clustering | Open |
git clone https://github.com/drumtorben/polars-ts.git
cd polars-ts
uv sync
uv pip install -e .
uv run pytestPre-commit hooks run via prek (Rust reimplementation of pre-commit) or standard pre-commit — both read .pre-commit-config.yaml:
# Option A: prek (faster)
uv tool install prek
prek run --all-files
# Option B: standard pre-commit
pre-commit run --all-files# mypy (authoritative)
uv run mypy polars_ts/
# ty (fast, informational — beta)
uvx ty check polars_ts/MIT
