High-performance feature store built on Delta Lake and Ray
Gamma Lake is a feature store built on Delta Lake and Ray, designed for efficient storage, versioning, and retrieval of time-indexed features backed by Polars DataFrames.
It was built to solve the real-world pain points that teams encounter when using flat Parquet files for ML feature storage: expensive column additions, no versioning, and painful cross-team collaboration.
The most common alternative — storing features in per-day Parquet files — breaks down quickly:
| Problem | Gamma Lake's answer |
|---|---|
| Adding a new feature group rewrites all existing files | Gamma Lake writes one new Delta table per feature group — O(1) cost regardless of existing data |
| Multiple teams can't write features in parallel | Each feature group is an independent Delta table — concurrent writes don't conflict |
| No versioning or audit trail | Delta Lake's transaction log provides full time-travel and version history |
| Cross-group reads require expensive joins | Gamma Lake maintains a master index; reads are aligned horizontal concatenations |
| Experimental features pollute production data | Features carry owner and version metadata; read() accepts a filtered metadata frame |
See performance for benchmark data comparing Gamma Lake against per-day Parquet at scale.
- Sortable index — a master index table ensures all feature groups stay aligned for efficient range-filtered reads
- Parallel reads and writes — Ray remote functions parallelise across feature groups for large-scale workloads
- Feature versioning — every
add_features()call records owner and version; read any historical version by filtering the metadata frame - Multiple signal types — standard features, as-of features, sparse features, and runtime-computed features
- Native Delta Lake storage — ACID transactions, schema enforcement, and time-travel out of the box
- Local or cloud —
base_pathaccepts a local directory, an on-prem object store path, or ans3://URI - Configurable compression —
zstdby default; pluggable viacompressionfield
import tempfile
from datetime import datetime, timezone
import polars as pl
import pyarrow as pa
from ccflow import ArrowSchema
from gammalake import GammaFeatureLake
# 1. Create and initialise a feature lake
with tempfile.TemporaryDirectory() as tmp:
lake = GammaFeatureLake(base_path=tmp, run_on_ray_cluster=False)
lake.initialize(
ArrowSchema.make(
pa.schema([("timestamp", pa.timestamp("us", tz="UTC")), ("symbol", pa.large_string())])
)
)
# 2. Build a toy feature DataFrame (timestamp × symbol × features)
now = datetime(2024, 1, 1, tzinfo=timezone.utc)
df = pl.DataFrame({
"timestamp": pl.Series([now]).cast(pl.Datetime("us", "UTC")),
"symbol": ["AAPL"],
"momentum": [0.42],
"vol_20d": [0.18],
})
# 3. Write features
lake.add_features(df, owner="my-team")
# 4. Read features back
result = lake.read(["momentum", "vol_20d"])
print(result)
# 5. Inspect metadata
print(lake.feature_metadata_frame().collect())Run the full annotated demo: examples/quickstart.py
from gammalake import GammaFeatureLake
lake = GammaFeatureLake(base_path="s3://my-bucket/features")
# One-time setup
lake.initialize(schema)
# Writing
lake.add_features(df, owner="team-a") # standard features
lake.add_targets(df, owner="team-a") # target/label columns
lake.add_as_of_features(df, params, owner=...) # point-in-time safe features
lake.add_sparse_features(df, owner=...) # sparse / infrequently-updated features
lake.add_runtime_computed_features(exprs, ...) # features computed at read time
# Reading
lake.read(["feature_a", "feature_b"]) # latest versions
lake.read(["feature_a"], start="2023-01-01", end="2024-01-01") # date range
lake.read(lake.feature_metadata_frame() # specific owner/version
.filter(pl.col("owner") == "team-a").collect())
# Metadata inspection
lake.feature_metadata_frame().collect() # features, versions, owners
lake.table_metadata_frame().collect() # Delta table details
lake.index_frame().collect() # master indexgamma-lake can be installed via pip or conda, the two primary package managers for the Python ecosystem.
To install gamma-lake via pip, run this command in your terminal:
pip install gamma-lakeTo install gamma-lake via conda, run this command in your terminal:
conda install gamma-lake -c conda-forgeSee our wiki!
Check out the contribution guide for more information.
Apache-2.0 — see LICENSE for details.