OPEN STREAM STORAGE FOR THE LAKEHOUSE
Streams on object storage.
Tables by materialization.
Lakestream is an open API and specification for stream storage on object storage, with a stream materialization framework that defines how a stream becomes a lakehouse table. Ursa implements it as a storage engine; Ursa for Apache Kafka is a Kafka distribution built on it. Both are open source under Apache 2.0.
WHERE LAKESTREAM FITS
One open foundation for streaming and analytics.
Lakestream is an open standard for stream storage on object storage, and it defines how a stream becomes a table. Open table formats gave the analytical side a layout several engines can read; the log side had no open equivalent, so brokers kept their own formats and exported through connectors. With both layers open, streaming and analytical systems share one foundation — a Streamhouse and a lakehouse on the same object storage.
Streamhouse is an open, vendor-neutral category maintained by the Streamhouse Working Group. Lakestream is one way to build its stream-storage layer.
LAKESTREAM AND THE STREAMHOUSE →STREAM–TABLE DUALITY
One retained log. Two outputs.
One compaction pass over the retained log serves both paths. The compacted objects it writes are Parquet: a stream reader replays them, and they can be registered as a lakehouse table and queried. Materializing into an external table is a separate output of the same pass — not a second pipeline reading the first.
WRITE — A STREAM
append-ordered records · WAL objects
ONE LIFECYCLE
OBJECT STORAGE
READ — A TABLE
columnar Parquet · Iceberg / Delta
topic → WAL objects → compacted Parquet → committed tables
PRINCIPLES
What the storage layer does.
STREAM–TABLE DUALITY
One log, read two ways
Tail it as a stream or query it as a table. The storage layer produces both from the same log.
LEARN MORE →DISKLESS & LEADERLESS
Brokers keep no partition data
Durability lives in object storage, not on broker disks. Replacing or scaling a broker moves no data.
LEARN MORE →ZERO-ETL
No connector to run
Tables are written by the same compaction that stores the stream, not by a separate pipeline you deploy and operate.
LEARN MORE →ARCHITECTURE
The layers, from protocol to object storage.
Each layer depends only on the one beneath it. A protocol layer speaks Kafka; Lakestream defines the API it calls and the format its records land in on object storage; Ursa implements both and runs the compaction that rewrites WAL objects into Parquet for a lakehouse table catalog to commit.
PROTOCOL LAYER
LAKESTREAM — API + STORAGE FORMAT
URSA — THE IMPLEMENTATION
OBJECT STORAGE & LAKEHOUSE
Lakestream specifies the API and the storage format; Ursa implements both.
THE SPECIFICATION
A storage format, a materialization framework, and an API.
The Lakestream Storage Spec defines how a stream's logs are laid out in object storage and indexed in a metadata store, at format version 3. The Materialization Spec defines how a stream becomes a lakehouse table. lakestream-api is the Java API that applications and protocol layers call.
lakestream-api · embedded in Ursa 1.0.0
IMPLEMENTATIONS
Two runtimes implement the specification.
Both are open source under Apache 2.0, with their limits stated in the docs. The verification harness and protocol models are listed with them on the projects page.
Ursa
STORAGE ENGINE · 1.0StreamNative's implementation of the spec: WAL objects on object storage, compaction to Parquet, materialization to Iceberg and Delta. Java, embeddable, and the subject of the VLDB 2025 Best Industry Paper.
Ursa for Apache Kafka (UFK)
KAFKA DISTRIBUTION · 4.3.1A Kafka distribution built on the Lakestream API and specification, whose diskless topics store records through Ursa: brokers hold no partition data, any broker is eligible to serve a partition, and compaction writes Parquet that can be queried as a lakehouse table. One config flag per topic.
RESEARCH · VLDB 2025 Best Industry Paper
Ursa: A Lakehouse-Native Data Streaming Engine for Kafka
PVLDB 18(12). The design behind Ursa.
QUICKSTART
Diskless Kafka on your laptop.
Three brokers, Oxia, MinIO, and the compactor from one Compose stack. Create a topic whose records live in object storage, not on a broker disk, then query it as an Iceberg table.