Lakestream

OPEN STREAM STORAGE FOR THE LAKEHOUSE

Streams on object storage. Tables by materialization.

Lakestream is an open API and specification for stream storage on object storage, with a stream materialization framework that defines how a stream becomes a lakehouse table. Ursa implements it as a storage engine; Ursa for Apache Kafka is a Kafka distribution built on it. Both are open source under Apache 2.0.

WHERE LAKESTREAM FITS

One open foundation for streaming and analytics.

Lakestream is an open standard for stream storage on object storage, and it defines how a stream becomes a table. Open table formats gave the analytical side a layout several engines can read; the log side had no open equivalent, so brokers kept their own formats and exported through connectors. With both layers open, streaming and analytical systems share one foundation — a Streamhouse and a lakehouse on the same object storage.

A Streamhouse, holding data to run the business, and a lakehouse, holding data to analyze the business, share one governance layer and one open infrastructure layer on object storage, where open stream storage (Lakestream) sits beside open table formats (Iceberg and Delta) and streams materialize into tables.

Streamhouse is an open, vendor-neutral category maintained by the Streamhouse Working Group. Lakestream is one way to build its stream-storage layer.

LAKESTREAM AND THE STREAMHOUSE →

STREAM–TABLE DUALITY

One retained log. Two outputs.

One compaction pass over the retained log serves both paths. The compacted objects it writes are Parquet: a stream reader replays them, and they can be registered as a lakehouse table and queried. Materializing into an external table is a separate output of the same pass — not a second pipeline reading the first.

WRITE — A STREAM

append-ordered records · WAL objects

ONE LIFECYCLE

OBJECT STORAGE

READ — A TABLE

columnar Parquet · Iceberg / Delta

topic → WAL objects → compacted Parquet → committed tables

PRINCIPLES

What the storage layer does.

STREAM–TABLE DUALITY

One log, read two ways

Tail it as a stream or query it as a table. The storage layer produces both from the same log.

LEARN MORE →

DISKLESS & LEADERLESS

Brokers keep no partition data

Durability lives in object storage, not on broker disks. Replacing or scaling a broker moves no data.

LEARN MORE →

ZERO-ETL

No connector to run

Tables are written by the same compaction that stores the stream, not by a separate pipeline you deploy and operate.

LEARN MORE →

ARCHITECTURE

The layers, from protocol to object storage.

Each layer depends only on the one beneath it. A protocol layer speaks Kafka; Lakestream defines the API it calls and the format its records land in on object storage; Ursa implements both and runs the compaction that rewrites WAL objects into Parquet for a lakehouse table catalog to commit.

PROTOCOL LAYER

Kafka

LAKESTREAM — API + STORAGE FORMAT

StreamCatalogLogLogCursorTableMaterializationPolicyFormat v3

URSA — THE IMPLEMENTATION

WAL objectsCompactionMaterializersOxia metadata

OBJECT STORAGE & LAKEHOUSE

S3GCSAzureIcebergDelta

Lakestream specifies the API and the storage format; Ursa implements both.

THE SPECIFICATION

A storage format, a materialization framework, and an API.

The Lakestream Storage Spec defines how a stream's logs are laid out in object storage and indexed in a metadata store, at format version 3. The Materialization Spec defines how a stream becomes a lakehouse table. lakestream-api is the Java API that applications and protocol layers call.

STORAGE SPECWAL objects, compacted objects, offset index
MATERIALIZATIONPolicy, table catalogs, resolution
JAVA APIStreamCatalog, Log, LogCursor

lakestream-api · embedded in Ursa 1.0.0

StreamCatalog catalog =
new StreamCatalogService().open(uri, props);
StreamIdentifier id =
StreamIdentifier.of("default", "orders");
 
catalog.openWriter(id).thenCompose(w ->
w.write(RoutingKey.roundRobin(), 1, payload));
// durable in a WAL object, readable as a stream
 
catalog.openReader(id).thenCompose(r ->
r.read(logId, offset, 100, 1_000_000));
// the same records, in order
 
// compaction materializes them as an Iceberg or Delta table

IMPLEMENTATIONS

Two runtimes implement the specification.

Both are open source under Apache 2.0, with their limits stated in the docs. The verification harness and protocol models are listed with them on the projects page.

Ursa

STORAGE ENGINE · 1.0

StreamNative's implementation of the spec: WAL objects on object storage, compaction to Parquet, materialization to Iceberg and Delta. Java, embeddable, and the subject of the VLDB 2025 Best Industry Paper.

Ursa for Apache Kafka (UFK)

KAFKA DISTRIBUTION · 4.3.1

A Kafka distribution built on the Lakestream API and specification, whose diskless topics store records through Ursa: brokers hold no partition data, any broker is eligible to serve a partition, and compaction writes Parquet that can be queried as a lakehouse table. One config flag per topic.

ALL PROJECTS →

RESEARCH · VLDB 2025 Best Industry Paper

Ursa: A Lakehouse-Native Data Streaming Engine for Kafka

PVLDB 18(12). The design behind Ursa.

READ THE PAPER

QUICKSTART

Diskless Kafka on your laptop.

Three brokers, Oxia, MinIO, and the compactor from one Compose stack. Create a topic whose records live in object storage, not on a broker disk, then query it as an Iceberg table.

$ make up
$ make create-topic
$ make produce && make consume
$ make destroy && make lakehouse-demo
# Avro orders → Iceberg → DuckDB