Why Lakestream
Lakestream addresses two forms of coupling: brokers that must store the partition data they serve, and lakehouse tables that require a separate export pipeline.
Lakestream addresses two forms of coupling: brokers that must store the partition data they serve, and lakehouse tables that require a separate export pipeline from those brokers.
Open table formats gave the lakehouse a standard for the table: a published layout for files, metadata, and snapshots that several engines can read. The log side had no open equivalent. Brokers kept their own storage formats and exported into the lakehouse through connectors, so the same events were written twice and governed twice. Lakestream defines the log side as an open API and spec, and defines the materialization from that log into tables as part of the same standard.
The cost of separate storage stacks
A conventional replicated broker keeps partition logs on leaders and followers. Scaling or replacing those nodes can involve moving retained data, and replication across availability zones can add network cost.
Exporting the same events to a lakehouse introduces another retained representation and another write path. A connector must handle offsets, schemas, retries, and table commits while the broker independently manages its log. This can be the right architecture, but it has an operational cost beyond the storage itself.
Put table production inside the storage lifecycle
A lakestream can replace that export path with storage-managed compaction and materialization. The compaction pass that writes the compacted objects a stream reader replays also materializes the destination table, so producing the table is part of the storage lifecycle rather than a second pipeline reading the log again. Shared durable storage also allows serving brokers to change without copying local partition replicas.
These are two separate benefits:
- Stream–table duality can remove a separately operated export pipeline for retained log data.
- Diskless & leaderless removes Kafka follower-payload replication from the diskless data path and decouples serving-node selection from file placement.
Neither removes the need for durable metadata, storage redundancy, or background processing.
Why standardize the model?
The shared vocabulary and API contract let clients distinguish streams, logs, offsets, layouts, and materialization policies without depending on a broker's local segment layout. Open table formats let compatible query engines consume the resulting table snapshots without becoming streaming clients.
The specification and the API reference describe that contract; Ursa and Ursa for Apache Kafka (UFK) show concrete implementations. A protocol integration still needs its own entry encoding, decoder, and reader support. An API type or an open file format alone does not prove arbitrary cross-implementation byte compatibility.
Trade-offs that remain
Table freshness still depends on materialization and catalog commit. WAL and compacted files can overlap during migration and cleanup. The delivered table is a second representation with its own files and lifecycle. Shared storage and Oxia become availability dependencies, and replacement brokers may need cache warming and producer-state recovery.
The goal is fewer independently operated storage and export paths, not a system with no coordination or maintenance. Zero-ETL explains what is removed and what remains configurable.
What is Lakestream?
Lakestream is an open API and specification for stream storage on object storage, with a stream materialization framework that defines how a stream becomes a lakehouse table.
Lakestream and the Streamhouse
Streamhouse names a category of data architecture; Lakestream is a standard for one layer inside it, and this page states the relationship and its limits.