What is Lakestream?
Lakestream is an open API and specification for stream storage on object storage, with a stream materialization framework that defines how a stream becomes a lakehouse table.
Lakestream is an open API and specification for stream storage on object storage, with a stream materialization framework that defines how a stream becomes a lakehouse table. Ursa implements it as a storage engine; Ursa for Apache Kafka is a Kafka distribution built on it. Both are open source under Apache 2.0.
Its central idea is to produce the table from retained log data inside the storage lifecycle instead of exporting it through a downstream connector.
Naming convention
Capitalized "Lakestream" names the standard. Lowercase "a lakestream" names a system that conforms to it. Ursa is the storage engine that implements Lakestream; Ursa for Apache Kafka (UFK) is the Kafka distribution built on it.
The problem
A conventional streaming stack keeps an ordered log in brokers and exports the same events into a lakehouse. The connector becomes another system to deploy, recover, and keep aligned with schema changes. The broker log and the analytical table have separate storage lifecycles and visibility boundaries.
Lakestream moves table materialization into the storage lifecycle: one compaction pass over the retained log writes the compacted objects a stream reader replays and materializes the destination table. See why Lakestream for the motivation and trade-offs.
Streams are composed of logs
The API distinguishes a stream from a log:
- A stream is a named catalog entity with configuration, schema, lifecycle state, and a layout.
- Its layout maps writes to one or more logs. Ursa currently implements an indexed layout: an ordered list of logs addressed by partition index.
- Each log is an independently ordered, append-only sequence of entries. An entry carries one or more records and occupies a range of record offsets.
Ordering and offsets belong to a log, not to the stream as a whole. A multi-log stream has no single global offset. A Kafka topic maps naturally to a stream whose logs represent its partitions.
One storage lifecycle, two access paths
New entries first become durable in row-oriented write-ahead log (WAL) objects. Background compaction rewrites WAL ranges into per-log compacted objects that the stream offset index points to and, where a materialization policy resolves, decodes the same range into destination-table files and commits them to a table catalog. The stream offset index directs readers to raw or compacted data; query engines discover table files through table metadata instead.
The table side is produced by the stream materialization framework: the standard's policy model states which streams become which tables and how, and the implementation's materializers write and commit the files.
One lifecycle does not mean immediate table visibility
An acknowledged append can be readable from the log before it is visible in a table. Table access requires configured materialization, successful file generation, and a table commit. During conversion and cleanup, WAL and compacted files can coexist.
Three concepts explain the model:
- Stream–table duality — one retained log is consumed as a stream and materialized as a table by the same storage lifecycle. The two outputs are separate files; there is no shared-file guarantee for any destination.
- Diskless & leaderless — serving brokers do not need local durable partition replicas. Shared metadata still coordinates ordering, routing, and lifecycle safety.
- Zero-ETL — storage-managed materialization removes a separately operated stream-to-table connector, not decoding, rewriting, configuration, or commit latency.
Standard and implementation
The specification defines the storage format and the materialization policy model; the API reference documents the Java contract. These Concepts pages explain the model and identify implementation-specific behavior in the local Ursa and Kafka code, rather than treating every possible architecture or API option as an implemented feature.
Continue with architecture for the write, read, and materialization paths, or the glossary for terminology.