Lakestream
Concepts

Architecture

A lakestream's write, read and materialization paths run through stream metadata, shared WAL storage, offset indexes and background compaction.

A lakestream separates stream metadata, durable log storage, and background processing. In Ursa these are cooperating libraries and services, not a requirement to deploy exactly three standalone servers.

Stream catalog and layout

The StreamCatalog API manages namespaces, streams, layouts, and materialization policies. Creating or loading a stream returns a StreamMetadata snapshot; it does not open a reader or writer. Data-plane handles are opened explicitly.

A stream has a StreamIdentifier and a layout of LogId values. Ursa's IndexedLayout selects a log by explicit index or round-robin routing. Each log has its own offset space. The older storage internals also call a numeric log ID a "stream ID"; that identifier must not be confused with the named, multi-log stream in the public API.

Ursa uses Oxia for durable metadata, including storage indexes and offset sequencing. The public catalog API and the storage index are related parts of the system, but catalog lookup itself is not the operation that allocates an offset for each record.

Append path and WAL objects

In Ursa's object-backed WAL path:

  1. A write selects a log and submits an entry with its record count and payload.
  2. The WAL buffers entries from multiple logs and flushes them into a shared raw object through FileStorage.
  3. After the object flush succeeds, the WAL publishes index entries in Oxia. For ordinary appends, Oxia sequence-key deltas allocate offset ranges using the number of records, while also tracking cumulative size.
  4. The append completes with the assigned entry metadata after the flush and index result have been processed.

Buffer admission alone is not a durable append acknowledgment. Offset allocation coordinates concurrent appends through shared metadata; it does not rely on a broker replicating a local partition log to followers.

The current WalStorageFactory constructs ObjectWalStorageImpl. Ursa includes S3, GCS, Azure, and local-file storage implementations. A local-file deployment is not equivalent to a shared, durable object-store deployment, and alternative WAL systems should not be inferred from the general pattern.

Read path and the Stream Offset Index

The Stream Offset Index maps a logical log-offset range to physical file information. Ursa's unified reader inspects the indexed file type:

  • RAW reads use the storage API.
  • PARQUET reads use a compacted-object reader capable of reconstructing stream entries.

A compacted-object reader must be installed for the relevant data format; Parquet support is not implied by opening an arbitrary low-level raw storage handle. Ursa's Kafka integration includes a Kafka lakehouse reader for this purpose, currently limited to its V2 layout, which records a CompactedObjectFileIndex in each entry index, and Kafka batched-raw Parquet encoding; it rejects the V1 read path. This is not an arbitrary-Parquet-to-Kafka adapter.

A cursor tracks consumption and acknowledgment separately from this index. The index locates data; it does not itself store every consumer's progress. Likewise, a SQL engine discovers data through Iceberg or Delta table metadata, not by consulting the Stream Offset Index.

Compaction, materialization, and table commit

Compaction consolidates data from WAL objects into per-log compacted objects, columnar Parquet in the Kafka path. External materialization decodes the same range and writes destination files using the configured schema and destination policy.

File production and table visibility are distinct stages. In the lakehouse path, LakehouseTableMaterializer.commit() flushes its writer and collects file results — despite its name, it does not commit the external table catalog. Task completion persists a COMPACTED task, and the downstream group-commit pipeline commits the table metadata.

For internal compacted objects, task completion records each file's offset range and per-file index, and the corresponding storage-index range is replaced with the compacted-file location. These are recoverable steps, not one atomic transaction across Oxia and the table catalog.

An append flushes a raw WAL object, then publishes its index in Oxia, which assigns the offset range and completes the append, and stream reads follow from there. One compaction and materialization pass over the same WAL range then splits into two lanes: internal compacted objects whose offset ranges replace the storage-index entries for stream reads, and materialized columnar files that are persisted as a COMPACTED task and become table scans only after the table catalog commit.

Stream–table duality explains how the two read paths relate. Zero-ETL explains the configuration and visibility boundaries.

Retention and lifecycle safety

Retention is not just a compaction timer. LogStorage distinguishes an inclusive soft trim, which logically removes an offset range, from an exclusive hard trim, which requests physical removal. Shared-file cleanup and table lifecycle processing must respect the data still referenced by other logs or tables.

The Kafka integration also schedules retention checks from Kafka topic metadata. Stream deletion, open-handle leases, table ownership, and eventual file cleanup are separate concerns; compaction is not the sole owner of every deletion policy.