Lakestream
Ursa for Kafka

Limitations

The current diskless implementation has compatibility and operational boundaries, and this page states them.

These limits describe diskless-topic behavior, not a roadmap or a guarantee of compatibility with every Kafka workload. Classic topics retain their separate local-log path.

Transactions and producer recovery

Transactional record batches are rejected with INVALID_REQUEST. Non-transactional producers, including idempotent producers, are supported. The other three transaction requests are rejected too, each with its own error code; see Kafka compatibility for the full table. Idempotence is not a claim of transactional exactly-once processing; applications that depend on Kafka transactions cannot use diskless topics as a substitute.

Producer-state snapshots are scoped by topic incarnation, partition, and zone. Keep an idempotent producer's routing zone stable. Recovery normally loads a snapshot and replays the log tail, but the implementation skips replay when there is no snapshot and the record-offset span exceeds the positive ursa.storage.producer.state.snapshot.record.threshold. Do not assume unlimited deduplication recovery after losing snapshot state; preserve Oxia and monitor snapshot/recovery failures.

Storage mode and internal topics

Diskless topics use RF=1 and cannot be converted in place to classic topics, or vice versa. Use a new topic and migrate data when changing storage mode.

Internal topics, including __consumer_offsets and __transaction_state, stay classic. KRaft metadata and any classic topics still need local storage. Diskless message storage does not make the entire Kafka deployment stateless.

The LOCAL backend does not make independent local directories a shared store. Multi-broker failover needs shared remote storage and available metadata services.

No Kafka key-based log compaction

Diskless topics do not implement cleanup.policy=compact. The controller accepts the setting, at creation or on a later config change, without an error, but nothing collapses each key's history. Ursa's WAL-to-compacted-object compaction changes the physical layout and keeps records readable; it is not Kafka's key-based compaction.

The policy does not exempt a diskless topic from retention either. On a classic topic, compact without delete turns off time- and size-based deletion; the diskless retention path does not read cleanup.policy, so retention.ms and retention.bytes trim the topic whatever its policy. A topic created with cleanup.policy=compact therefore does not keep each key's latest value: its records are trimmed like any other topic's, and Kafka's default broker retention is 7 days.

Keep topics that depend on compaction classic. When ursa.storage.topic.default.enable is on, a new topic is diskless unless its create request sets ursa.storage.enable=false, and that includes compacted topics applications create for themselves, such as Kafka Streams changelog topics and Kafka Connect's config, offset and status topics.

retention.ms and retention.bytes produce logical soft trims. Physical WAL reclamation waits for compaction progress and cleanup. A compactor outage or a lagging stream can hold back reclamation even after Kafka consumers can no longer read the trimmed offsets. External Iceberg table retention is separate. See operating the compactor for why the delete watermark is a global minimum across streams.

Long polling is implemented, but not full fetch batching parity

The diskless reader now waits when a fetch is empty and caught up, waking on a local append or the request deadline before re-reading. It does not always return immediately as older versions did.

However, any available records cause a response without waiting to reach fetch.min.bytes. Append notifications are local to the serving writer; another zone owner's writes are discovered through storage refresh rather than those notifications. The offset window is cached for 100 ms, and a caught-up fetch can wait until fetch.max.wait.ms before re-reading. Do not treat an idle fetch or a cached latest offset as an instantaneous cluster-wide view.

Timestamp lookups are bounded

Timestamp ListOffsets searches entry write-time headers and scans at most 256 entries for a matching record timestamp. On scan exhaustion it can return an earlier candidate entry's base offset and write time, so a consumer may need to read forward. The search assumes record timestamps are not newer than their entry write-time headers.

MAX_TIMESTAMP is answered from the last entry's header, not a full scan for the maximum producer-supplied record timestamp. Earliest/latest offsets and timestamp-based seeking should therefore not be conflated with exact event-time indexing for arbitrary timestamps.

Failover and analytics are not instantaneous

Owner remapping avoids copying message logs, but failure detection, metadata propagation, client retries, and producer-state recovery still take time. Removing Kafka ISR replication for diskless records also does not eliminate object-store, Oxia, controller, or internal-topic traffic across zones.

Iceberg visibility is asynchronous and needs both a configured external materialization sink and a catalog. The default Docker stack writes compacted objects without exposing an external table. The demo's Polaris catalog is in-memory and uses fixed development credentials.

Topic deletion cleans up Kafka-owned storage state asynchronously; it is not an instruction to drop an external table. The default external naming uses the logical topic name, so recreating the same Kafka topic name can append to the existing Iceberg table.

Where next

Issue tracking

Track changes and report issues at openlakestream/kafka. For the storage and compaction implementation, see openlakestream/ursa.