Lakestream

WAL settings

Write buffering, the read cache, the write-ahead log backend, and object naming.

These are the keys the source declares under the wal category. They govern how appends are batched into write-ahead log (WAL) objects, how much memory the write and read paths may hold, and which backend and object-naming scheme the WAL uses.

Defaults are the values in force when the key is absent from a properties file. See how properties are loaded for what that means.

Memory sizing comes from the JVM

Several defaults are derived from the JVM's maximum direct memory rather than being fixed numbers. Ursa reads it once at class-load time through Netty's estimate, which follows -XX:MaxDirectMemorySize and can also be influenced by the io.netty.maxDirectMemory system property.

The practical consequence: -XX:MaxDirectMemorySize is the primary capacity knob. Raising it raises the write-buffer segment count, the pending-append budget, the read cache and the entry-index cache together, without touching a single Ursa property.

Write buffering

PropertyTypeDefaultPurpose
writeBufferSizeint4194304 (4 MiB)Size of one write-buffer segment. Also the unit the concurrency default divides by.
writeBufferSegmentint-1Number of write-buffer segments. -1 means derive: maximum direct memory x 0.25 / writeBufferSize.
writeBufferFlushSizelong268435456 (256 MiB)Accumulated pending bytes that trigger a flush.
writeBufferFlushIntervalMslong250Time-based flush trigger. With a light append rate this is the floor on how long an append waits before it is written.
writeBufferMaxStreamIdslong4Maximum number of distinct streams packed into one WAL object.
maxPendingAddRequestsUsedByteslongmaximum direct memory x 0.15Admission-control budget for un-flushed appends, measured in bytes. Appends beyond it are rejected rather than queued.
writeCacheEnabledbooleantrueRetain just-flushed segments in memory so a tailing read is served without a round trip to the object store.

Read cache

PropertyTypeDefaultPurpose
readCacheMemorySizelongmaximum direct memory x 0.15Memory budget for WAL objects fetched back from the object store.
defaultReadBatchContextInitializeSizeint32Initial capacity of the per-batch index list a read allocates before fetching indexes.

Backend and object naming

PropertyTypeDefaultPurpose
backendStorageTypeStringS3WAL backend. One of LOCAL, S3, GCS, AZUREBLOB.
storagePathStringunsetBackend storage path, used by the local backend and as the Iceberg warehouse fallback.
compactionBackendStorageTypeStringunsetBackend for compacted objects. When unset, falls back to backendStorageType. Accepts LOCAL, S3, GCS, AZUREBLOB, and additionally AZUREDFS and AZURELOCAL.
idGeneratorTypeStringDATEUUIDNaming strategy for WAL objects. One of MEMORY, RANDOM, OXIA, UUID, DATEUUID.
indexSerializeFormatVersionint3Format version used when writing offset-index entries to the metadata store. One of 1, 2, 3.

Where the object-store connection settings live

Buckets, prefixes, regions, endpoints and connection limits are declared under the storage category, not wal. See object storage settings, and storage backends for how each cloud is wired.

The format version written here is the storage format the Storage Spec defines. Which versions an implementation reads and writes is recorded on implementation status.