Lakestream

Object storage settings

Buckets, regions, endpoints, connection limits, and the lakehouse reader's prefetch cache.

These are the keys the source declares under the storage category: where objects live, how Ursa connects to the object store, and how the lakehouse reader caches what it fetches.

For the backend type and for WAL buffering, see WAL settings. For how a particular cloud is wired, including credentials and permissions, see storage backends.

Location

PropertyTypeDefaultPurpose
bucketStringunsetBucket or container holding WAL objects. Falls back to s3Bucket when unset.
prefixStringunsetObject-name prefix for WAL objects. Falls back to s3Prefix when unset.
regionStringunsetRegion of the WAL bucket.
cloudStorageEndpointStringunsetEndpoint override, for an S3-compatible service such as MinIO.

bucket, prefix and region supersede s3Bucket, s3Prefix and s3Region. The older names remain readable as fallbacks so existing configurations keep working; prefer the generic names for new deployments.

Superseded propertyUse instead
s3Bucketbucket
s3Prefixprefix
s3Regionregion

Credentials

PropertyTypeDefaultPurpose
s3AccessKeyIdStringunsetStatic access key. Leave unset to resolve credentials from the ambient chain.
s3SecretAccessKeyStringunsetStatic secret key.
disableS3ExpressSessionAuthbooleanfalseTurn off S3 Express session authentication.

When both key properties are unset, Ursa resolves credentials from a shared profile or a projected web-identity token, which is what makes IRSA and workload identity work without static secrets. s3AccessKeyId, s3SecretAccessKey, unityCatalogToken, unityCatalogClientSecret and iceberg.credential are masked when the compactor logs its configuration at startup.

Connection limits

PropertyTypeDefaultPurpose
cloudStorageMaxConcurrencyRequestintmaximum direct memory x 0.25 / 4 MiBMaximum concurrent object-store requests. Tracks the write-buffer segment count by default.
cloudStorageMaxPendingConnectionAcquiresint-1Queue depth for connection acquisition. -1 leaves the SDK's own default in force.
cloudStorageMaxPendingAcquireTimeoutInMsint-1Timeout for acquiring a connection. -1 leaves the SDK's default in force.
cloudStorageOpsRateLimitPerSecondint-1Client-side request rate limit. Values of 0 or below leave requests unlimited.
s3OpsMaxRetriesint-1Retry attempts for S3 operations. -1 leaves the SDK's default in force.

Three further S3-specific keys are superseded but still read. Each is consulted only when set to a positive value, and otherwise defers to its generic counterpart:

Superseded propertyConsulted when positive, else falls back to
s3OpsRateLimitPerSecondno rate limit
s3MaxPendingConnectionAcquirescloudStorageMaxPendingConnectionAcquires
s3ConnectionAcquisitionTimeoutMscloudStorageMaxPendingAcquireTimeoutInMs

Lakehouse reader prefetch cache

The reader that serves compacted objects keeps its own prefetch cache, separate from the WAL read cache.

PropertyTypeDefaultPurpose
customExpireTimeMsint60000Expiry for entries the reader marks as short-lived.
defaultExpireTimeMsint120000Expiry for ordinary entries.
cacheEvictionWatermarkdouble0.9Occupancy fraction at which eviction begins.
maxInflightReadingTasksint20 x available processorsConcurrent prefetch tasks.
lakehouseIOThreadNumintavailable processorsThreads serving lakehouse reads.

Entry index cache

PropertyTypeDefaultPurpose
maxEntryIndexCacheSizeintmaximum direct memory x 0.01 / 1024Maximum cached offset-index entries.
entryIndexCacheTTLInSecsint600Time to live for a cached offset-index entry.

WAL reclamation

PropertyTypeDefaultPurpose
cleanupJobIntervalInHoursint12Interval between runs of the WAL reclamation job.

Opening a catalog starts this job automatically for every persistent backend. It is skipped when backendStorageType is local, because prefix-based expiry has no local-filesystem equivalent.

Each run takes a cluster-wide lock in the metadata store, so only one node reclaims at a time however many are running. On object stores, reclamation marks whole date prefixes for expiry through the bucket's own lifecycle configuration rather than deleting objects one by one — which means the bucket's lifecycle policy must be readable and writable by Ursa. See storage backends for the permissions this requires and how the rules are named.