Object storage settings
Buckets, regions, endpoints, connection limits, and the lakehouse reader's prefetch cache.
These are the keys the source declares under the storage category: where objects live, how Ursa connects to the object store, and how the lakehouse reader caches what it fetches.
For the backend type and for WAL buffering, see WAL settings. For how a particular cloud is wired, including credentials and permissions, see storage backends.
Location
| Property | Type | Default | Purpose |
|---|---|---|---|
bucket | String | unset | Bucket or container holding WAL objects. Falls back to s3Bucket when unset. |
prefix | String | unset | Object-name prefix for WAL objects. Falls back to s3Prefix when unset. |
region | String | unset | Region of the WAL bucket. |
cloudStorageEndpoint | String | unset | Endpoint override, for an S3-compatible service such as MinIO. |
bucket, prefix and region supersede s3Bucket, s3Prefix and s3Region. The older names remain readable as fallbacks so existing configurations keep working; prefer the generic names for new deployments.
| Superseded property | Use instead |
|---|---|
s3Bucket | bucket |
s3Prefix | prefix |
s3Region | region |
Credentials
| Property | Type | Default | Purpose |
|---|---|---|---|
s3AccessKeyId | String | unset | Static access key. Leave unset to resolve credentials from the ambient chain. |
s3SecretAccessKey | String | unset | Static secret key. |
disableS3ExpressSessionAuth | boolean | false | Turn off S3 Express session authentication. |
When both key properties are unset, Ursa resolves credentials from a shared profile or a projected web-identity token, which is what makes IRSA and workload identity work without static secrets. s3AccessKeyId, s3SecretAccessKey, unityCatalogToken, unityCatalogClientSecret and iceberg.credential are masked when the compactor logs its configuration at startup.
Connection limits
| Property | Type | Default | Purpose |
|---|---|---|---|
cloudStorageMaxConcurrencyRequest | int | maximum direct memory x 0.25 / 4 MiB | Maximum concurrent object-store requests. Tracks the write-buffer segment count by default. |
cloudStorageMaxPendingConnectionAcquires | int | -1 | Queue depth for connection acquisition. -1 leaves the SDK's own default in force. |
cloudStorageMaxPendingAcquireTimeoutInMs | int | -1 | Timeout for acquiring a connection. -1 leaves the SDK's default in force. |
cloudStorageOpsRateLimitPerSecond | int | -1 | Client-side request rate limit. Values of 0 or below leave requests unlimited. |
s3OpsMaxRetries | int | -1 | Retry attempts for S3 operations. -1 leaves the SDK's default in force. |
Three further S3-specific keys are superseded but still read. Each is consulted only when set to a positive value, and otherwise defers to its generic counterpart:
| Superseded property | Consulted when positive, else falls back to |
|---|---|
s3OpsRateLimitPerSecond | no rate limit |
s3MaxPendingConnectionAcquires | cloudStorageMaxPendingConnectionAcquires |
s3ConnectionAcquisitionTimeoutMs | cloudStorageMaxPendingAcquireTimeoutInMs |
Lakehouse reader prefetch cache
The reader that serves compacted objects keeps its own prefetch cache, separate from the WAL read cache.
| Property | Type | Default | Purpose |
|---|---|---|---|
customExpireTimeMs | int | 60000 | Expiry for entries the reader marks as short-lived. |
defaultExpireTimeMs | int | 120000 | Expiry for ordinary entries. |
cacheEvictionWatermark | double | 0.9 | Occupancy fraction at which eviction begins. |
maxInflightReadingTasks | int | 20 x available processors | Concurrent prefetch tasks. |
lakehouseIOThreadNum | int | available processors | Threads serving lakehouse reads. |
Entry index cache
| Property | Type | Default | Purpose |
|---|---|---|---|
maxEntryIndexCacheSize | int | maximum direct memory x 0.01 / 1024 | Maximum cached offset-index entries. |
entryIndexCacheTTLInSecs | int | 600 | Time to live for a cached offset-index entry. |
WAL reclamation
| Property | Type | Default | Purpose |
|---|---|---|---|
cleanupJobIntervalInHours | int | 12 | Interval between runs of the WAL reclamation job. |
Opening a catalog starts this job automatically for every persistent backend. It is skipped when backendStorageType is local, because prefix-based expiry has no local-filesystem equivalent.
Each run takes a cluster-wide lock in the metadata store, so only one node reclaims at a time however many are running. On object stores, reclamation marks whole date prefixes for expiry through the bucket's own lifecycle configuration rather than deleting objects one by one — which means the bucket's lifecycle policy must be readable and writable by Ursa. See storage backends for the permissions this requires and how the rules are named.