Lakestream
Ursa

Storage backends

S3, GCS, Azure and local storage: how each is selected and addressed, how credentials resolve, and what permissions Ursa needs.

Ursa writes two kinds of object: write-ahead log (WAL) objects on the append path, and compacted objects plus table files on the compaction path. They are configured separately and can live in different buckets, and the set of backends each supports is not identical.

For the property names, see object storage settings.

Two backend settings

SettingSelects the backend forAccepted values
backendStorageTypeWAL objectsLOCAL, S3, GCS, AZUREBLOB
compactionBackendStorageTypeCompacted objects and table filesLOCAL, S3, GCS, AZUREBLOB, AZUREDFS, AZURELOCAL

compactionBackendStorageType falls back to backendStorageType when unset, which is the right setting for every backend except Azure — where the two sides accept different names. See Azure below.

Addressing

The path Ursa builds for table output depends on the backend:

BackendSchemeLocation is built from
S3s3a://bucket and prefix
GCSgs://bucket and prefix
AZUREBLOBwasbs://<container>@<account>.blob.core.windows.net
AZUREDFSabfss://<container>@<account>.dfs.core.windows.net
AZURELOCALwasb://emulator address
LOCALplain pathstoragePath

A storagePath given with an s3:// scheme is rewritten to s3a://, because the Hadoop filesystem Ursa uses for table output is the S3A connector.

S3 and S3-compatible stores

Set backendStorageType=S3 with region, bucket and prefix. For MinIO, LocalStack or another S3-compatible service, also set cloudStorageEndpoint.

Credentials resolve in one of two ways:

  • Static, when s3AccessKeyId and s3SecretAccessKey are both set.
  • From the environment, when they are not. Ursa builds a chain of the shared-profile provider and the web-identity-token-file provider, which is what makes IAM Roles for Service Accounts on EKS and comparable workload-identity mechanisms work without static secrets.

Prefer the second in a cluster. It avoids a long-lived secret in a properties file, and the token is rotated by the platform.

disableS3ExpressSessionAuth turns off S3 Express session authentication where the store advertises but does not implement it.

Permissions

Beyond reading and writing objects under the configured prefixes, Ursa manages the WAL bucket's lifecycle configuration to reclaim expired WAL objects. That requires:

  • s3:GetLifecycleConfiguration
  • s3:PutLifecycleConfiguration

on the WAL bucket. Reclamation is a read-modify-write of the whole lifecycle configuration: Ursa adds rules whose IDs begin with ursa-wal-delete-, preserves any rule it did not add, and drops its own rules once their date prefix has aged out. Objects are removed by the bucket's lifecycle processing rather than by Ursa deleting them, so reclamation is not immediate.

Without those permissions the rest of Ursa works and WAL objects accumulate.

Google Cloud Storage

Set backendStorageType=GCS with bucket and prefix. Table output is addressed as gs://.

PropertyPurpose
googleCloudProjectIDProject owning the bucket.
googleCloudServiceAccountFileService-account key file, when not using ambient credentials.

The Hadoop connector is configured with a 64 MiB block size and 8 MiB buffers, tuned for the large sequential reads compaction performs.

Azure

Azure is the one backend where the WAL side and the compaction side take different names, because they use different Hadoop connectors.

UseSettingValue
WAL objectsbackendStorageTypeAZUREBLOB
Compacted objects and table filescompactionBackendStorageTypeAZUREDFS

Whether you have to set this yourself depends on how the catalog is opened:

  • Through the Kafka runtime, the provider fills it in: when compactionBackendStorageType is unset and the WAL backend is AZUREBLOB, it sets the compaction backend to AZUREDFS for you. It also derives compactionBucketRegion from region when that is unset.
  • Through StreamCatalogService directly, or in a standalone compactor's properties file, nothing translates it. Set compactionBackendStorageType=AZUREDFS explicitly, or the generic fallback yields AZUREBLOB and the compacted-object reader rejects it at startup:
AZUREBLOB is not supported because Hadoop 3.5 removed the WASB connector; use AZUREDFS

AZUREDFS requires a storage account with hierarchical namespace enabled. The bucket must be written as account@container; any other form is rejected.

Credentials come from workload identity, read from the standard environment variables AZURE_AUTHORITY_HOST, AZURE_CLIENT_ID, AZURE_TENANT_ID and AZURE_FEDERATED_TOKEN_FILE.

Local files

backendStorageType=LOCAL with storagePath writes to the local filesystem. It is a development backend: objects are not shared between processes, so it gives none of the properties that make storage on object storage useful.

WAL reclamation is skipped on this backend, because prefix expiry has no local equivalent.

Passing options to Hadoop

Any property prefixed hadoop. is passed to the Hadoop configuration with the prefix stripped, which is how connector settings Ursa does not itself expose are set. For example:

hadoop.fs.gs.outputstream.direct.upload.enable=true

iceberg.hadoop. does the same for the Iceberg catalog's own Hadoop configuration, and iceberg.hadoop-conf-dir loads configuration files from a directory.

Which backends the compacted-object reader accepts

The reader that serves compacted objects back to a Kafka integration supports local files, S3 and S3A, GCS, and Azure Data Lake Storage Gen2 through AZUREDFS. AZUREBLOB and AZURELOCAL are rejected when it is configured.

It reads the V2 KAFKA_BATCHED_RAW_PARQUET format only. An entry index without a compacted-object file index, or a Parquet file written with a different serde, fails explicitly rather than being skipped.