Lakestream

Lakehouse and table settings

Table format, compacted-object output, catalog bootstrap properties, and Iceberg and Delta options.

These keys are read by the lakehouse module rather than declared on the storage configuration, so they follow different naming conventions — some are camelCase, some hyphenated, and several are prefixes under which arbitrary properties are passed through to a catalog or table.

They are read from the same properties file as the compaction settings.

Output selection

One compaction pass can produce two outputs: internal compacted objects that keep the stream readable, and files in an external table. Each is switched on independently.

PropertyTypeDefaultPurpose
compactedObjectEnabledbooleantrueWrite internal compacted objects for stream replay.
lakehouseTypeStringNONETable format to materialize into. One of ICEBERG, DELTA, DELTA_AND_ICEBERG, NONE.

With lakehouseType left at NONE, compaction produces internal compacted objects alone. Those objects are Parquet files that keep the log readable past compaction, and Ursa does not register them in a table catalog. See lakehouse tables.

Table location and naming

PropertyTypeDefaultPurpose
storagePathStringa data directory resolved to an absolute pathBase path for table output, and the Iceberg warehouse fallback.
directExternalStoragePathStringunsetExplicit path for external table output.
catalog.nameStringunsetNamed catalog this stream materializes into.
catalog.defaultStringunsetCatalog used when a stream names none.
catalog-backendStringhadoopIceberg catalog backend when the catalog does not declare one.

Table shape

PropertyTypeDefaultPurpose
partitionKeyStringnonePartition column. The value __partition is treated as none.
identifierFieldsSet<String>emptyComma-separated columns forming the row identity, required for upsert writes.
upsertModebooleanfalseWrite rows as upserts rather than appends.
compressTypeStringZSTDParquet compression codec.
rowGroupSizelong10485760 (10 MiB)Parquet row-group size.
kafka.compression.typeStringLZ4Compression applied to stored Kafka record batches.

partitionKey and identifierFields describe one table's shape, so they are set per namespace or per topic rather than cluster-wide. See per-namespace and per-topic overrides.

Schema handling

PropertyTypeDefaultPurpose
tableEvolveSchemaEnabledbooleantrueApply schema changes to the destination table.
check-orderingbooleanfalseRequire column order to match during schema comparison.
check-nullabilitybooleantrueRequire nullability to match during schema comparison.
make-new-fields-optionalbooleanfalseAdd newly seen fields as optional columns.
variantTypeEnabledbooleanfalseAllow variant columns.
allowIcebergV3booleanfalseAllow writing Iceberg format version 3.
persistExtraMetadatabooleanfalsePersist additional record metadata as columns.
persistKeybooleanfalsePersist the record key as a column.

Commit behaviour

PropertyTypeDefaultPurpose
lakehouseCommitMaxRetryTimesint3Retries for a failed table commit.
catalogOpsRetryMaxAttemptsint3Attempts for a failed catalog operation.
catalogOpsRetryDelayMslong100Delay between catalog retries.
deltaKernelWriteBatchSizeint1000Rows per Delta write batch.
deltaSupportManagedCommitbooleanfalseUse Delta managed commits.
icebergSnapshotExpirationIntervalInSecondsint-1Interval for expiring Iceberg snapshots. -1 leaves snapshots in place.

Dead-letter tables

Records that cannot be written to the destination table are routed to a companion table rather than dropped.

PropertyTypeDefaultPurpose
delta.dlt.enabledbooleantrueRoute failed records to a dead-letter table.
dlt.suffixString_dltSuffix appended to the table name to form the dead-letter table's name.

Property prefixes

Everything under these prefixes is passed through with the prefix stripped, so catalog and table options that Ursa does not itself interpret can still be set.

PrefixPassed to
iceberg.catalog.<name>.Connection properties for a named Iceberg catalog
delta.catalog.<name>.Connection properties for a named Delta catalog
clickhouse.catalog.<name>.Connection properties for a named ClickHouse destination
iceberg.hadoop.Hadoop configuration for Iceberg
iceberg.write-props.Iceberg write properties
iceberg.table-props.Iceberg table properties
hadoop.Hadoop configuration for the object-store filesystem

A catalog is declared by setting properties under its prefix; the name in the middle segment becomes the catalog's name, which a stream then selects through catalog.name.

Iceberg table options

PropertyTypeDefaultPurpose
iceberg.table.evolve-schema-enabledbooleantrueApply schema evolution to the Iceberg table.
iceberg.table.schema-force-optionalbooleanfalseCreate all columns as optional.
iceberg.table.schema-case-insensitivebooleanfalseMatch column names case-insensitively.
iceberg.table.schema-update-retriesint2 (3 attempts)Retries for a schema update.
iceberg.table.create-table-retriesint2 (3 attempts)Retries for table creation.
iceberg.table.upsert-mode-enabledbooleanfalseUpsert writes for this table.
iceberg.table.cdc-fieldStringunsetColumn carrying the change-type indicator.
iceberg.hadoop-conf-dirStringunsetDirectory of Hadoop configuration files to load.
commitBranchStringunsetIceberg branch to commit to.

Credentials from files

PropertyTypeDefaultPurpose
iceberg.credentialFileStringunsetFile whose contents become the Iceberg catalog credential.
unityCatalogTokenFileStringunsetFile whose contents become the Unity Catalog token.

These are read once at startup. catalogMaxOpenTimeInSeconds bounds how long a catalog handle built from them is kept before being reopened.

Cloud project identity

PropertyTypeDefaultPurpose
googleCloudProjectIDStringunsetGoogle Cloud project for GCS and BigQuery-backed catalogs.
googleCloudServiceAccountFileStringunsetService-account key file.