Lakehouse and table settings
Table format, compacted-object output, catalog bootstrap properties, and Iceberg and Delta options.
These keys are read by the lakehouse module rather than declared on the storage configuration, so they follow different naming conventions — some are camelCase, some hyphenated, and several are prefixes under which arbitrary properties are passed through to a catalog or table.
They are read from the same properties file as the compaction settings.
Output selection
One compaction pass can produce two outputs: internal compacted objects that keep the stream readable, and files in an external table. Each is switched on independently.
| Property | Type | Default | Purpose |
|---|---|---|---|
compactedObjectEnabled | boolean | true | Write internal compacted objects for stream replay. |
lakehouseType | String | NONE | Table format to materialize into. One of ICEBERG, DELTA, DELTA_AND_ICEBERG, NONE. |
With lakehouseType left at NONE, compaction produces internal compacted objects alone. Those objects are Parquet files that keep the log readable past compaction, and Ursa does not register them in a table catalog. See lakehouse tables.
Table location and naming
| Property | Type | Default | Purpose |
|---|---|---|---|
storagePath | String | a data directory resolved to an absolute path | Base path for table output, and the Iceberg warehouse fallback. |
directExternalStoragePath | String | unset | Explicit path for external table output. |
catalog.name | String | unset | Named catalog this stream materializes into. |
catalog.default | String | unset | Catalog used when a stream names none. |
catalog-backend | String | hadoop | Iceberg catalog backend when the catalog does not declare one. |
Table shape
| Property | Type | Default | Purpose |
|---|---|---|---|
partitionKey | String | none | Partition column. The value __partition is treated as none. |
identifierFields | Set<String> | empty | Comma-separated columns forming the row identity, required for upsert writes. |
upsertMode | boolean | false | Write rows as upserts rather than appends. |
compressType | String | ZSTD | Parquet compression codec. |
rowGroupSize | long | 10485760 (10 MiB) | Parquet row-group size. |
kafka.compression.type | String | LZ4 | Compression applied to stored Kafka record batches. |
partitionKey and identifierFields describe one table's shape, so they are set per namespace or per topic rather than cluster-wide. See per-namespace and per-topic overrides.
Schema handling
| Property | Type | Default | Purpose |
|---|---|---|---|
tableEvolveSchemaEnabled | boolean | true | Apply schema changes to the destination table. |
check-ordering | boolean | false | Require column order to match during schema comparison. |
check-nullability | boolean | true | Require nullability to match during schema comparison. |
make-new-fields-optional | boolean | false | Add newly seen fields as optional columns. |
variantTypeEnabled | boolean | false | Allow variant columns. |
allowIcebergV3 | boolean | false | Allow writing Iceberg format version 3. |
persistExtraMetadata | boolean | false | Persist additional record metadata as columns. |
persistKey | boolean | false | Persist the record key as a column. |
Commit behaviour
| Property | Type | Default | Purpose |
|---|---|---|---|
lakehouseCommitMaxRetryTimes | int | 3 | Retries for a failed table commit. |
catalogOpsRetryMaxAttempts | int | 3 | Attempts for a failed catalog operation. |
catalogOpsRetryDelayMs | long | 100 | Delay between catalog retries. |
deltaKernelWriteBatchSize | int | 1000 | Rows per Delta write batch. |
deltaSupportManagedCommit | boolean | false | Use Delta managed commits. |
icebergSnapshotExpirationIntervalInSeconds | int | -1 | Interval for expiring Iceberg snapshots. -1 leaves snapshots in place. |
Dead-letter tables
Records that cannot be written to the destination table are routed to a companion table rather than dropped.
| Property | Type | Default | Purpose |
|---|---|---|---|
delta.dlt.enabled | boolean | true | Route failed records to a dead-letter table. |
dlt.suffix | String | _dlt | Suffix appended to the table name to form the dead-letter table's name. |
Property prefixes
Everything under these prefixes is passed through with the prefix stripped, so catalog and table options that Ursa does not itself interpret can still be set.
| Prefix | Passed to |
|---|---|
iceberg.catalog.<name>. | Connection properties for a named Iceberg catalog |
delta.catalog.<name>. | Connection properties for a named Delta catalog |
clickhouse.catalog.<name>. | Connection properties for a named ClickHouse destination |
iceberg.hadoop. | Hadoop configuration for Iceberg |
iceberg.write-props. | Iceberg write properties |
iceberg.table-props. | Iceberg table properties |
hadoop. | Hadoop configuration for the object-store filesystem |
A catalog is declared by setting properties under its prefix; the name in the middle segment becomes the catalog's name, which a stream then selects through catalog.name.
Iceberg table options
| Property | Type | Default | Purpose |
|---|---|---|---|
iceberg.table.evolve-schema-enabled | boolean | true | Apply schema evolution to the Iceberg table. |
iceberg.table.schema-force-optional | boolean | false | Create all columns as optional. |
iceberg.table.schema-case-insensitive | boolean | false | Match column names case-insensitively. |
iceberg.table.schema-update-retries | int | 2 (3 attempts) | Retries for a schema update. |
iceberg.table.create-table-retries | int | 2 (3 attempts) | Retries for table creation. |
iceberg.table.upsert-mode-enabled | boolean | false | Upsert writes for this table. |
iceberg.table.cdc-field | String | unset | Column carrying the change-type indicator. |
iceberg.hadoop-conf-dir | String | unset | Directory of Hadoop configuration files to load. |
commitBranch | String | unset | Iceberg branch to commit to. |
Credentials from files
| Property | Type | Default | Purpose |
|---|---|---|---|
iceberg.credentialFile | String | unset | File whose contents become the Iceberg catalog credential. |
unityCatalogTokenFile | String | unset | File whose contents become the Unity Catalog token. |
These are read once at startup. catalogMaxOpenTimeInSeconds bounds how long a catalog handle built from them is kept before being reopened.
Cloud project identity
| Property | Type | Default | Purpose |
|---|---|---|---|
googleCloudProjectID | String | unset | Google Cloud project for GCS and BigQuery-backed catalogs. |
googleCloudServiceAccountFile | String | unset | Service-account key file. |