Compaction settings
Scheduling, commit batching, quarantine, metadata and catalog connections, and compacted-object cleanup.
These are the keys the source declares under the compact category. They configure the compaction service — the process that rewrites log ranges into compacted objects and, when enabled, materializes streams into tables.
For how to run that process, see operations. For the table-format settings it hands to a sink, see lakehouse and table settings.
Metadata and coordination
The compaction service connects to two metadata stores, and they are configured separately because they serve different roles: cluster coordination uses notifications and locks, while the storage store holds the offset indexes.
| Property | Type | Default | Purpose |
|---|---|---|---|
metadataStoreUrl | String | unset | Metadata store for catalog metadata, locks and leader election. |
metadataStoreConfig | String | unset | Client configuration JSON for that connection. |
oxiaStorageUrl | String | unset | Metadata store holding WAL offset indexes. |
oxiaStorageConfig | String | unset | Client configuration JSON for that connection. |
The configuration JSON is passed through to the Oxia client unchanged; Ursa checks only that it parses. This is where transport security and authentication for the metadata connection are set, and neither is on by default.
When running alongside Kafka, match metadataStoreUrl to the broker's ursa.catalog.oxia.service.url and oxiaStorageUrl to ursa.oxia.service.url.
Compacted-object location
| Property | Type | Default | Purpose |
|---|---|---|---|
compactionBucket | String | unset | Bucket for compacted objects. Falls back to s3CompactionBucket. |
compactionPrefix | String | unset | Object-name prefix for compacted objects. Falls back to s3CompactionPrefix. |
compactionBucketRegion | String | unset | Region of the compacted-object bucket. |
compactionBucket and compactionPrefix supersede s3CompactionBucket and s3CompactionPrefix, which remain readable as fallbacks. The compacted-object location is independent of the WAL location; both must be set consistently on every process that reads or writes them.
What triggers a compaction task
A task is published when the uncompacted range grows past a size threshold, or failing that, when its oldest uncompacted record is old enough.
| Property | Type | Default | Purpose |
|---|---|---|---|
compactedFileSizeLimit | long | 268435456 (256 MiB) | Accumulated uncompacted bytes that publish a task immediately. |
tailCompactDataVisibilityIntervalInSeconds | int | 180 | Age of the oldest uncompacted record that publishes a task when the size threshold has not been reached. This is the upper bound on how long a record waits before it is visible in a table. |
checkCompactMessageStepLength | int | 10000 | Offset-index stride used while measuring the uncompacted range. |
internalCompactionTaskPublisherEnabled | boolean | true | Whether this process publishes tasks. Set to false when an external scheduler owns publication. |
Scheduling and threads
| Property | Type | Default | Purpose |
|---|---|---|---|
compactedThreadNum | int | max(1, availableProcessors() - 1) | Worker threads running compaction tasks. |
publishThreadNum | int | available processors | Threads publishing compaction tasks. |
commitThreadNum | int | available processors | Threads committing compacted output. |
refreshLocalTopicInternalInSeconds | long | 60 | Interval for refreshing the local view of topics. |
refreshLocalTaskIntervalInSeconds | long | 30 | Minimum interval at which a worker polls for new tasks. |
metastoreRequestRateLimitPerSecond | int | 500 | Per-second ceiling on compaction-task reads from the metadata store while a process fetches new tasks. |
compactionMaintenanceIntervalInSeconds | long | 300 | Interval for periodic maintenance, including cache cleanup. |
walReadRateLimitInBytesPerSecond | long | 52428800 (50 MiB) | Read-rate ceiling while a task reads the WAL. 0 or below removes the limit. |
Spelling matters
refreshLocalTopicInternalInSeconds contains Internal, not Interval. Field names are matched exactly and unrecognized keys are retained without complaint, so a "corrected" spelling is silently ignored.
Commit batching
| Property | Type | Default | Purpose |
|---|---|---|---|
maxCommitIntervalInSeconds | int | 180 | Interval at which the commit runner flushes accumulated tasks. |
maxTaskCombineSize | int | 250 | Maximum tasks combined into one commit. |
catalogMaxOpenTimeInSeconds | long | 1200 | How long a catalog handle is kept before being closed and reopened. Bounds the lifetime of credentials read at open time. |
Failure handling
A task that fails is held back rather than retried immediately, for a period that depends on whether the failure is classified as retryable.
| Property | Type | Default | Purpose |
|---|---|---|---|
retryableQuarantineInSeconds | int | 30 | Hold-off after a retryable failure. |
nonRetryableQuarantineInSeconds | int | 300 | Hold-off after a non-retryable failure. |
replayDLQTasksEnabled | boolean | true | Replay dead-lettered commit tasks on the first round after startup. |
recordNonCommittableTaskThreshold | int | 500 | Count of non-committable tasks at which the condition is recorded. |
Excluding streams
| Property | Type | Default | Purpose |
|---|---|---|---|
blackNamespaceOfCompact | Set<String> | empty | Namespaces excluded from compaction, comma-separated. For example public/__system, public/default. |
blackTopicOfCompact | Set<String> | empty | Individual streams excluded from compaction, named without a partition suffix. For example public/default/my-topic. |
Materialization dispatch
| Property | Type | Default | Purpose |
|---|---|---|---|
materializationEnabled | boolean | false | Dispatch resolved streams through the materialization SPI. |
materializationServiceClass | String | io.lakestream.ursa.lakehouse.compact.LakehouseMaterializationService | Implementation of the materialization service. |
compactionStorageBindingsClass | String | io.lakestream.ursa.lakehouse.compact.LakehouseCompactionStorageBindings | Implementation supplying the long-running publish, commit and cleaner runners. |
upsertModeEnabled | boolean | false | Write in upsert mode where the sink supports it. Can be overridden per namespace or topic. |
Materialization is off by default
With materializationEnabled left at false, the worker does not dispatch through the materialization SPI and the committer runners own the table commit. Setting it to true moves catalog commits onto the SPI path. Set it deliberately — it changes which code path writes your tables.
compactionServiceClass is superseded by materializationServiceClass. When both are set, materializationServiceClass wins; when only the older key is set, its value is used and a deprecation warning is logged.
Unity Catalog connection
| Property | Type | Default | Purpose |
|---|---|---|---|
unityCatalogUri | String | unset | Unity Catalog server URI. |
unityCatalogName | String | unset | Catalog name. |
unityCatalogToken | String | unset | Access token. Masked in logged configuration. |
These flat properties are an alternative to declaring the same catalog under the delta.catalog.<name>. prefix.
Compacted-object cleanup
| Property | Type | Default | Purpose |
|---|---|---|---|
compactedDataCleanupJobIntervalInSecs | int | 43200 (12 hours) | Interval between cleanup scans for compacted objects no longer needed. |
compactedDataCleanupThreadNum | int | available processors | Threads processing cleanup work in parallel. |
compactedDataCleanupPendingTasks | int | 100 | Maximum queued cleanup tasks. |
This job runs on the elected leader only. It removes compacted objects that have fallen behind each stream's mark-deleted offset; it does not touch files written to an external table, which the table catalog owns.