Lakestream

Compaction settings

Scheduling, commit batching, quarantine, metadata and catalog connections, and compacted-object cleanup.

These are the keys the source declares under the compact category. They configure the compaction service — the process that rewrites log ranges into compacted objects and, when enabled, materializes streams into tables.

For how to run that process, see operations. For the table-format settings it hands to a sink, see lakehouse and table settings.

Metadata and coordination

The compaction service connects to two metadata stores, and they are configured separately because they serve different roles: cluster coordination uses notifications and locks, while the storage store holds the offset indexes.

PropertyTypeDefaultPurpose
metadataStoreUrlStringunsetMetadata store for catalog metadata, locks and leader election.
metadataStoreConfigStringunsetClient configuration JSON for that connection.
oxiaStorageUrlStringunsetMetadata store holding WAL offset indexes.
oxiaStorageConfigStringunsetClient configuration JSON for that connection.

The configuration JSON is passed through to the Oxia client unchanged; Ursa checks only that it parses. This is where transport security and authentication for the metadata connection are set, and neither is on by default.

When running alongside Kafka, match metadataStoreUrl to the broker's ursa.catalog.oxia.service.url and oxiaStorageUrl to ursa.oxia.service.url.

Compacted-object location

PropertyTypeDefaultPurpose
compactionBucketStringunsetBucket for compacted objects. Falls back to s3CompactionBucket.
compactionPrefixStringunsetObject-name prefix for compacted objects. Falls back to s3CompactionPrefix.
compactionBucketRegionStringunsetRegion of the compacted-object bucket.

compactionBucket and compactionPrefix supersede s3CompactionBucket and s3CompactionPrefix, which remain readable as fallbacks. The compacted-object location is independent of the WAL location; both must be set consistently on every process that reads or writes them.

What triggers a compaction task

A task is published when the uncompacted range grows past a size threshold, or failing that, when its oldest uncompacted record is old enough.

PropertyTypeDefaultPurpose
compactedFileSizeLimitlong268435456 (256 MiB)Accumulated uncompacted bytes that publish a task immediately.
tailCompactDataVisibilityIntervalInSecondsint180Age of the oldest uncompacted record that publishes a task when the size threshold has not been reached. This is the upper bound on how long a record waits before it is visible in a table.
checkCompactMessageStepLengthint10000Offset-index stride used while measuring the uncompacted range.
internalCompactionTaskPublisherEnabledbooleantrueWhether this process publishes tasks. Set to false when an external scheduler owns publication.

Scheduling and threads

PropertyTypeDefaultPurpose
compactedThreadNumintmax(1, availableProcessors() - 1)Worker threads running compaction tasks.
publishThreadNumintavailable processorsThreads publishing compaction tasks.
commitThreadNumintavailable processorsThreads committing compacted output.
refreshLocalTopicInternalInSecondslong60Interval for refreshing the local view of topics.
refreshLocalTaskIntervalInSecondslong30Minimum interval at which a worker polls for new tasks.
metastoreRequestRateLimitPerSecondint500Per-second ceiling on compaction-task reads from the metadata store while a process fetches new tasks.
compactionMaintenanceIntervalInSecondslong300Interval for periodic maintenance, including cache cleanup.
walReadRateLimitInBytesPerSecondlong52428800 (50 MiB)Read-rate ceiling while a task reads the WAL. 0 or below removes the limit.

Spelling matters

refreshLocalTopicInternalInSeconds contains Internal, not Interval. Field names are matched exactly and unrecognized keys are retained without complaint, so a "corrected" spelling is silently ignored.

Commit batching

PropertyTypeDefaultPurpose
maxCommitIntervalInSecondsint180Interval at which the commit runner flushes accumulated tasks.
maxTaskCombineSizeint250Maximum tasks combined into one commit.
catalogMaxOpenTimeInSecondslong1200How long a catalog handle is kept before being closed and reopened. Bounds the lifetime of credentials read at open time.

Failure handling

A task that fails is held back rather than retried immediately, for a period that depends on whether the failure is classified as retryable.

PropertyTypeDefaultPurpose
retryableQuarantineInSecondsint30Hold-off after a retryable failure.
nonRetryableQuarantineInSecondsint300Hold-off after a non-retryable failure.
replayDLQTasksEnabledbooleantrueReplay dead-lettered commit tasks on the first round after startup.
recordNonCommittableTaskThresholdint500Count of non-committable tasks at which the condition is recorded.

Excluding streams

PropertyTypeDefaultPurpose
blackNamespaceOfCompactSet<String>emptyNamespaces excluded from compaction, comma-separated. For example public/__system, public/default.
blackTopicOfCompactSet<String>emptyIndividual streams excluded from compaction, named without a partition suffix. For example public/default/my-topic.

Materialization dispatch

PropertyTypeDefaultPurpose
materializationEnabledbooleanfalseDispatch resolved streams through the materialization SPI.
materializationServiceClassStringio.lakestream.ursa.lakehouse.compact.LakehouseMaterializationServiceImplementation of the materialization service.
compactionStorageBindingsClassStringio.lakestream.ursa.lakehouse.compact.LakehouseCompactionStorageBindingsImplementation supplying the long-running publish, commit and cleaner runners.
upsertModeEnabledbooleanfalseWrite in upsert mode where the sink supports it. Can be overridden per namespace or topic.

Materialization is off by default

With materializationEnabled left at false, the worker does not dispatch through the materialization SPI and the committer runners own the table commit. Setting it to true moves catalog commits onto the SPI path. Set it deliberately — it changes which code path writes your tables.

compactionServiceClass is superseded by materializationServiceClass. When both are set, materializationServiceClass wins; when only the older key is set, its value is used and a deprecation warning is logged.

Unity Catalog connection

PropertyTypeDefaultPurpose
unityCatalogUriStringunsetUnity Catalog server URI.
unityCatalogNameStringunsetCatalog name.
unityCatalogTokenStringunsetAccess token. Masked in logged configuration.

These flat properties are an alternative to declaring the same catalog under the delta.catalog.<name>. prefix.

Compacted-object cleanup

PropertyTypeDefaultPurpose
compactedDataCleanupJobIntervalInSecsint43200 (12 hours)Interval between cleanup scans for compacted objects no longer needed.
compactedDataCleanupThreadNumintavailable processorsThreads processing cleanup work in parallel.
compactedDataCleanupPendingTasksint100Maximum queued cleanup tasks.

This job runs on the elected leader only. It removes compacted objects that have fallen behind each stream's mark-deleted offset; it does not touch files written to an external table, which the table catalog owns.