Lakestream
Ursa for Kafka

Operating the compactor

Why the Ursa compactor is not optional, how it scales, and the settings that matter when running it next to Ursa for Kafka.

The compactor is a separate Ursa process that consolidates the write-ahead log into compacted objects. Ursa for Apache Kafka (UFK) does not run it inside the broker, so it never competes with the request path — and so it is something you have to deploy and keep running.

This page is the operating model. For the manifests that start one on Kubernetes, see deployment; for the full property reference, see Ursa configuration.

It is not a lakehouse add-on

The common mistake is to treat the compactor as something you only need when you want an Iceberg table. Retention is the reason it always runs:

  1. retention.ms and retention.bytes on a diskless topic issue a soft trim. The log's first offset moves, consumers stop seeing the records, and the objects holding them stay exactly where they are.
  2. WAL objects are physically deleted only up to a delete watermark.
  3. That watermark is the oldest un-compacted position across every stream, and only compaction advances it.

A cluster with no compactor running has a WAL that only grows. Retention hides records without freeing a byte.

The watermark is a global minimum

Because the watermark is the minimum across every stream, one stream that stops being compacted pins reclamation for all of them. A quarantined partition or a compaction backlog on one topic shows up as storage that never shrinks on topics that are compacting normally.

Registering an Iceberg table is a second, optional sink on the same compaction pass. Compaction runs with no catalog at all — this is TableCatalogType.NONE, storage-only compaction, and it is the default posture of the Compose stack.

Running one

The compactor ships inside the lakestream/kafka image, alongside the brokers and the Kafka CLI tools. There is no separate compactor image.

docker run lakestream/kafka \
  /opt/kafka/bin/ursa-compactor.sh --conf /path/to/ursa-storage.properties

--conf is the only flag; it points at an Ursa properties file. URSA_JAVA_OPTS sets the JVM flags and defaults to -Xmx1024M -XX:+UseZGC.

Not in the Strimzi image

lakestream/kafka-strimzi carries the broker distribution in Strimzi's layout and does not include ursa-compactor.sh. On Kubernetes, run the compactor from lakestream/kafka as its own Deployment — Strimzi does not manage it. See deployment.

The compactor reads the same storage and metadata that the brokers write. Its metadataStoreUrl, oxiaStorageUrl, bucket, prefix, compactionBucket and compactionPrefix must resolve to exactly the same locations as the brokers' corresponding ursa.* settings. Pointing them at different buckets or prefixes does not fail loudly; the two sides stop seeing each other's data, and the first symptom is a WAL that never shrinks.

Scaling and leader election

Compactor instances elect a leader through Oxia. The leader publishes compaction tasks; every instance runs compaction workers that claim and execute them. There is no split-brain risk in adding instances, and no external coordinator to deploy.

Start with one instance. Adding instances adds worker throughput for executing tasks — it does not add a second task publisher, and it does not shard the delete watermark, which stays a single global minimum.

The flag that must be right

materializationEnabled must be true, whether or not you want an external table.

Setting it to false selects a legacy path that fails on Ursa 1.0.0 with Unsupported lakehouse type: NONE. The effect is not a clean shutdown: it quarantines every partition and never reclaims WAL objects, which is the failure this page exists to prevent. Storage-only compaction is expressed by leaving lakehouseType and the catalog settings unset while materializationEnabled stays true.

Turning the external Iceberg sink on adds lakehouseType, catalog.name, materializationDefaultNamespace, clusterSdtEnabled and the iceberg.catalog.<name>.* group. Those are covered on deployment.

Three keys the demo entrypoint still sets

lakehouse/compactor-entrypoint.sh in the Kafka repository also sets streamTableMode, clusterSbtEnabled and managedTableSchemaEvolutionEnabled. Ursa 1.0.0 no longer defines any of them; they are leftovers from earlier releases and setting them changes nothing. In particular, every compacted-object write now records the per-file index Kafka needs to read a range after compaction, with no switch to turn it on.

Throughput and freshness

These properties trade compaction latency against the number of objects and requests produced. Defaults and value ranges are in Ursa compaction settings.

PropertyWhat it controls
compactedFileSizeLimitTarget size of a compacted file; smaller files appear sooner and cost more requests
tailCompactDataVisibilityIntervalInSecondsHow quickly freshly written records become visible in compacted output
refreshLocalTopicInternalInSecondsHow often the compactor rediscovers streams to compact
refreshLocalTaskIntervalInSecondsHow often a worker looks for newly published tasks
compactedThreadNumWorker threads performing compaction
publishThreadNumThreads publishing compaction tasks
commitThreadNumThreads committing compaction results
maxCommitIntervalInSecondsUpper bound on how long results are batched before commit
metastoreRequestRateLimitPerSecondCeiling on compaction-task reads from the metadata store while fetching new tasks

Demo values are not defaults

The values in the Compose stack and in run-lakehouse-demo.sh are tuned to make a laptop demo produce visible output within seconds — a compactedFileSizeLimit of 1, second-scale refresh intervals, and a lowered WAL flush interval and threshold. They are verifier overrides. Do not carry them into a cluster you care about.

What to watch

The compactor's own log is the primary signal: task errors, quarantined partitions, and whether ranges are being published and committed at all. In the Compose stack, make compaction-logs follows it.

Storage is the second signal. If WAL object count or bucket size grows without bound while Kafka retention is expiring records, compaction is not advancing the delete watermark — check the compactor first, and check it for a single stalled stream before checking the topic whose storage you noticed. See monitoring.

Next steps