Lakehouse tables
Ursa produces two distinct outputs: internal compacted objects that serve stream reads, and external tables committed to a named table catalog.
Ursa distinguishes the compacted data used for streaming reads from delivery into an external table. One compaction pass can produce both, but their ownership and storage costs differ.
Internal compacted objects versus external tables
| Internal compacted objects (COs) | External / stream-delivered table (SDT) | |
|---|---|---|
| Data | Per-log Parquet files under the compaction bucket and prefix | Separate sink output |
| Offset index | Per-file index recorded in Ursa's stream offset index | Not indexed for Ursa streaming reads |
| Table catalog | Ursa does not register them | Registered and committed in the destination catalog |
| Streaming replay | Reads the internal COs | Continues from the internal COs, not the external table |
| Analytical access | Parquet files, so they can be registered as a lakehouse table and queried; Ursa does not do so itself | Table managed through the external catalog/sink |
| Lifecycle | Ursa retention, compaction, and cleanup | External system's retention and maintenance |
| Typical use | Stream replay after WAL reclamation | Independently governed or transformed analytical data |
External materialization is not zero-copy
The Kafka Compose lakehouse profile turns on materializationEnabled and clusterSdtEnabled and registers an Iceberg REST catalog named polaris. Its entrypoint also still sets clusterSbtEnabled and streamTableMode, keys that Ursa 1.0.0 no longer defines. The compactor keeps writing internal COs for Kafka fetch and writes a second output into Polaris/Iceberg. Sharing a compaction pass does not make those outputs the same physical files.
Internal compaction runs with materializationEnabled=false and no external catalog, and it is still needed to advance WAL reclamation. compactedObjectEnabled (default true) in the compactor's properties controls whether internal COs are written; keep it enabled for Ursa for Apache Kafka (UFK), which reads them for replay. Enabling table materialization is additional work, not a replacement for this storage lifecycle.
Sinks and schemas
The materialization API represents sinks using TableCatalogType: ICEBERG, DELTA, DELTA_UC, CLICKHOUSE, and NONE. Iceberg and Delta support table-format output; ClickHouse is a separate delivery sink, not another compacted-object format. NONE selects storage-only compaction: internal COs and no table sink. See table catalogs for declaring each of them.
Kafka materialization has Avro, JSON Schema, and Protobuf codecs. Decoding, schema compatibility, and sink evolution policy determine whether a record can be written. A registered catalog alone does not guarantee that arbitrary payloads can be materialized. See schemas and the feature matrix.
Named catalogs and policy resolution
Use the StreamCatalog API to:
- Register a named
TableCatalogwithregisterTableCatalog(...). - Set a namespace baseline with
setNamespaceMaterialization(...). - Apply sparse stream overrides with
setStreamMaterialization(...). - Inspect the effective result with
resolveMaterialization(streamId).
Resolution deep-merges a namespace baseline and sparse stream overrides field by field; see resolution for the rules every resolver follows. Ursa additionally falls back to a cluster-wide default policy when a namespace sets none — see Appendix B: implementation notes.
The compactor can also bootstrap catalog records from operator properties rather than through the API. Table catalogs covers those property forms, the Iceberg backends Ursa packages, and how a policy selects among registered catalogs. Catalog properties are not interchangeable with Kafka broker configuration.
For a working local REST-catalog setup, use the Kafka lakehouse demo; do not copy its static demo credentials into a production deployment.
Table names and recreated topics
Without an explicit table identifier or namespace naming template, the table name is the source logical name (lakestream.source.logical.name), then the legacy Kafka topic-name property, then the storage stream name. Ursa 1.0.0 removed the TableMode distinction under which Ursa-owned tables used the storage stream name instead.
Kafka records the logical topic name in system-owned stream metadata. Recreating a Kafka topic creates a new storage incarnation, but can still append to the same topic-named external table. Deleting a topic does not imply deleting that external table or its history.
Namespace TableNaming templates support ${stream.namespace}, ${stream.name}, ${stream.logicalName}, and ${stream.property.<key>}. Unknown variables and missing or blank referenced properties fail resolution rather than silently selecting a default. The table namespace prefix is literal, not interpolated.
Changing a destination
Several catalogs can be registered for different warehouses or access boundaries, chosen through namespace baselines and stream overrides; see table catalogs. The compactor needs access to every catalog it might route to, and to that catalog's object store.
Before changing a catalog or table name, decide what happens to existing snapshots and files. Policy reassignment is not a data migration. For each output, verify decoded schema, partitioning, table commits, streaming reads after compaction, and retention independently.
Try it locally
With the Kafka image pulled or built as described in the Kafka quickstart, run from Kafka's Compose stack directory:
make lakehouse-demoThe script exercises Kafka → internal compaction plus external Iceberg → DuckDB and asserts the result. It is a disposable demo, not a production deployment recipe.