Lakestream
Ursa

Lakehouse tables

Ursa produces two distinct outputs: internal compacted objects that serve stream reads, and external tables committed to a named table catalog.

Ursa distinguishes the compacted data used for streaming reads from delivery into an external table. One compaction pass can produce both, but their ownership and storage costs differ.

Internal compacted objects versus external tables

Internal compacted objects (COs)External / stream-delivered table (SDT)
DataPer-log Parquet files under the compaction bucket and prefixSeparate sink output
Offset indexPer-file index recorded in Ursa's stream offset indexNot indexed for Ursa streaming reads
Table catalogUrsa does not register themRegistered and committed in the destination catalog
Streaming replayReads the internal COsContinues from the internal COs, not the external table
Analytical accessParquet files, so they can be registered as a lakehouse table and queried; Ursa does not do so itselfTable managed through the external catalog/sink
LifecycleUrsa retention, compaction, and cleanupExternal system's retention and maintenance
Typical useStream replay after WAL reclamationIndependently governed or transformed analytical data

External materialization is not zero-copy

The Kafka Compose lakehouse profile turns on materializationEnabled and clusterSdtEnabled and registers an Iceberg REST catalog named polaris. Its entrypoint also still sets clusterSbtEnabled and streamTableMode, keys that Ursa 1.0.0 no longer defines. The compactor keeps writing internal COs for Kafka fetch and writes a second output into Polaris/Iceberg. Sharing a compaction pass does not make those outputs the same physical files.

Internal compaction runs with materializationEnabled=false and no external catalog, and it is still needed to advance WAL reclamation. compactedObjectEnabled (default true) in the compactor's properties controls whether internal COs are written; keep it enabled for Ursa for Apache Kafka (UFK), which reads them for replay. Enabling table materialization is additional work, not a replacement for this storage lifecycle.

Sinks and schemas

The materialization API represents sinks using TableCatalogType: ICEBERG, DELTA, DELTA_UC, CLICKHOUSE, and NONE. Iceberg and Delta support table-format output; ClickHouse is a separate delivery sink, not another compacted-object format. NONE selects storage-only compaction: internal COs and no table sink. See table catalogs for declaring each of them.

Kafka materialization has Avro, JSON Schema, and Protobuf codecs. Decoding, schema compatibility, and sink evolution policy determine whether a record can be written. A registered catalog alone does not guarantee that arbitrary payloads can be materialized. See schemas and the feature matrix.

Named catalogs and policy resolution

Use the StreamCatalog API to:

  1. Register a named TableCatalog with registerTableCatalog(...).
  2. Set a namespace baseline with setNamespaceMaterialization(...).
  3. Apply sparse stream overrides with setStreamMaterialization(...).
  4. Inspect the effective result with resolveMaterialization(streamId).

Resolution deep-merges a namespace baseline and sparse stream overrides field by field; see resolution for the rules every resolver follows. Ursa additionally falls back to a cluster-wide default policy when a namespace sets none — see Appendix B: implementation notes.

The compactor can also bootstrap catalog records from operator properties rather than through the API. Table catalogs covers those property forms, the Iceberg backends Ursa packages, and how a policy selects among registered catalogs. Catalog properties are not interchangeable with Kafka broker configuration.

For a working local REST-catalog setup, use the Kafka lakehouse demo; do not copy its static demo credentials into a production deployment.

Table names and recreated topics

Without an explicit table identifier or namespace naming template, the table name is the source logical name (lakestream.source.logical.name), then the legacy Kafka topic-name property, then the storage stream name. Ursa 1.0.0 removed the TableMode distinction under which Ursa-owned tables used the storage stream name instead.

Kafka records the logical topic name in system-owned stream metadata. Recreating a Kafka topic creates a new storage incarnation, but can still append to the same topic-named external table. Deleting a topic does not imply deleting that external table or its history.

Namespace TableNaming templates support ${stream.namespace}, ${stream.name}, ${stream.logicalName}, and ${stream.property.<key>}. Unknown variables and missing or blank referenced properties fail resolution rather than silently selecting a default. The table namespace prefix is literal, not interpolated.

Changing a destination

Several catalogs can be registered for different warehouses or access boundaries, chosen through namespace baselines and stream overrides; see table catalogs. The compactor needs access to every catalog it might route to, and to that catalog's object store.

Before changing a catalog or table name, decide what happens to existing snapshots and files. Policy reassignment is not a data migration. For each output, verify decoded schema, partitioning, table commits, streaming reads after compaction, and retention independently.

Try it locally

With the Kafka image pulled or built as described in the Kafka quickstart, run from Kafka's Compose stack directory:

make lakehouse-demo

The script exercises Kafka → internal compaction plus external Iceberg → DuckDB and asserts the result. It is a disposable demo, not a production deployment recipe.