Quickstart
Run diskless Kafka, then produce Avro orders with a schema registry and query Iceberg through DuckDB.
The Kafka repository's Docker Compose stack runs Oxia, MinIO, three combined broker/controller nodes, and the Ursa compactor. The optional lakehouse profile adds Polaris, a Karapace Schema Registry, and a Karapace Kafka REST proxy. Demo workloads are separate opt-in profiles.
Local evaluation only
The stack uses plaintext listeners, fixed development credentials, and single-instance Oxia and MinIO. Published ports bind to 127.0.0.1. It is not a production deployment template.
Prerequisites
- Docker >= 20.10.4 and Docker Compose >= 2.20.
- Git and Make; curl for the optional HTTP examples.
- JDK 17 and Python >= 3.7, only if you build the image yourself instead of pulling it.
The commands use the 4.3-ursa branch and the lakestream/kafka:latest image, which bundles Ursa 1.0.0. No Ursa source checkout is needed: Kafka resolves the org.openlakestream artifacts from Maven Central.
Get the image
One image, lakestream/kafka, carries the brokers, the Kafka CLI tools, and the standalone Ursa compactor at /opt/kafka/bin/ursa-compactor.sh. Releases are pushed to Docker Hub from the Kafka repository's v* tags; latest tracks the newest one. Clone the repository for the Compose files and pull the image:
git clone --branch 4.3-ursa https://github.com/openlakestream/kafka.git
cd kafka/docker/examples/docker-compose-files/cluster/ursa
docker pull lakestream/kafka:latestThe Compose file reads the broker and compactor image from IMAGE and defaults to lakestream/kafka:latest. Set IMAGE (for example export IMAGE=lakestream/kafka:4.3.1.3) to pin a release; keep it exported in every shell that runs make or docker compose here.
To build the image from the checkout instead, run:
./build-image.shbuild-image.sh runs ./gradlew releaseTarGz, then resolves the compactor's classpath from org.openlakestream:ursa-storage-compact on Maven Central at the Ursa version the tarball ships. The result is lakestream/kafka:latest; the --amd64 option targets linux/amd64. There is no separate compactor image and no local Maven install step.
Run all remaining commands from the Compose directory above. Kafka CLI examples use tools inside a running container, so they do not depend on a nonexistent ./bin directory here.
Start, produce, and consume
make up
docker compose up -d --wait
make ps
make create-topic
make list-topics
make produce
make consumemake create-topic creates test-diskless with 12 partitions and RF=1. Override it with make create-topic TOPIC=my-topic PARTITIONS=6. Pass the same TOPIC override to producer/consumer targets when using another topic.
For interactive records, run make console-producer and make console-consumer in separate terminals. To create a topic directly:
docker compose exec kafka-1 /opt/kafka/bin/kafka-topics.sh \
--bootstrap-server kafka-1:19092 \
--create \
--if-not-exists \
--config ursa.storage.enable=true \
--topic my-diskless-topic \
--partitions 6 \
--replication-factor 1Inspect storage and services
| Service | Host port | Purpose |
|---|---|---|
| kafka-1 / kafka-2 / kafka-3 | 29092 / 39092 / 49092 | Host Kafka clients |
| Oxia | 6648 | Metadata |
| MinIO | 19000 | S3 API |
| MinIO console | 19001 | Object browser |
| Polaris (optional) | 18181 / 18182 | Catalog API / health and metrics |
| Schema Registry (optional) | 18081 | Karapace's Confluent-compatible schema REST API |
| Kafka REST proxy (optional) | 18082 | Schema-aware HTTP produce and consume |
Open http://localhost:19001 with username and password minioadmin. The kafka-ursa bucket holds the WAL under ursa/wal and compacted objects under ursa/compacted. The separate lakehouse bucket holds the external Iceberg warehouse.
make logs
make compaction-logsThe compactor is part of the default stack because WAL reclamation depends on compaction. Kafka retention alone advances a logical offset; it does not free WAL objects immediately.
Try broker failover
After producing records to test-diskless:
docker compose stop kafka-1
docker compose exec kafka-2 /opt/kafka/bin/kafka-console-consumer.sh \
--bootstrap-server kafka-2:19092 \
--topic test-diskless \
--from-beginningAllow time for broker failure detection, metadata refresh, and client retries. Remaining brokers can serve the shared logs without copying partition data. Stop the console consumer with Ctrl+C, then restart the broker:
docker compose start kafka-1This demonstrates serving-owner remapping, not uninterrupted availability or a test of Oxia/MinIO redundancy.
Add an external Iceberg table
URSA_MATERIALIZATION_ENABLED=true docker compose --profile lakehouse up -d --wait
make schemasThe environment variable enables the compactor's external sink. The profile adds Polaris and its setup job, plus Karapace running as a Schema Registry and a Kafka REST proxy. The registry stores schemas in a classic _schemas topic on the same brokers. The REST proxy lets curl produce schema-encoded records without a schema-aware console producer.
Compose recreates the compactor when its environment changes. Without the variable, starting the profile leaves the compactor writing only the compacted objects Kafka reads. The external table is a separate output, not the files used for Kafka streaming reads. See Ursa lakehouse tables.
The external table uses the logical Kafka topic name, while the internal stream includes the topic UUID. Recreating a Kafka topic with the same name can append to an existing external table; a new Kafka topic is not an automatic reset of that table.
Polaris uses an in-memory development metastore here. Its catalog is not a durable production catalog even though MinIO and the Kafka nodes have data volumes.
Produce your own schema-backed records
With the lakehouse profile running, create a new diskless topic and send Avro records through the REST proxy:
make create-topic TOPIC=events PARTITIONS=1
curl -X POST localhost:18082/topics/events \
-H 'Content-Type: application/vnd.kafka.avro.v2+json' \
-d '{"value_schema": "{\"type\":\"record\",\"name\":\"Event\",\"fields\":[{\"name\":\"id\",\"type\":\"long\"},{\"name\":\"kind\",\"type\":\"string\"}]}",
"records": [{"value": {"id": 1, "kind": "click"}}, {"value": {"id": 2, "kind": "view"}}]}'
make schemasThe proxy registers the schema under events-value and encodes the values in Confluent wire format. The compactor resolves the <topic>-value subject to materialize typed Iceberg columns. Plain JSON text sent with make console-producer is not equivalent to these schema-encoded records.
| Value schema | External table shape |
|---|---|
Avro, JSON Schema, or Protobuf registered under <topic>-value, with matching wire-format records | Typed columns from the schema |
| No registered subject | A single binary payload column containing the raw record value |
Register schemas before producing
The compactor caches a missing subject and treats that topic as raw bytes until it restarts. Register the schema, or use a schema-aware producer, before the first record reaches the compactor. If the subject was already cached as missing, restart the compactor after registration; this does not retroactively convert an existing binary table into a typed table.
Only the external Iceberg sink uses the registry. Compacted objects preserve the original Kafka record batches, and brokers do not need the registry to serve them. A schema is not required for a diskless topic to function as a Kafka topic.
Use make duckdb to open a SQL shell once the compactor has materialized records. Tables appear asynchronously; use make compaction-logs in another terminal to follow progress.
End-to-end verification demo
Use a fresh Compose project for this verifier. The first command deletes this demo stack's containers and volumes, including any records from the earlier steps:
make destroy
make lakehouse-demoThe verifier enables the lakehouse services and materialization, then:
- Creates
ursa-lakehouse-e2eand registers anOrderAvro schema underursa-lakehouse-e2e-value. - Produces 100 orders through the REST proxy, with
order_idvalues from 1 to 100. - Decodes all orders through the proxy's Avro consumer, waits for compacted Parquet objects, and reads them again with a fresh consumer group to check the compacted read path.
- Polls DuckDB until the external table has the expected
count(*)andsum(order_id). For 100 orders, these are 100 and 5050. It prints the columns, the first five orders, and a per-region aggregate. - Checks that the range was materialized once and that the compactor did not report task errors.
The typed table contains order_id, customer, region, quantity, amount, and order_ts_ms. The timestamp is epoch milliseconds stored as a plain integer, not an Avro logical timestamp, to work with the REST proxy's JSON input.
On success, the verifier removes containers and volumes unless KEEP_RUNNING=true; on failure it retains the stack for debugging. Do not run it against data you need to retain. A retained or failed run must be cleaned up with make destroy before retrying.
For inspection, run this instead of make lakehouse-demo, starting with a fresh project:
KEEP_RUNNING=true ./run-lakehouse-demo.sh
make schemas
curl -s localhost:18081/subjects/ursa-lakehouse-e2e-value/versions/latest
make duckdbInside DuckDB:
SHOW ALL TABLES;
DESCRIBE lakehouse.default."ursa-lakehouse-e2e";
SELECT region, count(*) AS orders, round(sum(amount), 2) AS revenue,
min(epoch_ms(order_ts_ms)) AS first_order
FROM lakehouse.default."ursa-lakehouse-e2e"
GROUP BY region ORDER BY region;Use NUM_RECORDS=500 ./run-lakehouse-demo.sh to change the order count. The verifier lowers the WAL flush interval to 100 ms and the flush threshold to 4096 bytes. These are verifier overrides, not normal broker defaults.
The compactor's entrypoint also sets managedTableSchemaEvolutionEnabled=true, a key from earlier Ursa releases that Ursa 1.0.0 no longer defines: every compacted-object write now records the per-file index Kafka needs to read a range after compaction, with no switch. The second consumer check helps catch a missing index.
Other demos and cleanup
make demo runs the producer/consumer performance profile; make share-demo runs the share-group profile. Both tear down the stack and remove volumes on exit, including Ctrl+C. Use them only with disposable demo data.
# Stop every profile's services, retaining volumes
make down
# Stop services and delete the demo volumes
make destroyIf services are not ready, inspect docker compose ps, docker compose logs oxia, and docker compose logs minio. If the compactor container is missing, check that IMAGE names a pulled or locally built lakestream/kafka image: the compactor runs from the same image as the brokers and does not require the lakehouse profile. Use make compaction-logs when retention is not reclaiming storage or an expected table does not appear.
For schema registration or HTTP produce/consume failures, inspect docker compose logs schema-registry kafka-rest and confirm the lakehouse profile is healthy. If an expected typed table contains only payload, check that the <topic>-value subject existed before compaction and that the producer used the matching wire format.
Next steps
- Deployment — run the same setup on Kubernetes with Strimzi.
- Configuration — exact property names and defaults.
- Diskless architecture — storage and compaction paths.
- Limitations — workload and operational boundaries.