Lakestream
Ursa

Embedding Ursa

Dependencies, the two ways to open a catalog, classpath isolation, and close ordering.

The quickstart opens a catalog, writes to a stream and reads it back in a single program. This page covers what that example leaves out: which artifacts to depend on, the difference between the two ways to open a catalog, and the ordering rules a long-running embedder has to respect.

For the interface contracts themselves, see the API reference. This page describes Ursa; that one describes what any implementation promises.

Dependencies

All artifacts are published to Maven Central under org.openlakestream at 1.0.0. The Java packages are io.lakestream.* — the group and the package names deliberately differ.

ArtifactRole
lakestream-apiThe interfaces you compile against. No implementation, no runtime dependencies.
ursa-storage-coreThe storage engine: WAL, object-store backends, caches.
ursa-storage-lakestreamUrsa's implementation of the Lakestream API.

Compile against lakestream-api alone and keep the other two on the runtime classpath. That is what lets an integration swap implementations without recompiling, and it is how Ursa for Apache Kafka (UFK) is built.

Ursa requires JDK 17. Later JDKs are not usable for building the project itself.

A metadata store is required

There is no embedded or in-memory metadata mode. Offsets are assigned by the metadata store, so a single-process embedding against local disk still needs Oxia running. A "local only, no dependencies" setup is not possible.

Two ways to open a catalog

The choice determines which capabilities you get, so make it deliberately.

Through the ServiceLoader

StreamCatalog catalog = StreamCatalogLoader.open(catalogMetadataUri, properties);

StreamCatalogLoader discovers a single StreamCatalogProvider on the class loader, and fails if it finds none or more than one. A class-loader overload exists for runtimes that isolate the implementation.

The provider registered by ursa-storage-kafka-runtime assembles more than the catalog: it wires the compacted-object reader, applies backend-name normalization, and bootstraps OpenTelemetry. Properties are defensively copied before the provider sees them.

This is the path to prefer. It is the one an integration is expected to use, and it produces a fully wired runtime.

Directly

IndexedStreamCatalog catalog = new StreamCatalogService()
        .open(catalogMetadataUri, new DefaultCatalogPaths(), properties, otel);

StreamCatalogService constructs the catalog itself. It is the lower-level path, and it does less for you.

The difference that matters: the compacted-object reader is only installed when you supply one. With no compactedObjectReaderFactoryClass property and no factory passed in, reads of data that compaction has already rewritten fail. A program that writes and reads recent data works; one that reads back far enough to cross a compaction boundary does not.

Supply the reader factory if you open the catalog this way and intend to read compacted ranges.

Classpath isolation

lakestream-api has no runtime dependencies, which makes it safe on a host application's main classpath. The implementation is a different matter: ursa-storage-core and ursa-storage-lakehouse bring Netty, the cloud SDKs, Hadoop and the table-format libraries.

UFK loads the implementation on an isolated runtime classpath for exactly this reason. Do the same in any host that has its own opinions about those libraries.

Lifecycle and close ordering

Handles are owned by the catalog that produced them, and the ordering is not advisory:

  1. Close writers, readers and cursors.
  2. Close logs.
  3. Close the catalog.

A log opened through the catalog holds a durable write lease. Closing it rejects new operations, waits for operations already in flight and for child cursors to finish, closes the delegate, and only then releases the lease. Closing the catalog first cuts that sequence short.

A write lease is per holder rather than per log, so any number of holders may append to and trim the same log concurrently; each gets its own lease.

Startup is lazy in the other direction. Constructing the storage engine does metadata work only; the object-store client, the WAL and the direct-memory caches are created on the first data-plane operation. A program that creates a catalog and never writes never allocates them.

Buffer ownership

Entries carry reference-counted buffers. Ownership transfers on the call — a write takes ownership of the buffer you hand it, including when it rejects the write, and a read hands you entries you must close exactly once.

This is the one rule whose violation shows up as a slow leak rather than an immediate error. See buffer ownership in the API reference.

Configuration

An embedder passes a Properties object rather than a file, but the keys are the same ones the configuration reference documents. Unrecognized keys are retained rather than rejected, which is how integrations pass their own settings through the same object.