Schemas
Record formats Ursa decodes, schema registry configuration, the Protobuf constraint, and Iceberg variant columns.
Ursa stores record batches as bytes and does not need to understand them to append or serve them. It needs a schema only when materializing a stream into a table, because a table has typed columns.
That makes schema configuration a materialization concern: a deployment that compacts without writing tables needs none of it.
Supported formats
| Format | Supported |
|---|---|
| Avro | Yes |
| JSON Schema | Yes |
| Protobuf | Yes |
Schemas are resolved from a Confluent-compatible schema registry, keyed by the subject derived from the stream's logical topic name.
No Confluent-licensed artifacts are distributed
JSON Schema and Protobuf are decoded without shipping the Confluent Community License provider artifacts. Those are used at test scope only, so nothing under that license is redistributed.
Registry configuration
| Property | Purpose |
|---|---|
schemaRegistryUrl | Registry endpoint. Required for table materialization. |
schemaRegistryConfig<Name> | Passed through to the registry client with the prefix stripped. |
schemaRegistryHttpHeader<Name> | Sent as an HTTP header on every registry request; the name is lowercased. |
schemaRegistryHttpHeaderAuthorizationFile | File whose contents become the Authorization header, so the credential is not in the properties file. |
The authorization file is read once at startup and not re-read. For a rotating credential, catalogMaxOpenTimeInSeconds bounds how long a built catalog handle survives before being rebuilt.
Protobuf: one message type per topic
Protobuf allows many message types in one file, and Kafka records carry a message index identifying which was used. That index can only be read by reading the data.
Ursa resolves a stream's schema from the registry without reading records, so it requires that a topic carries exactly one message type, fixed when the topic is first written.
Nesting is unaffected. A message may reference others:
syntax = "proto3";
package io.lakestream.ursa.test;
message A {
B b = 1;
}
message B {
...
}Writing A to test-topic is fine, and B appears as a nested field. What is not allowed is also writing B directly to another topic from this same file — that topic needs B declared in a file of its own.
Plan the proto layout around this before producing data. Changing a topic's message type later is not a schema evolution Ursa can follow.
Schema evolution
When the source schema advances, the destination table schema is evolved to match, subject to what the sink supports.
| Property | Default | Effect |
|---|---|---|
tableEvolveSchemaEnabled | true | Apply schema changes to the table. |
make-new-fields-optional | false | Add newly seen fields as optional columns. |
check-nullability | true | Require nullability to match when comparing schemas. |
check-ordering | false | Require column order to match when comparing schemas. |
Records that do not fit the table schema are routed through the failure policy rather than dropped — by default to a dead-letter table named with the _dlt suffix. See lakehouse and table settings.
Which evolutions each sink accepts is listed per materializer on implementation status.
Iceberg variant columns
A variant column stores heterogeneous or semi-structured values in one column. Ursa recognizes logicalType: "variant" in Avro and JSON schemas and maps the value to an Iceberg variant column.
Three things must hold: the Iceberg catalog and the query engine must support variant, variantTypeEnabled must be on, and the registry must preserve the logical-type annotation — some serializers discard unknown logical-type properties, so check the registered schema before producing data.
Supported values are primitives (string, integer, long, float, double, boolean, bytes), maps, lists, arrays and sets, nested records, and whole records.
Declaring a variant column
In Avro:
{
"name": "attributes",
"type": {
"type": "map",
"values": "string",
"logicalType": "variant",
"variant-metadata-fields": "[\"region\", \"source\"]"
}
}In JSON, with reflection-based schema generation:
@JsonPropertyDescription("logicalType: variant")
private Map<String, String> attributes;What can change
Adding a variant field and removing one are both supported. Converting an existing non-variant field to variant, or a variant field to another type, is not.
variant-metadata-fields promotes nested fields for filtering and projection. Each one adds write and metadata cost, so list only the fields that are actually queried, and measure against the Iceberg engine version you run.