Lakestream
Ursa for Kafka

Monitoring

The metrics a diskless partition reports, how to tell diskless from classic in one scrape, and the log lines worth alerting on.

Ursa for Apache Kafka (UFK) reports diskless partitions through Kafka's existing metrics registry rather than a parallel one, so an existing JMX or Prometheus scrape picks them up without new wiring. What changes is which metrics a diskless partition carries and how you tell it apart from a classic one.

Telling diskless partitions apart

Diskless partition gauges are registered in the kafka.log group with type Log, the same place classic partition log metrics live, and carry one extra tag:

TagValue
topicThe topic name
partitionThe partition index
storagediskless

Classic partitions have no storage tag. That tag is how you split the two modes in a cluster running both — filter on it to chart diskless storage separately, and remember that a dashboard summing across kafka.log type Log without grouping by it is mixing the two.

Gauges on a diskless partition

GaugeMeaning
SizeBytes between the log's first and last entry
LogStartOffsetThe current readable start, which a retention soft trim advances
LogEndOffsetOne past the last durable record

LogStartOffset is the one to watch against retention: a soft trim moves it while the underlying objects remain until compaction advances the delete watermark. A growing gap between what this gauge reports and the storage a bucket actually holds is the signature of a compaction problem, not a retention problem. See operating the compactor.

Two gauges that a classic partition carries are not registered for a diskless partition, because neither has a meaning without a local log: NumLogSegments and RetentionSizeInPercent. A panel that plots them per topic goes blank for diskless topics rather than reporting zero.

Gauges lag by up to a second

The values behind these three gauges are refreshed from the underlying log at most once per second, rather than computed on each scrape. Scraping faster than that returns a repeated value, and a gauge can trail a burst by about a second. Treat them as trends, not as an instantaneous read.

Produce rates

A successful append to a diskless partition marks the standard broker topic stats, both per topic and cluster-wide:

  • bytesInRate — bytes accepted, measured on the record batches as received
  • messagesInRate — records accepted

These are the same meters classic topics use, so per-topic produce dashboards keep working across both storage modes.

Dependencies

A diskless deployment has two dependencies the broker cannot report on, and both need their own monitoring:

  • Metadata store. Oxia holds the stream catalog, storage metadata and idempotent-producer snapshots. Losing it is not a degraded read path, it is a stopped one. In the Compose stack, Oxia exposes Prometheus metrics at /metrics on port 8080, which is also what its health check probes.
  • Object storage. WAL and compacted object counts and bucket size are where a compaction backlog becomes visible before it becomes an incident. The Compose stack runs MinIO, whose liveness endpoint is /minio/health/live on port 9000.

The compactor's own progress and task errors come from its log. In the Compose stack, make compaction-logs follows it.

Log lines worth alerting on

The diskless path logs failures with distinctive text. These are broker-side, at ERROR unless noted:

Log textWhat it means
Diskless storage append future failedA produce to a diskless partition failed in the storage layer
Diskless storage fetch future failedA fetch from a diskless partition failed in the storage layer
Diskless storage listOffsets future failedAn offset lookup failed in the storage layer
[Diskless]Prefix on share-partition initialization failures for diskless topics

Four more are logged at WARN when a client attempts something diskless topics reject: Attempt to produce transactional records to diskless topic, Attempt to call WriteTxnMarkersRequest with diskless topic, Attempt to call AddPartitionsToTxnRequest with diskless topic, and Attempt to call TxnOffsetCommitRequest with diskless topic. A steady rate of these is an application pointed at the wrong storage mode rather than a broker fault — see Kafka compatibility.

Next steps