Monitoring
The metrics a diskless partition reports, how to tell diskless from classic in one scrape, and the log lines worth alerting on.
Ursa for Apache Kafka (UFK) reports diskless partitions through Kafka's existing metrics registry rather than a parallel one, so an existing JMX or Prometheus scrape picks them up without new wiring. What changes is which metrics a diskless partition carries and how you tell it apart from a classic one.
Telling diskless partitions apart
Diskless partition gauges are registered in the kafka.log group with type Log, the same place classic partition log metrics live, and carry one extra tag:
| Tag | Value |
|---|---|
topic | The topic name |
partition | The partition index |
storage | diskless |
Classic partitions have no storage tag. That tag is how you split the two modes in a cluster running both — filter on it to chart diskless storage separately, and remember that a dashboard summing across kafka.log type Log without grouping by it is mixing the two.
Gauges on a diskless partition
| Gauge | Meaning |
|---|---|
Size | Bytes between the log's first and last entry |
LogStartOffset | The current readable start, which a retention soft trim advances |
LogEndOffset | One past the last durable record |
LogStartOffset is the one to watch against retention: a soft trim moves it while the underlying objects remain until compaction advances the delete watermark. A growing gap between what this gauge reports and the storage a bucket actually holds is the signature of a compaction problem, not a retention problem. See operating the compactor.
Two gauges that a classic partition carries are not registered for a diskless partition, because neither has a meaning without a local log: NumLogSegments and RetentionSizeInPercent. A panel that plots them per topic goes blank for diskless topics rather than reporting zero.
Gauges lag by up to a second
The values behind these three gauges are refreshed from the underlying log at most once per second, rather than computed on each scrape. Scraping faster than that returns a repeated value, and a gauge can trail a burst by about a second. Treat them as trends, not as an instantaneous read.
Produce rates
A successful append to a diskless partition marks the standard broker topic stats, both per topic and cluster-wide:
bytesInRate— bytes accepted, measured on the record batches as receivedmessagesInRate— records accepted
These are the same meters classic topics use, so per-topic produce dashboards keep working across both storage modes.
Dependencies
A diskless deployment has two dependencies the broker cannot report on, and both need their own monitoring:
- Metadata store. Oxia holds the stream catalog, storage metadata and idempotent-producer snapshots. Losing it is not a degraded read path, it is a stopped one. In the Compose stack, Oxia exposes Prometheus metrics at
/metricson port8080, which is also what its health check probes. - Object storage. WAL and compacted object counts and bucket size are where a compaction backlog becomes visible before it becomes an incident. The Compose stack runs MinIO, whose liveness endpoint is
/minio/health/liveon port9000.
The compactor's own progress and task errors come from its log. In the Compose stack, make compaction-logs follows it.
Log lines worth alerting on
The diskless path logs failures with distinctive text. These are broker-side, at ERROR unless noted:
| Log text | What it means |
|---|---|
Diskless storage append future failed | A produce to a diskless partition failed in the storage layer |
Diskless storage fetch future failed | A fetch from a diskless partition failed in the storage layer |
Diskless storage listOffsets future failed | An offset lookup failed in the storage layer |
[Diskless] | Prefix on share-partition initialization failures for diskless topics |
Four more are logged at WARN when a client attempts something diskless topics reject: Attempt to produce transactional records to diskless topic, Attempt to call WriteTxnMarkersRequest with diskless topic, Attempt to call AddPartitionsToTxnRequest with diskless topic, and Attempt to call TxnOffsetCommitRequest with diskless topic. A steady rate of these is an application pointed at the wrong storage mode rather than a broker fault — see Kafka compatibility.
Next steps
- Operating the compactor — what a stalled compaction does to storage.
- Limitations — the behaviors these signals are measuring against.
- Configuration — the settings that move these numbers.