Lakestream and the Streamhouse
Streamhouse names a category of data architecture; Lakestream is a standard for one layer inside it, and this page states the relationship and its limits.
"Streamhouse" names a category of data architecture. Lakestream is a standard for one layer inside such an architecture. The two are easy to conflate, so this page states the relationship and its limits.
The Streamhouse definition
The Streamhouse Working Group maintains an open, vendor-neutral definition of the category at streamhouse.com:
Streamhouse architectures empower organizations to capture, transport, transform, govern, and serve the current state of their business continuously, so that production applications and agents can act on it.
The definition gives the category three attributes:
- Real-time — data remains continuously current as business events occur, rather than being refreshed only through periodic batch processes.
- Production-native — the architecture is engineered to production service levels, because business-critical applications, analytics, and agents depend on it continuously.
- Decentralized — it meets data where it already lives rather than requiring it to be consolidated first.
The Working Group's page carries the authoritative text; the summary above is not a substitute for it.
Same ground
A Streamhouse and a lakehouse stand on the same foundation: object storage for bytes, open table formats for tabular data, and catalogs for discovery and governance. Neither replaces that foundation.
What differs is what the architecture is asked to do with it. A lakehouse analyzes the business: its tables are read by query engines, and freshness is measured against a reporting cycle. A Streamhouse runs the business: applications and agents act on the current state, so that state has to stay available continuously and at production service levels. Because the foundation is shared, building one does not require replacing the other.
What each house needs open
A lakehouse becomes open when the table format is open. Once file layout, metadata, and snapshots are published, more than one engine can read the same tables, and a table stops being a property of whichever engine wrote it.
A Streamhouse needs that, and it needs the log open as well. The current state of a business arrives as a stream before it lands as a table. When the log's storage format and API stay internal to a broker, the stream is a property of that broker: consuming it elsewhere means exporting it, and the table produced from it is a second retained copy on a separate lifecycle. Opening the table format alone does not reach the part of the architecture that carries the current state. See why Lakestream for the costs that follow from that split.
Where Lakestream sits
Lakestream is an open standard for the stream-storage layer. Two parts of it map onto the architecture:
- The stream-storage API and spec is the open-stream-storage block: a protocol-agnostic contract for namespaces, streams, logs, offsets, layouts, and durable appends on object storage. See the specification and the API reference.
- The stream materialization framework is how that layer feeds the lakehouse's tables instead of duplicating them. Its policy model states which stream becomes which table and under what write, partition, evolution, and commit settings, and materialization runs inside the storage lifecycle rather than in a separately operated export path. Stream–table duality describes what that does and does not guarantee.
A Streamhouse can use Lakestream for that layer, or other stream storage. The category is defined by what the architecture does, not by which implementation sits underneath it.
What Lakestream does not claim
- Lakestream is not a Streamhouse. It specifies one layer. An architecture is more than a storage contract, and the definition above covers transport, transformation, governance, and serving as well.
- Lakestream is not required for a Streamhouse. Nothing in the definition names a storage standard, and an architecture built on other stream storage is not excluded by it.
- Lakestream is not a Working Group deliverable. It is a separate open source project with its own repositories, specification, and release history. The Working Group's definition is cited here, not co-authored, and citing it is not a statement of endorsement in either direction.
For what Lakestream specifies, start with the specification and the API reference. For the model behind it, see what is Lakestream.
Why Lakestream
Lakestream addresses two forms of coupling: brokers that must store the partition data they serve, and lakehouse tables that require a separate export pipeline.
Stream–table duality
How one retained log is consumed as a stream and materialized as a table, and where the two outputs diverge.