Lakehouse interoperability: files, tables, catalogs and writers
Lakehouse interoperability is the ability of different engines to perform specified operations on a shared logical table while preserving the table’s defined behavior. Reading its data files establishes a narrower result than reading the correct snapshot or safely committing an update.
Files and the logical table
A data file contains encoded values. A table definition identifies which files and metadata constitute a logical dataset at a particular state. The logical table can include schema information, partition descriptions and a history of committed snapshots.
Iceberg documents snapshots, schema evolution and concurrent table changes. Those features make table state more than a directory of readable files. An engine needs the metadata interpretation required for the operation it performs.
Take two snapshots of a table, an earlier state and a corrected state. Some earlier files remain stored so the earlier snapshot can still be read. Scanning every file in the storage location does not select either snapshot correctly. The reader needs the metadata that identifies the intended state.
Reader compatibility and feature support
A compatible reader must implement the features used by the selected table version. File decoding is one requirement. Correct schema interpretation and membership of the selected snapshot are additional requirements.
Take a hypothetical table format that represents a later removal through separate metadata. A reader that decodes the original files but ignores that metadata can return a record that the logical table excludes. The example identifies a feature-support boundary; it does not assert that a particular named engine has that defect.
A read-compatibility claim therefore names the table format and version, the features in use and the selected operation. Support for one reader or table configuration does not establish support for every configuration carrying the same format name.
Catalogs and coordinated writes
A catalog participates in locating and managing table metadata according to its implementation. Writers need to agree on how a proposed table change becomes the committed state. Shared access to storage is not itself a commit protocol.
Take two writers that both start from the same table state. One adds newly received records while the other applies a correction. Preserving the intended result requires the table’s supported coordination and conflict behavior. Allowing both writers to replace the current metadata without that coordination can lose one change or produce an unintended combination.
Read access therefore establishes less than safe write access. An exit or integration test needs to exercise the operations the replacement engine will actually perform, including concurrent changes and recovery after an interrupted attempt.
Permissions outside the query service
A principal analytics service can enforce its own query permissions. Another engine with access to underlying storage creates an additional access path. Whether equivalent restrictions apply depends on the storage, identity and engine policies along that path.
The fact that an authorized query returned only permitted rows does not establish that a separately authorized file reader is subject to the same filtering. Governance needs the actual set of reading and writing paths.
Scope of portability
A portable workload needs compatible files, table metadata, catalog behavior, operation semantics and permissions. Financial replay adds the required snapshots, identifier history and calculation inputs.
What changes between implementations is the supported set of operations and dependencies. An open format can improve substitution options, but it does not establish universal multi-engine writes, equivalent access controls or a complete financial migration.
Questions about table format
If two engines read Parquet, can they safely write the same table?
No. Safe table writes also require compatible metadata features and commit coordination.
Does access to all stored files identify the current table?
No. The current logical state depends on the table metadata and selected snapshot.
Does an open table format guarantee a simple vendor exit?
No. Catalogs, permissions, calculations and operational dependencies can still require migration work.
Sources and method
- Introduction Apache Software Foundation
- Snowflake key concepts and architecture Snowflake
- Delta Lake Delta Lake project
- BigQuery overview Google Cloud
- About OpenLineage OpenLineage Project
Read next
- Transaction databases, tick stores and search indexes
Ledgers, tick stores and search indexes preserve different facts. Learn why one database rarely serves every financial workload.
- Vendor exit costs depend on reconstructing financial state
Exported files are only part of a vendor exit. Reconstructing balances, history and operating state can determine the real migration cost.
- Data lineage, quality checks and reconciliation establish different facts
Knowing where data came from does not prove it is complete or correct. Compare lineage, quality checks and reconciliation.
