Tag: parquet
All the articles with the tag "parquet".
- 31 MIN READ•Aug 24, 2026
Variant Shredding Explained: How Iceberg Gets Columnar Performance From Messy JSON
Variant shredding turns messy JSON into Parquet columns with statistics. How the layout works, how readers reassemble values, and why some queries prune.
Apache IcebergVariantParquet - 21 MIN READ•Aug 4, 2026
A Migration Playbook for Moving Legacy Warehouses onto Apache Iceberg
A dependency-first playbook for migrating legacy warehouses onto Apache Iceberg: snapshot vs migrate vs add_files, four-level parity validation, and federation-based cutover.
Apache IcebergMigrationData Warehouse - 31 MIN READ•Jul 28, 2026
The Parquet Versioning Problem, and Why Iceberg Cares About It
Parquet files have a version field that doesn't reliably signal feature requirements. A new versioning discipline is coming, borrowing from Iceberg's format version model.
ParquetApache IcebergData Engineering - 27 MIN READ•Jul 6, 2026
A Deep Dive Into File Compression: How Data Gets Smaller, Why Codecs Differ, and What to Actually Use in the Lakehouse
Somewhere in your data platform right now, a single configuration property is quietly deciding a meaningful percentage of your storage bill, your q...
file compressioncodecsparquet - 27 MIN READ•Jul 6, 2026
The File Format Renaissance: Parquet, Lance, Vortex, Nimble, BtrBlocks, and the New Physics of Columnar Storage
For a decade, the file format layer was the most settled real estate in data. Apache Parquet held the analytical world, ORC held the Hive legacy es...
parquetlancevortex