An independent resource forthe open lakehouse
Reference material on Apache Iceberg, lakehouse catalogs, the agentic lakehouse, and modern data architecture. It covers what table formats are, how to deploy Apache Polaris, and how to connect query engines to Iceberg tables. Written by a practitioner, free to read.
Not an Apache project. This is a personal site by Alex Merced. It is not affiliated with, endorsed by, or sponsored by the Apache Software Foundation or the Apache Iceberg project, whose official site is iceberg.apache.org.
- 500+
- Articles
- 200+
- Reference entries
- 14
- Subject areas
- Free
- No paywall, ever
Browse by topic
Nine long-form pillar guides that cover the lakehouse stack from the file format up to the agents querying it.
- Pillar guide
Apache Iceberg
Covers the metadata tree, snapshots, hidden partitioning, and the catalog API from end to end.
Read the guide - Deep dive
Iceberg Architecture
How manifest lists, manifest files, and data files fit together at query time.
Read the guide - Catalogs
The REST Catalog
What the Iceberg REST Catalog spec standardizes, and how engines authenticate against it.
Read the guide - Deep dive
Snapshots & Time Travel
Atomic commits, snapshot expiration, rollback, and querying a table as of any point in time.
Read the guide - Deep dive
Schema Evolution
Add, drop, rename, and reorder columns safely, and why Iceberg's field IDs make it work.
Read the guide - Comparison
Iceberg vs Delta Lake vs Hudi
A neutral comparison of the three open table formats across design, features, and ecosystem.
Read the guide - Pillar guide
The Data Lakehouse
What a lakehouse actually is, the layers it is built from, and how it differs from a warehouse.
Read the guide - Agentic AI
The Agentic Lakehouse
Semantic layers, MCP, and the architecture AI agents need to query your data reliably.
Read the guide - Foundations
Open Table Formats
Why table formats exist at all, and the problems they solved for data lakes.
Read the guide
Recent Posts
- 30 MIN READ•Sep 2, 2026
Data Quality Tooling Compared: Great Expectations, Soda, dbt Tests, and Anomaly Detection
A comparison of Great Expectations, Soda, dbt tests, and anomaly detection, and a layered design that uses each where it fits.
Great ExpectationsSodadbt - 25 MIN READ•Sep 2, 2026
The Data Team of the Agentic Era: Generalists Owning End-to-End Workflows
The case for generalists owning end-to-end data workflows with agents, the counterargument, and how to make the transition work.
Data TeamAIGeneralists - 30 MIN READ•Sep 2, 2026
dbt on Iceberg: Incremental Models on Open Tables
How dbt incremental materializations map to Iceberg operations, and the configuration, predicates, and maintenance that keep them healthy.
dbtApache IcebergIncremental Models - 30 MIN READ•Sep 2, 2026
Disaster Recovery for Iceberg Tables: Replication, Backup, and Restore
Disaster recovery for Iceberg across four tiers: snapshots, object versioning, catalog backup, and cross-region replication.
Apache IcebergDisaster RecoveryReplication - 30 MIN READ•Sep 2, 2026
Deleting User Data From an Immutable Lakehouse: GDPR Hard Deletes on Iceberg
How to turn a logical delete on immutable Iceberg into a physical erasure across snapshots, versions, replicas, and downstream copies.
GDPRApache IcebergData Privacy - 32 MIN READ•Sep 2, 2026
Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet
How Iceberg v3 geometry and geography types, bounding boxes, and native Parquet types give spatial data first-class standing.
Apache IcebergGeospatialGeoParquet
Learn alongside the rest of the community
A Slack workspace for lakehouse practitioners, plus a shared calendar of meetups, webinars, and Lakehouse Linkups.
Must reads on Iceberg, agentic AI, and the lakehouse
The Definitive Guide to the Semantic Layer
Understand what a semantic layer is, why it matters for modern data architectures, and how it creates a consistent, governed layer between raw data and business consumers.
Read articleApache Polaris: The Catalog Standard for Lakehouses and AI
A deep dive into Apache Polaris, the open-source catalog that is emerging as the standard for managing Iceberg tables across multi-engine Lakehouses and AI workloads.
Read articleWhat Are Table Formats and Why Were They Needed?
Explore the history and motivations behind open table formats like Apache Iceberg, Delta Lake, and Apache Hudi, and why they solved critical problems in big data engineering.
Read articleWhat is Dremio?
An overview of Dremio's Lakehouse platform: how it unifies data access, accelerates queries, and powers self-service analytics across cloud and on-premise sources.
Read articleWhat Apache Iceberg Native Actually Means
Not all Iceberg integrations are equal. This article breaks down what it truly means for a platform to be 'Apache Iceberg native' and why the distinction matters for your architecture.
Read articleOpen Source and the Data Lakehouse
A survey of the open source ecosystem powering modern Data Lakehouses, from Apache Iceberg and Nessie to Apache Arrow and Spark, and how they work together.
Read articleWhat is Agentic Analytics?
How AI agents are changing analytics pipelines by querying data, generating insights, and taking actions on their own, and what that means for the Lakehouse.
Read article
This site is an independent publication by Alex Merced. It is not affiliated with, endorsed by, or sponsored by the Apache Software Foundation or the Apache Iceberg project, whose official home is iceberg.apache.org.