Posts - Page 3
All the articles I've posted, page 3.
- 31 MIN READ•Sep 10, 2026
Kafka Connect to Iceberg: How the Commit Actually Works
Exactly-once semantics in the Iceberg sink connector: the coordinator, the control topic, offsets stored inside Iceberg snapshots, and where duplicates still get in.
Apache KafkaKafka ConnectApache Iceberg - 31 MIN READ•Sep 10, 2026
What a Query Costs
Four meters, their real proportions, and how to attribute compute to a query, a table, and a team so a platform can answer what a dashboard costs to run.
lakehousecost optimizationunit economics - 31 MIN READ•Sep 10, 2026
The Open Lakehouse Explained, Then Built on Your Laptop with Dremio and MinIO
The five layers of the open lakehouse explained, then a lab: Parquet, Iceberg, Polaris, Arrow, and Ossie running in two containers on your own machine.
open lakehouseApache IcebergApache Polaris - 31 MIN READ•Sep 10, 2026
Where Lock-In Went
The format war ended and exit cost did not: where lock-in relocated after open tables won, how to measure it, and which costs are worth keeping down.
vendor lock-inopen formatsApache Iceberg - 30 MIN READ•Sep 2, 2026
Data Quality Tooling Compared: Great Expectations, Soda, dbt Tests, and Anomaly Detection
A comparison of Great Expectations, Soda, dbt tests, and anomaly detection, and a layered design that uses each where it fits.
Great ExpectationsSodadbt - 25 MIN READ•Sep 2, 2026
The Data Team of the Agentic Era: Generalists Owning End-to-End Workflows
The case for generalists owning end-to-end data workflows with agents, the counterargument, and how to make the transition work.
Data TeamAIGeneralists - 30 MIN READ•Sep 2, 2026
dbt on Iceberg: Incremental Models on Open Tables
How dbt incremental materializations map to Iceberg operations, and the configuration, predicates, and maintenance that keep them healthy.
dbtApache IcebergIncremental Models - 30 MIN READ•Sep 2, 2026
Disaster Recovery for Iceberg Tables: Replication, Backup, and Restore
Disaster recovery for Iceberg across four tiers: snapshots, object versioning, catalog backup, and cross-region replication.
Apache IcebergDisaster RecoveryReplication - 30 MIN READ•Sep 2, 2026
Deleting User Data From an Immutable Lakehouse: GDPR Hard Deletes on Iceberg
How to turn a logical delete on immutable Iceberg into a physical erasure across snapshots, versions, replicas, and downstream copies.
GDPRApache IcebergData Privacy - 32 MIN READ•Sep 2, 2026
Geospatial Data in Apache Iceberg: Geometry, Geography, and GeoParquet
How Iceberg v3 geometry and geography types, bounding boxes, and native Parquet types give spatial data first-class standing.
Apache IcebergGeospatialGeoParquet - 31 MIN READ•Sep 2, 2026
Default Column Values and Field IDs: How Iceberg Schema Evolution Works at the Spec Level
How field IDs and initial and write defaults let Iceberg change schemas on large tables without rewriting data, at the spec level.
Apache IcebergSchema EvolutionField IDs - 30 MIN READ•Sep 2, 2026
The Iceberg Table Properties That Actually Matter
The Iceberg table properties that decide file count, pruning, write amplification, retention, and metadata growth, by workload.
Apache IcebergTable PropertiesTuning