The data platform bill climbs every quarter, and pipelines break whenever a source system changes.
We design and build AWS-native data lakes and lakehouses on S3, Apache Iceberg, Spark, EMR, Glue and Lake Formation, with Snowflake where it fits. Ingestion absorbs schema drift on its own, orchestration runs on events instead of an always-on scheduler, and compute scales to real usage.
Proof. University of Wisconsin: a consolidated data lake on EMR and Iceberg, with an auto-relationalization engine that turns Workday and PeopleSoft exports into queryable tables. Led by Robin Tanner.
Field briefs: the lean data lake · auto-relationalization · event-driven orchestration · geospatial lookup