A complete, governed data platform, ready to deploy
The architecture of Databricks, Snowflake and Microsoft Fabric, built entirely on open source. It runs on a laptop or your own Kubernetes cluster, with no licences, no cloud bill and no vendor lock-in.
Data flows from sources through Bronze, Silver and Gold layers to SQL, dashboards and AI, orchestrated by Airflow and governed by Polaris.
Everything a modern data platform needs
Runs anywhere
One command on a laptop with Docker Compose, or a full multi-user deployment on Kubernetes with Helm and Ansible.
Governed by design
Every table sits in an Apache Polaris catalog with users, roles and per-table grants, and each user works in an isolated workspace.
Open table formats
Apache Iceberg tables on S3-compatible storage. Spark writes, Trino reads, and your data never sits in a proprietary format.
Batch to dashboard
Ingestion, Bronze/Silver/Gold pipelines, MERGE upserts, data quality checks and SCD2 history, all scheduled by Airflow.
Notebooks, SQL and BI
Jupyter with Spark Connect, SQLPad on Trino and Superset dashboards, all behind one portal.
Self-service management
A management console to create, reset and remove users, and a portal where users browse files and share catalogs.
Where clients use it
Proofs of concept
Validate a lakehouse design in days, before committing to a cloud contract.
Migration rehearsal
Prototype and test pipelines before moving to Databricks, Snowflake or Fabric.
Dev & test environments
Give every engineer a full platform without growing the cloud bill.
On-premises analytics
A governed lakehouse where data must stay inside your own network.
Every component maps to the major platforms
What you build on the Lakehouse Stack carries over. Patterns, SQL and pipelines translate directly when you move to the cloud.
| Layer | Lakehouse Stack | Databricks | Snowflake | Microsoft Fabric | AWS |
|---|---|---|---|---|---|
| Object storage | S3-compatible storage | S3 / ADLS | Managed storage | OneLake | Amazon S3 |
| Table format | Apache Iceberg | Delta Lake (+ Iceberg) | Iceberg tables | Delta Lake | Iceberg / S3 Tables |
| Processing | Apache Spark | Databricks Spark | Snowpark | Fabric Spark | EMR / Glue |
| SQL engine | Trino | Databricks SQL | Virtual warehouses | Fabric Warehouse | Athena |
| Orchestration | Apache Airflow | Lakeflow Jobs | Tasks | Data Factory | MWAA |
| Catalog & governance | Apache Polaris | Unity Catalog | Horizon / Open Catalog | OneLake catalog | Glue + Lake Formation |
| Notebooks | Jupyter | Databricks notebooks | Snowflake Notebooks | Fabric notebooks | SageMaker Studio |
| Dashboards | Apache Superset | AI/BI dashboards | Streamlit | Power BI | QuickSight |
Two ways to run it
Single machine · Docker Compose
The whole platform on one laptop or server, ideal for proofs of concept, demos and individual development. Up and running with a handful of commands.
Multi-user · Kubernetes
A shared deployment with Helm and Ansible, isolated per-user workspaces, generated secrets and a management console, for teams and on-premises use.
- Apache Spark
- Apache Iceberg
- Trino
- Apache Airflow
- Apache Polaris
- Jupyter
- Apache Superset
- SQLPad
- PostgreSQL
- S3 object storage
- Kubernetes
- Helm
- Ansible
- Docker
See the Lakehouse Stack in action
Book a demo and we'll walk you through a live deployment, and how it could fit your data.