Home / Our Solutions / Lakehouse Stack
Epireum Lakehouse Stack

A complete, governed data platform, ready to deploy

The architecture of Databricks, Snowflake and Microsoft Fabric, built entirely on open source. It runs on a laptop or your own Kubernetes cluster, with no licences, no cloud bill and no vendor lock-in.

Apache Airflow · orchestrates and schedules every stepLAKEHOUSE · APACHE SPARK + APACHE ICEBERG ON S3-COMPATIBLE STORAGESourcesDatabasesFiles & ExcelAPIs & eventsBronzeraw, as ingestedSilvercleaned, joined,quality-checkedGoldbusiness-readydata productsConsumeSQL (Trino)DashboardsNotebooks & AIApache Polaris · catalog, users, roles and per-table access

Data flows from sources through Bronze, Silver and Gold layers to SQL, dashboards and AI, orchestrated by Airflow and governed by Polaris.

Capabilities

Everything a modern data platform needs

Runs anywhere

One command on a laptop with Docker Compose, or a full multi-user deployment on Kubernetes with Helm and Ansible.

Governed by design

Every table sits in an Apache Polaris catalog with users, roles and per-table grants, and each user works in an isolated workspace.

Open table formats

Apache Iceberg tables on S3-compatible storage. Spark writes, Trino reads, and your data never sits in a proprietary format.

Batch to dashboard

Ingestion, Bronze/Silver/Gold pipelines, MERGE upserts, data quality checks and SCD2 history, all scheduled by Airflow.

Notebooks, SQL and BI

Jupyter with Spark Connect, SQLPad on Trino and Superset dashboards, all behind one portal.

Self-service management

A management console to create, reset and remove users, and a portal where users browse files and share catalogs.

Use cases

Where clients use it

Proofs of concept

Validate a lakehouse design in days, before committing to a cloud contract.

Migration rehearsal

Prototype and test pipelines before moving to Databricks, Snowflake or Fabric.

Dev & test environments

Give every engineer a full platform without growing the cloud bill.

On-premises analytics

A governed lakehouse where data must stay inside your own network.

Cloud-ready

Every component maps to the major platforms

What you build on the Lakehouse Stack carries over. Patterns, SQL and pipelines translate directly when you move to the cloud.

LayerLakehouse StackDatabricksSnowflakeMicrosoft FabricAWS
Object storageS3-compatible storageS3 / ADLSManaged storageOneLakeAmazon S3
Table formatApache IcebergDelta Lake (+ Iceberg)Iceberg tablesDelta LakeIceberg / S3 Tables
ProcessingApache SparkDatabricks SparkSnowparkFabric SparkEMR / Glue
SQL engineTrinoDatabricks SQLVirtual warehousesFabric WarehouseAthena
OrchestrationApache AirflowLakeflow JobsTasksData FactoryMWAA
Catalog & governanceApache PolarisUnity CatalogHorizon / Open CatalogOneLake catalogGlue + Lake Formation
NotebooksJupyterDatabricks notebooksSnowflake NotebooksFabric notebooksSageMaker Studio
DashboardsApache SupersetAI/BI dashboardsStreamlitPower BIQuickSight
Deployment

Two ways to run it

Single machine · Docker Compose

The whole platform on one laptop or server, ideal for proofs of concept, demos and individual development. Up and running with a handful of commands.

Multi-user · Kubernetes

A shared deployment with Helm and Ansible, isolated per-user workspaces, generated secrets and a management console, for teams and on-premises use.

  • Apache Spark
  • Apache Iceberg
  • Trino
  • Apache Airflow
  • Apache Polaris
  • Jupyter
  • Apache Superset
  • SQLPad
  • PostgreSQL
  • S3 object storage
  • Kubernetes
  • Helm
  • Ansible
  • Docker

See the Lakehouse Stack in action

Book a demo and we'll walk you through a live deployment, and how it could fit your data.