The Data Hub Blueprint: A Field Guide to Databricks, Snowflake, Fabric, AWS, and Google Cloud
"Catalog" on Databricks is "Catalog" on Snowflake is "OneLake catalog" on Fabric is "Glue Data Catalog" on AWS is "Dataplex Universal Catalog" on Google Cloud — same job, five menu labels. I built a working map so I stop re-deriving that translation every time a client, a candidate, or my own memory needs it. Here's the map, how it's built, and where the maturity gap matters more than the naming gap.
1. Why a translation layer, not a scorecard
Multi-platform data estates are the default now — a Databricks lakehouse next to a Snowflake warehouse next to an Azure tenant running Fabric, a BigQuery project somewhere in the mix, all on top of AWS infrastructure. Nobody designed it that way; it accreted through acquisitions and cloud commitments made years apart.
The cost is constant small friction, not one big failure: a new hire spends their first week hunting for "the Databricks version of Worksheets" instead of doing the work; an architecture review compares two features with different names and calls it a gap; an RFP asks for "a marketplace" when what it actually needs is Delta Sharing-style live sharing, not a storefront. None of that is fixed by picking a favorite platform — it's fixed by a shared index: point at a job, know its name in all four dialects, get an honest read on who's ahead at it.
2. How the map is built
Everything sits in six domains, each carrying a maturity tier (defined below). Two candidates didn't make the final cut: Navigation & Workspace, which measures whether a vendor has one consolidated front door rather than an actual capability; and Notebooks & Apps, folded into AI/ML since a notebook is mostly where AI/ML work actually happens.
Three limits worth flagging up front — the parts most likely to be wrong or age badly:
- Tiers are a snapshot, not a benchmark — an opinion as of September 2026, not a formal ranking; a vendor is routinely a Leader in one domain and Developing in the next.
- AI/ML reorders fastest. Snowflake's Cortex AI is closing ground quickly enough that a fair call in September could read dated by December.
- Fabric, AWS, and GCP being inferred is architectural, not just a documentation gap. AWS is a set of independently mature, modular services (IAM, Glue, Athena, S3) that were never meant to share one nav; Fabric is an integrated SaaS layer on OneLake; Google Cloud sits in between — one console, but BigQuery, Vertex AI, and Dataplex still read as separate products stitched under it. Forcing all three into the same table as Databricks/Snowflake's native navigation is a real compromise.
3. The six domains, at a glance
Reading down any column tells you where a given vendor's weight sits. Reading across any row tells you the translation. A few patterns jump out immediately:
Furthest along; the reference implementation other vendors get compared against for this job.
Solid and production-proven, but not the category-setter.
Real and shipping, but visibly newer or still consolidating.
| Domain | Databricks | AWS | Snowflake | Google Cloud | Microsoft Fabric |
|---|---|---|---|---|---|
| Ingestion & Engineering | Mature | Leader | Mature | Mature | Mature |
| Compute | Mature | Leader | Mature | Mature | Developing |
| Catalog & Governance | Leader | Mature | Mature | Mature | Developing |
| SQL & BI | Mature | Mature | Leader | Leader | Leader |
| AI / ML | Leader | Leader | Developing | Leader | Mature |
| Marketplace & Sharing | Mature | Mature | Leader | Mature | Developing |
No vendor sweeps the board. Databricks leads catalog and AI/ML because Unity Catalog and MLflow set the reference pattern. Snowflake, Fabric, and GCP split SQL & BI leadership three ways for three different reasons — Snowflake on warehouse pedigree, Fabric on Power BI's install base, GCP on BigQuery's serverless architecture plus Looker's semantic layer. AWS leads raw ingestion and compute breadth on the strength of Glue, DMS, and EMR's long track record.
4. The full capability map
The summary above is the maturity read. This is the actual row-by-row translation — every capability, in each platform's own words.
| Capability | Databricks | AWS | Snowflake | Google Cloud | Microsoft Fabric |
|---|---|---|---|---|---|
| Ingestion & Engineering | |||||
| Data ingestion | Data Ingestion | Glue / DMS / AppFlow | Snowpipe / Openflow | Dataflow / Datastream | Data Factory |
| Pipeline orchestration | Jobs & Pipelines | Step Functions / MWAA | Tasks / Dynamic Tables | Cloud Composer | Data Factory pipelines |
| Visual data prep | Visual Data Prep | Glue DataBrew | Snowpark (code-first) | Cloud Data Fusion | Dataflows Gen2 |
| Compute | |||||
| General compute | Compute | EMR / Glue | Compute | Dataproc / serverless | Spark capacities |
| Catalog & Governance | |||||
| Data catalog | Unity Catalog | Glue Data Catalog | Catalog | Dataplex Catalog | OneLake / Purview |
| Governance & security | UC governance | Lake Formation + IAM | Governance & Security | Dataplex + IAM | Purview |
| SQL & BI | |||||
| SQL editor | SQL Editor | Athena / Redshift QE v2 | Worksheets | BigQuery Studio | Warehouse SQL editor |
| BI dashboards | Dashboards (AI/BI) | QuickSight | Dashboards | Looker / Looker Studio | Power BI |
| Natural-language Q&A | Genie | Amazon Q | Cortex Analyst | Gemini in BigQuery | Copilot in Fabric |
| Warehouse compute | SQL Warehouses | Redshift / Athena | Virtual warehouses | BigQuery slots | Capacity units |
| AI/ML | |||||
| Model playground | Playground | Bedrock Playground | Cortex AI Playground | Vertex AI Studio | AI Foundry Playground |
| Agent building | Agents | Bedrock Agents | Cortex Agents | Vertex AI Agent Builder | Copilot Studio |
| LLM gateway | AI Gateway | Bedrock + API Gateway | Cortex AI governance | Model Garden + Apigee | AI Foundry router |
| Experiment tracking | Experiments | SageMaker Experiments | Snowflake ML tracking | Vertex AI Experiments | Fabric Data Science |
| Feature store | Features | SageMaker Feature Store | Snowflake Feature Store | Vertex AI Feature Store | Azure ML Feature Store |
| Model registry | Models | SageMaker Registry | Snowflake Model Registry | Vertex AI Model Registry | Azure ML registry |
| Model/agent serving | Serving | SageMaker / Bedrock | Cortex serving | Vertex AI Endpoints | Azure ML endpoints |
| Notebooks & apps | Notebooks + Apps | SageMaker Studio | Notebooks + Streamlit | BigQuery Studio / Vertex Workbench | Fabric Notebooks |
| Marketplace & Sharing | |||||
| Data marketplace | Marketplace | AWS Data Exchange | Marketplace | Analytics Hub | Azure Marketplace |
| Secure data sharing | Delta Sharing | Redshift Data Sharing | Data Sharing | Analytics Hub | OneLake shortcuts |
5. Three spotlights worth the extra detail
Catalog & Governance: "open" is a narrowing gap, not a moat
Unity Catalog's lead isn't feature count — Horizon, Glue Catalog + Lake Formation, and Dataplex Universal Catalog all cover tags, lineage, and access policy. The original case was that Unity Catalog went open and cross-cloud from day one while Snowflake's catalog sat inside a tighter boundary. That's dating fast: Snowflake co-drove Apache Polaris to a top-level Apache project and matured Horizon's Iceberg REST catalog support, including bidirectional external writes; Google's BigLake Metastore now speaks Iceberg REST too. The openness gap is narrower than it was a year ago, across the board. What Databricks still holds is being first to a unified data-and-AI governance model, tables and models under one catalog — a harder lead to close than "open" alone.
AI/ML: three leaders, three different bets
Databricks (MLflow's home turf), AWS (SageMaker plus Bedrock), and GCP (Vertex AI, BigQuery ML, Gemini) all read as Leaders here, for different reasons — one set the open tracking standard, one has the longest end-to-end track record, and GCP has the deepest native data-to-AI integration plus a real cost edge on inference. Cortex AI is shipping real capability fast, but it's a recent entrant next to a decade of MLflow/SageMaker/Vertex maturity — which is why it's Developing, not Mature, and the domain most likely to reorder first.
Natural-language Q&A: one row in the map, five different bets
The capability map above lists five products for one job, and each picked a different entry point: Genie starts from the dataset, Cortex Analyst from a governed semantic model, Gemini in BigQuery from the warehouse's own query surface, Copilot in Fabric from the report, Amazon Q from the dashboard. That entry point tells you more about onboarding a team than the maturity tier does.
6. How I actually use this
- Onboarding across platforms: point a Snowflake-native engineer at the Databricks column for what they'll look for first — SQL Editor, Catalog, Jobs, Workspace, Query History.
- Vendor comparisons that don't strawman: when a comparison claims "Platform X doesn't have Y," check whether Y just lives under a different name, not a missing capability.
- Architecture reviews: Developing is where a specialized third-party tool still earns its keep over the platform's native option.
- RFP language: write requirements in domain terms ("governed cross-account data sharing without copying"), not vendor terms ("a Marketplace"), so they survive a platform switch.
This is a running reference — I expect to revise it as AI/ML keeps moving fastest.