Skip to content
BlogData Architecture

The Data Hub Blueprint: A Field Guide to Databricks, Snowflake, Fabric, AWS, and Google Cloud

8 min read By Kishore Namburi

"Catalog" on Databricks is "Catalog" on Snowflake is "OneLake catalog" on Fabric is "Glue Data Catalog" on AWS is "Dataplex Universal Catalog" on Google Cloud — same job, five menu labels. I built a working map so I stop re-deriving that translation every time a client, a candidate, or my own memory needs it. Here's the map, how it's built, and where the maturity gap matters more than the naming gap.

1. Why a translation layer, not a scorecard

Multi-platform data estates are the default now — a Databricks lakehouse next to a Snowflake warehouse next to an Azure tenant running Fabric, a BigQuery project somewhere in the mix, all on top of AWS infrastructure. Nobody designed it that way; it accreted through acquisitions and cloud commitments made years apart.

The cost is constant small friction, not one big failure: a new hire spends their first week hunting for "the Databricks version of Worksheets" instead of doing the work; an architecture review compares two features with different names and calls it a gap; an RFP asks for "a marketplace" when what it actually needs is Delta Sharing-style live sharing, not a storefront. None of that is fixed by picking a favorite platform — it's fixed by a shared index: point at a job, know its name in all four dialects, get an honest read on who's ahead at it.

2. How the map is built

Everything sits in six domains, each carrying a maturity tier (defined below). Two candidates didn't make the final cut: Navigation & Workspace, which measures whether a vendor has one consolidated front door rather than an actual capability; and Notebooks & Apps, folded into AI/ML since a notebook is mostly where AI/ML work actually happens.

Three limits worth flagging up front — the parts most likely to be wrong or age badly:

  • Tiers are a snapshot, not a benchmark — an opinion as of September 2026, not a formal ranking; a vendor is routinely a Leader in one domain and Developing in the next.
  • AI/ML reorders fastest. Snowflake's Cortex AI is closing ground quickly enough that a fair call in September could read dated by December.
  • Fabric, AWS, and GCP being inferred is architectural, not just a documentation gap. AWS is a set of independently mature, modular services (IAM, Glue, Athena, S3) that were never meant to share one nav; Fabric is an integrated SaaS layer on OneLake; Google Cloud sits in between — one console, but BigQuery, Vertex AI, and Dataplex still read as separate products stitched under it. Forcing all three into the same table as Databricks/Snowflake's native navigation is a real compromise.

3. The six domains, at a glance

Reading down any column tells you where a given vendor's weight sits. Reading across any row tells you the translation. A few patterns jump out immediately:

Leader

Furthest along; the reference implementation other vendors get compared against for this job.

Mature

Solid and production-proven, but not the category-setter.

Developing

Real and shipping, but visibly newer or still consolidating.

DomainDatabricksAWSSnowflakeGoogle CloudMicrosoft Fabric
Ingestion & EngineeringMatureLeaderMatureMatureMature
ComputeMatureLeaderMatureMatureDeveloping
Catalog & GovernanceLeaderMatureMatureMatureDeveloping
SQL & BIMatureMatureLeaderLeaderLeader
AI / MLLeaderLeaderDevelopingLeaderMature
Marketplace & SharingMatureMatureLeaderMatureDeveloping

No vendor sweeps the board. Databricks leads catalog and AI/ML because Unity Catalog and MLflow set the reference pattern. Snowflake, Fabric, and GCP split SQL & BI leadership three ways for three different reasons — Snowflake on warehouse pedigree, Fabric on Power BI's install base, GCP on BigQuery's serverless architecture plus Looker's semantic layer. AWS leads raw ingestion and compute breadth on the strength of Glue, DMS, and EMR's long track record.

4. The full capability map

The summary above is the maturity read. This is the actual row-by-row translation — every capability, in each platform's own words.

CapabilityDatabricksAWSSnowflakeGoogle CloudMicrosoft Fabric
Ingestion & Engineering
Data ingestionData IngestionGlue / DMS / AppFlowSnowpipe / OpenflowDataflow / DatastreamData Factory
Pipeline orchestrationJobs & PipelinesStep Functions / MWAATasks / Dynamic TablesCloud ComposerData Factory pipelines
Visual data prepVisual Data PrepGlue DataBrewSnowpark (code-first)Cloud Data FusionDataflows Gen2
Compute
General computeComputeEMR / GlueComputeDataproc / serverlessSpark capacities
Catalog & Governance
Data catalogUnity CatalogGlue Data CatalogCatalogDataplex CatalogOneLake / Purview
Governance & securityUC governanceLake Formation + IAMGovernance & SecurityDataplex + IAMPurview
SQL & BI
SQL editorSQL EditorAthena / Redshift QE v2WorksheetsBigQuery StudioWarehouse SQL editor
BI dashboardsDashboards (AI/BI)QuickSightDashboardsLooker / Looker StudioPower BI
Natural-language Q&AGenieAmazon QCortex AnalystGemini in BigQueryCopilot in Fabric
Warehouse computeSQL WarehousesRedshift / AthenaVirtual warehousesBigQuery slotsCapacity units
AI/ML
Model playgroundPlaygroundBedrock PlaygroundCortex AI PlaygroundVertex AI StudioAI Foundry Playground
Agent buildingAgentsBedrock AgentsCortex AgentsVertex AI Agent BuilderCopilot Studio
LLM gatewayAI GatewayBedrock + API GatewayCortex AI governanceModel Garden + ApigeeAI Foundry router
Experiment trackingExperimentsSageMaker ExperimentsSnowflake ML trackingVertex AI ExperimentsFabric Data Science
Feature storeFeaturesSageMaker Feature StoreSnowflake Feature StoreVertex AI Feature StoreAzure ML Feature Store
Model registryModelsSageMaker RegistrySnowflake Model RegistryVertex AI Model RegistryAzure ML registry
Model/agent servingServingSageMaker / BedrockCortex servingVertex AI EndpointsAzure ML endpoints
Notebooks & appsNotebooks + AppsSageMaker StudioNotebooks + StreamlitBigQuery Studio / Vertex WorkbenchFabric Notebooks
Marketplace & Sharing
Data marketplaceMarketplaceAWS Data ExchangeMarketplaceAnalytics HubAzure Marketplace
Secure data sharingDelta SharingRedshift Data SharingData SharingAnalytics HubOneLake shortcuts

5. Three spotlights worth the extra detail

Catalog & Governance: "open" is a narrowing gap, not a moat

Unity Catalog's lead isn't feature count — Horizon, Glue Catalog + Lake Formation, and Dataplex Universal Catalog all cover tags, lineage, and access policy. The original case was that Unity Catalog went open and cross-cloud from day one while Snowflake's catalog sat inside a tighter boundary. That's dating fast: Snowflake co-drove Apache Polaris to a top-level Apache project and matured Horizon's Iceberg REST catalog support, including bidirectional external writes; Google's BigLake Metastore now speaks Iceberg REST too. The openness gap is narrower than it was a year ago, across the board. What Databricks still holds is being first to a unified data-and-AI governance model, tables and models under one catalog — a harder lead to close than "open" alone.

AI/ML: three leaders, three different bets

Databricks (MLflow's home turf), AWS (SageMaker plus Bedrock), and GCP (Vertex AI, BigQuery ML, Gemini) all read as Leaders here, for different reasons — one set the open tracking standard, one has the longest end-to-end track record, and GCP has the deepest native data-to-AI integration plus a real cost edge on inference. Cortex AI is shipping real capability fast, but it's a recent entrant next to a decade of MLflow/SageMaker/Vertex maturity — which is why it's Developing, not Mature, and the domain most likely to reorder first.

Natural-language Q&A: one row in the map, five different bets

The capability map above lists five products for one job, and each picked a different entry point: Genie starts from the dataset, Cortex Analyst from a governed semantic model, Gemini in BigQuery from the warehouse's own query surface, Copilot in Fabric from the report, Amazon Q from the dashboard. That entry point tells you more about onboarding a team than the maturity tier does.

6. How I actually use this

  • Onboarding across platforms: point a Snowflake-native engineer at the Databricks column for what they'll look for first — SQL Editor, Catalog, Jobs, Workspace, Query History.
  • Vendor comparisons that don't strawman: when a comparison claims "Platform X doesn't have Y," check whether Y just lives under a different name, not a missing capability.
  • Architecture reviews: Developing is where a specialized third-party tool still earns its keep over the platform's native option.
  • RFP language: write requirements in domain terms ("governed cross-account data sharing without copying"), not vendor terms ("a Marketplace"), so they survive a platform switch.

This is a running reference — I expect to revise it as AI/ML keeps moving fastest.