Favicon of Google Cloud Knowledge Catalog

Google Cloud Knowledge Catalog

Knowledge Catalog — previously Dataplex Universal Catalog, and the replacement for the retired Data Catalog — harvests metadata from BigQuery, Vertex AI, Pub/Sub, Bigtable, Cloud SQL, and AlloyDB into entries described by aspects, with natural-language search, lineage, and glossaries over the result.

LicenseCommercial
DeploymentManagedServerless
PricingSubscription
Operational complexityLow
WorkloadInteractiveBatch

Use it when

  • The platform is BigQuery-centred: metadata is harvested automatically, and free, from BigQuery, Vertex AI, Pub/Sub, Bigtable, Cloud SQL, and AlloyDB.
  • You want serverless cataloguing with no capacity to size, governed by the IAM model the platform enforces anyway.
  • Natural-language semantic search over entries and aspects serves analysts better than a table browser.
  • Data quality scans on BigQuery tables, with rules recommended from profiles, cover your checking needs.

Think twice when

  • The estate is multi-cloud: outside Google Cloud, coverage depends on what Google harvests, and non-GCP connectors are new and few.
  • Name churn is a concern: this is the fourth name and second architecture since Data Catalog, and the API, CLI, and IAM names still say dataplex.
  • Lineage coverage must be uniform; it is automatic only for some services and has documented gaps.

How it runs

Fully managed and serverless in Google Cloud. Automatic ingestion feeds entries described by aspects, structured and column-level; billing meters metadata storage by the GiB and processing in DCU-hours, with automatically ingested Google Cloud technical metadata free. Access requests on data products, in preview, grant IAM roles themselves on approval.

Details

Compare

How Google Cloud Knowledge Catalog answers the questions Data Catalogs turns on.

Data Catalogs
How it works
Metadata modelEntries and aspects: an entry is a data asset, and aspects are structured metadata sets attached to it or to an individual column, with entry groups containing entries and carrying their access control, aspect types as reusable templates and entry types demanding particular aspects be present. It replaces the flat tag model of the Data Catalog it succeeded
LineageAutomatic from BigQuery, Dataflow, Managed Airflow, Managed Service for Apache Spark, Data Fusion, Vertex AI and Looker in preview, and column-level, with real edges: not collected for BigQuery load jobs or routines, top-level columns only, and switched off above 1,500 column links in a job. Everything else is reported through the Data Lineage API against a custom entry
Search and discoverySemantic search in natural language across every ingested source and the aspects added on top, with a defined search syntax underneath it; entry links then relate assets to one another as synonyms, related items, schema joins or definitions
Business glossaryCategories nested up to three levels holding terms, linked to entries and to each other, so data can be found by business concept rather than technical name, exportable to a Google Sheet, with a documented migration path from the old Data Catalog glossaries
Quality checksRow-level rules for range, null, set, regex and uniqueness, aggregate rules, and custom SQL assertions, with rules recommended from a profile scan; scans run full or incremental, filtered and sampled to control cost, scored at job, column and dimension level, alerting through Cloud Logging or email. The limit is the point, rules run only on BigQuery and Iceberg REST Catalog tables, capped at 1,000 per scan
Custom metadataYour own aspect types and entry types, custom entries for systems Google does not harvest, custom IAM roles, and REST, gcloud and Terraform surfaces over all of it
Connections
ConnectorsAutomatic ingestion from across Google Cloud (BigQuery datasets, tables, views and models, Dataform, Dataproc Metastore, Vertex AI models, datasets and feature groups, Cloud SQL, AlloyDB, Spanner, Bigtable, Pub/Sub topics, Cloud Storage and Looker) with SQL Server and PostgreSQL connectors added in preview in July 2026; anything else arrives through a Knowledge Catalog connector or a managed connectivity pipeline
Access
Policy and complianceIAM rather than a policy language of its own, predefined and custom roles granted on entry groups and entries, with VPC Service Controls drawing the perimeter around them, and governance workflows layering approval on top
Access requestsStronger than the IAM framing suggests, and newer, a consumer requests access to a data product inside the catalog and, on approval, the governance workflow grants the IAM roles and group memberships itself rather than handing somebody a task to do it. In preview since July 2026
Cost
Billing unitMetadata storage by the GiB as a monthly average (but technical metadata ingested automatically from Google Cloud services is free, and the first 1 MiB besides) plus processing in DCU-hours billed by the second with a one-minute minimum and the first 100 DCU-hours a month free. Data quality anomaly detection is the exception, billed as ordinary BigQuery usage instead

Share:

Alternatives to Google Cloud Knowledge Catalog

Favicon

 

  
  
Favicon

 

  
  
Favicon