Knowledge Catalog — previously Dataplex Universal Catalog, and the replacement for the retired Data Catalog — harvests metadata from BigQuery, Vertex AI, Pub/Sub, Bigtable, Cloud SQL, and AlloyDB into entries described by aspects, with natural-language search, lineage, and glossaries over the result.
Fully managed and serverless in Google Cloud. Automatic ingestion feeds entries described by aspects, structured and column-level; billing meters metadata storage by the GiB and processing in DCU-hours, with automatically ingested Google Cloud technical metadata free. Access requests on data products, in preview, grant IAM roles themselves on approval.
How Google Cloud Knowledge Catalog answers the questions Data Catalogs turns on.
| How it works | |
| Metadata model | Entries and aspects: an entry is a data asset, and aspects are structured metadata sets attached to it or to an individual column, with entry groups containing entries and carrying their access control, aspect types as reusable templates and entry types demanding particular aspects be present. It replaces the flat tag model of the Data Catalog it succeeded |
| Lineage | Automatic from BigQuery, Dataflow, Managed Airflow, Managed Service for Apache Spark, Data Fusion, Vertex AI and Looker in preview, and column-level, with real edges: not collected for BigQuery load jobs or routines, top-level columns only, and switched off above 1,500 column links in a job. Everything else is reported through the Data Lineage API against a custom entry |
| Search and discovery | Semantic search in natural language across every ingested source and the aspects added on top, with a defined search syntax underneath it; entry links then relate assets to one another as synonyms, related items, schema joins or definitions |
| Business glossary | Categories nested up to three levels holding terms, linked to entries and to each other, so data can be found by business concept rather than technical name, exportable to a Google Sheet, with a documented migration path from the old Data Catalog glossaries |
| Quality checks | Row-level rules for range, null, set, regex and uniqueness, aggregate rules, and custom SQL assertions, with rules recommended from a profile scan; scans run full or incremental, filtered and sampled to control cost, scored at job, column and dimension level, alerting through Cloud Logging or email. The limit is the point, rules run only on BigQuery and Iceberg REST Catalog tables, capped at 1,000 per scan |
| Custom metadata | Your own aspect types and entry types, custom entries for systems Google does not harvest, custom IAM roles, and REST, gcloud and Terraform surfaces over all of it |
| Connections | |
| Connectors | Automatic ingestion from across Google Cloud (BigQuery datasets, tables, views and models, Dataform, Dataproc Metastore, Vertex AI models, datasets and feature groups, Cloud SQL, AlloyDB, Spanner, Bigtable, Pub/Sub topics, Cloud Storage and Looker) with SQL Server and PostgreSQL connectors added in preview in July 2026; anything else arrives through a Knowledge Catalog connector or a managed connectivity pipeline |
| Access | |
| Policy and compliance | IAM rather than a policy language of its own, predefined and custom roles granted on entry groups and entries, with VPC Service Controls drawing the perimeter around them, and governance workflows layering approval on top |
| Access requests | Stronger than the IAM framing suggests, and newer, a consumer requests access to a data product inside the catalog and, on approval, the governance workflow grants the IAM roles and group memberships itself rather than handing somebody a task to do it. In preview since July 2026 |
| Cost | |
| Billing unit | Metadata storage by the GiB as a monthly average (but technical metadata ingested automatically from Google Cloud services is free, and the first 1 MiB besides) plus processing in DCU-hours billed by the second with a one-minute minimum and the first 100 DCU-hours a month free. Data quality anomaly detection is the exception, billed as ordinary BigQuery usage instead |
vs Google Cloud Knowledge Catalog: Self-hosted · Managed · Operational complexity: Medium
Same headline facts as Google Cloud Knowledge Catalog
vs Google Cloud Knowledge Catalog: Open source (permissive) · Self-hosted · Free · Operational complexity: High · Java