~/tools/datahub
DataHub
tool

DataHub

DataHub is an open source data catalog and metadata platform, originally built at LinkedIn — a self-hostable Collibra and Alation alternative for data discovery, column-level lineage, and governance across your whole data stack.

What is DataHub?

DataHub is an open source data catalog and metadata management platform that helps teams find, understand, and govern their data. Originally built and open-sourced by LinkedIn in 2020, it collects metadata from across your stack — warehouses, dashboards, pipelines, and ML models — into one searchable graph with lineage and ownership.

What is DataHub best for?

Engineering-heavy data teams that want a central place to discover data assets and trace where they come from, without paying enterprise catalog licensing. It fits organizations that already run infrastructure like Kafka and Elasticsearch and have a platform engineer who can own the deployment and extend the metadata model.

What can DataHub do?

  • Search across datasets, tables, dashboards, and ML models in one place
  • Track column-level lineage from source system to final consumption
  • Ingest metadata from 70+ sources (Snowflake, BigQuery, dbt, Airflow, Looker, Kafka, and more)
  • Enforce governance with ownership, tags, glossary terms, and PII classification
  • Monitor data quality — freshness, schema changes, and assertion results
  • Automate metadata with GraphQL and REST APIs plus a streaming-first Kafka architecture
  • Provide context to AI agents through Model Context Protocol (MCP) integration

Where does DataHub fall short?

  • It’s engineering-first, not analyst-first. Compared to Collibra and Alation, it ships less out-of-the-box governance automation and fewer no-click workflows, so non-technical stewards face a steeper learning curve.
  • The metadata backend depends on Kafka and Elasticsearch alongside a JVM-based service, which is resource-hungry and needs real tuning for large deployments — more moving parts than a single-container app.
  • Governance features like policy workflows and stewardship are lighter than a dedicated enterprise suite, so heavily regulated teams may find the formal-controls story thinner.

Is DataHub free?

Yes — DataHub Core is fully open source under Apache 2.0, with no seat limits or feature gates, and free to self-host. The paid option is DataHub Cloud (formerly Acryl Data), a managed service from the project’s original creators that runs the infrastructure for you and adds automation and enterprise support; managed plans typically start in the low tens of thousands per year.

What does DataHub replace?

DataHub is a self-hosted alternative to proprietary data catalogs like Collibra, Alation, Informatica, and Atlan. It covers the core discovery, lineage, and governance job those platforms do, but you run it on your own infrastructure without per-seat licensing. Among open source options, its closest peer is OpenMetadata.

FAQ

Is DataHub open source? Yes — DataHub Core is licensed under Apache 2.0, a permissive OSI-approved license. The code is public and free to use, modify, and self-host commercially.

Can I self-host DataHub for free? Yes. The open source project is free to self-host with no feature gates; you only pay for the servers it runs on. DataHub Cloud is the paid managed alternative.

Is DataHub a good Collibra alternative? For technical teams, yes — it delivers discovery, lineage, and governance without licensing cost. Regulated enterprises that need turnkey stewardship workflows and formal policy controls may still prefer Collibra or Alation.

What do I need to run DataHub? A container or Kubernetes environment with the supporting services it depends on — Kafka, Elasticsearch, and a metadata store — plus engineering time for setup and tuning. Docker Compose works for evaluation; production runs on Kubernetes with Helm.