10 Best Open Source Data Quality Tools in 2026

    Discover the best 10 open-source data quality tools you can self-host and run for free, across data testing, data profiling, and data observability.

    By Ari Bajo - Data Engineer turned Writer.

    Updated on August 12, 2026

    Great Expectations

    Open-source Python library with declarative expectations to validate data in files, SQL databases, data warehouses, and in-memory DataFrames.

    OSSdata testingpython-native
    My Opinion

    Best for data engineering teams looking for a code-first OSS data testing library with a large built-in expectation library and Python extensibility.

    Deequ

    Open-source Scala library built on Apache Spark to define and verify data quality constraints and profile large datasets at scale.

    My Opinion

    Best for data engineering teams using Apache Spark looking for a code-first OSS library to define data quality constraints programmatically in Scala or Python.

    Google CloudDQ

    Cloud-native data validation CLI with YAML-based data quality checks for BigQuery tables and GCS structured data.

    OSSdata testingbigquery-nativepython-native
    My Opinion

    Best for data teams looking for a BigQuery-native solution to write reusable SQL checks and consume data quality outputs programmatically.

    DQX by Databricks

    Data quality framework for Apache Spark with data quality rule generation from profiling results, and YAML and Python-based data validation checks.

    My Opinion

    Best for Databricks users looking to validate PySpark DataFrames and Tables across Spark Core, Spark Structured Streaming, and Lakeflow Pipelines / DLT.

    DQOps

    Open-source data quality testing and observability platform with data quality checks, monitors, data lineage with Marquez, and data quality dashboards.

    My Opinion

    Best for data teams looking to customize built-in data quality checks and data quality dashboards with Looker Studio to monitor data quality KPIs.

    DataKitchen

    Open-source data testing and observability platform with automated test generation, data profiling, and anomaly detection.

    My Opinion

    Best for data teams looking for a cost-effective data testing and observability solution that prices per database connection and user.

    Elementary OSS

    Open-source dbt package to add data observability to dbt projects with anomaly detection tests and a local data observability report generated via CLI.

    My Opinion

    Best for data analytics teams using dbt looking to add anomaly detection monitors to their existing dbt codebase without a cloud account.

    Soda Core

    Open-source Python library and CLI to write and run data contracts in YAML using SodaCL with integrations for data warehouses, databases and query engines.

    My Opinion

    Best for data engineering teams looking for a YAML-based OSS data testing library that embeds directly in pipelines and CI/CD workflows.

    Recce

    Open-source dbt validation toolkit and managed platform with data-diff, data impact reports and column-level data lineage.

    OSSdata diffdata lineagedbt-native
    My Opinion

    Best for data analytics teams using dbt looking to validate code changes with data impact reports during PR reviews.

    OpenMetadata

    Open-source unified metadata platform with data discovery, data quality checks, observability metrics, column-level lineage, and governance workflows.

    OSSdata testingdata observabilitydata lineagedata catalog
    My Opinion

    Best for data teams looking for a self-hosted open-source platform covering data discovery, observability, and governance with a wide range of integrations.

    Evaluating data quality tools?

    Market Guide (7,000 words)

    Feature Matrix (73 features)

    Integration Matrix (227 integrations)

    Join the Newsletter

    One email a month — a new tool list, comparison matrix, and market guide, straight to your inbox. Next up: Data Governance, LLMOps, Data Orchestration.

    By Ari Bajo - Data Engineer turned Writer.

    Frequently Asked Questions

    What is an open source data quality tool?
    An open source data quality tool is a freely available, self-hostable tool for validating, testing, or profiling data, with source code you can inspect, modify, and run without a vendor subscription. Examples include Great Expectations and Soda Core for data testing, and Elementary OSS for data observability. Many open-source data quality tools also offer a paid managed or cloud version with additional features like hosting, alerting, and support. Read more on my data quality tool market guide.
    Why create yet another tools list?
    I found no comprehensive, actionable, and up-to-date list of data quality tools. The MAD Landscape misclassifies 3 out of 19 data quality and observability tools. The Gartner Magic Quadrant for augmented data quality solutions lists 13 tools, half of which are enterprise data platforms, and I need to enter my professional email on a featured tool's website to get access to a reprint. Other lists by vendors contain a random sample of less than 10 tools, are written by AI, are highly biased, or are never updated.
    How can I edit this tools list?
    If you think a tool belongs here or you want to suggest an edit, I would love to hear from you. You can fill up the feedback form or DM on LinkedIn.