Available for Enterprise Edition only.
- Understanding the success of your pipeline migrations.
- Catching discrepancies before deploying pipelines to production.
- Detecting unintended schema changes in datasets over time.
- Identifying data mismatches when troubleshooting data transformation issues.
Prerequisites
To compute a data diff, you need:- Prophecy 4.0 or later.
- A Prophecy Python project. Data diffs are not supported for Scala or SQL projects.
- ProphecyLibsPython 1.9.37 or later as a dependency for your project. For more information, see Prophecy libraries.
- An expected dataset. It must be in Parquet or Databricks Catalog Table format.
Data diff cannot be computed on Databricks
Serverless.
What is data diff?
Data diffs are outputs of Target gems that show you differences between your target table and an expected table. Similar to a data sample, you can explore the data diff after interactively running a pipeline. The data diff has four views:- Overview: A summary that displays various high-level comparisons of the generated (target) and expected datasets. Review the following section to understand each statistic in more detail.
- Column differences: The dataset schemas and the number of matching values for each column in both datasets.
- Values differences: A table that displays side-by-side differences of every value in both datasets. This will show a sample of the data.
- Data samples: Samples of the generated and expected datasets for data exploration.
Data diffs only temporarily appear in the pipeline. They are not persisted in your project.
Overview
The following table provide in-depth descriptions of each statistic in the Overview tab of the data diff.Unique and duplicated primary keys
Prophecy only calculates the data diff on rows with unique primary keys. Why is that? Assume you have a the following table, wherefirst_name and *last_name` are the primary keys:
If you try to compute the data diff John Smith, how will you know which row is the correct match? It is impossible to match the rows with 100% confidence. Because of this ambiguity, Prophecy ignores rows with duplicated primary keys in the data diff.
Configuration
Data diffs are configured in Target gems.- Open a Target gem in your pipeline.
- In the top right of the gem dialog, click the Options (ellipses) menu.
- Select Data Diff. This adds the Data Diff step to your Target gem configuration.
- Open the Data Diff step.
- Fill in the required parameters and Save the gem.
The row order of the generated and expected dataset does not matter, as the rows are joined by
keys, rather than row order.

