Phase 4: Data Quality & Governance

OpenLineage, Marquez & Monte Carlo

Intermediate ~3 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're baking your absolute favorite cookies! You gather flour, sugar, chocolate chips, and eggs, mix them all up, pop them in the oven, and out come delicious cookies. But what if you wanted to remember exactly where those chocolate chips came from, or how much sugar you added last time to make them extra perfect? Today we're talking about three cool names for tools that help grown-ups keep track of digital stuff: OpenLineage, Marquez, and Monte Carlo.

OpenLineage is like a super special, universal recipe card that everyone agrees to use. Think about it: when you bake, you follow steps, right? "Take 2 cups of flour, add 1 cup of sugar, mix well." OpenLineage is a way for all the different cooking machines in a big digital kitchen to write down their recipes in the exact same way. So, whether one machine is making bread, another is making a cake, or a third is just chopping vegetables, they all use the same kind of recipe card to say what ingredients went in, what they made, and how they made it. This means any chef, no matter which machine they're using, can read and understand any other chef's recipe.

Now, where do all these perfectly written recipe cards go? That's where Marquez comes in! Marquez is like a giant, super-organized digital recipe book for your whole kitchen. Every time one of your cooking machines finishes a task and uses an OpenLineage recipe card, it sends a copy to Marquez. Marquez collects all these recipe cards. It doesn't just stack them up; it uses them to draw a big picture, like a map, showing how all your ingredients became different dishes. It knows that the flour from this pantry shelf went into that cookie dough, and that cookie dough became those delicious cookies. It creates a complete story of your food, from individual ingredients to the final yummy dish.

So, why is having this amazing digital recipe book (Marquez) filled with universal recipe cards (OpenLineage) so useful? Imagine you run a huge bakery. If a batch of muffins doesn't turn out quite right, you can instantly look up its recipe in Marquez and see exactly which ingredients were used, when they were added, and by which machine. Or, if you want to invent a new cake, you can easily look at all the recipes that use blueberries, helping you get new ideas! This means you can always tell the full story of your digital "food," making sure everything is perfect and easy to understand for everyone. So, when grown-ups build amazing digital data "dishes," they use OpenLineage and Marquez to keep track of every single ingredient and step!

Data lineage is crucial for understanding your data's journey, from its origin to its current state, including all transformations and dependencies. OpenLineage is an open standard that provides a common language and specification for collecting and exchanging this lineage metadata. Think of it as a universal protocol that allows different data systems (like Spark, Airflow, dbt) to describe their operations and the datasets they interact with in a consistent, interoperable way. This standardization is key for building a comprehensive view of your data pipelines without being locked into a proprietary format.

Marquez is an open-source metadata service that acts as a central repository for your OpenLineage events. When your data jobs (e.g., an Apache Spark job or an Airflow DAG) run, they can emit OpenLineage events describing their inputs, outputs, schemas, and processing logic. Marquez ingests these events, stores them, and uses them to construct a detailed, graph-based view of your data lineage. It provides APIs for programmatic access and a user interface to visualize your data pipelines, helping you trace data provenance, understand impact analysis, and improve data governance by seeing exactly how data flows and transforms across your entire ecosystem.

While OpenLineage and Marquez focus on building and storing the lineage graph, Monte Carlo is a commercial end-to-end data observability platform that leverages this understanding to ensure data quality and reliability. Monte Carlo monitors your data warehouses, lakes, and other data sources for freshness, volume anomalies, schema changes, and other quality issues. By understanding data dependencies (often through its own lineage capabilities or integrations with systems like Marquez), it can pinpoint the root cause of data problems, understand their blast radius, and proactively alert data engineers. It moves beyond just seeing lineage to actively monitoring, detecting, and helping resolve data incidents across your stack.

Key Takeaways

  • OpenLineage is an open standard for consistent data lineage metadata collection.
  • Marquez is an open-source metadata service that stores and visualizes OpenLineage events, acting as your central lineage hub.
  • Monte Carlo is a commercial data observability platform focused on data quality and incident management, often leveraging lineage to detect and alert on data issues.
  • Together, these tools help you build a clear lineage graph, monitor data health, and react quickly to data quality problems.

Code Example

python
from openlineage.client import OpenLineageClient, set_producer
from openlineage.client.utils import get_hostname
from openlineage.client.run import Dataset, Job, Run, RunEvent, RunState

# Configure OpenLineage client to send events to Marquez
set_producer("my_simple_etl", get_hostname())
client = OpenLineageClient(url="http://marquez:5000/api/v1", timeout=5.0)

# Define job, run, and datasets
job = Job("my_namespace", "users_transform_job")
run = Run(runId="unique-run-id-456")
input_dataset = Dataset("my_source_db", "public.raw_users")
output_dataset = Dataset("my_warehouse", "public.clean_users")

# Emit START event before processing
client.emit(RunEvent(RunState.START, run, job, [], [input_dataset]))

# Simulate ETL processing here
print("Processing data from raw_users to clean_users...")

# Emit COMPLETE event after successful processing
client.emit(RunEvent(RunState.COMPLETE, run, job, [output_dataset], [input_dataset]))
print("OpenLineage events emitted to Marquez.")

How this code works

This code demonstrates how to use OpenLineage to track a simple data transformation job, specifically an ETL (Extract, Transform, Load) process that moves data from raw_users to clean_users. It configures an OpenLineageClient to send events to a Marquez server, identified by its URL. The set_producer function helps identify the source system sending these events, in this case, a 'my_simple_etl' process running on the current get_hostname.

The code defines the job name, a unique run identifier for this specific execution, and describes the input_dataset (public.raw_users) and output_dataset (public.clean_users). It then uses client.emit to send a RunEvent with RunState.START before the simulated data processing begins, signaling the job's initiation and identifying its inputs. After the processing, another client.emit sends a RunEvent with RunState.COMPLETE. A subtle but important detail is including the input_dataset again in the COMPLETE event, along with the output_dataset; this ensures Marquez fully understands the entire lineage, linking both sources and destinations to the completed job. This way, Marquez builds a comprehensive view of how clean_users was derived from raw_users.