Phase 1: Foundations

Git hooks & CI triggers for data validation

Beginner ~3 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you're baking a super special cake for a big party. You want it to be perfect and yummy, right? Just like baking, when grown-ups work with important information (we call it 'data'), they want to make sure it's correct and safe to use. If the data is wrong, it's like serving a cake with salt instead of sugar – big problem! We need ways to check our work.

This is where some clever helpers come in. The first helper is like having a little checklist or a quick sniff-test you do before you even start mixing your cake batter. You quickly check your ingredients: are the eggs cracked? Is the flour fresh and not lumpy? Did you remember to get enough sugar? This quick check happens right in your kitchen, before you combine anything. It helps you catch simple mistakes, like grabbing the wrong spice jar, right away. If something is wrong, you fix it immediately, and you haven't even messed up your mixing bowl yet!

Now, let's say your batter is all mixed, and it looks pretty good. But what if there's a tiny problem you missed, like a weird ingredient that only shows up when baked? This is where the second helper comes in, like sending your cake batter to a special, super-smart kitchen lab that automatically checks everything when you're ready to put it in the oven. This lab has fancy machines and trained experts that can test if the cake will rise properly, if it’s safe for allergies, or if it has the perfect amount of sweetness. It’s a bigger, more thorough check that happens automatically as soon as you say, "Okay, this is ready for the oven!"

So, when data engineers are building programs that use important information, they use these two types of helpers. The 'local checks' (like your ingredient sniff-test) prevent small errors from getting into their work early. The 'automatic lab checks' (the oven quality control) then give everything a final, more powerful check before anyone else sees or uses the data. This means they can be super confident that the information they're working with is always top-notch and trustworthy.

As a Data Engineer, ensuring data quality is paramount. Git hooks and CI triggers are powerful tools that help you enforce data validation right from your local machine to your shared data pipelines. Git hooks are scripts that run automatically before or after certain Git events, like committing or pushing. For data validation, a common use is a pre-commit hook that runs checks on your local machine before your changes are even saved to your Git history. This could involve validating a CSV file's schema, checking for missing values in a dataset, linting your SQL or Python scripts, or ensuring no sensitive data is accidentally being committed. These local checks provide immediate feedback, catching simple errors early and preventing low-quality data or code from entering your repository's history.

While Git hooks handle local, lightweight checks, CI triggers take data validation to the next level in an automated, shared environment. CI (Continuous Integration) systems like GitHub Actions, GitLab CI, or Jenkins, can be configured to automatically start a series of tasks whenever you push your changes to a remote repository. A CI trigger for data validation might spin up a test environment, run more extensive schema validation against a large dataset, execute complex data quality rules, or even perform data profiling on newly added data. This ensures that any new data transformations, models, or data sources integrate correctly and maintain the required quality standards across your entire data platform, not just on your local machine.

Together, Git hooks and CI triggers form a robust, multi-layered defense for data quality. Git hooks act as your personal guard, preventing basic errors from leaving your local workspace. CI triggers then serve as the team's quality assurance, running comprehensive automated tests in a shared environment to catch more complex issues. By integrating these tools into your data engineering workflow, you ensure that only high-quality, validated data and code make it into your production systems, significantly improving data reliability and reducing downstream errors.

Key Takeaways

  • Git hooks perform local, immediate data validation checks (e.g., pre-commit).
  • CI triggers initiate automated, comprehensive data validation in a shared environment (e.g., after git push).
  • They create a two-stage defense: local prevention and remote assurance for data quality.
  • Help prevent bad data, incorrect schemas, or faulty code from entering main branches.
  • Essential for maintaining data integrity and building reliable data pipelines.

Code Example

bash
# .git/hooks/pre-commit
#!/bin/bash

echo "Running pre-commit data validation checks..."

# Check if any staged file is larger than 10MB (10240 KB)
# This is a basic example to prevent committing excessively large data files.
MAX_FILE_SIZE_KB=10240
for file in $(git diff --cached --name-only --diff-filter=AM);
do
  FILE_SIZE_KB=$(du -k "$file" | cut -f1)
  if (( FILE_SIZE_KB > MAX_FILE_SIZE_KB )); then
    echo "Error: File '$file' ($FILE_SIZE_KB KB) exceeds max allowed size of ${MAX_FILE_SIZE_KB}KB."
    exit 1 # Abort commit if check fails
  fi
done

echo "Pre-commit checks passed."
exit 0

How this code works

This script acts as a .git/hooks/pre-commit hook, running automatically before a git commit. Its primary role is to validate data files by ensuring no excessively large files are included in the commit, thereby helping maintain a clean and manageable repository. If any staged file exceeds a predefined size limit, the commit operation is aborted, and an informative error message is displayed.

The script starts by defining MAX_FILE_SIZE_KB for the allowed limit. It then iterates through files using git diff --cached --name-only --diff-filter=AM. This command is crucial because git diff --cached specifically targets files currently staged for commit, and the --diff-filter=AM option subtly ensures only newly Added or Modified files are checked, preventing unnecessary processing of files being deleted. For each file, du -k measures its size, and if it exceeds MAX_FILE_SIZE_KB, exit 1 stops the commit. Otherwise, exit 0 allows the commit to proceed.