I2 min
Two equal numbers
The sneakiest mistake in checking data is to relax when two numbers match. Equal numbers only say the two are the same today; whether they measure the same thing, they can’t tell you.
This summer I built a data collection pipeline: every day it fetches data from several platforms and stores it in a read-only database for the analysis on top.
It started as an RPA flow, a robot clicking through pages, waiting for them to load and exporting the tables. It was slow and fragile, ran only on the one machine set up for it, and when it failed it was hard to tell which step had broken. I rewrote it in Python. Fetching no longer opens a browser; it sends the requests behind the page directly. The browser is used only to log in and to work out the endpoints the first time, and neither is on the daily path.
What took the time wasn’t writing the requests. It was checking.
For a while, if the numbers I fetched matched what the platform’s page showed, I relaxed. Later I saw that was wrong: two fields can be equal on a given day and still measure different things. The rule now is that the data can only say “these two numbers are equal today”; it can’t tell you what they measure. That has to be found in two places: the metrics, dimensions and grouping parameters written into the request, and the field structure of the platform’s raw response. Conclusions rest on those two, never on equal values worked backwards.
Two more kinds of checking came later. One compared the data column by column with the platform’s own exported reports, and found three errors that raised no error: the program ran normally, the numbers looked reasonable, and they were wrong. The other took screenshots back to the pages and went through them line by line, and found two real problems: an empty result returned under rate limiting had been read as “no data today”, and one kind of data had been filed under the wrong entry.
Then there was logging in. The platforms treat frequent logins as suspicious, and one did trigger a security check. Now each day starts by checking whether the session is still alive, and logs in again only if it isn’t.
Last was delivery. One line per platform, one folder per line, and the folder has to be something you can zip up and send, which the other person unzips and runs on their own machine. The test comes down to one sentence: no file in the folder needs anything outside it. To hold to that, I wrote tests that check the folder’s structure, and a pre-release check that really unzips it somewhere else, really creates an empty virtual environment, and really runs the whole thing there. Running fine on my own machine proves nothing about anyone else’s.
All of this is tedious. But the people using these numbers compare them week on week and watch them for anomalies. If one definition underneath is wrong, every chart built on it is wrong too, and nothing anywhere raises an error.