Measuring a defect before fixing it, and being wrong about the measurement
Reduced the estimated blast radius of a customer-facing metric defect from ~29k cycles to 445 that actually showed a wrong number to a customer, turning a proposed historical-correction project into an audited script over a few hundred rows.
impactBlast radius corrected from ~29,000 records to 445 that actually displayed a wrong percentage to a customer, across three customers rather than the twenty implied by the raw count.A whole class of defect found that the first analysis had missed entirely, 4,174 records, almost all customer-facing, because my initial filter assumed a condition that did not hold.Turned a proposed historical-correction project into an audited one-off script over a few hundred rows, reviewable by hand.Every defect reproduced in a controlled environment with a written prediction before any fix was merged; each fix shipped with a regression test verified to fail without it.Two independent reproductions, fifteen hours apart, of the defect whose cause had been described incorrectly in the tracker.
01
Context
This is where I stopped treating "how big is this?" as a preamble to the real work and started treating it as the work. I produced a number, acted on it, then found it wrong by an order of magnitude, twice, and each correction changed what the team should do next.
It is my strongest evidence of investigating under ambiguity, of separating a defect from behaviour that merely looks like one, and of being the person who corrects their own analysis in front of stakeholders rather than defending it. It also shows a discipline I now consider non-negotiable: reproduce first, fix second.
02
Problem
A compliance percentage, derived from how many inspection items were completed against how many were due in a time window, was visibly wrong for several customers. Support had escalated individual cases; a product analyst had extracted a list of affected records; nobody knew the real extent.
The metric was not wrong in one way. It was wrong in several independent ways that produced similar-looking symptoms: a denominator that silently dropped items, a denominator that included items that were not due, counters that summed the same inspection twice, and a repair script that had itself written bad values. Percentages above 100% existed in production, up to 279%.
03Constraints
+
The symptoms did not map one-to-one to causes. A record showing 0% could be a real defect, or a route that legitimately had nothing to inspect that period. A record showing "more inspections than items" could come from duplicate answers or from a dropped item. Counting symptoms would have produced a number, just not a true one.
The correct value was not always recoverable. The system does not keep a history of what a route contained at a past moment. For a large share of records it was possible to prove the stored value was wrong, and impossible to say what it should have been.
Most of the data was noise. The production database also hosts internal and homologation tenants whose routes generate records continuously and are never inspected. Any naive count is dominated by them.
The investigation ran against a read-only production replica, so every question had to be answered with a SELECT, and expensive ones had to be shaped to finish at all.
04
Decision
I framed one question: *for each affected record, can I prove it is wrong, and can I compute what it should be?* Those are two different questions, and the answer to the second decides whether a fix is even possible.
Classify before counting. I built a taxonomy where each record falls in exactly one bucket: correct; provably wrong with a recoverable value; provably wrong with an unrecoverable value. The third bucket only exists because route composition history is not kept, and naming it early stopped me from promising a correction I could not deliver.
Validate the method against the data. The recomputation agreed with the stored value in 95% of records. That agreement is what made the 5% disagreement trustworthy: if my rule were wrong, it would have disagreed everywhere.
Separate defect from expected behaviour. The largest bucket, ~24,000 records showing 0%, turned out to be routes that genuinely had nothing to inspect. Real, but not a counter defect: a display decision for product, not a correction for engineering. This is where the first order-of-magnitude correction came from.
Separate customers from test data. Of the records that were both wrong and correctable, 93% belonged to internal and homologation tenants. Reporting the raw number would have overstated customer impact by more than an order of magnitude.
Then reproduce, one defect at a time. Each defect got a named route in staging, a written prediction of the expected numbers before the test ran, and a control route designed to stay unchanged. The control is what let me claim a defect was the null handling and not the periodicity rule itself: a route without the triggering condition behaved correctly across twenty consecutive cycles while the affected one failed in nineteen of twenty.
Only then, the fixes. Small, single-cause changes, each with a regression test verified to fail without the fix.
05Trade-offs
+
—Exact counts for the small, decidable populations; sampling for the large one. A full recomputation across every record was measured at roughly twelve hours of database work. I ran exact counts where the population was small enough to enumerate, and a 1% sample where it was not, then narrowed the exact scope by restricting to the routes that could exhibit the defect at all. Precision where it changed the decision, estimates where it did not.
—A documented "unrecoverable" bucket over an estimated correction. A previous repair script had guessed at values and introduced a defect that persisted for eleven consecutive periods for one customer. That precedent is why I preferred publishing "we can fix this half and not that half" over a heuristic that would look complete.
—Declining seven stacked pull requests to rebuild one problem at a time. The work had grown into dependent branches spanning multiple defects, which made any single one impossible to validate alone. Discarding open work is expensive and looks like backtracking; I argued against it at first, then followed it once the decision was made, and the rebuilt sequence was genuinely easier to review and to test. My initial objection was about sunk cost, and the decision was about reviewability.
—Reproducing before fixing, even when the cause was already visible in the code. For several defects I could point at the line from reading alone. Reproducing anyway caught two cases where my explanation was wrong: a code path I believed was live had been replaced by another service, and a scenario I had written could not occur because a client-side guard prevented it.
06
Impact
—Blast radius corrected from ~29,000 records to 445 that actually displayed a wrong percentage to a customer, across three customers rather than the twenty implied by the raw count.
—A whole class of defect found that the first analysis had missed entirely, 4,174 records, almost all customer-facing, because my initial filter assumed a condition that did not hold.
—Turned a proposed historical-correction project into an audited one-off script over a few hundred rows, reviewable by hand.
—Every defect reproduced in a controlled environment with a written prediction before any fix was merged; each fix shipped with a regression test verified to fail without it.
—Two independent reproductions, fifteen hours apart, of the defect whose cause had been described incorrectly in the tracker.
07Lessons Learned
+
—"How big is this?" is an engineering task, not a preamble. The first number I produced was wrong by ten times, and it was the number the team would have planned around. Measurement deserves the same scepticism as code.
—Separate "provably wrong" from "fixable". They are different questions, and only the second one decides whether a correction is possible. Conflating them leads to promising a repair for data whose correct value no longer exists.
—A control case is worth more than another failing case. The route that behaved correctly across twenty cycles is what made the failing route's nineteen failures attributable to one specific cause instead of a general suspicion.
—Test data in a production database will dominate any naive count. Filtering it out changed the recommendation completely; reporting without filtering would have cost credibility the first time someone checked.
—Reproduce before fixing, even when the cause looks obvious. Reading code tells you what a path does, not whether that path is the one running. Two of my confident explanations were wrong for exactly that reason.
—Correcting your own published number early is cheaper than defending it. Each correction changed the plan, and each was easier to make before the plan had been committed to than after.
08Evidence
+
—Production measurement, read-only replica, 120-day window: ~888,000 records analysed; 95% verified correct by an independent recomputation; defect population classified into correctable and unrecoverable buckets with exact counts for the former.
—Staging reproduction, named scenarios with written predictions: one route exhibited the defect in 19 of 20 consecutive cycles with a constant signature, while its control route was correct in all 20. The same route isolated a second, independent defect with the opposite signature, making both visible in a single dataset.
—Fixes shipped as small, single-cause pull requests, each with regression tests confirmed to fail without the change, across two services.
—Tracker restructured to one task per defect, each carrying the reproduction, the measured volume, the product decision quoted verbatim, and the acceptance scenarios.
—Two published numbers corrected by me before release, once in a stakeholder-facing document and once in a pull request description.