Accepting inspection answers instantly and processing them durably, with an outbox
Redesigned how field inspectors' answers reach the platform, so the mobile app gets an immediate confirmation after validation while a background consumer processes the answers durably, with retries and traceable failures; outcome qualitative, not measured.
impactThe mobile app receives a confirmation as soon as validation passes, instead of waiting for every answer to be processed.Accepted answers are stored durably before processing, so a downstream failure no longer risks losing them; failed groups are retried and recorded with enough context to reprocess.Heavy processing moved off the request path, so peak-hour submissions no longer do all their work at arrival.Every group is traceable end to end by its identifier, and time-to-integration is logged, which gives the team a way to measure the flow in production.
01
Context
It is the clearest example of me designing a flow across a boundary I don't control: a mobile client on an unreliable connection, talking to a backend that does expensive work. It is evidence for system design and reliability reasoning: what must be synchronous, what can be deferred, and what has to be durable in between.
02
Problem
At the end of an inspection route, the mobile app sends the whole batch of answers: questions answered, measurements, justifications, "no anomalies" confirmations, plus who sent them and for which route and checklist. The backend handled the entire batch inside the request. That had three consequences:
—Large submissions were slow, because the user waited for every answer to be fully processed.
—Failures could be silent. If the connection dropped mid-processing, the app had no reliable signal of what had or hadn't been saved.
—Peak hours concentrated load, since many inspectors finish routes around the same times and every submission did its full work on arrival.
03Constraints
+
—The client is a mobile app in the field. Connectivity is unreliable, and the inspector needs a clear "your answers are safe" signal before moving on. Any design that leaves that ambiguous pushes the problem onto the person holding the phone.
—Not everything can be deferred. Some failures must reach the inspector while they can still act on them (wrong permissions, a route or checklist that doesn't exist, an asset that isn't on that route). Others (a processing error deep in the pipeline) should never be the inspector's problem.
—One submission mixes several kinds of work. Checklist answers, measurement-point answers, justifications and "not applicable" results each need different processing and different downstream events.
—Duplicates are expected. A retrying client on a flaky connection will sometimes send the same answers twice.
04
Decision
I split the flow at the point where the backend can honestly promise "your answers are safe", and made everything before that point fast and everything after it durable.
—Validate synchronously, completely. Authentication, role, required fields, existence of routes, checklists and measurement points, that each asset belongs to the route it was answered under, and that the inspector may access those routes. All of it runs before anything is accepted, so every error the inspector can fix comes back immediately, with details. Accidental duplicates are removed at this stage too.
—Persist to an outbox, grouped and traceable. Valid answers are written to an outbox table, grouped by what they belong to (a checklist on a route, or a measurement point on a route). Each group gets a unique identifier and records its kind, its operation mode, the full payload, the inspector, and whether it was sent by the inspector or on their behalf by support.
—Confirm to the app at that point. The response lists the groups created and their identifiers. From the inspector's perspective the submission is done; the answers are durably stored even if nothing downstream has run yet.
—Process in the background. A consumer receives the outbox entries through Kafka, validates each message, routes it to the processor for its mode, writes the answers to their final tables, and publishes events so reports, dashboards, alerts and inspection history update.
—Make failures recoverable, not just visible. A processing failure is recorded with the Kafka topic, partition and offset, the mode, the group identifier, the answer count and the stack trace, and the message is kept so it can be reprocessed. Each processed group also logs the time between submission and integration.
05Trade-offs
+
—Eventual consistency over a synchronous "fully saved" answer. The confirmation now means "safely stored", not "visible everywhere". The inspector gets a fast, reliable answer; the cost is a short window in which reports don't yet reflect the submission, which is why time-to-integration is logged per group.
—Strict validation at the edge over validating in the background. Doing all validation in the request makes it heavier than a bare "store and acknowledge", but it keeps every user-fixable error in front of the user and keeps the background consumer from ever needing to report back to a phone.
—Grouping by checklist or measurement point over one record per submission. Groups give each unit of work its own identifier, mode and failure record, so one bad group can be traced and reprocessed without replaying the whole submission.
06
Impact
Qualitative; no before/after metric was captured.
—The mobile app receives a confirmation as soon as validation passes, instead of waiting for every answer to be processed.
—Accepted answers are stored durably before processing, so a downstream failure no longer risks losing them; failed groups are retried and recorded with enough context to reprocess.
—Heavy processing moved off the request path, so peak-hour submissions no longer do all their work at arrival.
—Every group is traceable end to end by its identifier, and time-to-integration is logged, which gives the team a way to measure the flow in production.
07Lessons Learned
+
—Put the promise where you can keep it. A synchronous boundary should confirm only what is already guaranteed. Here that was "validated and durably stored", not "processed".
—Errors belong to whoever can fix them. User-fixable errors go back in the request; system errors go to a durable, reprocessable record. Mixing the two is what makes failures feel silent.
—The unit of retry is a design decision. Choosing the group (not the request, not the single answer) as the unit made identifiers, logging and reprocessing all line up.
08Evidence
+
—Feature functional documentation (internal), explaining the flow for a non-engineering audience.
—Structured logs per submission and per processed group, including time-to-integration.