Sound familiar?
You run operations for a retailer with more than one shop. Stock lives in a warehouse system; the storefronts want it in their own format, on their own API, right now. So you built the bridge yourself, on Bolt, in an afternoon. It reads a stock change, calls three APIs, done. It has worked every day since. Nobody has looked at it since either, because it has never shown an error. That is the part that should worry you. A tool that fails loudly gets fixed. A tool that fails quietly gets trusted.
What they sent us
An internal tool, built on Bolt by the operations manager of a multi-store retailer. One Node process, one 1,800-line server file, one Postgres table of stock levels. Roughly 4,000 SKUs, three storefronts, four countries. Every time the warehouse export changed a quantity, the tool called each storefront API with the new number.
The call path was three lines long. Try the request. If it fails, try again, up to three times, one second apart. If it still fails, move on. No queue, no log line, no alert. The code did not have a bug. It had an opinion: a failed push is not worth remembering.
One Saturday night the storefront API answered 502 Bad Gateway for four hours. Thirteen stock changes hit that window. Each got its three tries, three seconds of patience, and was gone. The tool never crashed, the process never restarted, the dashboard stayed green. On Monday the warehouse held fewer units than the shops were selling.
At Discovery we scored the app 1.9 out of 5 on our readiness scale. The build was fine. The operations layer, the part that tells you what happened while you slept and makes a tool production-ready, did not exist.
What would have happened
What actually happens
Saturday 23:40, the storefront API starts returning 502. By 03:50 it recovers. Thirteen stock changes are missing from three shops, and the tool has no record that they ever existed. Sunday is the busiest day of the week. By Monday morning, 546 units that are not in the warehouse are on sale and 41 of them are already ordered. On day 3 the picking list comes back short and customer service starts writing cancellation emails. On day 5 the marketplace seller score drops below the threshold and the listings get buried. On day 12 the finance lead asks why refunds went up and the answer is "the sync must have missed something", because that is all anyone can say. The outage that costs money is never the one you saw. It is the one you did not.
It never crashed. It just quietly stopped being true.
Forty-eight hours
Hour 0–6. We did not touch the push code first. We put a table in front of it. Every stock change now inserts one row into stock_outbox: the SKU, the storefront, a next_attempt_at and an attempt counter. That is all. The row is a key, not a value. The current quantity is read from the stock table at send time, so if the same SKU changes five times during an outage, the shop receives the latest number once, and order never matters.
Hour 6–18. The sender. A loop wakes every 2 seconds, takes up to 50 due rows per storefront, sends them, and reschedules or clears them. The backoff schedule came from a real overnight outage, not a textbook. Five seconds, doubling, capped at five minutes, for eight attempts. That covers about fifteen minutes of trouble. Then three more attempts an hour apart. Eleven attempts in total, roughly three and a quarter hours of patience, then the row is moved to stock_dead_letter and someone gets a message.
// attempt n: 5 s, 10 s, 20 s ... capped at 5 min; after 8, hourly
const backoffMs = (n: number) =>
n > 8 ? 3_600_000 : Math.min(5000 * 2 ** (n - 1), 300_000);
const isPermanent = (e) => [400, 401, 404, 422].includes(e.status);
if (isPermanent(err) || row.attempts + 1 >= 11) {
await markDead(row, err); // dead-letter + alert
} else {
await reschedule(row, Date.now() + backoffMs(row.attempts + 1));
}
Two kinds of failure, two different answers. A 502, a 504 or an ECONNRESET is the network having a bad night; you wait and try again. A 400 or a 404 is your request being wrong; retrying it eleven times is just eleven more wrong requests. Those go straight to dead-letter with the response body attached, so the person who opens the ticket can read why.
A four-hour outage still outlasts eleven attempts. The difference is where the row ends up. It sits in dead-letter with its history, an alert has gone out, and replaying it is one command once the API is back. Lost is not a state the system has any more.
Hour 18–36. The things around the sender. The original tool ran its push on a one-minute cron, and when the API was slow, runs stacked on top of each other; in our test against a stalled endpoint we counted six concurrent runs fighting over the same rows. One flag fixes that.
if (this.running) { log.warn('sender: previous run still going, skipping'); return; }
this.running = true;
try {
const due = await db.select().from(outbox)
.where(lte(outbox.nextAttemptAt, new Date())).limit(50);
for (const row of due) await send(row);
} finally {
this.running = false;
}
If the run is still going, the next tick skips and says so. On boot, any row still marked processing is set to failed with the reason "server restarted during processing", so a deploy in the middle of a batch cannot strand a row forever.
Then the health endpoint. The old one returned 200 OK as long as the process was alive, which is exactly what it did during the outage. The new one checks the database and the outbox backlog, and answers 503 when either is wrong. An uptime monitor polls it every 60 seconds, and 503 is the signal it understands.
app.get('/api/health', async (_req, res) => {
try {
await db.execute(sql`select 1`);
const stuck = await outboxRowsOlderThan('15 minutes');
const dead = await deadLetterCount();
const ok = stuck === 0 && dead === 0;
res.status(ok ? 200 : 503).json({ ok, stuck, dead, version });
} catch {
res.status(503).json({ ok: false, error: 'db' });
}
});
The version field comes from a version.json written on every build, so a screenshot of the health page tells you which build was running when it went wrong. Builds are locked to a dependency lockfile and produced by one script; nobody deploys from a laptop.
Hour 36–48. Backups, and the proof that they work. A nightly pg_dump at 02:10, fourteen days kept. Then the part most teams skip: we restored last night's dump into a scratch database and compared row counts, table by table. Six minutes. That is now a monthly calendar entry, and the result is a line in the report.
# nightly at 02:10, keep 14 days
pg_dump -Fc "$DATABASE_URL" > /backups/stock-$(date +%F).dump
find /backups -name 'stock-*.dump' -mtime +14 -delete
# restore drill, monthly, into a scratch database
createdb stock_drill
pg_restore --no-owner --no-acl -d stock_drill /backups/stock-$(date +%F).dump
psql stock_drill -c "select count(*) from stock_levels" # compare to prod
dropdb stock_drill
Last, the Monday report. Every Monday at 07:00 the tool sends a short message: rows in dead-letter, oldest pending row, how many pushes needed more than one attempt, restarts since last week, health flaps, last backup, last restore drill. Six lines. If every line says zero, the weekend was quiet. If it does not, the operations manager knows before the first customer does.
From our desk
While reading the settings code we found a second, quieter bug. The tool decided whether to re-sync the catalogue by comparing a hash of the settings object with the last saved one. Save the same settings from a different screen and the keys came out in a different order, the hash changed, and the whole catalogue, every SKU for every storefront, was queued for a push. On the old code that was 4,000 pointless API calls and a good way to get rate-limited during business hours. The fix is one line: sort the keys before hashing. The test for it is four lines. Both are in the repository now, next to a test that feeds the sender the literal HTML of a 502 error page and checks it gets classified as transient.
What it looks like now
The same tool, the same Bolt code for the parts that were fine, the same three storefronts. Between the stock table and the outside world there is now a queue that survives restarts, a sender that knows the difference between "later" and "never", and a small set of instruments that answer one question: what went wrong while nobody was looking.
| Question | Before | After |
|---|---|---|
| A storefront API returns 502 for four hours | 3 tries, 3 seconds, update gone | 11 attempts over ~3 h, then dead-letter and an alert |
| The process restarts mid-batch | In-flight pushes vanish | Rows survive; stuck ones marked failed on boot |
| Two runs overlap on a slow API | They fight over the same rows | Second run skips and logs a warning |
| Is the tool healthy? | 200 as long as the process is alive | 503 when the database or the backlog is wrong |
| Which build was running when it broke? | Nobody knows | version.json in every health response |
| The database is lost | No backup, no plan | Nightly dump, 14 days, restore proven monthly |
| What went wrong over the weekend? | Find out from customers | Six-line report, Monday 07:00 |
Readiness went from 1.9 to 3.9. The remaining gap is deliberate: the tool still runs as one process on one server, and the team decided that a second instance is a problem for next year, not this quarter. The Monday report has said zero silent failures every week since.
Do this tonight
Do this tonight
1. Break the API on purpose. Point your tool's outbound URL at a port where nothing listens (http://127.0.0.1:9 works), change one stock level, wait five minutes, then put the URL back. Did the update arrive? Did anything log or alert while it was failing? If the answer is no and no, you have the same bug.
2. Ask your health endpoint a hard question. Run curl -i https://your-app/api/health, then stop your database (or revoke the app's password for a minute) and run it again. If it still says 200, your uptime monitor is watching a light that is always on.
3. Restore a backup somewhere. Find the most recent dump. If you cannot find one, that is the result. If you can, run pg_restore into a fresh database and count the rows in your biggest table. Write the date and the number in a file in the repository.
These three checks find the gaps. They do not build the outbox, the classifier that tells 502 from 400, the boot recovery, the report, or the tests that keep all of it true after the next feature. That last mile is a different job, and it is the one prodready does in 48 hours.
The rule
The rule
A failed write that nobody recorded is not a failure, it is a lie the system tells you on Monday. Never drop a message without writing down that you dropped it.
It never crashed. It just quietly stopped being true.— Operations manager