Success story · Retail · Integrations

Thirteen lost stock updates in one night, and the Monday-morning test.

An internal tool built on Bolt pushed stock levels to three storefronts. One night the storefront API returned 502 for four hours. The tool retried three times, gave up, and logged nothing. On Monday, 546 phantom units were on sale and nobody knew.

Multi-store retail operationsBuilt with BoltOperations managerReadiness 1.9 → 3.948 hours
Success story · Retail · Integrations
13updates lost in 24 h→0silent failures since
Thirteen lost stock updates in one night, and the Monday-morning test.
prodreadyBuilt with Bolt · production-ready in 48 hours
13
updates lost in 24 h
546
phantom units on sale
11
retry attempts, then dead-letter
0
silent failures since

Sound familiar?

You run operations for a retailer with more than one shop. Stock lives in a warehouse system; the storefronts want it in their own format, on their own API, right now. So you built the bridge yourself, on Bolt, in an afternoon. It reads a stock change, calls three APIs, done. It has worked every day since. Nobody has looked at it since either, because it has never shown an error. That is the part that should worry you. A tool that fails loudly gets fixed. A tool that fails quietly gets trusted.

What they sent us

An internal tool, built on Bolt by the operations manager of a multi-store retailer. One Node process, one 1,800-line server file, one Postgres table of stock levels. Roughly 4,000 SKUs, three storefronts, four countries. Every time the warehouse export changed a quantity, the tool called each storefront API with the new number.

The call path was three lines long. Try the request. If it fails, try again, up to three times, one second apart. If it still fails, move on. No queue, no log line, no alert. The code did not have a bug. It had an opinion: a failed push is not worth remembering.

One Saturday night the storefront API answered 502 Bad Gateway for four hours. Thirteen stock changes hit that window. Each got its three tries, three seconds of patience, and was gone. The tool never crashed, the process never restarted, the dashboard stayed green. On Monday the warehouse held fewer units than the shops were selling.

At Discovery we scored the app 1.9 out of 5 on our readiness scale. The build was fine. The operations layer, the part that tells you what happened while you slept and makes a tool production-ready, did not exist.

Warehouse export stock change, any hour Bolt tool push now, in memory 3 retries, 1 s apart then: forget it Storefront API 1 502 for 4 hours Storefront API 2 marketplace Storefront API 3 marketplace Log: nothing Alert: nobody. 13 updates gone.
BEFORE: every stock change went straight from memory to three APIs. Three tries, then the update was dropped. Nothing wrote it down.

What would have happened

What actually happens

Saturday 23:40, the storefront API starts returning 502. By 03:50 it recovers. Thirteen stock changes are missing from three shops, and the tool has no record that they ever existed. Sunday is the busiest day of the week. By Monday morning, 546 units that are not in the warehouse are on sale and 41 of them are already ordered. On day 3 the picking list comes back short and customer service starts writing cancellation emails. On day 5 the marketplace seller score drops below the threshold and the listings get buried. On day 12 the finance lead asks why refunds went up and the answer is "the sync must have missed something", because that is all anyone can say. The outage that costs money is never the one you saw. It is the one you did not.

It never crashed. It just quietly stopped being true.

Forty-eight hours

Hour 0–6. We did not touch the push code first. We put a table in front of it. Every stock change now inserts one row into stock_outbox: the SKU, the storefront, a next_attempt_at and an attempt counter. That is all. The row is a key, not a value. The current quantity is read from the stock table at send time, so if the same SKU changes five times during an outage, the shop receives the latest number once, and order never matters.

Hour 6–18. The sender. A loop wakes every 2 seconds, takes up to 50 due rows per storefront, sends them, and reschedules or clears them. The backoff schedule came from a real overnight outage, not a textbook. Five seconds, doubling, capped at five minutes, for eight attempts. That covers about fifteen minutes of trouble. Then three more attempts an hour apart. Eleven attempts in total, roughly three and a quarter hours of patience, then the row is moved to stock_dead_letter and someone gets a message.

// attempt n: 5 s, 10 s, 20 s ... capped at 5 min; after 8, hourly
const backoffMs = (n: number) =>
  n > 8 ? 3_600_000 : Math.min(5000 * 2 ** (n - 1), 300_000);

const isPermanent = (e) => [400, 401, 404, 422].includes(e.status);

if (isPermanent(err) || row.attempts + 1 >= 11) {
  await markDead(row, err);          // dead-letter + alert
} else {
  await reschedule(row, Date.now() + backoffMs(row.attempts + 1));
}

Two kinds of failure, two different answers. A 502, a 504 or an ECONNRESET is the network having a bad night; you wait and try again. A 400 or a 404 is your request being wrong; retrying it eleven times is just eleven more wrong requests. Those go straight to dead-letter with the response body attached, so the person who opens the ticket can read why.

A four-hour outage still outlasts eleven attempts. The difference is where the row ends up. It sits in dead-letter with its history, an alert has gone out, and replaying it is one command once the API is back. Lost is not a state the system has any more.

Hour 18–36. The things around the sender. The original tool ran its push on a one-minute cron, and when the API was slow, runs stacked on top of each other; in our test against a stalled endpoint we counted six concurrent runs fighting over the same rows. One flag fixes that.

if (this.running) { log.warn('sender: previous run still going, skipping'); return; }
this.running = true;
try {
  const due = await db.select().from(outbox)
    .where(lte(outbox.nextAttemptAt, new Date())).limit(50);
  for (const row of due) await send(row);
} finally {
  this.running = false;
}

If the run is still going, the next tick skips and says so. On boot, any row still marked processing is set to failed with the reason "server restarted during processing", so a deploy in the middle of a batch cannot strand a row forever.

Then the health endpoint. The old one returned 200 OK as long as the process was alive, which is exactly what it did during the outage. The new one checks the database and the outbox backlog, and answers 503 when either is wrong. An uptime monitor polls it every 60 seconds, and 503 is the signal it understands.

app.get('/api/health', async (_req, res) => {
  try {
    await db.execute(sql`select 1`);
    const stuck = await outboxRowsOlderThan('15 minutes');
    const dead  = await deadLetterCount();
    const ok = stuck === 0 && dead === 0;
    res.status(ok ? 200 : 503).json({ ok, stuck, dead, version });
  } catch {
    res.status(503).json({ ok: false, error: 'db' });
  }
});

The version field comes from a version.json written on every build, so a screenshot of the health page tells you which build was running when it went wrong. Builds are locked to a dependency lockfile and produced by one script; nobody deploys from a laptop.

Hour 36–48. Backups, and the proof that they work. A nightly pg_dump at 02:10, fourteen days kept. Then the part most teams skip: we restored last night's dump into a scratch database and compared row counts, table by table. Six minutes. That is now a monthly calendar entry, and the result is a line in the report.

# nightly at 02:10, keep 14 days
pg_dump -Fc "$DATABASE_URL" > /backups/stock-$(date +%F).dump
find /backups -name 'stock-*.dump' -mtime +14 -delete

# restore drill, monthly, into a scratch database
createdb stock_drill
pg_restore --no-owner --no-acl -d stock_drill /backups/stock-$(date +%F).dump
psql stock_drill -c "select count(*) from stock_levels"   # compare to prod
dropdb stock_drill

Last, the Monday report. Every Monday at 07:00 the tool sends a short message: rows in dead-letter, oldest pending row, how many pushes needed more than one attempt, restarts since last week, health flaps, last backup, last restore drill. Six lines. If every line says zero, the weekend was quiet. If it does not, the operations manager knows before the first customer does.

Discovery before hour 0 · customer traffic after hour 48 0–6 6–18 18–36 36–48 0 h 6 h 18 h 36 h 48 h Outbox table rows are keys only Sender + backoff 11 attempts, then dead-letter + alert Guards, health, version cron guard · boot recovery · 503 · version.json Backups + report nightly pg_dump · restore drill · Monday 07:00 13 lost pushes replayed 502 test page as fixture fingerprint bug found and fixed restore in 6 min, handover Readiness 1.9 → 3.9
TIMELINE: the 48-hour engineering window. The outbox table went in first, so every later step had somewhere durable to stand.

From our desk

While reading the settings code we found a second, quieter bug. The tool decided whether to re-sync the catalogue by comparing a hash of the settings object with the last saved one. Save the same settings from a different screen and the keys came out in a different order, the hash changed, and the whole catalogue, every SKU for every storefront, was queued for a push. On the old code that was 4,000 pointless API calls and a good way to get rate-limited during business hours. The fix is one line: sort the keys before hashing. The test for it is four lines. Both are in the repository now, next to a test that feeds the sender the literal HTML of a 502 error page and checks it gets classified as transient.

What it looks like now

The same tool, the same Bolt code for the parts that were fine, the same three storefronts. Between the stock table and the outside world there is now a queue that survives restarts, a sender that knows the difference between "later" and "never", and a small set of instruments that answer one question: what went wrong while nobody was looking.

Stock change warehouse export Outbox table durable, key only value read at send time Sender, every 2 s 5 s → 5 min × 8 then hourly × 3 one run at a time 3 storefront APIs 502 = try later Dead-letter + alert after attempt 11 or a 4xx · replayable GET /api/health 503 if DB or backlog is bad Nightly pg_dump 14 days kept · monthly restore drill Monday report 07:00, six lines On boot: rows stuck in "processing" are marked failed · version.json in every health response 0 silent failures since
AFTER: a durable outbox in front of the sender, a backoff schedule with an end, and three instruments that report what happened. Nothing is dropped without a record.
QuestionBeforeAfter
A storefront API returns 502 for four hours3 tries, 3 seconds, update gone11 attempts over ~3 h, then dead-letter and an alert
The process restarts mid-batchIn-flight pushes vanishRows survive; stuck ones marked failed on boot
Two runs overlap on a slow APIThey fight over the same rowsSecond run skips and logs a warning
Is the tool healthy?200 as long as the process is alive503 when the database or the backlog is wrong
Which build was running when it broke?Nobody knowsversion.json in every health response
The database is lostNo backup, no planNightly dump, 14 days, restore proven monthly
What went wrong over the weekend?Find out from customersSix-line report, Monday 07:00

Readiness went from 1.9 to 3.9. The remaining gap is deliberate: the tool still runs as one process on one server, and the team decided that a second instance is a problem for next year, not this quarter. The Monday report has said zero silent failures every week since.


Do this tonight

Do this tonight

1. Break the API on purpose. Point your tool's outbound URL at a port where nothing listens (http://127.0.0.1:9 works), change one stock level, wait five minutes, then put the URL back. Did the update arrive? Did anything log or alert while it was failing? If the answer is no and no, you have the same bug.

2. Ask your health endpoint a hard question. Run curl -i https://your-app/api/health, then stop your database (or revoke the app's password for a minute) and run it again. If it still says 200, your uptime monitor is watching a light that is always on.

3. Restore a backup somewhere. Find the most recent dump. If you cannot find one, that is the result. If you can, run pg_restore into a fresh database and count the rows in your biggest table. Write the date and the number in a file in the repository.

These three checks find the gaps. They do not build the outbox, the classifier that tells 502 from 400, the boot recovery, the report, or the tests that keep all of it true after the next feature. That last mile is a different job, and it is the one prodready does in 48 hours.

The rule

The rule

A failed write that nobody recorded is not a failure, it is a lie the system tells you on Monday. Never drop a message without writing down that you dropped it.

It never crashed. It just quietly stopped being true.— Operations manager

Send us what you have.

The readiness scan is free — install our MCP and your own coding agent runs it. When you want it fixed, send us what you have: fixed scope, fixed price, agreed before the clock starts. 48 hours later, it's production-ready — on your own cloud, with the tests to prove it.

Feed the machineWatch the 49-second film