Sound familiar?
You built it on Replit because the spreadsheet broke. One screen, one CSV upload, a free Postgres. It worked, so the supplier feed got wired in. Then stock started flowing from it to the shops. Then a colleague in another country asked for a login. Nobody decided this was the product catalogue; it just is now. The schema still changes through a tool that asks you yes/no questions at 11 p.m. You are not sure what happens when a feed sends garbage, because so far it mostly has not. You are the only person who knows how any of it works.
What they sent us
The head of e-commerce operations at a fashion brand selling in four countries built the first version in a weekend. By Discovery it was the system of record: 3,782 product models, a 106 MB Postgres dump, four currencies, seven languages. Feeding it, one supplier XML feed pulled over FTPS every hour and one stock event stream on Kafka. Out of it, stock pushes to four storefronts, all from one Node process.
The code was still the weekend. The schema was applied with an interactive push tool that diffs a schema file against the live database; the file had drifted, so every run asked was table X renamed from Y? The wrong answer drops a production table. Nobody had answered wrong yet.
The import loop ran hourly with no overlap guard, so a slow file and the next run started on top of each other. A restart left jobs marked processing forever. Identifiers were merged with incoming || stored, which is how a supplier's internal counter had already overwritten 193 real barcodes and left the warehouse unable to match picking lists.
Stock went to the storefronts as a bare HTTP call; when a shop's proxy answered 502 the update was gone, thirteen of them in one 24-hour window. Nothing read stock back from the shops, so the catalogue and the storefronts had drifted apart by 78,482 units in two months and nobody could see it.
The rest was the usual prototype residue: sessions in memory, a session secret that fell back to a placeholder string, cookies with secure:false and httpOnly:false, bearer tokens from Math.random held in a Map that emptied on every deploy, database credentials committed in a scripts folder. Our readiness scan put it at 1.8 out of 5.
What would have happened
Peak season was weeks away and the team knew the numbers were wrong, so they had already tried to fix the drift themselves. The script matched catalogue rows to shop rows on EAN, got a false reading for one category, and took 11,778 real units off sale for two hours on a normal trading afternoon. That is the trap with a system that was never built to be checked: the first check you write is as dangerous as the bug.
What actually happens
Peak season starts on a Thursday. On day 3 the supplier feed sends its counter in the barcode field again; the warehouse stops trusting picking lists and hand-checks every parcel. On day 9 a storefront proxy returns 502 overnight; a night of stock updates never arrives and 546 phantom units sell against sizes that do not exist. On day 12 someone needs one more column, runs the schema push, answers yes to a rename prompt, and the size-variant table is gone. The last plain dump is six days old. Drift by then: 78,482 units, spread over four countries. The site goes to a holding page in the biggest week of the year, the ops lead rebuilds sizes from a spreadsheet export, and the board asks who signed off on a catalogue built in a weekend.
Forty-eight hours
The point is not that the prototype was bad. It did its job and taught the team exactly what the product needed to be. The last mile is a different job, and it starts with taking the sharp objects away.
Hour 0–6: foundations
We stopped the app, took a compressed dump, restored it onto managed Postgres with --no-owner --no-acl, and compared row counts table by table before anything else moved; the dump stayed on disk as the rollback. Then the push tool left the deploy path for good. In its place: hand-written, forward-only, additive SQL files numbered 0001 to 0004, run by a small runner that records the date each one was applied.
-- 0003_reconcile_runs.sql · additive, re-runnable, never a DROP
CREATE TABLE IF NOT EXISTS reconcile_runs (
id SERIAL PRIMARY KEY,
ran_at TIMESTAMP NOT NULL DEFAULT now(),
country TEXT NOT NULL,
mismatched INT NOT NULL DEFAULT 0,
error TEXT
);
CREATE INDEX IF NOT EXISTS idx_runs_country
ON reconcile_runs (country, ran_at DESC);
Every statement is safe to run twice, so a migration that half-applied during a deploy can simply be re-run. There is no rename prompt because there are no renames: a rename is a new column, a backfill and a later drop, in three separate files.
The same six hours closed the prototype residue. The app now refuses to start with a placeholder session secret. Sessions and bearer tokens live in the database, cookies are secure and httpOnly, the committed credentials were rotated, and the connection pool got sane limits: ten connections, thirty-second idle timeout, each connection recycled after 7,500 uses, with retries only on connection-class errors.
Hour 6–18: the data model
Fashion stock has three levels and the prototype had one and a half. We made it explicit: model, then model-colour as the primary key, then model-colour-size. Every feed, every storefront push and every report now speaks the same three keys.
Then the cleanup scripts, all dry-run by default. Some size SKUs carried _ where the shops expected -, because the feed's internal id uses underscores; fixing them in place collided with correct rows that already existed (23505, unique violation), so on collision the script deletes the bad row and its country rows, in batches of 200, printing an ETA.
Sizes that differed only in spelling, "ONE SIZE" against "ONESIZE", were merged; that alone removed 546 phantom units from sale. Lowercase country codes were folded into uppercase at startup. Barcodes were separated from the supplier's counter by GS1 prefix, each in its own column, and incoming || stored left every identifier path.
Two smaller things earned their own helpers. A static VAT table per country returns null for an unknown country rather than guessing. And every date-time in the codebase goes through one fixed-timezone function, after we found a promotion that had started two hours early because a zone-less timestamp was read as UTC.
Hour 18–36: ingest and outbox
This is where the prototype and the product part ways. The Kafka consumer now writes raw events to the database first and only then commits the offset. If the insert fails, nothing is committed and the broker redelivers. The client retries eight times with exponential backoff capped at 30 seconds, sends a heartbeat every 100 messages and warns after two minutes of silence.
eachBatch: async ({ batch, resolveOffset, commitOffsetsIfNecessary }) => {
try { await saveRawEvents(batch.messages); } // durable first
catch { return; } // no commit → redelivery
await applyChanges(batch.messages);
resolveOffset(batch.messages.at(-1)!.offset);
await commitOffsetsIfNecessary();
}
Order events became immutable rows with a unique key on (topic, partition, offset), so a redelivery is ignored. The order aggregate is upserted by order number and only updated when the incoming timestamp is newer, so an event that arrives late cannot overwrite a newer one. A message that fails validation is saved with status='error' and the consumer keeps going. Validation asks for two required fields and passes everything else through, because a strict schema on someone else's feed is a self-inflicted outage.
The FTP importer got a processed-files table so only new or failed files are touched, a running flag so overlapping runs skip, and a boot step that marks anything still "processing" as failed with the reason "server restarted during processing". Bad XML returns a typed error; invalid products are skipped with a warning and a count, not a crash.
Outbound, every stock change to a storefront is now a row in a durable outbox. The row holds only a key; the current figures are read from the database at send time, so order does not matter and a burst of changes collapses into one send. The outbox ticks every two seconds, 50 rows per country. Transient errors (502, 504, ECONNRESET) are retried; a 400 or a not-found fails at once.
// 5 s doubling to 5 min for 8 attempts, then 3 hourly, then dead-letter
const backoffMs = (n: number) =>
n > 8 ? 3_600_000 : Math.min(5000 * 2 ** (n - 1), 300_000);
if (isPermanent(err) || row.attempts + 1 >= 11) markDead(row);
else reschedule(row, Date.now() + backoffMs(row.attempts + 1));
Eleven attempts over about three and a quarter hours, then a dead-letter row a person can see and replay. The update is never dropped on the floor; the worst case is that it waits.
From our desk
The three hourly attempts were not in the first version. On the ninth night after go-live a storefront's proxy returned 502 for most of the night, longer than the fast retries lasted, and 41 rows went to dead-letter. The morning shift replayed them in one click; nothing was lost, but a person had been needed, so we added the hourly tail the same day. The literal HTML of that 502 page is now a test fixture.
Hour 36–48: reconciliation, tests, gate
Reconciliation reads stock back from each storefront four times a day and compares it to the catalogue on the same key the sync uses, not on EAN. Exactly one class of mismatch is fixed automatically: the shop is behind and the catalogue has a newer, complete figure. Every other class goes into a report with a proposed action and waits for a person. That is the difference between a check and a script that takes 11,778 units off sale.
The tests were written from the incident list: 40 files, 717 cases across API, services and client, plus one browser end-to-end run. SKU normalisation is tested against non-breaking spaces and typographic dashes from spreadsheet exports; retry classification against the real 502 page. The order tests check that an older event never overwrites newer data and that a duplicate offset is ignored.
it('an older event never overwrites newer order data', async () => {
await ingest(orderEvent({ id: 'A', total: 90, at: '2026-03-01T10:00Z' }));
await ingest(orderEvent({ id: 'A', total: 40, at: '2026-03-01T09:00Z' })); // late
expect((await orders.get('A')).total).toBe(90);
});
it('a redelivered offset is ignored', async () => {
const e = rawEvent({ topic: 'orders', partition: 0, offset: 17 });
await ingest(e); await ingest(e);
expect(await events.count()).toBe(1);
});
Each test is one incident that cannot recur; that is the whole strategy, not coverage but a list of things that already went wrong once.
Then the gate. Three independent reviewers score every change from 0 to 5 on architecture, security and functional completeness; any score under 3 is a rejection and each rejection finding becomes a mandatory task. /api/health returns 503 when the database is unhealthy, and the Kafka consumer and the outbox lane sit behind environment flags, so they were switched on one at a time without a redeploy. Readiness at the gate: 4.1.
What it looks like now
The same codebase, with the same name, now runs product data for one of Europe's largest fashion e-commerce brands: 18,162 product models, 42,024 colour variants, 311,497 size variants across 62 tables and a 222 MB dump, six countries with a VAT rate for each, seven languages. In one two-week window the stock consumer took in 725,132 events; 713,509 carried a real barcode, and the 11,623 that did not were stored, flagged and never applied. Peak season went through it with zero lost updates.
| At Discovery | Now | |
|---|---|---|
| Schema changes | Interactive push; a wrong answer drops a live table | Numbered forward-only SQL, re-runnable, dated log |
| Stock push fails with 502 | Update lost; 13 in one day | Retried 11 times over 3 h, then dead-letter; 0 lost |
| Duplicate or late event | Overwrote newer data | Ignored by (topic, partition, offset); upsert only if newer |
| Identifier from a feed | incoming || stored; 193 barcodes overwritten | Separated by GS1 prefix; identifiers never overwritten |
| Catalogue vs storefronts | Never compared; 78,482 units of drift | Read back 4× a day; one class auto-fixed, rest to a person |
| Import runs | Overlapping; stuck "processing" after a restart | Seen-files table, run guard, boot cleanup |
| Sessions and secrets | Memory, Math.random, placeholder secret | Database-backed, fail-closed config, secure cookies |
| Tests | None that touched a feed | 717 cases in 40 files, each one a past incident |
I built the first version myself in a weekend. I did not build the version that runs the business. Turns out those are two different products with the same name.Head of e-commerce operations
Those two products still share a repository and most of their screens. What changed is everything underneath, and the ops lead who built the first one now reads a mismatch report over coffee instead of a spreadsheet at midnight.
Do this tonight
Do this tonight
1. Find out how your schema reaches production. Run grep -rn "db push\|drizzle-kit push\|synchronize: true\|create_all" package.json scripts deploy README.md. If it hits anything a deploy runs, your tables are one prompt away from gone. Take a pg_dump -Fc right now and put it somewhere that is not the same server.
2. Break a downstream on purpose. Point the storefront URL in your env at a port with nothing on it, http://127.0.0.1:9, change one product's stock, wait one minute, put the URL back. Does the change arrive, or is it gone? If your logs cannot tell you, that is also your answer.
3. Sample fifty SKUs. SELECT sku, stock FROM variants ORDER BY random() LIMIT 50, then look up the same fifty in the shop's admin. Any difference is drift you did not know you had, and it only grows.
These three tell you whether you have the problem. They do not give you the store-first consumer, the outbox, the migration runner or the reconciliation that knows which mismatches it may fix alone. That is the 48 hours, and it is what prodready does.
The rule
The rule
A feed is not an input, it is a stranger with a schedule. Store it before you trust it, retry it before you lose it, and read it back before you believe it.
I built the first version myself in a weekend. I did not build the version that runs the business. Turns out those are two different products with the same name.— Head of e-commerce operations