Success story · E-commerce · PIM

A product catalogue built on Replit now runs one of Europe's largest fashion e-commerce brands.

It started as a weekend prototype: one screen, one CSV, one very tired ops lead. Today it manages 311,000 size variants across six countries, eats 725,000 stock events a fortnight and survives a broken supplier feed at 3 a.m. Here is what changed between those two sentences.

Fashion e-commerce, 6 countriesBuilt with ReplitHead of e-commerce operationsReadiness 1.8 → 4.148 hours
Success story · E-commerce · PIM
18kproduct models→0lost updates since
A product catalogue built on Replit now runs one of Europe's largest fashion e-commerce brands.
prodreadyBuilt with Replit · production-ready in 48 hours
18k
product models
311k
size variants
725k
stock events / 2 weeks
0
lost updates since

Sound familiar?

You built it on Replit because the spreadsheet broke. One screen, one CSV upload, a free Postgres. It worked, so the supplier feed got wired in. Then stock started flowing from it to the shops. Then a colleague in another country asked for a login. Nobody decided this was the product catalogue; it just is now. The schema still changes through a tool that asks you yes/no questions at 11 p.m. You are not sure what happens when a feed sends garbage, because so far it mostly has not. You are the only person who knows how any of it works.

What they sent us

The head of e-commerce operations at a fashion brand selling in four countries built the first version in a weekend. By Discovery it was the system of record: 3,782 product models, a 106 MB Postgres dump, four currencies, seven languages. Feeding it, one supplier XML feed pulled over FTPS every hour and one stock event stream on Kafka. Out of it, stock pushes to four storefronts, all from one Node process.

The code was still the weekend. The schema was applied with an interactive push tool that diffs a schema file against the live database; the file had drifted, so every run asked was table X renamed from Y? The wrong answer drops a production table. Nobody had answered wrong yet.

The import loop ran hourly with no overlap guard, so a slow file and the next run started on top of each other. A restart left jobs marked processing forever. Identifiers were merged with incoming || stored, which is how a supplier's internal counter had already overwritten 193 real barcodes and left the warehouse unable to match picking lists.

Stock went to the storefronts as a bare HTTP call; when a shop's proxy answered 502 the update was gone, thirteen of them in one 24-hour window. Nothing read stock back from the shops, so the catalogue and the storefronts had drifted apart by 78,482 units in two months and nobody could see it.

The rest was the usual prototype residue: sessions in memory, a session secret that fell back to a placeholder string, cookies with secure:false and httpOnly:false, bearer tokens from Math.random held in a Map that emptied on every deploy, database credentials committed in a scripts folder. Our readiness scan put it at 1.8 out of 5.

BEFORE · at Discovery · readiness 1.8 Supplier XML FTPS, polled hourly Stock events Kafka, applied as read One import loop no overlap guard, no checks incoming || stored on ids 193 barcodes overwritten Postgres schema via interactive push wrong answer drops a table HTTP, no retry Storefronts × 4 a 502 loses the update 13 pushes lost in 24 h nothing reads stock back drift: 78,482 units in two months
BEFORE · One loop, one push, no way back. Every red box is an incident that had already happened at least once.

What would have happened

Peak season was weeks away and the team knew the numbers were wrong, so they had already tried to fix the drift themselves. The script matched catalogue rows to shop rows on EAN, got a false reading for one category, and took 11,778 real units off sale for two hours on a normal trading afternoon. That is the trap with a system that was never built to be checked: the first check you write is as dangerous as the bug.

What actually happens

Peak season starts on a Thursday. On day 3 the supplier feed sends its counter in the barcode field again; the warehouse stops trusting picking lists and hand-checks every parcel. On day 9 a storefront proxy returns 502 overnight; a night of stock updates never arrives and 546 phantom units sell against sizes that do not exist. On day 12 someone needs one more column, runs the schema push, answers yes to a rename prompt, and the size-variant table is gone. The last plain dump is six days old. Drift by then: 78,482 units, spread over four countries. The site goes to a holding page in the biggest week of the year, the ops lead rebuilds sizes from a spreadsheet export, and the board asks who signed off on a catalogue built in a weekend.

Forty-eight hours

The point is not that the prototype was bad. It did its job and taught the team exactly what the product needed to be. The last mile is a different job, and it starts with taking the sharp objects away.

TIMELINE · 48 engineering hours, after Discovery, before live traffic 0 h 6 h 18 h 36 h 48 h 6 h 12 h 18 h 12 h · go-live Hour 0–6 Foundations dump, restore, row counts forward-only migrations fail-closed config, secrets sessions and tokens to DB Hour 6–18 Data model model → colour → size keys SKU and size cleanups barcodes split by prefix VAT table, timezone helper Hour 18–36 Ingest and outbox Kafka: store, then commit FTP: seen-files, run guard outbox: retry, dead-letter late events never overwrite Hour 36–48 Reconcile and gate read-back reconciliation 717 tests, three reviewers health check, env flags go-live, readiness 4.1 Discovery before hour 0 · the brand's own traffic after hour 48
TIMELINE · Foundations first, data model second, feeds third, proof last. Nothing touched a storefront until hour 36.

Hour 0–6: foundations

We stopped the app, took a compressed dump, restored it onto managed Postgres with --no-owner --no-acl, and compared row counts table by table before anything else moved; the dump stayed on disk as the rollback. Then the push tool left the deploy path for good. In its place: hand-written, forward-only, additive SQL files numbered 0001 to 0004, run by a small runner that records the date each one was applied.

-- 0003_reconcile_runs.sql · additive, re-runnable, never a DROP
CREATE TABLE IF NOT EXISTS reconcile_runs (
  id         SERIAL PRIMARY KEY,
  ran_at     TIMESTAMP NOT NULL DEFAULT now(),
  country    TEXT NOT NULL,
  mismatched INT NOT NULL DEFAULT 0,
  error      TEXT
);
CREATE INDEX IF NOT EXISTS idx_runs_country
  ON reconcile_runs (country, ran_at DESC);

Every statement is safe to run twice, so a migration that half-applied during a deploy can simply be re-run. There is no rename prompt because there are no renames: a rename is a new column, a backfill and a later drop, in three separate files.

The same six hours closed the prototype residue. The app now refuses to start with a placeholder session secret. Sessions and bearer tokens live in the database, cookies are secure and httpOnly, the committed credentials were rotated, and the connection pool got sane limits: ten connections, thirty-second idle timeout, each connection recycled after 7,500 uses, with retries only on connection-class errors.

Hour 6–18: the data model

Fashion stock has three levels and the prototype had one and a half. We made it explicit: model, then model-colour as the primary key, then model-colour-size. Every feed, every storefront push and every report now speaks the same three keys.

Then the cleanup scripts, all dry-run by default. Some size SKUs carried _ where the shops expected -, because the feed's internal id uses underscores; fixing them in place collided with correct rows that already existed (23505, unique violation), so on collision the script deletes the bad row and its country rows, in batches of 200, printing an ETA.

Sizes that differed only in spelling, "ONE SIZE" against "ONESIZE", were merged; that alone removed 546 phantom units from sale. Lowercase country codes were folded into uppercase at startup. Barcodes were separated from the supplier's counter by GS1 prefix, each in its own column, and incoming || stored left every identifier path.

Two smaller things earned their own helpers. A static VAT table per country returns null for an unknown country rather than guessing. And every date-time in the codebase goes through one fixed-timezone function, after we found a promotion that had started two hours early because a zone-less timestamp was read as UTC.

Hour 18–36: ingest and outbox

This is where the prototype and the product part ways. The Kafka consumer now writes raw events to the database first and only then commits the offset. If the insert fails, nothing is committed and the broker redelivers. The client retries eight times with exponential backoff capped at 30 seconds, sends a heartbeat every 100 messages and warns after two minutes of silence.

eachBatch: async ({ batch, resolveOffset, commitOffsetsIfNecessary }) => {
  try { await saveRawEvents(batch.messages); }   // durable first
  catch { return; }                              // no commit → redelivery
  await applyChanges(batch.messages);
  resolveOffset(batch.messages.at(-1)!.offset);
  await commitOffsetsIfNecessary();
}

Order events became immutable rows with a unique key on (topic, partition, offset), so a redelivery is ignored. The order aggregate is upserted by order number and only updated when the incoming timestamp is newer, so an event that arrives late cannot overwrite a newer one. A message that fails validation is saved with status='error' and the consumer keeps going. Validation asks for two required fields and passes everything else through, because a strict schema on someone else's feed is a self-inflicted outage.

The FTP importer got a processed-files table so only new or failed files are touched, a running flag so overlapping runs skip, and a boot step that marks anything still "processing" as failed with the reason "server restarted during processing". Bad XML returns a typed error; invalid products are skipped with a warning and a count, not a crash.

Outbound, every stock change to a storefront is now a row in a durable outbox. The row holds only a key; the current figures are read from the database at send time, so order does not matter and a burst of changes collapses into one send. The outbox ticks every two seconds, 50 rows per country. Transient errors (502, 504, ECONNRESET) are retried; a 400 or a not-found fails at once.

// 5 s doubling to 5 min for 8 attempts, then 3 hourly, then dead-letter
const backoffMs = (n: number) =>
  n > 8 ? 3_600_000 : Math.min(5000 * 2 ** (n - 1), 300_000);

if (isPermanent(err) || row.attempts + 1 >= 11) markDead(row);
else reschedule(row, Date.now() + backoffMs(row.attempts + 1));

Eleven attempts over about three and a quarter hours, then a dead-letter row a person can see and replay. The update is never dropped on the floor; the worst case is that it waits.

From our desk

The three hourly attempts were not in the first version. On the ninth night after go-live a storefront's proxy returned 502 for most of the night, longer than the fast retries lasted, and 41 rows went to dead-letter. The morning shift replayed them in one click; nothing was lost, but a person had been needed, so we added the hourly tail the same day. The literal HTML of that 502 page is now a test fixture.

Hour 36–48: reconciliation, tests, gate

Reconciliation reads stock back from each storefront four times a day and compares it to the catalogue on the same key the sync uses, not on EAN. Exactly one class of mismatch is fixed automatically: the shop is behind and the catalogue has a newer, complete figure. Every other class goes into a report with a proposed action and waits for a person. That is the difference between a check and a script that takes 11,778 units off sale.

The tests were written from the incident list: 40 files, 717 cases across API, services and client, plus one browser end-to-end run. SKU normalisation is tested against non-breaking spaces and typographic dashes from spreadsheet exports; retry classification against the real 502 page. The order tests check that an older event never overwrites newer data and that a duplicate offset is ignored.

it('an older event never overwrites newer order data', async () => {
  await ingest(orderEvent({ id: 'A', total: 90, at: '2026-03-01T10:00Z' }));
  await ingest(orderEvent({ id: 'A', total: 40, at: '2026-03-01T09:00Z' })); // late
  expect((await orders.get('A')).total).toBe(90);
});

it('a redelivered offset is ignored', async () => {
  const e = rawEvent({ topic: 'orders', partition: 0, offset: 17 });
  await ingest(e); await ingest(e);
  expect(await events.count()).toBe(1);
});

Each test is one incident that cannot recur; that is the whole strategy, not coverage but a list of things that already went wrong once.

Then the gate. Three independent reviewers score every change from 0 to 5 on architecture, security and functional completeness; any score under 3 is a rejection and each rejection finding becomes a mandatory task. /api/health returns 503 when the database is unhealthy, and the Kafka consumer and the outbox lane sit behind environment flags, so they were switched on one at a time without a redeploy. Readiness at the gate: 4.1.

What it looks like now

The same codebase, with the same name, now runs product data for one of Europe's largest fashion e-commerce brands: 18,162 product models, 42,024 colour variants, 311,497 size variants across 62 tables and a 222 MB dump, six countries with a VAT rate for each, seven languages. In one two-week window the stock consumer took in 725,132 events; 713,509 carried a real barcode, and the 11,623 that did not were stored, flagged and never applied. Peak season went through it with zero lost updates.

AFTER · in production · readiness 4.1 Supplier XML FTPS hourly, seen-files Stock events Kafka, 725k / 2 weeks Validated ingest store raw, then apply late or duplicate: skip bad rows flagged, held Postgres forward-only migrations model → colour → size 62 tables, 311k sizes Durable outbox key only, read at send retry ×8, then 3 hourly 2 s tick, 50 per country after 11 attempts Dead-letter a person replays 502: retry · 400: fail Storefronts × 6 6 countries, 6 VAT rates 0 lost updates since Reconciliation reads shops back 4×/day matches on the sync key one class auto-fixed A person all other mismatches gets a report, decides Operations health 503 on bad DB 717 tests, env flags
AFTER · Store first, send from an outbox, read it back, and hand the ambiguous cases to a human. Six storefronts, zero lost updates.
At DiscoveryNow
Schema changesInteractive push; a wrong answer drops a live tableNumbered forward-only SQL, re-runnable, dated log
Stock push fails with 502Update lost; 13 in one dayRetried 11 times over 3 h, then dead-letter; 0 lost
Duplicate or late eventOverwrote newer dataIgnored by (topic, partition, offset); upsert only if newer
Identifier from a feedincoming || stored; 193 barcodes overwrittenSeparated by GS1 prefix; identifiers never overwritten
Catalogue vs storefrontsNever compared; 78,482 units of driftRead back 4× a day; one class auto-fixed, rest to a person
Import runsOverlapping; stuck "processing" after a restartSeen-files table, run guard, boot cleanup
Sessions and secretsMemory, Math.random, placeholder secretDatabase-backed, fail-closed config, secure cookies
TestsNone that touched a feed717 cases in 40 files, each one a past incident
I built the first version myself in a weekend. I did not build the version that runs the business. Turns out those are two different products with the same name.Head of e-commerce operations

Those two products still share a repository and most of their screens. What changed is everything underneath, and the ops lead who built the first one now reads a mismatch report over coffee instead of a spreadsheet at midnight.


Do this tonight

Do this tonight

1. Find out how your schema reaches production. Run grep -rn "db push\|drizzle-kit push\|synchronize: true\|create_all" package.json scripts deploy README.md. If it hits anything a deploy runs, your tables are one prompt away from gone. Take a pg_dump -Fc right now and put it somewhere that is not the same server.

2. Break a downstream on purpose. Point the storefront URL in your env at a port with nothing on it, http://127.0.0.1:9, change one product's stock, wait one minute, put the URL back. Does the change arrive, or is it gone? If your logs cannot tell you, that is also your answer.

3. Sample fifty SKUs. SELECT sku, stock FROM variants ORDER BY random() LIMIT 50, then look up the same fifty in the shop's admin. Any difference is drift you did not know you had, and it only grows.

These three tell you whether you have the problem. They do not give you the store-first consumer, the outbox, the migration runner or the reconciliation that knows which mismatches it may fix alone. That is the 48 hours, and it is what prodready does.

The rule

The rule

A feed is not an input, it is a stranger with a schedule. Store it before you trust it, retry it before you lose it, and read it back before you believe it.

I built the first version myself in a weekend. I did not build the version that runs the business. Turns out those are two different products with the same name.— Head of e-commerce operations

Send us what you have.

The readiness scan is free — install our MCP and your own coding agent runs it. When you want it fixed, send us what you have: fixed scope, fixed price, agreed before the clock starts. 48 hours later, it's production-ready — on your own cloud, with the tests to prove it.

Feed the machineWatch the 49-second film