Sound familiar?
You built the customer data platform yourself because the agency quote had six digits and Claude Code had a working version in a fortnight. Imports work. Segments work. The nightly job runs, the emails go out, the board is pleased. Then the profile count went from 4,000 to 560,000 and everything started taking longer. The nightly job is at 45 minutes and climbing. Last week a segment came back with people who had unsubscribed. Somewhere in the config there is a line that reads process.env.API_KEY || 'default_key' and you are not sure it is the only one. You have not had a double send yet. You also cannot prove you will not.
What they sent us
The head of CRM at a retail group with a loyalty programme and a newsletter sent us a repository, a read-only connection string and a one-line brief: "We send 50,000 emails a day off this and I would like to sleep." Discovery took two days. Then the clock started.
The numbers at Discovery: 560,000 customer profiles, about half a million from the loyalty programme and 60,000 newsletter subscribers, all in one collection under one schema. 129 indexes on that collection, 70-odd compound and the rest single-field. When we asked which query each index served, the answer for roughly forty of them was "the performance advisor suggested it".
Four scheduled jobs, all UTC. Profile cache at 03:00, about 20 minutes. Loyalty profile generation at 05:00, 30 to 45 minutes depending on the night. Segment snapshots at 05:45, 15 minutes. Newsletter regeneration every five days. The snapshot job started at 05:45 whether or not generation had finished, so on a slow night it read yesterday's profiles and called them today's segments. Each job's trigger endpoint called the service directly with no check for a run already in progress. The diagnostic script meant to catch that queried status: 'running' on a model whose value was 'processing', so it always reported "no running jobs". Comforting, and wrong.
The campaign sender worked in batches of 500 with a one-second pause. When the email provider answered 429 Too Many Requests the job stopped. The resume path was a TODO comment that said, honestly, "will start from beginning".
Then the config. The ingest API key fell back to the string 'default_key' when the environment variable was missing. An isAuthenticated middleware was imported and never applied to a single route. The webhook signature check from the email provider was commented out with the word "temporarily". The internal-access check treated a missing Origin header as same-origin, so anything sent from a terminal walked in. The request logger wrote response bodies, including customer emails, to plain console.log. In the project root sat a sample JSON of real customers, two CSV exports and a SQL dump with password hashes. The .gitignore covered .env and nothing else.
Readiness score at Discovery: 2.1.
What would have happened
What actually happens
On day 12 the autumn campaign goes to 480,000 loyalty profiles. At batch 300 the provider returns 429. The job dies. The dashboard says "no running jobs". Someone presses Send again and it starts from the beginning. 150,000 people receive the same email twice before anyone notices. The segment it went to was snapshotted at 05:45 from a profile set that finished generating at 05:52, so it includes 1,100 people who unsubscribed the day before. On day 13 the complaints start. On day 19 one of them goes to the data protection authority instead of to support, with a screenshot of both emails. The letter asks for the consent record for that address, with a timestamp. The record is a boolean. Nobody remembers who set it or when. That is not a bug ticket. It is a case number.
Forty-eight hours
Hour 0–6. Lock the doors.
Nothing about data models matters while the ingest endpoint accepts 'default_key'. We rotated the key and made every secret fail closed: if the variable is unset, the server refuses the request with a 500 and a log line, instead of quietly accepting whatever a stranger sends.
// Fail closed on a missing secret
const key = process.env.INGEST_API_KEY;
if (!key) {
return res.status(500).json({ error: 'Server misconfigured' });
}
if (req.header('x-api-key') !== key) {
return res.status(401).end();
}
Three lines, and the difference between "misconfigured" and "open". The same six hours: the isAuthenticated middleware went onto every route except /api/health and the tracking pixel, the webhook signature check came back on, and a missing Origin header stopped meaning "trusted". The customer files left the repository, the history was rewritten, and .gitignore learned about .csv, .sql and data dumps.
Hour 6–18. Unique indexes are the idempotency layer.
We split the one collection into two with the same schema, loyalty and newsletter, because their query patterns differ and one index set for both was where the 129 had come from. In front of both sits a dedup layer with one rule: if an email exists in both, loyalty wins and the newsletter record is not merged into it. An opt-out collection with a unique, lowercased email is checked at send time by every campaign, no exceptions.
Then the rule we apply to every prototype that ingests anything: if a row must exist once, the database must enforce it, not the code. Orders got a unique index on {orderId, source}. Campaign recipients got {campaign_id, email}. Webhook events from the email provider got a unique index on the provider's event id, so a redelivered webhook is a no-op rather than a second "opened" record. Workflow runs got the interesting one:
// One active run per person, case-insensitive
RunSchema.index(
{ flowId: 1, email: 1 },
{ unique: true,
partialFilterExpression: { status: 'active' },
collation: { locale: 'en', strength: 2 } }
);
The partial filter means a person can have a hundred finished runs but only one active one. The collation means Anna@ and anna@ are the same person, which the code previously assumed and the database previously did not. The job trigger got the same treatment: instead of "check if running, then start", one atomic upsert that either claims the slot or tells you it was taken.
// Atomic "only one running job" guard
const job = await Job.findOneAndUpdate(
{ type, status: { $in: ['pending', 'processing'] } },
{ $setOnInsert: { type, status: 'processing', startedAt: new Date() } },
{ upsert: true, new: true, rawResult: true }
);
if (job.lastErrorObject?.updatedExisting) {
return res.status(409).json({ error: 'Already running' });
}
Two people pressing the button in the same second now get one job and one 409. The diagnostic script now queries the field names the model actually uses.
Hour 18–36. Jobs and the sender.
The segment snapshot no longer starts at 05:45. It starts when profile generation reports done, and if generation fails it does not run at all, because a segment built from half a profile set is worse than yesterday's segment. Every job now has a status machine (pending, processing, paused, cancelled), a progress figure, a log array and a checkpoint. "Will start from beginning" became "resumes from the last checkpoint". Routes refuse to delete a job that is active.
The campaign sender got three changes. The recipient list and its count are frozen the moment sending starts, so a profile that changes mid-campaign cannot fall in or out of the run. The sender records last_processed_offset after every batch of 500. And a 429 from the provider pauses the campaign rather than killing it; a person resumes it from the offset. If someone resumes from the wrong offset anyway, the unique index on recipients turns the second insert into error code 11000 and the row is skipped. That is what "zero double sends" rests on: not the offset, the index.
The tracking pixel was the other volume path. It had been writing cart events straight into the database on every request. Now the endpoint writes to a queue, returns 202 Accepted in a few milliseconds, and a worker drains the queue every five seconds in batches of 500, with an attempts counter, a retry endpoint and a cleanup of anything older than seven days.
// Queue fast, process in batches
await Queue.create({ type: 'cart', payload, status: 'pending', attempts: 0 });
res.status(202).json({ queued: true });
// worker, every 5 s
const batch = await Queue.find({ status: 'pending' })
.sort({ createdAt: 1 })
.limit(500);
Anything that talks to an external service got a retry with exponential backoff capped at ten seconds. The workflow exit grace period, documented as 10 minutes in one file and 60 in another, became one number, written down once.
Hour 36–48. Consent and proof.
The old model had consent_email, consent_sms and request_delete as flat booleans. A boolean answers "is it true"; a regulator asks "since when, on what basis, from where". Consent is now an array of records: type, legal basis (terms acceptance, contract, legitimate interest, unsubscribe), status, the sources it came from and a timestamp. The old booleans still exist as derived, read-only mirrors so nothing that reads them breaks, but the record is the truth. Newsletter subscribers carry a double opt-in flag. request_delete stopped being a flag someone might notice and became a job with a log.
Then the tidy-up that separates a demo from a deployment. Response bodies no longer go to the logs; the pixel logs an event id, not an email. CORS reflects an allowlist rather than any origin with credentials. The global error handler stopped sending a response and then re-throwing. The in-memory profile map is capped at 100,000 entries with a 24-hour cache instead of trying to hold half a million users in RAM. And we went through all 129 indexes and wrote next to each one the query it serves.
The last four hours were a load test against a copy of the full 560,000 profiles: the nightly chain end to end, a 200,000-recipient campaign with a forced 429 at batch 120, resumed twice from the wrong offset on purpose. Recipients emailed once: 200,000. Twice: none. Readiness re-scored at 3.9.
We had a platform that could email half a million people and no one could tell me, on paper, that it would never email one of them twice. Now I can.
From our desk
The comment above the disabled webhook signature check said "temporarily". We asked git when temporarily had started. Four months. In that time the platform had accepted every "delivered", "opened" and "bounced" event anyone cared to post to it, and the head of CRM's dashboards were built on those numbers.
What it looks like now
The platform still sends 50,000-plus emails a day. The sender's ceiling at 500 recipients a second is about 100 seconds of that, so the daily volume is not what limits it. What changed is what can be proved.
| Before | After | |
|---|---|---|
| Profiles | 560,000 in one collection, duplicates possible across programmes | 500k loyalty + 60k newsletter, dedup layer, loyalty wins |
| Indexes | 129, about 40 with no known query | 129, each annotated with the query it serves |
| Double sends | Prevented by an offset in memory and a TODO | Prevented by a unique index on {campaign_id, email} |
| Nightly chain | Fixed times; segments could read yesterday's data | Snapshot runs on completion; one running job per type, enforced by upsert |
| Provider 429 | Job dies; restart begins at recipient 0 | Campaign pauses; resumes from offset; index catches any overlap |
| Consent | Three booleans | Records with type, legal basis, sources and timestamp; booleans derived |
| Ingest key | Falls back to 'default_key' | Fails closed; auth middleware on every route |
| Customer data in repo | JSON sample, CSV exports, SQL dump with hashes | Gone, history rewritten, ignored by pattern |
Double sends since go-live: zero. Not because nobody has pressed Send twice. Because the database will not let it matter.
Do this tonight
Do this tonight
- Find every secret with a default. In your repo root run
grep -rnE "process\.env\.[A-Z_]+ *(\|\||\?\?) *['\"]" src/. Every hit is a place your app works without the secret, which means it works for anyone. Each one should become "unset means refuse". - Send yourself the same webhook twice. Take one real payload from your email or payment provider and post it to your endpoint two times with
curl. Then count the rows. If you have two, the fix is a unique index on the provider's event id, and it takes one line. - Look for customer data in git. Run
git ls-files | grep -iE "\.(csv|sql|xlsx|json)$"and open anything that is not config. If you find a real email address, it is in every clone of the repo and in every laptop that ever pulled it.
These three take half an hour and will each find something. What they will not tell you is whether your nightly job can run twice at once, whether your segments were built from yesterday's data, or what your sender does at batch 300 when the provider says 429. That is a different evening, and it is the one prodready sells.
The rule
The rule
A platform that can email half a million people must be able to prove, on paper, that it will email each of them once. If the proof is a unique index, you have one. If the proof is a person remembering to check, you do not.
We had a platform that could email half a million people and no one could tell me, on paper, that it would never email one of them twice. Now I can.— Head of CRM