Sound familiar?
You built it with Claude Code. It runs on your laptop with a .env file that has every key in it. The repo has a .env.example full of replace_with_ placeholders, and a config.ts that says process.env.X || 'something sensible' forty times. You deployed a Docker image, the container came up, the homepage loaded, and you posted the launch. You have error monitoring. It says zero errors. You have not, at any point, opened a shell in the production container and typed printenv. You are fairly sure you don't need to.
What they sent us
Two co-founders, pre-seed, a marketplace for workshop equipment: used lathes, CNC spindles, dust extraction, sold by small shops to other small shops. Built with Claude Code in about seven weeks. Roughly 14,000 lines of TypeScript, a Node API, Postgres in the local compose file, a checkout backed by a payments provider, a seller payout flow and a tax-invoice generator. On launch day it worked. Orders came in from the first hour.
Eleven days later a seller emailed to ask when the payout for a $1,900 spindle would arrive. The founders opened the payments dashboard. It was empty. They flipped the dashboard into test mode. There were the 92 orders, every one of them a sandbox transaction, every one fulfilled, every one $0.
The cause took us about forty minutes to find at Discovery and about three seconds to understand. The production container had no environment variables set at all. Not APP_ENV, not the payment key, not the database URL, not the invoicing flag. And the app did not care, because every setting had a fallback. APP_ENV || 'development'. PAYMENT_SECRET_KEY || 'sk_sandbox_…', the sandbox key hard-coded in the file as its own default. INVOICING_ENABLED === 'true', so unset meant off. DATABASE_URL || 'sqlite:./dev.db', so every order the marketplace had ever taken lived in a SQLite file inside the container's writable layer, one --force-recreate away from not existing.
The demo seed ran on every start too. Twelve sample sellers and forty sample listings, which the founders had deleted by hand on day one. They came back after the first restart on day four. Nobody connected the two events.
Our Discovery scan covered 18 dimensions and produced 3 critical, 21 high, 44 medium and 17 low findings. Readiness 2.2 out of 5, "developing". The three criticals were the same fact seen from three sides: config fails open, secrets in source, schema created at start-up with no migration owner. There were 36 tests. Not one set an environment variable. The app was not broken. It was configured, by nobody, to be a demo.
We had monitoring for errors. There were no errors. The app was doing exactly what it was configured to do.Co-founder
What would have happened
What actually happens
On day 12 the accountant asks for the invoices behind 92 sales. There are none, because the flag that issues them was never set. On day 14 seller payouts fall due: the marketplace owes its sellers roughly $26,000 for goods already shipped, out of a pre-seed account that collected nothing. On day 20 the payments provider's risk team notices sandbox transactions coming from a live checkout on a public domain, which their terms forbid, and freezes the live account before it has processed a single dollar. On day 30 the quarter closes with 92 sales, no tax collected and no invoices, which is a filing problem, not a bug. Somewhere in there the founders email 92 customers to ask, politely, whether they would mind paying now for the machine they already own. About a third do not reply.
Forty-eight hours
Hour 0–6. We did not start with the payment key. We started with a grep. Every || or ?? in front of a string literal after process.env is a place where production can quietly become something else, and the count tells you how many decisions were never made.
# every env read that falls back to a literal (Node and Python forms)
grep -rnE "process\.env\.[A-Z0-9_]+\s*(\|\||\?\?)\s*['\"]" src/ | grep -v "\.test\." | wc -l
grep -rnE "os\.getenv\(\s*['\"][A-Z0-9_]+['\"]\s*," src/ | wc -l
The first command returned 47. Nine of them were load-bearing: environment, database URL, session secret, payment key, webhook signing key, invoicing flag, demo seed, base URL and the outbound email key. Three of the nine had a real secret as the fallback, committed to the repo. We replaced the whole file with a loader that has no defaults in production and refuses to start rather than guess.
const ENV = (process.env.APP_ENV ?? '').trim().toLowerCase(); // no default
if (!['production','staging','development','test'].includes(ENV)) fail('APP_ENV invalid');
const strict = ENV === 'production' || ENV === 'staging';
for (const name of ['SESSION_SECRET','PAYMENT_SECRET_KEY','WEBHOOK_SIGNING_KEY']) {
const v = process.env[name] ?? '';
const weak = v.length < 32 || new Set(v).size < 8;
if (!v || v.startsWith('replace_with_') || (strict && weak)) fail(`${name} missing/weak`);
}
if (strict && !process.env.DATABASE_URL?.startsWith('postgresql')) fail('Postgres required');
if (strict && process.env.PAYMENT_SECRET_KEY?.includes('_test_')) fail('sandbox key in production');
if (strict && process.env.DEMO_SEED === '1') fail('demo seed in production');
function fail(msg: string): never { console.error(msg); process.exit(78); } // names only, never values
Read it top to bottom and it is the whole policy. The environment must be one of four known words. In production or staging, every secret must exist, must not be a placeholder, must be at least 32 characters with at least 8 distinct ones (a reviewer rejected our first version for accepting aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa). The database must be Postgres, checked as an allow-list, not a SQLite deny-list. A payment key with the provider's test-mode marker is refused. The demo seed is refused. Exit code 78 is the standard "configuration error" code, so the process manager and the alerting can tell it apart from a crash. The loader logs which variable failed, never its value.
Then we wrote the matrix: nine ways to misconfigure the app, each one a test that starts the process with that single fault and asserts it exits 78 before opening a port.
✓ refuses APP_ENV unset exit 78, no port opened
✓ refuses APP_ENV=prod (unknown value) exit 78
✓ refuses SESSION_SECRET unset exit 78
✓ refuses SESSION_SECRET=replace_with_secret exit 78
✓ refuses SESSION_SECRET=changeme-changeme-... exit 78 (weak)
✓ refuses DATABASE_URL=sqlite:./app.db exit 78
✓ refuses PAYMENT_SECRET_KEY with _test_ marker exit 78
✓ refuses DEMO_SEED=1 exit 78
✓ refuses WEBHOOK_SIGNING_KEY unset exit 78
9 passing (2.1s)
At hour 9 the app refused to start on the production host for the first time. That was the first correct thing it had done there.
Hour 6–18. Secrets. We deleted the three committed keys from source and from git history, revoked them at the provider, and generated the replacements on the server itself with openssl rand -base64 48, into a file owned by the deploy user with mode 600. Nothing crosses a laptop, nothing lands in a chat thread, nothing is in the repo. The compose file reads that file with env_file. The live payment key and the webhook signing key were entered by the founders, on the server, once.
Then the schema. The app had been calling the ORM's create-tables-at-start on every boot, which is how the SQLite file came to exist. We froze the current schema as an explicit baseline migration, made migrations the only thing allowed to touch the schema, and moved the 92 orders and every account out of the container's SQLite file into Postgres with a one-off script that we ran twice and diffed. Migration became its own compose service that must finish before the web service is allowed to start.
migrate:
command: ["node", "dist/migrate.js"]
depends_on: { db: { condition: service_healthy } }
restart: "no"
web:
depends_on: { migrate: { condition: service_completed_successfully } }
healthcheck:
test: ["CMD", "curl", "-fs", "http://localhost:3000/ready"]
interval: 10s
retries: 3
Two lines carry the weight. service_completed_successfully means web does not start if the migration job exits non-zero, so a broken migration produces no web process rather than a web process pointed at a half-migrated schema. The healthcheck hits /ready, not /, so a container that is up but not fit is marked unhealthy and the proxy stops routing to it.
Hour 18–36. /ready itself returns 200 only when three things are true: the database answers a query, the migration head in the database matches the head shipped in the image, and the background worker has written a heartbeat in the last 60 seconds. Anything else is a 503 with the failing check named. This is the difference between "the process is running" and "the process can take an order".
The invoicing flag went. Invoicing is not a feature; in the countries they sell to it is a legal obligation of every sale, so the code path is now unconditional and the flag survives only as a test-environment switch that the loader refuses in production. Same treatment for the demo seed: it exists, it runs only when APP_ENV is development or test, and production refuses it at start. We then ran a reconciliation over the 92 orders that produced the 92 missing tax invoices, dated to the actual order dates and flagged as late-issued for the accountant, plus a payment link per order. Sending them was the founders' call and they sent them the same evening.
Hour 36–48. The nine-case matrix went into CI and runs on every push, alongside 1,140 tests that now include a full checkout journey against the provider's test environment with the webhook signature verified. The uptime monitor was repointed from the homepage to /ready, so a refused start, a failed migration or a dead worker pages the founders within two minutes rather than being discovered by a seller. Deploy, verify: /health 200, /ready 200, API docs 404, security headers present. At hour 46 one of the founders bought a $1 item from the marketplace with his own card. It charged. He refunded it. The final gate scored readiness at 3.9.
What it looks like now
The app can no longer start in a state it should not be in. That is the whole result, and the four boxes on the right of the diagram are how it is enforced: the loader refuses, the migrate job must finish, /ready must pass, and the proxy sends nothing until it does.
| Before | After | |
|---|---|---|
| Missing setting in production | Silent fallback to a dev value | Exit 78 before a port opens, alert fires |
| Payment key | Sandbox key committed as the default | Live key on the server, mode 600; test-mode marker refused |
| Invoicing | Flag, unset means off | Unconditional in production |
| Database | SQLite file inside the container | Postgres only, allow-listed; migrate job before web |
| Demo data | Seeded on every start | Refused outside development and test |
| "Is it up?" | Homepage returns 200 | /ready checks DB, migration heads, worker heartbeat |
| Tests | 36, none touching config | 1,140, nine of them start-up refusals |
| Readiness | 2.2 | 3.9 |
From our desk
Our first loader was rejected by our own security reviewer, twice. First because a 32-character secret made of one repeated character passed, so the distinct-character rule went in. Then because the database check was a SQLite deny-list, which passes a MySQL URL, an empty string and a typo. It became a Postgres allow-list. Neither would have failed a test the founders could have written, because neither was a bug. They were policies that did not exist yet.
Do this tonight
Do this tonight
1. Count your defaults. Run the grep from Hour 0–6 on your repo. Write the number down. Then read each hit and ask what production does if that variable is missing. Anything where the answer is "uses the dev value" or "turns a feature off" is not a default. It is an unmade decision.
2. Look at what production actually has. docker compose exec web printenv | grep -E 'ENV|KEY|SECRET|DATABASE' | cut -d= -f1, or the equivalent on your host. You get names only, which is all you need. If the list is shorter than your .env.example, you are running on fallbacks right now. Then open your payments provider's dashboard and flip the test-mode toggle: your last ten orders should be on the live side.
3. Break it on purpose. In staging, unset one required variable and start the app. If it comes up, it fails open. If it comes up with a green health check, your monitoring cannot see this class of failure at all.
These three checks find the defaults. They do not tell you whether your migrations run before your web process, whether your worker is alive, which of your 47 hits matter, or what else is in the same file that has never been tested. That is the part prodready does in 48 hours.
The rule
The rule
An app that starts with a missing setting is guessing. In production, the guess is the bug. Make it refuse, make the refusal loud, and test every way it can refuse.
We had monitoring for errors. There were no errors. The app was doing exactly what it was configured to do.— Co-founder