Sound familiar?
You built the app with Claude Code over a few weekends. Offices, sellers, leads, approvals: every screen works, every click does something. Investors saw the demo and asked for a second meeting. Then a real client asked a boring question: where is the data backed up? And you realised you had never checked what the browser actually sends to the server. You have 36 tests. They pass. You are not sure what they test.
What they sent us
The founder of a field-sales management startup in Italy sent us a repository and a deck. The app manages offices, outside sellers, daily leads, holds and manager approvals. He had built it with Claude Code in a few months, alone. The first serious client, an Italian company with sellers in several regions, wanted a rollout date.
Our discovery scan covers 19 dimensions and scores each from 1 to 5. This app scored 2.0, "developing". 125 findings: 4 critical, 32 high, 71 medium, 18 low. Thirty-six tests in total.
The four critical findings were one finding in four outfits. The browser held the data. Offices, sellers, leads, holds and approvals were seeded JavaScript arrays; the daily lead list came from a demo template. Of the whole API, the UI called two routes: login and bootstrap. Nothing a user did was saved anywhere.
The founder believed otherwise because a "UI regression gate" reported 155/155 functions preserved on every run. That gate stubbed window.fetch. It proved the buttons existed. It could not see that none of them talked to the server.
Below the UI it did not improve. The schema was created by create_all at start-up. Both migration "revisions" were snapshots of that same call, so a migration check would have dropped the other tree's tables. The pytest configuration dropped every table on whatever DATABASE_URL it found; run the suite with the staging URL exported and 13 tables go. Config failed open when an environment variable was missing. There was no way to create the first admin account in production mode.
Then the load test. With small data the API answered in 91.5 ms at p95 and took 205 requests a second. With 50,000 leads, a plausible year for the client, p95 went to 15.8 s, throughput to 2.55 req/s, and memory from 220 MB to 842 MB. The list and export endpoints had no pagination and ran on one synchronous worker.
The rest of the list is the usual one. No CI, no lockfile, no logs, no readiness probe, no metrics, no release process. Backups were manual. The business engine existed in three copies that nobody enforced, and one rule had two implementations that disagreed. A new profile returned HTTP 500. All 16 acceptance rows in the client's contract were marked NOT RUN. Actor identity came from request headers; the SSRF guard checked IP literals only; the PDF parser carried 41 advisories.
What would have happened
What actually happens
Day 1 of the pilot: twelve sellers in three offices log their morning visits, refresh, and see the demo template again. The account manager calls it a glitch. Day 3: a contractor exports DATABASE_URL on the staging box "to test it properly" and runs the suite; 13 tables are dropped before lunch. Day 12: the client's IT lead asks for the backup policy and an export of everything entered so far. There is no export because there was never anything to export. Day 20: the client's legal team opens the contract at the acceptance rows, 16 of them, none run. That pilot carried the company's whole first-year revenue plan. The investor update that month says "technical delays".
I thought I actually had a finished product. I had a very convincing screenshot of one.Founder & CEO
Forty-eight hours
Discovery produced 64 stories in six sprints. Every story took the same route: built in its own branch, verified by a separate person against every suite and benchmark, then reviewed by three independent reviewers scoring architecture, security and functional completeness from 0 to 5. Any score under 3 is a rejection. Over the 48 hours that meant 67 approvals and 31 rejections, and every rejection became a mandatory task before the story could close.
Hour 0–6: make it safe to touch
Nothing else could start until the test suite was physically unable to destroy a database. Tests may now touch only SQLite or a throwaway local Postgres whose name ends in _test, and hosted database hostnames sit on a denylist. The guard lives in conftest.py and runs before any fixture.
# conftest.py: refuse anything that is not disposable
url = os.getenv("TEST_DATABASE_URL") or f"sqlite:///{tempfile.mkstemp(suffix='.db')[1]}"
p = urlparse(url)
if url.startswith("postgres") and (
p.hostname not in {"localhost", "127.0.0.1"} or not p.path.endswith("_test")
):
raise SystemExit("refusing non-disposable test database")
os.environ["DATABASE_URL"] = url # force it; never setdefault
Eight lines. Note the last one: the suite overwrites DATABASE_URL instead of respecting whatever is exported. The original used setdefault, which is how staging was one shell away from deletion.
Then config went fail-closed. No default environment. Secrets are checked for length and for the number of distinct characters, and the app refuses placeholder values. In production and staging the database must be Postgres.
ENV = os.getenv("APP_ENV", "").strip().lower() # no default
if ENV not in {"production", "staging", "development", "test"}:
sys.exit("APP_ENV invalid")
strict = ENV in {"production", "staging"}
for name in ("SECRET_KEY", "SERVICE_API_KEY"):
v = os.getenv(name, "")
if not v or v.startswith("replace_with_") or (strict and (len(v) < 32 or len(set(v)) < 8)):
sys.exit(f"{name} missing/weak") # names only, never values
if strict and not os.getenv("DATABASE_URL", "").startswith("postgresql"):
sys.exit("Postgres required")
The rest of the first six hours: an audited command to create the first admin, migrations made the only owner of the schema with a frozen explicit-DDL baseline, create_all refused outside development and test, and migrations moved into a pre-deploy job. A /ready endpoint that checks the database, the migration heads and a worker heartbeat. JSON logs with request IDs. Backup and restore-drill scripts.
Hour 6–18: wire the UI to the API
Every screen was connected to the API one at a time, and the fetch stub was replaced with a Playwright journey per role: seller, manager, admin. Seed data went behind a demo flag that production refuses to start with. List endpoints got keyset pagination and the export became a stream.
This is the sprint where 15.8 s became 25 ms. Nothing clever happened. The endpoint stopped loading 50,000 rows into memory to return 50 of them.
Hour 18–36: the database learns the rules
A background worker using SELECT … FOR UPDATE SKIP LOCKED took over the work that used to run inside requests. Eight runtime migrations added 12 CHECK constraints and 42 foreign keys. Each foreign key went in as NOT VALID and was validated only after an orphan-row query returned zero; on a schema that already has rows you cannot assume the data agrees with your model. Money became NUMERIC(14,2). An append-only audit table got a trigger that rejects updates and deletes.
Then security. The SSRF guard now pins the DNS answer it checked to the connection it opens. A signed actor token replaced identity taken from headers. Login lockout, token revocation, spend caps on the paid providers. The security dimension moved from 2.4 to 4.0.
Hour 36–48: prove it and ship it
The CI file ended at 766 lines: hash-locked installs, digest-pinned images, an SBOM, a dependency audit, a migration check and the accessibility gate at 375 and 1440 px. The three lines that matter most are short.
# ci.yml (excerpt)
- run: pip install --require-hashes -r requirements.lock
- run: alembic upgrade head && alembic check
- run: pip-audit -r requirements.lock
Then the load test. k6, seeded with 50,000 leads, 30 s warm-up excluded, then 300 s at a constant 20 requests a second with 20 to 50 virtual users. Tokens come from setup() so no virtual user logs in on the hot path. The thresholds are the interesting part.
thresholds: {
"http_req_duration{gated:yes}": ["p(95)<500"],
"http_req_failed{gated:yes}": ["rate<0.01"],
functional_failures: ["count<1"] // a 200 that did no work
},
scenarios: { steady: { executor: "constant-arrival-rate", rate: 20, timeUnit: "1s",
duration: "300s", preAllocatedVUs: 20, maxVUs: 50 } }
functional_failures is a custom counter. A request that returns 200 but did not create the lead, did not move the hold, did not write the audit row, counts as a failure. A reviewer rejected the first version of this gate because it only failed on exit code. Result: 9,171 requests, p95 25.46 ms, p99 40 ms, zero errors, zero lockouts.
The deploy went to the client's own Debian host, not ours. Docker Compose with Postgres 16, a migrate job that must finish before the web containers start, two web workers that are not published to the network, a background worker, and a proxy with auto-renewing TLS certificates. Secrets are generated on the server with mode 600. We verified /health and /ready return 200, the API docs return 404, CSP is set, and the first login forces a password change.
From our desk
The 31 rejections read like a checklist we now carry to every job. Secrets need a minimum number of distinct characters, not just a length. The database check must be an allow-list (Postgres only), not a deny-list of SQLite. The login throttle had a check-then-act race and became reserve-then-check. A password passed in argv became PGPASSWORD. Unbounded concurrent exports got a cap that returns 429. An unsupported query parameter must be a 422 before any query runs. Raw exception text stopped reaching clients. And crond silently does nothing as a non-root user, so the scheduler became an in-process loop.
What it looks like now
The compose file encodes the order: web depends on migrate with condition: service_completed_successfully, and the proxy will not route to a web container until its /ready health check passes.
Readiness went from 2.0 to 3.4. Findings went from 4 critical, 32 high, 71 medium, 18 low to 0, 12, 53 and 49; the medium and low counts include the things we now know about and did not before. No dimension regressed. Testing, data, observability, dependencies, scalability and compliance each gained two full points.
| Test layer | Before | After |
|---|---|---|
| Runtime: API, config, migrations, worker | 36 tests | 2,882 tests |
| Business engine, frozen benchmarks | 3 unenforced copies, 0 tests | 591 tests, run twice (10/10 and 14/14 benchmarks) |
| End-to-end, per role | 155 "functions preserved" with fetch stubbed | 185 Playwright journeys against the real API |
| Load, 50k leads | none (p95 15.8 s when we ran it) | k6 gate: p95 < 500 ms, errors < 1 %, functional failures 0 (p95 25 ms) |
| Production start-up | fails open on missing config | fail-closed matrix: missing env, placeholder secrets, SQLite URL, demo seed, missing signing key all refused |
| Accessibility | none | axe-core gate at 375 and 1440 px |
| Acceptance rows | 16, all NOT RUN | 16, run by the verifier before every merge |
| Total | 36 | about 3,650 |
What we shipped switched off is as important as what we shipped. At hour 48, 62 of the 64 product-owner questions were still open, because we do not answer those for a founder. Every rule without an answer is built fail-closed or configurable. Twelve seller-flow handlers show "Not yet available". The duplicate-assignment guard is in the code with both of its constants set to None; setting only one stops the app from starting, and switching it on is a reviewed constant change plus an activation migration. Backup tooling exists but is not scheduled. The alert receiver is a placeholder. Manager views are capped at 20,000 leads. Each answer becomes a config change and a test flip, not a rebuild.
Do this tonight
Do this tonight
1. Watch the network. Open DevTools, Network tab, filter to Fetch/XHR. Use your app for two minutes: create a record, edit it, refresh the page. You should see a POST or PUT for each change, and the record should survive the refresh. If you see one login call and one bootstrap call and nothing else, your data lives in the browser.
2. Read what your tests are allowed to destroy. Run grep -rn "drop_all\|create_all\|window.fetch =" . in the repo. Open the test setup file and find which database URL it uses. If it reads DATABASE_URL from the environment and drops tables, anyone with the staging URL in their shell is one command from an empty database.
3. Start it wrong. Unset your main secret and start the app: env -u SECRET_KEY python -m app or the equivalent for your stack. If it comes up, config fails open. Then insert 50,000 rows with a loop and time your list endpoint with curl -s -o /dev/null -w '%{time_total}\n'. Anything over one second at 50k rows will be a minute at 500k.
Those three checks find the demo. They do not find the 42 missing foreign keys, the race in the login throttle, the export that falls over on the third concurrent caller, or whether the containers come up in the right order on a server you have never seen. That part is the 48 hours, with three reviewers who reject.
The rule
The rule
A demo proves the screens exist. Production-ready proves the data survives a refresh, a wrong config and a year of rows. Until something in your stack refuses to start when it is misconfigured, you have a screenshot that works.
The point isn't vibe coding bad; the last mile is a different job. The founder built a product one person could not have built two years ago. prodready took the last mile in 48 hours, on the client's own server, with the tests to show for it.
I thought I actually had a finished product. I had a very convincing screenshot of one.— Founder & CEO