Sound familiar?
You build apps for clients. The last one went well: a booking tool on Base44 for a business with several locations, delivered in a fortnight, and the client's staff actually use it. You have a maintenance retainer. What you don't have is an answer to a question nobody has asked yet: where, exactly, is the data. The prompt said "save appointments" and the app saves appointments, on screen, every time. You have not opened the generated backend since the demo. You would rather not be the person who finds out what's in it during a phone call with the client.
What they sent us
The customer was a two-person agency. They had built a booking and rota tool on Base44 for a chain of four outpatient clinics: receptionists book, practitioners see their day, the manager sees all four sites. For three weeks the client was happy, the staff were happy, and the agency invoiced the second milestone.
On day 21 a receptionist at the second clinic closed the tab. When she reopened it, the day view showed fourteen appointments with placeholder names, all dated the day of the original demo. Every booking made in three weeks was gone. The other three clinics were still fine, because nobody there had refreshed yet.
The agency owner opened the code to work out what had happened. That is the moment in the quote at the top of this page. He sent the repository to prodready the same evening; Discovery ran overnight.
Discovery scanned 19 dimensions and logged 118 findings: 5 critical, 28 high, 64 medium, 21 low. The critical ones were all the same finding in different clothes. The UI held patients, practitioners, rooms, appointments and the waitlist as JavaScript arrays in browser state, seeded from a file the tool had generated on day one. The API had two endpoints the UI ever called: login, and a bootstrap call that returned a demo template. The same JSON, every time, for every clinic.
A database existed. Thirteen tables, correctly named, with sensible columns. They were created by create_all every time the server started, and nothing ever wrote to them. Rows persisted across three weeks of use at four clinics: 0. Our estimate of what the staff believed they had booked in those three weeks: about 3,400 appointments.
The test setup was the part that made our reviewer put her coffee down. The pytest configuration dropped every table on whatever DATABASE_URL was exported in the shell. Point it at a hosted database by accident and it wipes 13 tables without asking. The two "migration revisions" in the repository were both snapshots of the same models; running a migration check would have tried to drop the other one's tables. The UI regression suite stubbed window.fetch, which is why its "all screens preserved" report had been green for the whole three weeks.
Readiness at Discovery: 1.4 out of 5. Data scored 1. So did testing.
What would have happened
What actually happens
Day 22: the laptop at clinic three freezes and gets restarted. Its bookings reset to the demo. Day 23: patients arrive for appointments that no longer exist at two sites; a practitioner is double-booked for the afternoon; the front desk goes back to a paper diary. Day 25: a patient complains about a missed procedure and the clinical director asks for the appointment record. There is none, and there never was. Day 30: the client's compliance lead asks how long booking records are retained. Medical appointment records must be kept for years in every European jurisdiction; the answer is "until someone presses F5". Day 45: the regulator's questionnaire arrives, and the agency's retainer, its reference, and its rebuild at its own cost all go on the same invoice. The four clinics run on paper for six weeks.
Forty-eight hours
Discovery had already told us the order. Nothing in the first six hours was about making the app better; it was about making it impossible to lose more.
Hour 0–6: stop anything from making it worse
First the test fence. Tests may touch a temporary SQLite file or a throwaway local Postgres whose name ends in _test, and nothing else. Any hosted hostname is refused before a single fixture runs. Then the schema: we froze the 13 tables into one explicit-DDL baseline migration, made create_all refuse to run outside development and test, and made the migration job the only thing allowed to own the schema.
# conftest.py, evaluated before any fixture
url = os.getenv("TEST_DATABASE_URL") or f"sqlite:///{tempfile.mkstemp(suffix='.db')[1]}"
p = urlparse(url)
if url.startswith("postgres") and (
p.hostname not in {"localhost", "127.0.0.1"}
or not p.path.endswith("_test")
):
raise SystemExit("refusing non-disposable test database")
os.environ["DATABASE_URL"] = url # force it; never setdefault
The guard is eight lines and you can paste it into your own conftest tonight. The last line matters: it overwrites the variable rather than filling it in when missing, so an exported staging URL in your shell can never leak into the suite.
"""0001_baseline: explicit DDL. Never imports live models. Forward-only."""
def upgrade():
op.execute("""
CREATE TABLE appointment (
id BIGSERIAL PRIMARY KEY,
clinic_id BIGINT NOT NULL,
patient_id BIGINT NOT NULL,
starts_at TIMESTAMPTZ NOT NULL,
deposit_eur NUMERIC(14,2) NOT NULL DEFAULT 0
)""")
def downgrade():
raise RuntimeError("baseline is irreversible; rollback = restore backup")
A baseline migration is the CREATE TABLE statements written out by hand, not generated from whatever the models say today. It never changes again. Rollback of a baseline is a restore from backup, and saying so in the file stops someone trying to be clever at 2 a.m.
Config went fail-closed in the same window: an unset environment, a placeholder secret, or a SQLite URL in production now stops the process from starting, with the variable name in the log and never the value.
Hour 6–18: make the app remember
We wired every screen to real endpoints. The day view, the booking form, the rota and the waitlist each got a read and a write route, and the bootstrap call now returns the clinic's own rows. The demo template survived, but only behind a DEMO_SEED flag that production refuses to start with. The day view got keyset pagination so a busy Monday at the largest clinic does not load four months of history.
The three weeks of bookings were gone and we did not pretend otherwise. The team re-keyed the coming fourteen days from paper diaries and calendar invites in one afternoon, while the four sites stayed on the demo build in read-only mode. By hour 18 there were Playwright journeys for each role: receptionist books, practitioner sees the day, manager sees all four sites, and every one of them asserts on a row in the database rather than on a screen.
Hour 18–36: make wrong data impossible
This is the part a tool does not do for you. Eight runtime migrations added 42 foreign keys and 12 CHECK constraints: an appointment must point at a real patient, a real practitioner, a real room in the same clinic; an end time must be after a start time; a deposit cannot be negative. Money moved from a JavaScript float to NUMERIC(14,2).
ALTER TABLE appointment
ADD CONSTRAINT fk_appointment_patient
FOREIGN KEY (patient_id) REFERENCES patient(id) NOT VALID;
-- count the orphans, fix them, and only then:
ALTER TABLE appointment VALIDATE CONSTRAINT fk_appointment_patient;
CREATE TRIGGER audit_no_rewrite
BEFORE UPDATE OR DELETE ON audit_log
FOR EACH ROW EXECUTE FUNCTION refuse_change();
Adding a constraint as NOT VALID means new rows are checked immediately while old rows are left alone, so the migration cannot fail halfway on a live table. You then count the orphan rows, repair them, and validate once the count is zero. We validated all 42 at zero orphans; on your app, expect to find some.
The second half of the snippet is the audit table. Every booking, move, cancellation and no-show writes a row that says who, what, when and from which clinic. The trigger refuses any UPDATE or DELETE on that table, for everyone, including the application's own database user. That is what "append-only" means when the compliance lead asks.
Hour 36–48: make loss recoverable
Persisting the data was half the job. The other half is knowing, on a bad morning, how old the newest copy is and whether it restores.
# /etc/cron.d/clinic-db (server time, root)
15 2 * * * app /srv/app/scripts/backup.sh # pg_dump -Fc, encrypted, 30 kept
45 2 * * * app /srv/app/scripts/restore_drill.sh # restore into scratch DB, count rows, alert
The first line takes the dump. The second restores it into a scratch database half an hour later, compares row counts with live and pages someone if they differ. A backup nobody has restored is a file, not a backup.
The same window added a /ready endpoint that checks the database, the migration head and the worker heartbeat; a CI pipeline that installs with hash-locked dependencies and runs alembic upgrade head && alembic check on every push; and a Compose deploy where the migrate job must finish before the web container is allowed to start. The suite went from 36 tests to 1,140, and every one of them runs against a database whose name ends in _test.
From our desk
The first version of the test fence was rejected by our own security reviewer. It was a deny-list: refuse SQLite in production, refuse a short list of hosted database hostnames in tests. The reviewer's note was one line: a deny-list is a list of the mistakes you have already thought of. The version that shipped is an allow-list. Tests get localhost and a name ending in _test; production gets Postgres; everything else fails to start. Thirty-one review rejections like that came back on this job; each became a task before its story could close.
What it looks like now
The client's four clinics have been on the production-ready build for five weeks. The app looks the same to the receptionists, which was the point. What changed is underneath.
| Question the client can ask | Before | After |
|---|---|---|
| Where is tomorrow's appointment list? | In whichever tab has not been closed | In Postgres, with an audit row for every change |
| What happens on refresh? | Reset to the demo day, 14 sample bookings | Nothing. The browser holds no state worth keeping |
| How is the schema changed? | create_all at every start-up; two snapshot "migrations" that would drop each other's tables | One explicit baseline plus 8 forward-only migrations, run as a pre-deploy job |
| Can a test delete production data? | Yes, on any exported DATABASE_URL, 13 tables | No. Tests refuse anything but localhost and a name ending in _test |
| Can a deposit be negative or a room double-booked? | Yes, floats and free text in an array | No. CHECK constraints and NUMERIC(14,2), enforced by the database |
| Who cancelled that appointment? | Unknown | Append-only audit row: who, what, when, from which clinic |
| How old is the newest backup, and does it restore? | No backup existed | Nightly, restore-tested at 02:45, row counts compared, someone paged if off |
| Readiness | 1.4 | 3.7 |
Two things stayed switched off, and the client knows it. Online self-booking by patients is built but disabled until the clinics decide the cancellation rules, and the SMS reminder route has no provider configured. Both ship as a config change plus a test flip when the answer arrives. That is what 3.7 means rather than 4.5.
Do this tonight
Do this tonight
1. The refresh test. Open your app, create something real, and note its name. Press F5. Then close the tab entirely, open a fresh one, and log in again. If it is not there, or if the list has gone back to sample names, you have what this agency had. Ten minutes.
2. Count the rows. Open your database console (on Base44 the Data tab; on Supabase the Table Editor; anywhere else psql) and run SELECT count(*) FROM appointments; or whatever your main table is. Use the app. Run it again. If the number does not move, the app is not writing. Five minutes.
3. Find the wipe. In the repository: grep -rn "drop_all\|create_all\|DATABASE_URL" tests/ conftest.py. Any hit that reads a URL from the environment without checking the hostname can empty whatever that URL points at. Fifteen minutes, including the fix.
Those three checks tell you whether the data exists. They do not tell you whether it is consistent, who changed it, or whether the backup you think you have will restore on the morning you need it. That is the other forty-seven hours.
The rule
The rule
If you cannot point at the row, you do not have the data. A screen that remembers is not a database, and a database nobody has restored is not a backup.
My client asked me where the appointments were stored. I opened the code to answer and realised I couldn't.— Agency owner