What this is
Every vibe-coded app reaches the same moment. It works. People are using it. And nobody in the room can answer the only question that matters: if this gets real traffic on Monday, what breaks, and who finds out first — us, or a customer?
The prodready MCP server answers that question with the same catalogue our reviewers use on paid work. It is the diagnosis we run in the first hours of every engagement, handed to you for nothing, inside the tool you already code in.
It is a remote server. It has no access to your disk, your database or your network. It cannot read a file, and it is not asked to. Instead it hands your agent a list of read-only commands to run locally — greps, file counts, git ls-files — and then reads back only what your agent chose to send: paths, line numbers, counts and matched identifiers. Never file contents. Never values.
The rule this server is built on
The report names the gap and its consequence. It never names the remedy.
“The row-level policy on invoices gates only on the tenant setting, with no operator branch” is a finding. So is “three tables carrying an owner column have row security switched off”. What you do about either one is the work we sell, and you will not find a patch, a migration, a config snippet or a step list anywhere in the output. That is deliberate, and we would rather say so on the box than have you discover it.
The trade is honest in both directions. A finding has to be specific enough and evidenced enough that you can walk to the file and see it for yourself — a vague warning is worth nothing and we treat it as a defect. What it will not do is fix it for you.
How it works
Your agent calls prodready_probe
The server returns the detector plan: the exact read-only commands and file globs for the stack it was told about, and what evidence each one needs.
Your machine runs them
Your agent executes the commands in your repository. Nothing is written, nothing is installed, no network call is made. You watch it happen in your own terminal.
Your agent calls prodready_assess
It sends the collected evidence — paths, counts, identifiers. The server maps it onto numbered rules, scores five dimensions and picks a grade.
You get the report
Verdict first, then what holds, then what does not, then the scope band for closing it. The evidence is discarded the moment the response is written.
The server is stateless. Probe, then assess, and nothing is remembered between the two — no session, no account, no history, nowhere for your evidence to sit. If you run it twice, the second run knows nothing about the first.
The two tools
prodready_probe(stack?)
Returns the read-only detector plan your host should run locally.
- Takes
- stack — optional hint, any of: postgres, supabase, prisma, sql-migrations, nextjs, queue, multi-tenant, i18n, email, hosted-runtime. Omit it and the plan covers every stack-agnostic detector and asks your agent to fingerprint the rest.
- Returns
- detectors[{ id, rule, group, severity, dimension, commands[], expect, collect[], redact[] }]
- Guarantee
- Every command is read-only: grep and rg over source globs; find, ls, wc, head, cat; awk, sort, uniq, xargs and shell loops over their output; git ls-files, git log, git shortlog; and one
node -eline that reads package.json. Four of the 118 checks carry a read-only query against the database's own catalogue —pg_roles,pg_class,pg_policies. The plan marks thoseSQL>rather than$, they run only against a database you hand it, and no check depends on one.
prodready_assess(evidence)
Turns what your agent collected into a graded report and a priced offer.
- Takes
- evidence — the output of the probe, keyed by detector id. Anything your agent could not run comes back as not evaluated, which is not the same as passed.
- Returns
- score, score_is_ceiling, grade, status, verdict, dimensions[{ id, score, max, unevidenced }], findings[{ id, rule_ref, severity, dimension, finding, consequence, evidence_summary }], not_evaluated[], offer
- Never returns
- remedy, fix, patch, snippet, migration, solution, steps. Those fields do not exist in the contract.
The catalogue behind them is 118 detectors drawn from the 25 numbered platform-hardening rules — wiring, configuration, multi-tenancy and row-level security, UI and internationalisation, operations, compliance and lifecycle — plus our own rules on data integrity, API design, state machines, testing, sessions, secrets and logging. Twenty-six of them are rated critical, which in this catalogue means one thing: a live exposure, or a failure that reaches your users with no gate in front of it.
Install
One line per host. The endpoint speaks MCP over streamable HTTP, needs no key and holds no account.
claude mcp add --transport http prodready https://mcp.prodready.si/mcp --scope user
Drop --scope user to add it to the current project only. Then ask: use prodready to check whether this app is production ready.
codex mcp add prodready --url https://mcp.prodready.si/mcp
Or, in ~/.codex/config.toml: [mcp_servers.prodready] then url = "https://mcp.prodready.si/mcp"
{"type":"mcp","server_label":"prodready","server_url":"https://mcp.prodready.si/mcp","require_approval":"never"}
Add it to the tools array of a Responses call. Leave require_approval off if you want to see each tool call before it runs.
{"mcpServers":{"prodready":{"url":"https://mcp.prodready.si/mcp"}}}
Project-local instead: .cursor/mcp.json in the repository root, same shape.
{"mcpServers":{"prodready":{"serverUrl":"https://mcp.prodready.si/mcp"}}}
Windsurf names the key serverUrl for remote servers, not url. Open it from Cascade → MCP servers → Manage → View raw config.
{"mcpServers":{"prodready":{"type":"streamableHttp","url":"https://mcp.prodready.si/mcp"}}}
Set type explicitly — Cline will otherwise guess the transport.
npx -y mcp-remote https://mcp.prodready.si/mcp
Older hosts without remote transport can bridge through this as a normal stdio command.
Once it is installed, you do not call the tools by name. Ask the question in plain language — is this secure? is this production ready? what breaks when we get real traffic? — and your host picks the tools up on its own.
What a report looks like
This is a real shape with a real worked example: a Lovable-built invoicing app on Supabase, forty paying customers, running for five months.
Not production-ready. The gaps reach into your data and who can reach it, and the scope has to be agreed before any clock starts.
criticaldatarule 11 · group C
Three tables carrying an owner or tenant column have no row-level security in force: invoices, invoice_lines, customers.
For those tables isolation rests entirely on every query in the codebase remembering to narrow by owner, so the single query that forgets — today's or next month's — returns every customer's rows to whoever asked.
criticaltestsSP-01
The repository holds no tests at all: no file matches *.test.*, *.spec.*, __tests__/ or tests/ anywhere in the tree, across 214 source files.
Nothing in this codebase is checked by anything except the last person who clicked through it, so a regression in sign-in, payment or tenant scoping reaches your users before anyone knows it exists.
criticalreleaseSP-02
No continuous integration exists: .github/workflows holds no workflow file and no .gitlab-ci.yml, .circleci/config.yml or equivalent is present anywhere in the repository.
Every commit reaches the deployed branch without a single automated gate in front of it, so the first environment where a broken build is discovered is the one your users are standing in.
highreviewrule 6 · group B
The application origin at src/lib/url.ts:14 falls back to a development default when NEXT_PUBLIC_APP_URL is unset, rather than failing.
A deploy that forgets one variable is not detectably broken — the service boots, answers, and quietly builds every absolute link, redirect and callback against the wrong host until a user reports it.
mediumreviewSP-12
Nine source files are over 500 lines — app/dashboard/page.tsx (1,340), lib/billing.ts (880), components/InvoiceTable.tsx (742) — holding 38% of the codebase in files no reader takes in at one sitting.
A change anywhere in one of these files has to be reasoned about against the whole file, and a defect introduced in one is the hardest kind to see in review: the reader is shown forty changed lines with no way to hold the surrounding thirteen hundred in their head.
Scope band: Heavy
14 remediation stories, 3 of them critical. Fits one 48-hour engagement. Fixed price, fixed scope, quoted after Discovery.
Discovery is $500, credited against the engagement if you go ahead.
Score 70 lands in band B. Two hard rules cap it: a critical finding in the data dimension caps at D, and a missing test suite or missing CI caps at C. The most severe cap wins, so the report says D and shows why. A cap never raises a grade.
What is missing from that report, on purpose
Not one line of it tells you what to type. No policy to create, no workflow file to add, no variable to set, no library to install. Five findings, five consequences, a number and a price. If that is all you ever take from us, you are still ahead of where you were — you now know what is wrong and can go fix it yourself.
How the score works
Five dimensions, twenty points each, one hundred in total. Each dimension starts at twenty and subtracts: eight points for a critical finding, five for a high, one for a medium. A dimension clamps at zero and never goes below it. Findings below medium are reported as observations and move nothing.
| Dimension | What it covers | Max |
|---|---|---|
| data | Schema constraints and foreign-key integrity, tenant scoping and row-level security, money and status typing, deletes, retention, query shape | 20 |
| safety | One session truth, role gates, rate limiting before side effects, audit trail, GDPR surfaces, secret handling | 20 |
| tests | Existence of a suite, coverage gating, integration and end-to-end probes of real flows | 20 |
| review | Wiring and orphan code, hardcoded config and locale, form and state contracts, dependency hygiene, unresolved decisions | 20 |
| release | CI gates, health and liveness, job reliability, deploy-time verification against the real origin | 20 |
Those are the same five areas the rest of this site talks about, and each one is also shown on the five-point scale used elsewhere — twenty points divided by four. The hundred-point total is the number of record.
| Score | Grade | Status | What it means |
|---|---|---|---|
| 85 – 100 | A | Green | Production-ready in everything we can see from the outside. What is left is finish, not structure. |
| 70 – 84 | B | Green | Close. A short and bounded list stands between this and production. |
| 55 – 69 | C | Amber | Not production-ready. The gaps are known and countable — this is the shape a 48-hour engagement is built for. |
| 40 – 54 | D | Amber | Not production-ready. The gaps reach into your data and who can reach it, and the scope has to be agreed before any clock starts. |
| 0 – 39 | F | Red | Do not put real users on this yet. What we found is not a punch list; it sits in the foundations. |
The hard rules that override the number
A good score does not survive a bad finding. Five caps are applied after the band is chosen, and the most severe one wins.
- A critical finding anywhere, on a score of 85 or more, caps the grade at C.
- No test suite, or no continuous integration, caps at C regardless of score.
- A critical finding in the data dimension — a live cross-tenant or integrity exposure — caps at D.
- Credential files tracked in version control cap at F. While a committed credential is live, nothing the report says about access control can be relied on.
- A total under 40 caps at F, and the report says plainly that the standing cannot be settled from a scan alone.
When the scan is incomplete, the report says so
A detector your agent could not run is recorded as not evaluated. It is never counted as passed. A dimension with no evidence behind it keeps its full twenty points and is labelled unverified — and if two or more of the five are unverified, the total is reported as a ceiling, “at best 74 out of 100”, never as a flat number. The number is the number. It is not rounded up.
What it reads, and what never leaves your machine
Runs locally, read-only
grep/rgover source globsfind,ls,wc,head,catgit ls-files,git log,git shortlogawk,sort,uniq,xargsand shell loops over their output- One
node -eline that readspackage.jsonand prints dependency versions - Four optional catalogue queries —
pg_roles,pg_class,pg_policies— and only against a database you hand it
Never, under any plan
- Any write, move, delete or install
- Any network call a probe command makes on its own — no fetch, no curl, no package download, no telemetry. The only connection a probe ever opens is the read-only database connection you hand it, and no check needs one: skip them and they come back as not evaluated
- Any command that prints a secret value
- Any command you did not see your own agent run
What crosses the wire is the evidence your agent sends to prodready_assess, and you can read it before it goes. That is file paths, line numbers, counts, table and column names, package names and versions, and matched identifiers. It is not your source code, and the server has no way to ask for it.
Credential detectors never print the line their match came from. They run with -l and -c — list the files, count the matches — or with -o against a pattern that matches only an identifier, so the output is path:line:VAR_NAME and stops there. Never bare -n, never -A/-B/-C. The evidence for a secrets finding is a filename, a line number, a count and the name of the variable. Never the value.
The report itself is assembled from counts, denominators and paths. A few findings also quote the short line your agent wrote about what a command returned — twelve words at most — and the server checks that line before it prints: key material, an email address, a phone number, a URL with parameters, a named person, anything that reads as an instruction, or anything longer than a quoted fragment, and the finding is dropped rather than softened. The evidence line under each finding never carries it, an entry in a file list that is not shaped like a path is left out rather than trusted, and the whole report is checked once more on the way out.
Storage, in one sentence
Evidence is never stored and never logged — the server holds it for the length of one request, writes the report, and forgets it, because there is no session, no database and nowhere else for it to go.
What it cannot see
A read-only scan of a repository is not a full Discovery, and pretending otherwise would be the same dishonesty as a vague finding. These are the things this server does not claim:
- Whether the app does what you meant. Requirements clarity and scope control are scored, in our engagements, from written stories and open questions. A repository holds neither. Nothing in a codebase substitutes for them.
- Anything that needs a running system. Row counts in the live database, a worker actually booting, a build manifest. A remote server that cannot read your disk certainly cannot query your database.
- Business rules nobody approved. The most expensive thing we find in vibe-coded apps is a decision the model made and the founder never saw. It reads as working code; no detector catches it. That is a conversation, not a grep.
- Your history with us. Delivery confidence normally leans on velocity from sprints already run. At first contact there are none, and its absence is not counted against you.
Each of these appears in the report by name, in the not_evaluated list, rather than being quietly folded into the score.
How the price is worked out
Every finding maps to a class of work with a base estimate, multiplied for the conditions that genuinely make work slower — a first story in an unfamiliar codebase, a file over 500 lines, a pattern that exists nowhere yet, a third-party integration — and capped so that no single story runs long without being split. Count the stories, count the criticals, and the assessment lands in a scope band.
| Scope band | When | 48 h |
|---|---|---|
| Contained | 6 stories or fewer, no critical findings | yes |
| Standard | 7 – 15 stories, at most one critical | yes |
| Heavy | 16 – 30 stories, or two or more criticals | yes |
| Beyond one engagement | More than 30 stories, or grade F | no |
The band is the size of the job, not a quote. The report names it, says how many stories and how many criticals put the assessment there, and then says the fixed price is quoted after Discovery — because the only price we publish is the one we can stand behind. When the work is genuinely larger than one 48-hour engagement, it says that instead. Discovery is $500 — we look at your product and tell you honestly whether it is anywhere near production — and it comes off the bill if you go ahead.
Short answers
Is it free?
Yes. No key, no account, no rate-limit tier to buy. The diagnosis is free because it is the honest half of a sales conversation: you learn what is wrong, and we earn the right to quote for fixing it.
Why won't it just fix the problems?
Because that is the product. Forty-eight hours of our delivery, on your own cloud, with the tests to prove it, is what you would be paying for. A free tool that also shipped the fix would be us doing the work and calling it marketing.
Can I run it on a client's codebase?
Yes, and agencies do. The server never sees the code — your agent runs the commands, you see the output, you decide what is sent.
Does it work on Lovable, Replit, Base44 or Bolt exports?
Yes. Export or clone the project so the files are on disk where your agent can read them. The detectors are written for what those tools actually generate — Supabase schemas without policies, environment variables with development fallbacks, handlers reachable only by side-effect import.
What if the scan comes back clean?
Then it says so, and you have a grade A with the evidence behind every dimension. We would rather publish that than manufacture a finding.
Where do the rules come from?
The same playbooks our reviewers work from: 25 numbered platform-hardening rules across wiring, configuration, multi-tenancy, UI and internationalisation, operations and compliance; a set of critical rules on data integrity, API design, state machines, testing, sessions, secrets and logging; and the scoring model that turns findings into a number. Nothing in the catalogue was invented for the demo.
Or skip the scan and send us what you have.
Discovery is $500: we look at your product and tell you, honestly, whether it is anywhere near production. If you go ahead, the $500 comes off the bill. 48 hours later, it's production-ready — on your own cloud, with the tests to prove it.