· 7 min read

Hardening an AI-built app: the 12-point checklist

The twelve checks I run first on an app built quickly with AI coding tools, in the order that finds the most serious problems soonest.

By

Hardening an AI-built app means finding and fixing the security and reliability gaps that fast, AI-assisted development tends to leave behind, before real users or an attacker find them. AI coding tools make it possible to ship a working product in weeks; they also reproduce the same handful of mistakes across thousands of projects, which makes those mistakes predictable and the checks for them short.

This is the checklist I start with when I assess an app built quickly, whether by a founder with an AI assistant, a freelancer or a small team. The order matters: it puts the problems that expose other people's data first.

Why do AI-built apps fail in the same places?

Because the tools optimize for "it works when I try it". A sign-in flow works when you try it; whether another user can read your records is a question nobody asked the tool. Defaults that are fine on a laptop, such as an open database rule, a key in a config file or no rate limit, survive into production because nothing failed. The OWASP Top 10:2025 puts broken access control first and security misconfiguration second, and in my experience those two cover most of what goes wrong in these apps.

The 12 checks

#CheckHow I test itTypical fix
1Secrets in code or bundlesScan the repository, its history and the built front endRotate the keys; move secrets to the platform's secret store
2Authorization on every endpointSign in as user A, request user B's recordsScope every query by owner or organization, server-side
3Database rules (Supabase, Firebase)Query the database directly with the public keyEnable and write row-level or security rules for every table
4Admin and internal routesRequest them as a normal user and signed outRole checks on the server, not hidden buttons
5Input validationSend wrong types, huge values, scripts, SQLValidate on the server with a schema; parameterized queries
6Rate limitsRepeat sign-in, sign-up, password reset and costly AI callsLimits per user and per IP; usage caps on paid APIs
7DependenciesAudit for known vulnerabilities and abandoned packagesUpdate, replace or remove; automate update checks
8Error handling and loggingTrigger errors; read what the user and the logs seeGeneric user errors; useful, scrubbed server logs
9Data modelRead the schema against the product's real rulesConstraints, foreign keys and migrations in version control
10Tests on the critical pathsLook for tests on money, permissions and data deletionAdd tests there first, and run them in CI
11Backups and restoreAsk when a restore was last tested; then try oneScheduled backups and a timed, documented restore
12Deploys and environmentsCheck how code reaches production, and with which keysCI/CD, separate staging and production credentials

1. Are there secrets in the code?

AI tools often paste API keys straight into source files, and front-end frameworks will happily ship any key the browser code references. I scan the repository, its full history and the built JavaScript. Any secret found is treated as leaked: rotated first, then moved to a secret store, as the OWASP Secrets Management Cheat Sheet recommends instead of hardcoding. GitHub secret scanning with push protection then stops the next one at the commit.

2. Can one user read another user's data?

This is the finding I see most and the one that matters most. The app checks that a request comes from a signed-in user but not that the record belongs to them. I test every endpoint that takes an identifier with two accounts. The fix is always server-side: load records through the current user's ownership, never trust an identifier from the client alone.

3. Are the database rules open?

Many AI-built apps talk to Supabase or Firebase directly from the browser, which is fine only if the database enforces access itself. Supabase's documentation is blunt: a table in an exposed schema without Row Level Security is readable and writable by any role with a grant on it, so RLS should be enabled on every such table. Firebase has the equivalent in its Security Rules. I test with the public key exactly as an attacker would, straight from a script.

4. Are admin routes protected on the server?

Hiding an admin button isn't access control. I request every admin and internal route as a normal user and as nobody at all.

5. Is input validated on the server?

Client-side validation is a convenience for honest users. The server validates types, lengths and formats against a schema, and every database query is parameterized.

6. Are there rate limits?

Without them, sign-in can be brute-forced, sign-up can be flooded, and an AI feature can run up a bill overnight. Limits go per user and per IP, with hard caps on paid APIs.

7. Are the dependencies current?

Fast builds pull in many packages. I audit for known vulnerabilities and for packages nobody maintains, and set up automated update checks so the list doesn't rot again.

8. What do errors reveal?

Stack traces in the browser tell an attacker your framework, file paths and sometimes queries. Users get a generic message; the server log gets the detail, with personal data and secrets scrubbed.

9. Does the data model hold?

AI-generated schemas often lack constraints: no foreign keys, nullable everything, uniqueness enforced only in code. That's how duplicate accounts and orphaned records appear. Constraints go into the database, and schema changes become migrations in version control.

10. Are the critical paths tested?

Not everything needs a test on day one. Payments, permissions and data deletion do. Those tests go in first and run in CI, so the next AI-assisted change can't silently break them.

11. When was a backup last restored?

A backup nobody has restored is a hope. I restore one into a separate environment, time it, and write the steps down.

12. How does code reach production?

If the answer is "from my laptop", that's the last fix: a pipeline that runs the tests, separate staging and production credentials, and a way to roll back. NIST's Secure Software Development Framework is a good reference for the practices to grow into once the basics are in place.

In what order should you fix things?

By who gets hurt. First anything that exposes other people's data or your keys (checks 1 to 4). Then anything that lets someone run up costs or take the service down (5, 6, 8). Then the foundations that make the next feature safe to build (7, 9 to 12). Most AI-built code is worth keeping; the aim is to separate what is dangerous from what is merely untidy, and fix them in that order. The Application Security Verification Standard is the fuller list to work towards once these twelve are done.

How do you keep it fixed?

With gates rather than good intentions: secret scanning with push protection, dependency checks, the critical-path tests and a type check in CI, and review rules for what an AI coding agent may change on its own. The same gates that catch an agent's mistakes catch a tired human's.

How I do this

A security audit and penetration test covers these checks and more, with findings ranked by severity, a fix for each and a retest. If your team is adopting AI coding agents and wants the gates set up before the problems appear, AI coding agents for engineering teams covers repository instructions, CI gates and review rules. Either starts with a short brief on the contact page.

Sources