· 7 min read
Hardening an AI-built app: the 12-point checklist
The twelve checks I run first on an app built quickly with AI coding tools, in the order that finds the most serious problems soonest.
Hardening an AI-built app means finding and fixing the security and reliability gaps that fast, AI-assisted development tends to leave behind, before real users or an attacker find them. AI coding tools make it possible to ship a working product in weeks; they also reproduce the same handful of mistakes across thousands of projects, which makes those mistakes predictable and the checks for them short.
This is the checklist I start with when I assess an app built quickly, whether by a founder with an AI assistant, a freelancer or a small team. The order matters: it puts the problems that expose other people's data first.
Why do AI-built apps fail in the same places?
Because the tools optimize for "it works when I try it". A sign-in flow works when you try it; whether another user can read your records is a question nobody asked the tool. Defaults that are fine on a laptop, such as an open database rule, a key in a config file or no rate limit, survive into production because nothing failed. The OWASP Top 10:2025 puts broken access control first and security misconfiguration second, and in my experience those two cover most of what goes wrong in these apps.
The 12 checks
| # | Check | How I test it | Typical fix |
|---|---|---|---|
| 1 | Secrets in code or bundles | Scan the repository, its history and the built front end | Rotate the keys; move secrets to the platform's secret store |
| 2 | Authorization on every endpoint | Sign in as user A, request user B's records | Scope every query by owner or organization, server-side |
| 3 | Database rules (Supabase, Firebase) | Query the database directly with the public key | Enable and write row-level or security rules for every table |
| 4 | Admin and internal routes | Request them as a normal user and signed out | Role checks on the server, not hidden buttons |
| 5 | Input validation | Send wrong types, huge values, scripts, SQL | Validate on the server with a schema; parameterized queries |
| 6 | Rate limits | Repeat sign-in, sign-up, password reset and costly AI calls | Limits per user and per IP; usage caps on paid APIs |
| 7 | Dependencies | Audit for known vulnerabilities and abandoned packages | Update, replace or remove; automate update checks |
| 8 | Error handling and logging | Trigger errors; read what the user and the logs see | Generic user errors; useful, scrubbed server logs |
| 9 | Data model | Read the schema against the product's real rules | Constraints, foreign keys and migrations in version control |
| 10 | Tests on the critical paths | Look for tests on money, permissions and data deletion | Add tests there first, and run them in CI |
| 11 | Backups and restore | Ask when a restore was last tested; then try one | Scheduled backups and a timed, documented restore |
| 12 | Deploys and environments | Check how code reaches production, and with which keys | CI/CD, separate staging and production credentials |
1. Are there secrets in the code?
AI tools often paste API keys straight into source files, and front-end frameworks will happily ship any key the browser code references. I scan the repository, its full history and the built JavaScript. Any secret found is treated as leaked: rotated first, then moved to a secret store, as the OWASP Secrets Management Cheat Sheet recommends instead of hardcoding. GitHub secret scanning with push protection then stops the next one at the commit.
2. Can one user read another user's data?
This is the finding I see most and the one that matters most. The app checks that a request comes from a signed-in user but not that the record belongs to them. I test every endpoint that takes an identifier with two accounts. The fix is always server-side: load records through the current user's ownership, never trust an identifier from the client alone.
3. Are the database rules open?
Many AI-built apps talk to Supabase or Firebase directly from the browser, which is fine only if the database enforces access itself. Supabase's documentation is blunt: a table in an exposed schema without Row Level Security is readable and writable by any role with a grant on it, so RLS should be enabled on every such table. Firebase has the equivalent in its Security Rules. I test with the public key exactly as an attacker would, straight from a script.
4. Are admin routes protected on the server?
Hiding an admin button isn't access control. I request every admin and internal route as a normal user and as nobody at all.
5. Is input validated on the server?
Client-side validation is a convenience for honest users. The server validates types, lengths and formats against a schema, and every database query is parameterized.
6. Are there rate limits?
Without them, sign-in can be brute-forced, sign-up can be flooded, and an AI feature can run up a bill overnight. Limits go per user and per IP, with hard caps on paid APIs.
7. Are the dependencies current?
Fast builds pull in many packages. I audit for known vulnerabilities and for packages nobody maintains, and set up automated update checks so the list doesn't rot again.
8. What do errors reveal?
Stack traces in the browser tell an attacker your framework, file paths and sometimes queries. Users get a generic message; the server log gets the detail, with personal data and secrets scrubbed.
9. Does the data model hold?
AI-generated schemas often lack constraints: no foreign keys, nullable everything, uniqueness enforced only in code. That's how duplicate accounts and orphaned records appear. Constraints go into the database, and schema changes become migrations in version control.
10. Are the critical paths tested?
Not everything needs a test on day one. Payments, permissions and data deletion do. Those tests go in first and run in CI, so the next AI-assisted change can't silently break them.
11. When was a backup last restored?
A backup nobody has restored is a hope. I restore one into a separate environment, time it, and write the steps down.
12. How does code reach production?
If the answer is "from my laptop", that's the last fix: a pipeline that runs the tests, separate staging and production credentials, and a way to roll back. NIST's Secure Software Development Framework is a good reference for the practices to grow into once the basics are in place.
In what order should you fix things?
By who gets hurt. First anything that exposes other people's data or your keys (checks 1 to 4). Then anything that lets someone run up costs or take the service down (5, 6, 8). Then the foundations that make the next feature safe to build (7, 9 to 12). Most AI-built code is worth keeping; the aim is to separate what is dangerous from what is merely untidy, and fix them in that order. The Application Security Verification Standard is the fuller list to work towards once these twelve are done.
How do you keep it fixed?
With gates rather than good intentions: secret scanning with push protection, dependency checks, the critical-path tests and a type check in CI, and review rules for what an AI coding agent may change on its own. The same gates that catch an agent's mistakes catch a tired human's.
How I do this
A security audit and penetration test covers these checks and more, with findings ranked by severity, a fix for each and a retest. If your team is adopting AI coding agents and wants the gates set up before the problems appear, AI coding agents for engineering teams covers repository instructions, CI gates and review rules. Either starts with a short brief on the contact page.
Sources
Related service
- Service: Security audit and penetration test · Security and compliance engineering
An attacker's view of your application and infrastructure, and a fix list you can act on.
- Service: AI coding agents for engineering teams · Applied AI engineering
Claude Code and similar agents set up to ship real work in your codebase, with guardrails, tests and review rather than unreviewed pull requests.