← GUIDES

The vibe coding security checklist we actually use

By Leon Mallett · Updated 8 October 2026 · 4 min read

AI agents write code that works. They're much less reliable at writing code that's safe, because "safe" is mostly about things that should not happen, and you rarely ask for those.

This is the checklist we use on Vibe Leagues, which is built almost entirely with an AI agent. Every item is here because we either got it wrong first or nearly did. It's written for a web app (ours is Next.js on Cloudflare), but most of it applies to anything with users.

Secrets

1. No secret ever appears in the repository, or in the chat. Why: a key pasted into a conversation with an agent is in that transcript for good, and the only real fix is to rotate it. A key committed to git is in the history even after you delete it. How: keep secrets in your OS keychain or your host's secret store, and reference them by name. Before every commit, look at the diff for anything shaped like a key.

Accounts and sign-in

2. Limit password guessing. Why: without a limit, anyone can try passwords against an account as fast as your server answers. Our first version had none. How: count failed sign-ins per email and per IP address over a time window, and refuse further attempts before checking the password. Test it: five wrong passwords should lock the account out even for the right one.

3. Limit every other thing that can be guessed. Why: invite codes, reset tokens, one-time links. If anything can be checked for validity, it can be brute-forced. Ours leaked validity just from loading a page with a code in the URL. How: list every place your app says "valid" or "invalid" to an anonymous visitor, and put the same limit on each.

4. Don't reveal which accounts exist. Why: "wrong password" versus "no such account" tells an attacker who's registered, and so does a faster response for unknown emails. How: return the same message either way, and do the same amount of work either way. We verify unknown emails against a dummy password hash so the timing matches.

5. Hash passwords with a slow, salted algorithm. How: Argon2id or bcrypt where available. On Cloudflare Workers, PBKDF2 through Web Crypto, at the platform's ceiling of 100,000 iterations. Never a plain hash.

Authorisation

6. Every action checks permission itself, not just the page it's on. Why: hiding a button isn't security. In Next.js, server actions and API routes can be called directly, without loading the page. Before we made our community publicly readable, we checked that every action verified the session on its own, rather than trusting the page's login gate. How: for each action, ask "what if a signed-out stranger calls this directly?"

7. API tokens: random, hashed, shown once, revocable, narrow. How: at least 256 bits of randomness; store only a hash; show it to the user once; let them revoke it; scope it to what the integration needs. Ours can post and delete the member's own posts, and nothing else: no reacting, no inviting.

8. Token APIs ignore cookies. Why: if your API accepts the browser's session cookie, any website can make a signed-in visitor's browser call it (cross-site request forgery). How: on bearer-token endpoints, read only the Authorization header. Test that a request carrying only a session cookie is refused.

Headers and the browser

9. Security headers on every response, including the cached and static ones. Why: on frameworks that serve some pages pre-built and others on demand, the headers often have to be set in two places. Set them in one and your homepage can quietly have none. How: fetch one URL of each kind (a static page, a dynamic page, a static file, an API response) and count the headers on each. We expect seven: HSTS, nosniff, referrer policy, frame protection, permissions policy, cross-origin opener policy, and a content security policy.

10. Start your Content Security Policy in report-only mode. Why: a wrong policy silently breaks pages in some browsers and not others. How: ship Content-Security-Policy-Report-Only, then load real pages in a real browser and listen for violations before enforcing it. When we added analytics, we updated the policy first, then loaded the real pages in a browser and confirmed zero violations before shipping.

Data and infrastructure

11. Read every database migration before it runs in production. Why: generated migrations can drop and rebuild tables. On our database that would have cascaded and deleted every reaction on the site. How: read the SQL. Look for DROP, table rebuilds and PRAGMA. Test it on a throwaway copy that has data in it, and compare row counts before and after.

12. Treat every new dependency as a security decision. Why: a freshly published package version is the favourite delivery route for supply-chain attacks. How: pin exact versions; be wary of anything released in the last 72 hours; check whether it runs install scripts; install from the lockfile.

13. Ship on your own domain, and check the live site after every deploy. Why: free platform subdomains (*.workers.dev, *.pages.dev) get blocked by corporate and home security filters. And "merged" isn't "deployed": one of our merges silently never built. How: use a custom domain, switch the platform subdomain off, and after each deploy re-run the checks above against the live site.


None of this is exotic, and most of it takes minutes. The reason it gets skipped with AI-written code is speed: the code arrives faster than anyone thinks to ask "what shouldn't this allow?". Building something with AI and learning this the hard way? That's a story worth sharing here.