Mihály Tari.

Sample report: an AI code audit of my own repo

A finished AI code audit of my own directory-engine repo. 50,000 lines, one week, mostly AI. Ten findings down to the file and line, with hour estimates and a fix order. Published as it came out.

This is a real audit, and the code it looks at is mine. directory-engine is a multi-tenant directory engine. I wrote it between 2026-08-27 and 2026-09-03, in 229 commits, mostly with AI. 50,384 lines of TypeScript, 169 test files against 279 source files. (I count as a test file anything matching *.test.* or *.spec.*, plus everything under a test/, tests/, __tests__/ or e2e/ directory.) It’s exactly the situation the AI code audit is for, so I picked it as the sample. I haven’t tidied it up. What came out is what’s here.

The layout is the same as a report I’d write for a client. First, one page for whoever doesn’t read code. Then the findings in two parts: what hurts now, and what will hurt in six months. At the end, the order to fix things in.

One-page summary

Verdict. Safe to keep live. Two things are worth fixing before the next deploy, but neither has an attacker behind it. Both are triggered by an error or an outage at the wrong moment. The thing multi-tenant systems usually get wrong, one tenant seeing another’s data, is fine here. Every database query filters on the tenant id, and so does the auth adapter.

Findings. 0 critical · 2 high · 6 medium · 2 low.

Time to fix. 9–15 hours for the two high items, 17–33 hours for the rest.

If nothing changes. The sign-in rate limit counts per instance, not globally. So the real ceiling isn’t three emails per ten minutes, it’s three times however many server instances happen to be running. If a publish dies between its two database writes, it leaves behind a listing that never expires, that nobody pays for, and that nothing goes looking for. During the first database outage, the catalogue page goes out to visitors and to Google with a 200 and an empty list, and nothing except a line in the Vercel log records that it happened.

What I looked at, and what I didn’t. Repository directory-engine, commit f5753ff, 2026-09-03. I read all 8 route handlers, the Auth.js adapter and how it scopes to a tenant, the ownership checks in the listing actions, both migrations with their indexes, the take-down token lifecycle, the upload signature and the Cloudinary namespace checks, the markdown sanitiser, the rate limiter, the block renderer, the nightly cron, and the whole git history, looking for secrets. I didn’t look at known vulnerabilities in dependencies (not part of this audit, and the report says nothing about them one way or the other), the running infrastructure (Vercel settings, DNS, the Postgres and Cloudinary account configuration), performance under load, what the Playwright e2e suite actually covers, accessibility, or GDPR compliance for the stored data.

One result that isn’t in the list. I went through all 229 commits, on every branch, merge diffs included. The only .env-style file ever added is apps/web/.env.example. It has variable names, comments, one public hostname, a tenant slug and two placeholder connection strings whose host is literally host. No secret values, and I found no secret-looking line anywhere else in the git history either. It’s not a finding because it’s fine, not because I didn’t look.

Method

I started with a scripted pass: a few scripts that flag where to look. Secrets across the whole git history, dependency advisories (whose output I didn’t process, see above), test ratio per workspace, oversized files, repeated code, AI-code tells, and a list of route handlers with an auth flag. The scripts don’t write findings. I checked every flag by hand, and only the ones where I could write down the evidence, a concrete failure scenario, the fix and the time stayed in the report. Whatever didn’t hold up came out. One example: the scripts flagged six route handlers because they found no auth idiom in them. Two genuinely needed reading. The other four (robots.txt, the two sitemap routes and Auth.js’s own endpoint) are public on purpose, and they’re not below.

Severity.

  • Critical: exploitable, or can lose data, today, with no precondition.
  • High: exploitable, or can lose data, given a likely condition (a signed-in user, a bad input, a concurrent request).
  • Medium: no immediate harm, but it will cause a defect at the next feature or the next developer.
  • Low: untidiness that costs time, not damage.

Part 1 — What hurts now

1. The rate limiter counts in one process’s memory, and the deploy runs on several — High

Evidence. packages/engine/src/rate-limit.ts:11: const buckets = new Map<string, Bucket>();. The module’s header (lines 1–3) says so:

// In-memory fixed-window rate limiter — best-effort defense-in-depth against floods/abuse, not a
// globally enforced quota (state lives in this process only; a multi-instance deployment needs a
// shared store instead). Ported from erdei-fahazak.hu's shared/rate-limit.ts, generalised to the

Both sign-in limits are built on it. apps/web/app/%5Fsites/[tenant]/auth/api/[...nextauth]/route.ts:63: limitByIp(tenant, "sign-in", req.headers, { limit: 5, windowMs: 60_000 }), and then on line 74, limitByKey(tenant, "sign-in-email", email.toLowerCase(), { limit: 3, windowMs: 600_000 }).

Scenario. Whoever wrote this knew what they were doing. The header says where the limit is and what should replace it (“a multi-instance deployment needs a shared store instead”). It’s just that the deploy became multi-instance in the meantime and the shared store never got built. In practice: someone types somebody else’s email address into the sign-in form and resubmits it every minute. Under load, Vercel runs the app on several independent instances, each with its own buckets Map, and every cold start begins with an empty one. “Three emails to this address per ten minutes” becomes three emails per instance. The victim’s inbox fills up with sign-in links, they mark them as spam, and the delivery reputation of the Postmark domain drops. Every tenant’s transactional mail goes through that domain. Listing notifications use the same limiter module, in a separate namespace (apps/web/lib/actions/listing-internal.ts:41: limitByKey(tenant, "listing-notify", listingId, { limit: 5, windowMs: 3_600_000 })), so more gets through there too than the limit promises.

Fix. Put the counter in shared storage: a rate_limits table in Postgres with an upsert, or Vercel KV / Upstash. The key shape (<tenant uuid>:<subject>:<route>) can stay, only the store underneath changes, so it’s one module to touch. 6–10 hours.

2. Publishing is two writes with no transaction, and if the second one is lost the listing stays free forever — High

Evidence. apps/web/lib/actions/listing-internal.ts:453: await storage.setListingStatus(listing.id, "LISTING_APPROVED");. Then, as a separate statement on line 455: await storage.setListingPlan(listing.id, transition.stamp.plan, transition.stamp.planUntil);.

Scenario. After line 453 the listing is live. If the request dies between the two writes (Postgres fails over, the Vercel function hits its timeout, the connection drops), what’s left is a LISTING_APPROVED row whose plan_until is NULL. The nightly cleanup cron only looks at rows that have an expiry (packages/engine/src/storage/listings.ts:450: isNotNull(listings.planUntil)), so it never archives this one. It stays live, for free, and there’s no query that would find “approved but with no plan” rows. There’s no transaction anywhere in the repo: Storage doesn’t open one, and the guarded database handle explicitly refuses a transaction call (packages/engine/src/db/guard.ts:35).

Fix. A Storage.publishListing(id, stamp) method that writes status and plan in a single UPDATE, plus a one-off query for the rows already in that state. 3–5 hours.

3. When a block fails, the page goes out with a 200 and content missing, and all that’s left is one log line — Medium

Evidence. apps/web/lib/blocks/render.tsx:34-35:

        console.error(`[tenant:${ctx.tenant.id}] block ${b.block} failed`, err);
        return { slot: b.slot, index, block: b.block, node: null };

A null node is simply left out by the renderer.

Scenario. The database is unreachable for half a minute (a cold start, a connection limit, a failover, whatever). The listing-grid loader throws, renderPage catches it, drops the block, and /hazikok, the listings page, goes out with a 200, a full header and footer, and zero listings. If Google’s crawler comes by in that window, it indexes an empty catalogue. A first-time visitor sees an empty list and moves on. There’s no alarm: none of the five package.json files has an error-monitoring dependency, so the only trace of the whole thing is a single console.error line in the Vercel log. Dropping the block is right, by the way. One broken block shouldn’t take the page down. What’s missing is anything that tells someone it happened.

Fix. Keep the drop, add a signal. An error reporter (Sentry), or at least a data-block-failed marker in the output and an external check that fails when the catalogue page renders without its listing-grid. 4–8 hours.

4. The take-down token points at a slug, not at the listing’s id — Medium

Evidence. packages/engine/src/storage/engagement.ts:269: return row ? { userId: row.userId, listingSlug: row.listingSlug } : undefined;. The handler uses it like this (apps/web/app/%5Fsites/[tenant]/api/takedown/route.ts:131): const listing = await storage.getListingBySlugAnyStatus(consumed.listingSlug);, and then on line 134, await storage.setListingStatus(listing.id, "SUSPENDED");.

Scenario. The token is valid for 30 days (apps/web/lib/takedown.ts:5). Owner A publishes a listing called “Csendes Tanya”, and the email to the admin carries the take-down link. A few days later A deletes the listing. The delete deliberately leaves the token in place, and says so at packages/engine/src/storage/listings.ts:548. Still within those thirty days, owner B posts a listing with the same name, which becomes the same slug. When the operator digs out the old email and clicks the link, the handler resolves the slug to B’s live listing, suspends that one, and sends B the take-down notice.

Fix. Put the listing’s id on the token (the slug can stay next to it, for display) and have the handler look the listing up by that. Or have deleteAction delete the tokens for that slug as well. 2–4 hours.

Part 2 — What will hurt in six months

5. The owner query calls lower() on the indexed column, so the index can’t be used — Medium

Evidence. packages/engine/src/storage/listings.ts:244:

    .where(this.scoped(listings, sql`lower(${listings.ownerEmail}) = lower(${trimmed})`))

The index, though, is on the raw column, packages/engine/src/schema/listings.ts:52: tenantOwner: index("listings_tenant_owner_idx").on(t.tenantId, t.ownerEmail),.

Scenario. The query calls lower() on the column, and the index is built on the raw column, so Postgres can’t use it. Every load of an owner’s profile page reads every listing the tenant has. Today that’s a few hundred rows and nobody would notice. After the old erdei-fahazak.hu catalogue is imported it’s tens of thousands, and the page slows down exactly when owners start using it. The lower() on the read side is redundant anyway, because both createListing (:282) and updateListing (:381) lowercase the email on write.

Fix. Either take lower() out of the query (normalising the input is enough), or build the index on (tenant_id, lower(owner_email)). One line plus one migration. 1–2 hours.

6. The accounts primary key is a natural key, and the tenant id isn’t in it — Medium

Evidence. packages/engine/migrations/0000_init.sql:15: CONSTRAINT "accounts_provider_providerAccountId_pk" PRIMARY KEY("provider","providerAccountId").

Four tables in the schema use a natural key as their primary key instead of a synthetic id: accounts (:15), sessions (:20), verification_tokens (:40) and listing_terms (:172). The key of listing_terms starts with tenant_id. The other three don’t have it in the key at all. All three do have a tenant_id column (accounts:3, sessions:19, verification_tokens:36), so the row is tied to a tenant, just not the key. For sessions and verification_tokens the key contains a random value, so two tenants can’t collide there. For accounts, though, the key is the provider’s account id, and that’s exactly the kind of value two tenants can share. Every other per-tenant unique constraint in the schema (users_tenant_email, listings_tenant_slug, taxonomy_terms_tenant_kind_slug) starts with tenant_id. This one doesn’t.

Scenario. Today there’s only email sign-in, nothing writes to accounts, and the bug is dormant. The day a Google button gets added: a user signs in on erdei-fahazak.hu with their Google account, and the (google, <account id>) row is created. The same person signs in on kapcsolodjki.hu with the same account, linkAccount tries to insert the same pair, the primary key rejects it, and sign-in ends on an error page. Only for people who use both sites, and only in production, where there really are two tenants. On a dev machine with one tenant it never shows up.

Fix. A migration that changes the primary key to (tenant_id, provider, providerAccountId), before the first OAuth row is written. 2–4 hours.

7. Two different rules decide whether an image belongs to a tenant — Medium

Evidence. The write side, apps/web/lib/cloudinary.ts:73:

  return parsed.pathname.startsWith(uploadPrefix) && parsed.pathname.includes(`/${folderFor(tenant)}/`);

The delete side, in the same file, on line 144: if (!publicId || !publicId.startsWith(prefix)) {.

Scenario. The write gate lets a path through if the tenant’s folder is anywhere in it. The delete gate expects the public id to start with it. Today they behave the same, because Cloudinary generates the public_id and the folder is always the first segment. They part ways as soon as a segment appears in front of the tenant folder. publicIdFromUrl (cloudinary.ts:82) keeps the whole path after a vNNN segment, so a subfolder under t-<id>/ still deletes fine, but with a segment in front of it, it doesn’t. That can happen two ways: a reorganisation to listings/t-<id>/…, or a transformation segment on a version-less delivery URL. /upload/w_500/t-<id>/photo.jpg has the public id w_500/t-<id>/photo, which the write gate accepts today and the delete gate quietly skips as skipped. An owner removes a photo from their listing, the database stops referencing it, and the image stays on the CDN, still fetchable by anyone with the URL. That’s the one thing the delete was supposed to prevent. The repo also documents a 2026-08-08 incident with a Cloudinary cleanup that went wrong (cloudinary.ts:116), so this area has caused trouble once before.

Fix. One exported predicate that both sides call. The rule is fine, there just shouldn’t be two of it. 2–4 hours.

8. The markdown sanitiser is 25 lines of regex, and only a comment guards it — Medium

Evidence. packages/ui/src/markdown.ts:77: const rendered = marked.parse(md, { async: false });, followed by a hand-written allow-list. The constraint is on lines 69–70, as a comment:

 * HARD CONSTRAINT: `md` must always be operator-authored (a tenant config value or a blog post
 * only an operator can write) — never a visitor-supplied string. The `sanitize()` allow-list above

Scenario. This is fine today. renderMarkdown gets either tenant configuration or a blog post, and the operator writes both. But only a comment holds that constraint. There’s no type, no test and no lint rule to stop the next caller. The obvious next feature is markdown in a listing’s description. The description column exists, the owner edits it, and the natural move is to call the repo’s own markdown helper on it. From then on, text written by a self-registered advertiser goes through a tag filter whose own comment says it wasn’t built to face an attacker.

Fix. Make the input a branded type (OperatorMarkdown) that only the two places allowed to produce it can create, so a plain string fails the type check. And if visitor-written text ever ends up here, swap in a maintained sanitiser. 3–6 hours.

9. The same parseArgs in four scripts, character for character — Low

Evidence. scripts/parity-cron.mjs:38, scripts/parity-from-source.mjs:47, scripts/parity-urls.mjs:71, scripts/preview-smoke.mjs:34. All four have export function parseArgs(argv) {, and all four bodies are byte-identical.

Scenario. The four scripts are four steps of one operator runbook. parity-urls.mjs is 642 lines, and it’s the one that gets extended. When somebody adds a valueless flag (--dry-run) there, the other three keep the old behaviour, which treats the value as mandatory and exits with code 2. The operator then runs into two different flag conventions inside one runbook, and the error message doesn’t say which script follows which.

Fix. One scripts/lib/args.mjs that all four import from. 1–2 hours.

10. The honeypot convention lives in eight places, held together by comments — Low

Evidence. The hidden field is in four components (packages/ui/src/pages/contact-form.tsx:49, packages/ui/src/blocks/inquiry-form.tsx:67, packages/ui/src/pages/plan-request-form.tsx:84, packages/ui/src/pages/sign-in.tsx:68), all name="website". The check is in four separate server actions, for example apps/web/lib/actions/contact.ts:41: const honeypot = formData.get("website");.

Scenario. All four pairs are in place today. Somebody will start the fifth form by copying the existing markup, and write its server action from scratch. The field will be on the form, and the check won’t. And nothing complains: the form works, the tests are green, the honeypot just does nothing. The first sign is spam, weeks later, and nobody will connect it to three missing lines.

Fix. A <Honeypot /> component and an isBot(formData) helper from the same module that exports the field name. That way both halves live in one place, and the fifth form can’t separate them. 2–3 hours.

Fix order

#ItemHoursWhat it unblocks
2Publish in one UPDATE3–5The cleanup sees every live listing from then on; no more invisible, permanent rows
1Rate limiter into shared storage6–10The sign-in and notification limits are worth what their docs claim
4Take-down token keyed on id2–4The moderation link cannot hit the wrong listing
3An alarm on a failing block4–8A database outage is not silent; missing content is not something the customer reports
6accounts key migration2–4Google / OAuth sign-in can be added
5Owner index1–2The profile page holds up after the erdei import
7One namespace predicate2–4The Cloudinary folder layout can change
8Branded markdown type3–6Visitor-written text can go anywhere in the system
10Honeypot into one module2–3The fifth form cannot quietly ship undefended
9Shared parseArgs1–2All four runbook steps read flags the same way

Item 2 first, because it’s cheap and it’s producing rows right now that will be hard to find later. Then item 1. The Part 2 items aren’t urgent, but each one is a precondition for a specific next feature, so the time to do each is right before that feature, not after.

What this sample shows

50,000 lines in a week, and the tenant separation, the authorisation checks, the token lifecycle and the handling of secrets are all fine. That’s not luck. 169 test files against 279 source files (counting *.test.* and *.spec.* plus everything under a test directory) is a better ratio than many codebases built over years can show. Most of what I read didn’t become a finding because it was fine.

What did come out isn’t the AI’s fault. Of the two high items, one comes from code written for one machine that runs on several (item 1). The other from something that looks like one database write and is actually two (item 2). Then there are five findings (item 1 among them again) where somebody thought through a constraint or a limit and wrote it down in a comment: the in-memory rate limiter, the slug-keyed token, the namespace check, the markdown constraint, the honeypot convention. None of them is enforced by a type, a test or an index. An AI won’t find these, because none of them is broken code. Each is a written-down rule that’s correct. Reality has already caught up with one of them (item 1), and the next feature will catch up with the rest. Finding them takes somebody reading the whole thing in one sitting and asking whether what was written down is still true.

If you want the same for your codebase: AI code audit, €970, 1–2 weeks, fixed price.

If you have something to build or fix, send me a few sentences and I'll reply within two business days.

Get in touch

← All writing