The 90 firms that fix AI-generated code
What the 90 firms say breaks in AI-built apps, how they package the work, what none of them sells, and a checklist to use on any of them, including us.
Chapters · 07
In September 2026 we went looking for companies that sell cleanup of AI-generated code: apps built with Lovable, Bolt, Cursor, Replit, v0, Claude Code and similar tools, by people who then got stuck. We found 90. This post is what those 90 say they fix, how they package the work, and what we think a buyer should demand from any of them.
Why this category exists
In our reading of the 90 service descriptions, the typical caller in 2026 is a business owner, a product manager or a COO who built something real with an AI tool and then met real users, a payment that fails, or a security email from a stranger.
The demand is visible in public. On r/ExperiencedDevs, the thread “Getting more calls to fix ai generated codebases than actual new builds lately” reached 405 upvotes and 102 comments; one comment: “Cleanup contracts are gonna be a whole market segment.” On r/SaaS, an MVP dev shop said it lost half its pipeline to Claude Code in 2025; of the prospects who tried it instead, about a third shipped, a third broke in production, and three came back for cleanups. On r/vibecoding, a post pointing to a vibe-code fix service, asking builders about getting their “mess” fixed, drew 1,086 upvotes and 125 comments.
The failure rate is measurable. Veracode’s 2025 report found that AI-generated code introduced security flaws in 45% of its tests. How Bad Is It?, a public URL scanner for Lovable, Bolt and v0 apps, publishes a running counter: 4,271 apps reviewed, median score 38 out of 100.
45%
of Veracode tests where AI-generated code introduced a security flaw. Source: Veracode, 202538
median score out of 100 across 4,271 Lovable, Bolt and v0 apps scanned. Source: How Bad Is It?
Who the 90 are
| Type of firm | Count |
|---|---|
| Established dev agencies that added a cleanup service line | 48 |
| Human audit specialists | 13 |
| Agencies formed for rescue work | 10 |
| Automated scanners | 10 |
| Solo fixers and freelance studios | 6 |
| Marketplaces and expert networks | 3 |
Twenty are in the USA, 18 in Europe outside the UK, 7 in India, 6 in the UK, 4 in Canada, 11 elsewhere, and 24 give no location. Size is mostly undisclosed (53 of 90); of the rest, 6 have 250 or more people, 10 are mid-sized and 21 are small, micro or one-person.
It is a young and thin market. We could open 75 of the 90 websites; 15 are known only from directory listings, and three domains (Vibe App Rescue, VibeCheckAudit, VibeFixLab) no longer resolve. Fifty-five name an AI builder in their service copy; Lovable leads (50), then Cursor (46), Bolt (44) and Replit (41).
What they say breaks
We counted which failure classes each firm’s service description names.
| Failure class named in the service description | Entries (of 90, scanners included) |
|---|---|
| Security in general (vulnerabilities, OWASP, hardening) | 48 |
| Missing or broken tests | 25 |
| Performance and scalability | 25 |
| No deploy pipeline, CI/CD or hosting setup | 23 |
| Authentication and access control (including Supabase RLS) | 18 |
| Data model or database problems | 15 |
| Secrets and API keys in code | 10 |
| Monitoring and observability | 9 |
| Documentation | 9 |
| Payments and Stripe | 8 |
| Duplicated logic and coupling | 2 |
| Hallucinated APIs | 1 |
| Runaway AI API spend | 0 |
Security is the sales pitch. Over half the market names it, and seven of the ten automated scanners are security scanners: Supabase row-level-security probes, leaked-key detection, public buckets. Beesoul, a US-listed audit shop (figures from its directory listing), reports that 10.3% of the Lovable apps it audited had critical RLS vulnerabilities, and that a typical MVP has 8 to 14 findings. Sidekick Interactive’s case study is typical of the genre: exposed API keys, database holes and broken payments in one app.
Structure is under-sold. Only two of the 90 entries, one of them a scanner, mention duplicated logic or unsafe coupling by name, though that is what makes an AI-generated codebase slow and risky to change. A security patch is usually local; untangling several copies of the same auth check, each slightly different, is where the work goes.
Nobody sells a spending cap. Not one of the 90 names uncontrolled AI API usage as something they fix; the two that mention “cost” at all mean something else. Yet on r/cursor, a thread about a Cursor bill that spiked within one hour, after a PM asked the agent to tag 87 tasks, reached 228 upvotes. An app that calls a model on every request with no budget, rate limit or per-user cap is a bill waiting to happen.
How they package it
The 80 human-service firms (the 90 minus the 10 scanners) use four patterns, often more than one at once.
| Packaging pattern | Firms (of 80) |
|---|---|
| Audit, assessment or review named as the first step | 61 |
| Audit or assessment promised within 48 hours | 10 |
| Work described in sprints | 9 |
| Rebuild or re-architecture offered as one option | 9 |
| Explicitly against rebuilding | 5 |
| Retainer, maintenance or monitoring mentioned | 16 |
Audit first. Valletta Software (Malta) is the cleanest version: a senior engineer reviews the repository across eight areas, delivers the audit in 48 hours with a debrief call, then scopes the cleanup, typically six to eight weeks. MetaCTO runs a 48-hour code audit, then promises production-ready in 30 days, with weekly demos.
Fixed-scope rescue. AxonBuild sells a 10-working-day production-engineering sprint covering auth, secrets, payments, CI/CD and monitoring, ending in a verified handover with two weeks of defect cover and 30 days of async support. Relux Works (Armenia) runs a one-week audit, then a two-to-three-week stabilisation sprint, then what it calls “de-vibe-coding”: moving the app to a production architecture.
Rebuild. Usually offered for the broken parts only, and five firms say outright that they avoid it. Bitnoise (Poznań) shows the whole sequence: audit in 3 to 7 business days, stabilisation sprint 2 to 4 weeks, re-architecture 6 to 12 weeks.
Retainer. Vibe-Audit (Barcelona, a one-person shop) runs a pre-launch security sweep delivered in 48 hours, done-for-you fixes, and then ongoing security monitoring.
Side by side, the market has settled on one sequence: read the code, stabilise what is dangerous, refactor or rebuild what will not hold, then watch it.
How cleanup is sold
- Read the code: audit or review first
- Stabilise: what is dangerous now
- Refactor or rebuild: what will not hold
- Watch it: retainer or monitoring
What a good cleanup engagement includes
Use this on anyone you are considering, including us.
- A written review before any commitment, with findings ranked by severity and a verdict: ship, patch, refactor or rewrite. If the verdict is “not worth fixing”, they should say so.
- An experienced engineer who has actually read the code. A scanner finds leaked keys; it does not find the auth check that is bypassed on one route.
- Security tested against the running app, not only the repository: authentication, per-row authorisation, secrets rotated rather than just deleted from git, public storage buckets closed.
- Characterisation tests written before refactoring, so current behaviour is pinned and the cleanup cannot silently change it.
- Duplicated logic collapsed and the data model reviewed, because that decides whether the next feature takes a day or a month.
- A deploy pipeline and monitoring at the end: staging, CI, error tracking, backups. Working on one person’s laptop does not count as done.
- AI API spend bounded: budgets, rate limits, per-user caps and an alert. Nobody in the 90 sells this; ask for it anyway.
- Handover: documentation, and you holding the repository, the cloud accounts and the keys. You own the code from day one.
- A fixed scope with a definition of done and a defect-cover window afterwards.
How we do it
Dravin AI is an AI software company. We build AI systems and software, and we clean up AI-built apps. Against the list above, this is what we commit to:
- A written report first. Read-only access to one repository, and within 48 hours a written report with findings ranked by severity: the file, the risk and the fix, and what to leave alone. If nothing needs fixing, it says so. There is no obligation to go further.
- Secrets. A secret we find in your repository is reported the same day, ahead of the report, and rotated within 24 hours.
- Auth and data access. Every route’s auth check, database access policies included, is traced in the code. Testing the running app is agreed on the call, because it needs access beyond the repository.
- Tests before refactoring. Tests on the paths that make money, run on every pull request, and structure changed in small, tested steps.
- Deploys. From the repository, with a rollback that has been tried.
- AI API spend. Timeouts, budgets and retry caps on every model call.
- Ownership. The code is yours from the first commit, we work on branches, never main, and our access is removed within seven days of handover, confirmed in writing.
- Defect cover. Anything that misses the statement of work, reported in writing within 30 days of handover, is corrected under it.
An engagement starts with a call about what you built, what it does in production and what worries you. If that is an app that got you to traction, book one or request a code review.
Method
Every count comes from the cleanup list in our research workbook: 90 firms, each with a segment, a region, a size band, a description of what it does, its delivery model and its stated turnaround. We read the list on 22 September 2026. The workbook itself is not published; these are the counting rules.
- Type, region, size: counts of each firm’s segment, region and size band. “Small, micro or one-person” is Small (10 to 49) plus Micro (2 to 9) plus Solo.
- Reachability: whether we could fetch the firm’s site or know it only from a directory listing, and whether its domain still resolves.
- Tool mentions: a case-insensitive search of each firm’s description for each builder’s name.
- Failure classes: a case-insensitive keyword search of the description only. Security: secur, vuln, OWASP, pentest. Tests: test, QA; Sherlock Forensics matches only through “pentest” and is excluded, so 25 rather than the mechanical 26. Performance: perf, scalab, bottleneck. Deploy: CI/CD, CI, deploy, pipeline, hosting, DevOps. Auth: auth, access control, RLS. Data model: data model, database, DB, schema, persistence, RLS. Secrets: secret, key(s), credential. Monitoring: monitor, observab, Sentry. Documentation: doc, docs, document or documentation as whole words. Payments: Stripe, payment. Duplication: duplic, coupling, dead code, untangle. Hallucination: hallucinat. AI spend: cost, spend, token, bill (the two hits, *instinctools’ “cost-benefit” and Vibe Code Rescue’s “perf & cost”, are not API spend). Scanner types: the descriptions of the 10 automated scanners.
- Packaging: over the 80 firms that are not automated scanners, searching the description, the delivery model and the stated turnaround unless a rule says otherwise. Audit-first: audit, assess, review, scan, diagnos or scorecard. 48 hours: the turnaround promises an audit, assessment, scan or check result in 48 hours or less; fix times and reply times excluded. The ten: MetaCTO, Pfaff Digital, Sonder, VibeAudits, Mitrix Technology, Valletta Software, ShipClarity, Vibe-Audit, VibeCodeBlue and one solo fixer. Sprints: “sprint” anywhere in those three fields. Rebuild offered: rebuild, re-architect, rewrite or full build (11 firms), minus the two of them that also say “without rebuild”, “no rewrites”, “not rebuild”, “diagnosis-over-rebuild” or “without starting from scratch”. Against rebuilding: one of those five phrases in those three fields, the firm’s hero line or the workbook’s notes on how each firm positions itself (5 firms). Retainer: retainer, maintenance, ongoing, monitoring, subscription, defect cover or post-rescue support (16 firms).
- Named offers, turnarounds and claims are quoted from what each firm publishes, unaudited.
- Reddit and studies: r/ExperiencedDevs (405 upvotes, 102 comments), r/SaaS (65 upvotes, 66 comments) and r/cursor (228 upvotes) from our AI-agency landscape research (48 searches, 2,837 posts, 17 September 2026); r/vibecoding (1,086 upvotes, 125 comments, 24 December 2025) from a separate Reddit sweep for the workbook; Veracode’s 45% from the same landscape research.