How a request becomes a deploy.
We agree the scope on a call and write it down. Then agents build in parallel, an engineer approves each change, and you see a demo every week.
Nothing starts until the scope is in writing.
01CALL
A scoping call about the problem, the users and the date that matters.02WRITTEN SCOPE
A statement of work names the problem, the scope, the engineers, the timeframes and the fee.03ACCESS
The least access that will do the job, on a named account of our own.04FIRST SLICE
A first working slice, then a demo every week and a preview link for work in progress.
Six stages, with an engineer at the centre.
Our engineers set the plan, direct the agents, read each change and approve it.
- Tests on the paths that carry money or data
- Quality and security checks that have to pass
- An engineer who can explain it before it merges
Intake You drop bugs and feature requests on one intake board.
Parallel agents Each works in its own isolated environment and branch.
Automated checks Tests, type checks, lint and security scans on every pull request.
Preview deploy Every branch gets a preview link on our domain.
Code review An engineer reads each change, with bug and security passes.
Merge and deploy Every merge builds, tests and deploys.
Many agents at once, each in its own worktree and branch.
- Our tooling turns each request into a PRD and a sprint plan.
- Developer and QA agents run in parallel, each in its own isolated environment, git worktree and branch.
- Guardrails on every run, and no two agents share a working copy.
- Everything comes back as a pull request on your repository. We write to branches, never to main, and your branch protection stays on.
Five questions each pull request answers before it merges.
We approve the pull request; your branch protection decides who merges it.
Does it do what the task asked, and nothing else?
Would it survive bad input, a slow network and a second user?
Is anything generated that nobody can explain?
Is the test testing the behaviour, or just passing?
Could the next engineer change it without breaking something else?
Bugs and security get a dedicated pass of their own, the same one we run in a code review.
Tests at four layers, run by CI.
Unit
Fast tests on the logic, run on every pull request.
Integration
Services, the database and the APIs they call, tested together.
End-to-end
Real user flows driven through the product.
Autonomous web testing
Our in-house platform, run on top of the three suites above.
A preview for every branch, and a deploy for every merge.
| Status | STEP | RESULT |
|---|---|---|
| Agent draft | agent-3/fix-n-plus-one opened pull request #214 | |
| Tests | 212 passed, 0 failed | |
| Type check | 0 errors | |
| Lint | 0 problems | |
| Dependency audit | clean | |
| Secrets scan | no secrets found | |
| Static security analysis | 1 finding: query built from user input, fixed in 3f9a2c1 | |
| Preview deploy | preview for PR #214 is live | |
| Code review | approved with 2 comments | |
| Merge | merged into main under branch protection | |
| Deploy | build, tests and deploy passed |
What has to pass before an engineer opens the pull request.
| CHECK | WHAT IT CATCHES |
|---|---|
| Tests | Behaviour that changed when it should not have |
| Type checks | Values that cannot be what the code assumes |
| Lint | Patterns known to cause bugs, and drift from the codebase’s style |
| Dependency audit | Packages with known vulnerabilities |
| Secrets scan | Keys, tokens and passwords committed to the code |
| Static security analysis | Injection, unsafe input handling and risky calls |
One loop from raw data to a model in production.
Training, fine-tuning and classical ML all follow it.
Data audit and a baseline
We read the data before we model it: where it comes from, how it is labelled and what is missing. Then we set a baseline with the simplest approach that could work, so every later result has something to beat.
An evaluation set built with you
Real cases chosen with your team and held out from training, with leakage checks so no test example, or a near copy of one, reaches the training data.
Tracked runs
Each experiment is logged with its data version, code, settings and scores, so any result can be traced and run again.
Ablations
We take changes out one at a time to find what actually moved the score, and keep only what earns its cost.
A promotion gate
A candidate goes live only when it beats the baseline on the held-out set, within the latency and cost you need.
Monitoring after launch
Once it is live, we watch the inputs for drift, score fresh samples of the outputs and track the cost of each request.
Bugs and requests on one board, shipped the same way.
Drop
Bugs, issues and feature requests, all on one board.
Triage
We sort them and agree with you what comes first.
Build
Agents build each one in its own isolated branch.
Ship
An engineer approves it, your branch protection governs the merge, then CI/CD ships it.
It worked in the demo. We make it hold up.
Every fix breaks something else.
Keys are sitting in the repository.
The AI assistant keeps looping on the same bug.
Nobody on the team can say how it works.
Built it with Lovable, Bolt, Cursor, Replit, v0 or Claude Code, and now it has real users?
We make it safe, tested and maintainable.
The in-house tools that run our agents and test their work.
Autonomous web-testing platform
The autonomous layer on top of our unit, integration and end-to-end suites.
Multi-agent engineering platform
Turns a prompt into a PRD, a sprint plan and parallel developer and QA agents on separate branches.
AI coding workspace
Runs many coding agents at once, each in its own git worktree, with a shared backlog and guardrails.
AI-native IDE
Our own desktop editor for working alongside coding agents, on Electron and Monaco.
See the system on your own code.
One repository, read-only, and a written review of it within 48 hours.
- Experience
- Two years serving customers, from startups to enterprise teams.
- Shipped
- 30+ production AI systems.