Agents that do the work: in your tools, in your content, on the phone.

Agents and automation

Agents that do the work, not just the talking: they plan, call your tools, remember, run on a schedule and ask a person before anything risky.

plangoalagenttoolsmemory

STACK

  • LangChain
  • MCP
  • PostgreSQL
  • Redis
  • Tool-calling and multi-agent systems, over MCP and A2A
  • Memory. Short- and long-term, so an agent picks up where it left off.
  • Background agents. On a schedule, or when a new email, ticket, deal or deploy arrives.
  • Coding agents. They take tickets, run the tests and open pull requests for your team to approve.
  • Personal and team agents. On OpenClaw, in WhatsApp, Slack or Teams, locked down and audited.
  • Approvals and audit. Risky steps wait for a named person, and every action is logged.

WHAT WE’VE BUILT

  • An agent platform driving 100+ agents from one chat window
  • A library of 70+ tool-using agents across search, travel, finance, media and productivity
  • An SME credit-underwriting platform: KYC, litigation and adverse-news checks, financial ratios and a cited document chat, as MCP tools

WE WORK WITH

  • A2A
  • OPENCLAW
  • CLAUDE CODE
  • CODEX
Everything we build in this area: Agents and automation

Agents we build

  • Copilots and chat products inside your own app
  • Browser agents that work portals with no API
  • Research agents that deliver cited reports overnight
  • Shared-inbox agents for orders, ops and support mailboxes
  • Runbook agents that follow your SOPs and pause on sensitive steps
  • Migration agents for upgrades across many repositories
  • Morning briefings, inbox triage and meeting prep for executives

OpenClaw, deployed properly

  • Private OpenClaw on your server, reachable only through your network
  • Team agents with shared sessions and per-person permissions
  • Custom skills for your CRM, ERP, helpdesk and internal APIs
  • Hardening of existing installs: sandboxing, vetted skills, secrets, backups
  • Local-model setups that keep sensitive data on your hardware

How we build them

  • MCP servers that give every agent one governed door to your systems
  • Untrusted email and web pages read by agents that cannot act
  • Budgets per agent, loop breakers, and easy steps routed to cheaper models
  • Tracing and evals that gate every release

Content and generative media

Content that ships every day in your voice: articles, posts, ads, images, video and dubbing, in any language, approved before it goes out.

storyboardtry-on+
cataloguemodelstry-onad videoimages

STACK

  • PyTorch
  • Hugging Face
  • NVIDIA
  • Content engines. Topics planned, drafted against your sources, checked and published to your CMS.
  • SEO and AI search. Programmatic pages, and answers that AI search tools cite.
  • Your brand voice. Trained once, kept across every writer, agent and channel.
  • Video. Presenter and avatar videos, ads, product b-roll and dubbing in the speaker’s own voice.
  • Images. Product photography at catalogue scale, ad creative and virtual try-on.
  • Every language. Copy, voice-over and on-image text in Hindi and regional languages as well as English, in 22+ scripts.

WHAT WE’VE BUILT

  • A content engine for D2C brands: brand DNA from a website, then ads, posts and avatar videos in Hindi, regional languages and English
  • A multi-agent writing platform with writer, editor, proofreader and researcher agents
  • An AI video platform routing jobs across diffusion backends, with a browser editor
  • Virtual try-on on a diffusion pipeline

WE WORK WITH

  • GEMINI OMNI
  • VEO 3.1
  • KLING 3.0
  • SEEDANCE 2.0
  • RUNWAY
  • HEYGEN
  • LTX-2
  • GPT IMAGE 2.5
  • NANO BANANA 2
  • FLUX.2
  • ELEVEN V3
Everything we build in this area: Content and generative media

Words

  • Newsletters assembled from the week’s news and releases
  • Product descriptions, FAQs and help articles at catalogue scale
  • Email and WhatsApp campaigns per segment
  • Brand and claims checks on every draft before it goes out

Social, ads and audio

  • A month of social posts drafted per platform, queued for your approval
  • One webinar or podcast turned into clips, articles, threads and carousels
  • Ad variants per audience and placement: hooks, headlines, images, short video
  • Multi-shot video ads with the product and presenter kept consistent across shots
  • Personalised videos for each prospect
  • Podcasts generated from your articles and data

Voice agents

Voice agents that answer every call, start speaking before the model has finished and stop when the caller cuts in.

STTLLMTTStokensspeechfirst audio

STACK

  • Whisper
  • Kokoro TTS
  • Triton
  • Twilio
  • LiveKit
  • Deepgram
  • ElevenLabs
  • Speech-to-speech assistants and streaming pipelines
  • Turn detection and barge-in
  • Telephony. Support lines, bookings, reminders and outbound calls within the limits you set.
  • Languages. Indian and European languages, with your vocabulary and accents.
  • Your voice. A brand voice, or a cloned voice used with consent.
  • Self-hosted. Speech models on your own GPUs when calls must stay private.

WHAT WE’VE BUILT

  • Whisper into a streaming LLM into Kokoro TTS on Triton, chunked so speech starts early
  • A speech-to-speech agent over WebSockets with voice-activity detection

WE WORK WITH

  • GPT-REALTIME-2
  • GEMINI LIVE
  • ELEVEN V3
  • SCRIBE V2
  • DEEPGRAM FLUX
  • CARTESIA SONIC
  • PARAKEET
  • VOXTRAL
  • INDICCONFORMER
  • KOKORO
  • LIVEKIT AGENTS
  • PIPECAT
Everything we build in this area: Voice agents

Voice agents we build

  • Warm transfers to a person, with the call so far
  • Voice pre-screens for high-volume hiring
  • Transcripts, summaries and CRM notes after every call
  • Voice interfaces inside apps, kiosks and devices
  • Narrated reports and audio summaries

Agents for every team.

The jobs agents take off each team, each with a person approving what matters.

Sales

  • Outbound agents that research each account and draft the first touch
  • Leads enriched, scored and sent to the right owner
  • A brief on the account before every call
  • Calls logged, next steps set and stale deals flagged in the CRM
  • Buying signals turned into a daily call list
  • Proposals and RFP answers drafted from your past wins

Support

  • Chat, email and WhatsApp requests resolved by acting in your systems
  • Every ticket classified, routed and given a drafted reply
  • Answers from your help centre that show their source
  • Every conversation checked for tone, policy and resolution
  • Churn signals raised before renewal

Operations

  • Order and shipment exceptions handled before the customer asks
  • Weekly business reports posted to your team channel
  • New clients set up the moment a deal closes
  • Access requests and resets for an IT helpdesk

Finance

  • Invoices matched to purchase orders, with mismatches chased
  • Daily reconciliation that leaves only the exceptions
  • Month-end drafts and plain-language variance notes
  • Every expense claim checked against policy
  • Polite, escalating reminders for overdue invoices

HR and hiring

  • Shortlists for every open role
  • Screening with a written reason, and a person making the call
  • Interviews scheduled across panels and time zones
  • Onboarding that makes day one work
  • Policy questions answered from your handbook

Engineering

  • Every pull request reviewed for bugs, security and your conventions
  • Tests written for legacy code before a refactor
  • Alerts triaged, with a likely cause and rollback proposed
  • Docs and runbooks kept in step with the code

Data and research

  • Questions answered from your warehouse in Slack, with the query shown
  • Metric digests that explain what changed and why
  • Competitor pricing, releases and hiring watched weekly
  • Reviews, tickets and calls grouped into themes with quotes

Your own AI, trained on your data and running from a phone to a GPU cluster.

Your own AI

Models trained on your data, tuned for your jobs and small enough to run where your data lives. You own the weights.

GPU nodeslosspretrainSFTDPO

STACK

  • PyTorch
  • Hugging Face
  • NVIDIA
  • vLLM
  • Your tasks first. An evaluation set from your real work, so every step is scored.
  • Continued pretraining, SFT and tool-use training on your APIs
  • RL post-training. Preference tuning, GRPO-family RL with checkable rewards, and agentic RL in a sandbox of your tools.
  • Distillation. A small model that does your job as well as a big one, for a fraction of the cost.
  • Pretraining 1B to 10B-class models, and adapting open MoE models with hundreds of billions of parameters
  • Indian-language models with extended tokenizers
WHERE IT RUNS
On your GPU cluster or in your cloud account, so training data can stay inside your network.
WHAT YOU OWN
The trained weights or adapters, the datasets and the pipelines that made them, the evaluation suite and the training code. When the work ends, we delete our copies and confirm it in writing.

WHAT WE’VE BUILT

  • A 3B function-calling model that beat a frontier API model on the client’s own benchmark
  • A 70B open-weight model fine-tuned with SFT and DPO
  • Reasoning models post-trained with GRPO, SFT and a process reward model, on 8-GPU runs
  • GRPO post-training of function-calling models, and an RL framework for vision-language models

WE WORK WITH

  • TORCHTITAN
  • MEGATRON
  • VERL
  • TRL
  • UNSLOTH
  • DEEPSPEED

BASE MODELS

  • QWEN3.8
  • DEEPSEEK V4
  • GEMMA 4
  • GPT-OSS
  • NEMOTRON 3
  • SARVAM
  • LFM2.5
Everything we build in this area: Your own AI

Models we train

  • Domain models on your documents, formats and jargon
  • Embeddings and rerankers trained on your corpus
  • Reward models and judges that encode your experts
  • Guard models for your own policies
  • Vision-language and document models
  • Speech models for your accents and vocabulary
  • Small language models for phones and laptops

How the training runs

  • Synthetic data generated from your documents
  • Multi-node GPU training with FSDP2, Megatron and TorchTitan
  • FP8 and FP4 training on Hopper and Blackwell GPUs
  • Specialists merged into one model
  • Red-teaming and safety tuning before release
  • Releases gated on your evaluation set, retrained from real usage

Inference and on-device

Models served where your data lives: on the phone, on one private GPU or across a cluster, at a lower cost per token where quality holds.

your networkdocumentspromptsgpu serveropen weightspublic api
routercache3B int4your GPUlargelatency, cost

STACK

  • vLLM
  • SGLang
  • TensorRT-LLM
  • Triton
  • Ollama
  • NVIDIA
  • On the device. iPhone, Android, Mac and Windows, and in the browser, offline.
  • On your GPUs. Private and air-gapped deployments inside your network.
  • At cluster scale. Large mixture-of-experts models with prefill and decode on separate GPUs.
  • Each request routed to the smallest model that passes your evaluation set
  • Quantisation to fit the hardware: FP8, NVFP4, INT4, GGUF
  • Measured first. Each technique is checked against your traffic today and kept only where quality holds.

WHAT WE’VE BUILT

  • A 70B open-weight model served with vLLM
  • Private deployments of open-weight LLMs and SLMs
  • An offline desktop app running OCR models locally on Windows, macOS and Linux

WE WORK WITH

  • GEMMA 4
  • LFM2.5
  • GPT-OSS
  • QWEN3.8
  • NEMOTRON 3
  • MISTRAL SMALL 4
  • DEEPSEEK V4
  • GLM-5.3
  • NVIDIA DYNAMO
  • LLAMA.CPP
  • MLX
  • CORE AI
  • EXECUTORCH
  • LITERT-LM
Everything we build in this area: Inference and on-device

On the device

  • Offline assistants and field tools on phones
  • Apps on Apple’s on-device models and Gemini Nano on Android
  • Redaction and search on the device, before anything leaves it
  • Edge boxes for cameras, kiosks and robots

On servers and clusters

  • GPU clusters set up with Slurm or Kubernetes
  • One base model with an adapter per team
  • KV cache tiered across GPU, CPU and SSD
  • Speculative decoding with trained draft models
  • Autoscaling and cost-per-token dashboards
  • OpenAI-compatible endpoints, so your apps switch without a rewrite
  • Guard models in front of every endpoint

Beyond LLMs

An LLM only where language is the job. For a label, a forecast or a finding in an image, a small model trained on your data is faster and easier to audit.

LLMnot needed+churnp 0.943 ms

STACK

  • scikit-learn
  • pandas
  • Optuna
  • SHAP
  • AutoGluon
  • Calibrated classifiers. Routing, triage and guardrails in milliseconds.
  • Forecasting: demand, pricing, inventory and hierarchical time series
  • Computer vision. Detection, segmentation and inspection, trained on your images.
  • Recommenders and ranking for feeds, catalogues and search
  • Gradient-boosted ensembles and tabular foundation models
  • World models, state-space and diffusion language models, where they fit

WHAT WE’VE BUILT

  • A demand-forecasting platform with 150+ engineered features, tuned with Optuna and explained with SHAP
  • A segmentation model trained from scratch that finds tampering in scanned identity documents
  • A microscopy classifier for materials research, validated with grouped cross-validation

WE WORK WITH

  • XGBOOST
  • LIGHTGBM
  • CATBOOST
  • TABPFN
  • CHRONOS-2
  • TIMESFM
  • SAM 3
  • V-JEPA 2
  • NVIDIA COSMOS
  • MAMBA
  • MERCURY 2.5
Everything we build in this area: Beyond LLMs

Models we build

  • Fraud, risk and churn scores that explain themselves
  • Anomaly detection on transactions, sensors and metrics
  • Quant research with cost-aware backtesting
  • Honest error bars on every forecast

The software your business runs on, and AI-built apps made safe to grow.

Apps and business software

Full products and the software your business runs on, with AI built in where it helps.

webmobiledesktopapideploylive

STACK

  • Next.js
  • React
  • Vite
  • Node.js
  • FastAPI
  • Fastify
  • Supabase
  • BullMQ
  • React Native
  • Expo
  • Electron
  • Web and SaaS. Frontend, backend, auth, payments, multi-tenant accounts and dashboards.
  • Mobile. iOS and Android, with our iOS apps live on the App Store.
  • Desktop. Mac, Windows and Linux apps, native or cross-platform, and developer tools.
  • Business software. CRMs, ERPs, inventory and order management, and project trackers you can self-host.
  • Platforms. Learning platforms, client portals, marketplaces and admin consoles.
  • Internal tools. Dashboards, approvals and workflows that replace spreadsheets.

WHAT WE’VE BUILT

  • An edtech platform launching to 50k students, with lesson planning, AI tutoring and generated simulations
  • A mobile AI app with 30k users
  • An AI-native IDE on Electron and Monaco
  • A native macOS command centre that supervises parallel coding-agent sessions
  • A desktop app for medical-document OCR, packaged for Windows, macOS and Linux
Everything we build in this area: Apps and business software

Business software we build

  • CRMs with pipelines, lead scoring, email sync and an agent that keeps them clean
  • ERPs and inventory: stock, orders, purchasing, warehouses and invoices
  • Project trackers you can self-host: boards, sprints, time logs and roles
  • HR and workforce tools: team dashboards, check-ins, goals and reviews
  • Booking, scheduling and field-service apps

Apps we build

  • Native macOS apps in Swift and SwiftUI
  • Desktop apps for Mac, Windows and Linux on Electron or Tauri
  • Offline-first apps that sync when they can
  • Browser extensions and developer tools

Fixing AI-built apps

We find what will break in production and fix it, including apps built with AI tools that now have real users.

beforeaftertestsall passing

STACK

  • Playwright
  • GitHub Actions
  • Code audits
  • Bug and security passes
  • Vibe-code cleanup. Apps from Lovable, Bolt, Cursor, Replit, v0 or Claude Code made safe, tested and maintainable, usually without a rewrite.
  • Test automation: unit, integration and end-to-end
  • Exposed secrets. Reported the same day, ahead of the written report.

WHAT WE’VE BUILT

  • Our own autonomous web-testing platform, run on top of your test suites
Everything we build in this area: Fixing AI-built apps

What we fix

  • Auth and database access rules
  • Slow queries, repeated calls and cold starts
  • CI/CD with previews and rollbacks
  • Moves off no-code backends when you outgrow them
  • Handover notes so your team can keep going

Search over your documents, scored evals, and the cloud underneath.

Retrieval, memory and document AI

Assistants that know your documents and remember past work, and pipelines that turn paperwork into clean, checked data.

documentsindexqueryanswer123memoryocr · extraction · knowledge graph

STACK

  • pgvector
  • Qdrant
  • Neo4j
  • MongoDB
  • Redis
  • Hybrid retrieval over documents, tables and images, reranked
  • Knowledge graphs
  • OCR, extraction and tables linked across Word and Excel
  • Document AI. Invoices, contracts, specifications, claims and clinical notes read and checked.
  • Contract pipelines with compliance checks
  • Answers that cite the page they came from

WHAT WE’VE BUILT

  • Hybrid search over enterprise documents with Qdrant, Neo4j and GPU-served embeddings
  • An infrastructure copilot that took solution documents from days to minutes
  • A medical coding engine that assigns diagnosis and procedure codes from clinical notes

WE WORK WITH

  • QWEN3-EMBEDDING
  • BGE-M3
  • PADDLEOCR-VL
  • DOCLING
Everything we build in this area: Retrieval, memory and document AI

Document pipelines we build

  • Insurance policies turned into section-mapped knowledge
  • Financial reports checked for numbers, units, currency and headings
  • A house style learned from reference documents and applied to every Word file
  • Annual reports matched page by page to their XBRL filing
  • Workbook versions compared cell by cell, without false alarms from inserted rows
  • Spreadsheets you can question in plain language
  • Enterprise search that respects each user’s permissions
  • Long-term memory for assistants

Evaluation, search and safety

Scored harnesses that show whether a change made the model better, safety checks before release, and search that answers with citations.

eval setabcdllm-as-judge · cost per runsearch1212plan · retrieve · rerank · cite

STACK

  • Qdrant
  • pgvector
  • PostgreSQL
  • Evaluation harnesses, rubric scoring and LLM-as-judge
  • Leaderboards and cost tracking
  • Before release. Red-teaming, guard models and regression suites.
  • In production. Tracing, quality monitoring and drift alerts.
  • Search with query planning, retrieval and reranking
  • Answers that cite their sources

WHAT WE’VE BUILT

  • An evaluation and leaderboard system with LLM-as-judge and cost tracking
  • Perplexity-style search with cited answers
  • A published multi-agent evaluation framework for language and vision-language outputs
  • A Swiss-tournament Elo leaderboard that ranks models on maths and penalises hallucination

WE WORK WITH

  • INSPECT AI
  • PROMPTFOO
  • RAGAS
  • LANGFUSE
  • LLAMA GUARD 4
  • QWEN3GUARD

Cloud, DevOps and architecture

System architecture and deployment on AWS, GCP, Azure or your own data centre.

ci/cdcommitbuildtestdeployclusternodenodegpu nodemetrics · logs · traces

STACK

  • AWS
  • Google Cloud
  • Azure
  • Docker
  • Kubernetes
  • Terraform
  • GitHub Actions
  • Containers and Kubernetes
  • Infrastructure as code
  • GPU clusters and serving
  • On-prem and air-gapped networks
  • CI/CD
  • Observability

WHAT WE’VE BUILT

  • Containerised FastAPI services and GPU serving on Triton and vLLM
  • AI services on AWS with Bedrock and Polly
  • Terraform, Kubernetes and Helm for a retrieval and model stack

Built for the way your industry works.

The same engineering, pointed at the paperwork, rules and customers of your field.

Healthcare

  • Coding that holds up to audit, straight from the clinical note
  • Dictation turned into structured clinical notes
  • Patient guidance on WhatsApp that knows when to hand over to a person
  • Records digitised on the desktop, with nothing leaving the machine

Finance and insurance

  • Credit files assembled for the analyst: checks, news and ratios on one screen
  • Policy and loan documents an underwriter can question
  • Reports and filings checked before they reach a regulator
  • Advisor desks with drill-down client reporting

Education

  • One school platform for classes, assessment and attendance, with a tutor for doubts
  • Lessons, quizzes and simulations generated from your own course material
  • Study workspaces with spaced repetition, quiz battles and podcast recaps

Retail and e-commerce

  • Catalogue content for every SKU, in every language
  • Demand, price and stock forecasts per store
  • Search that understands what shoppers mean
  • Returns and refunds handled in chat, inside your order system

Media and marketing

  • Newsrooms with research, drafting and fact-checking agents
  • Series and podcasts produced from your archive
  • Campaign reports that explain what worked

HR and workforce

  • Career and skills graphs from CVs and job descriptions
  • AI-disruption risk scored for every role
  • Skills and labour-demand forecasting

Manufacturing and industrial

  • Control-room displays migrated between SCADA platforms and repaired automatically
  • Quality inspection from camera feeds, trained on your own defects
  • Offline apps and local models for the plant floor
  • Copilots over equipment manuals and infrastructure documents

Build with us, join your team, or check the code first.

We agree the scope on a call.

  • Project

    • First slice in 2 to 3 weeks
    SHAPE
    A defined build with an end date.
    FITS WHEN
    You have one defined problem and a date.
  • Embedded engineering

    • YOUR REPOSITORY
    • YOUR PROCESS
    SHAPE
    We join your engineering team and work to your standards: your repository, CI, linters and code owners.
    FITS WHEN
    You have a team and need to ship faster now.
  • Ongoing engineering

    • INTAKE BOARD
    • WEEKLY DEMO
    SHAPE
    Steady engineering on a live product, with bugs and features dropped on one intake board.
    FITS WHEN
    A live product needs steady engineering.
  • Code audit

    • EVERY REPOSITORY IN SCOPE
    SHAPE
    Where the 48-hour report reads one repository, an audit covers every repository in scope and ends with a written plan for what to fix first.
    FITS WHEN
    You need to know what is in the code before you build on it.
  • Vibe-code cleanup

    • LOVABLE
    • BOLT
    • CURSOR
    • REPLIT
    • V0
    • CLAUDE CODE
    SHAPE
    An AI-built app made safe, tested and maintainable, usually without a rewrite.
    FITS WHEN
    An AI tool built most of it and nobody can change it safely.
  • Consulting

    • ARCHITECTURE
    • MODELS
    • DELIVERY
    SHAPE
    A second opinion on architecture, model choice and delivery for AI systems.
    FITS WHEN
    A decision is expensive to reverse.

How a cleanup runs

  1. 01 Read

    We read the codebase and write down what is sound, what is risky and what is broken.

  2. 02 Stabilise

    Secrets out of the repository, the crashes fixed, tests on the paths that matter.

  3. 03 Refactor

    Structure a new engineer can follow, changed in small, tested steps.

  4. 04 Guardrails and handover

    Automated checks switched on in your repository, documentation, and our access removed within seven days.

Most vibe-coded apps need surgery. Few need a rewrite, and we say which in writing.

We work on branches of your repository, never main, with your branch protection left on.

EVERY ENGAGEMENT

  • A demo every week
  • Defects corrected if reported within 30 days of handover
  • A README and runbooks at handover

Not sure where to start

Start with one repository. Give us read-only access and within 48 hours you get a written report: findings by severity, what we would build or fix first, and which of these ways of working fits. It stands on its own, with no obligation to go further.

Request a code review

Frontier APIs and open weights, tried on your task first.

MODELS WE WORK WITHAS OF SEP 2026

  • GPT-6
  • Claude Opus 5.5
  • Claude Fable 5.1
  • Claude Sonnet 5
  • Gemini 3.1 Pro
  • Gemini 3.8 Flash
  • Grok 4.6
  • DeepSeek V4
  • Qwen3.8
  • Kimi K3
  • GLM-5.3
  • Mistral Large 3
  • Gemma 4
  • Nemotron 3
  • Muse Glimmer
  • LFM2.5

Any other model, including one you trained, plugs in through a custom endpoint.

Bring the problem and we scope it with you on a call.

The repository helps if you have one. We agree the scope and the first milestone on the call.

Experience
Two years serving customers, from startups to enterprise teams.
Shipped
30+ production AI systems.