<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Dravin AI · Blog</title><description>Long reads from Dravin AI on building with current models and agents, fixing AI-generated code and the AI services market, with the data behind them.</description><link>https://dravin.ai/blog/</link><language>en-gb</language><image><url>https://dravin.ai/apple-touch-icon.png</url><title>Dravin AI · Blog</title><link>https://dravin.ai/blog/</link></image><lastBuildDate>Tue, 22 Sep 2026 18:30:00 GMT</lastBuildDate><atom:link href="https://dravin.ai/rss.xml" rel="self" type="application/rss+xml"/><item><title>More users per GPU: a playbook for self-hosted LLM serving</title><link>https://dravin.ai/blog/raising-llm-serving-throughput/</link><guid isPermaLink="true">https://dravin.ai/blog/raising-llm-serving-throughput/</guid><description>What to measure first, then the levers that raise LLM serving throughput on your own GPUs, in the order we pull them and what each one trades away.</description><pubDate>Tue, 22 Sep 2026 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Prefill keeps the GPU’s arithmetic busy while decode waits on memory, so most of the work is giving decode a bigger batch without letting prefill stall it.&lt;/li&gt;&lt;li&gt;Measure goodput first: the request rate each GPU serves while a target share of requests meet their latency limits.&lt;/li&gt;&lt;li&gt;Pull the levers in order of cost and risk: engine settings and prompt layout, then quantisation and speculative decoding, then parallelism and disaggregation.&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;A model that answers well in a notebook can be too slow and too expensive once hundreds of people use it at once. Before adding GPUs, we work through engine settings, prompt layout and a few serving techniques in a fixed order, and check each change against numbers we trust.&lt;/p&gt;
&lt;p&gt;One fact sits under everything here. Prefill, where the model reads the prompt, keeps the GPU’s arithmetic busy &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[1]&lt;/a&gt;, &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[2]&lt;/a&gt;. Decode, where it writes one token at a time per request, uses little of that arithmetic and mostly waits on memory bandwidth &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[3]&lt;/a&gt;. Most throughput work comes down to giving decode a bigger batch while stopping prefill from stalling it.&lt;/p&gt;
&lt;p&gt;We serve open-weight models privately, quantised, on vLLM. Versions are current as of 23 September 2026: vLLM 0.30.0, SGLang 0.5.20 and TensorRT-LLM 1.2.1, each the latest stable release &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[4]&lt;/a&gt;. Flags and metric names below are vLLM’s unless we name another engine.&lt;/p&gt;
&lt;p&gt;The order follows cost and risk. Batching, KV cache sizing, &lt;a href=&quot;https://dravin.ai/services/#inference&quot;&gt;prefix caching&lt;/a&gt; and chunk size are configuration on a current engine, and vLLM 0.30.0 turns most of them on by default, so the work is sizing and checking them. Quantisation and speculative decoding change the arithmetic the model runs, so they need your own evals and your own load. Parallelism and disaggregation change the topology, so they come last.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Lever&lt;/th&gt;
&lt;th&gt;What it buys&lt;/th&gt;
&lt;th&gt;What it costs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Batch and KV cache sizing&lt;/td&gt;
&lt;td&gt;More requests in flight&lt;/td&gt;
&lt;td&gt;Per-user latency rises with the batch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefix caching and routing&lt;/td&gt;
&lt;td&gt;Skipped prefill for shared prompts&lt;/td&gt;
&lt;td&gt;Cache memory, and uneven load across replicas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chunk size&lt;/td&gt;
&lt;td&gt;Steady streams while long prompts load&lt;/td&gt;
&lt;td&gt;A later first token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantisation&lt;/td&gt;
&lt;td&gt;Fewer bytes per token, room for more requests&lt;/td&gt;
&lt;td&gt;Accuracy risk; formats tied to GPU generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speculative decoding&lt;/td&gt;
&lt;td&gt;Faster streams at low load&lt;/td&gt;
&lt;td&gt;Spare compute; the gain fades as the batch fills&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Admission control&lt;/td&gt;
&lt;td&gt;Latency targets held under load&lt;/td&gt;
&lt;td&gt;Some requests deferred or rejected&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Replicas and expert parallelism&lt;/td&gt;
&lt;td&gt;Capacity beyond one GPU group&lt;/td&gt;
&lt;td&gt;More GPUs and fast interconnect&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disaggregation&lt;/td&gt;
&lt;td&gt;Separate control of TTFT and ITL&lt;/td&gt;
&lt;td&gt;Two pools to size, RDMA, experimental support&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&quot;measure-before-you-tune&quot;&gt;Measure before you tune&lt;/h2&gt;
&lt;p&gt;Four numbers describe a serving system.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;What it measures&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TTFT&lt;/td&gt;
&lt;td&gt;Time from sending a request to receiving the first token &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[5]&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;How long a user stares at a blank screen&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ITL and TPOT&lt;/td&gt;
&lt;td&gt;The gap between streamed outputs (ITL), or a request’s decode time spread over its output tokens (TPOT) &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[6]&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;How fast the answer streams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens per second&lt;/td&gt;
&lt;td&gt;Output tokens per second, summed over all concurrent requests &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[5]&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;What the hardware is worth to you&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Goodput&lt;/td&gt;
&lt;td&gt;The highest request rate per GPU served while a target share of requests meet their latency limits &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[7]&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;The number to optimise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Tools disagree on whether ITL counts the first token, so NVIDIA’s guide says to “compare results only when definitions align” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[5]&lt;/a&gt;. Pick one tool and keep it.&lt;/p&gt;
&lt;p&gt;DistServe defines goodput as “the maximum request rate that can be served adhering to the SLO attainment goal (say, 90%) for each GPU provisioned” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[7]&lt;/a&gt;, where the SLO (service-level objective) is the latency limit you commit to. It leads the dashboard because raw tokens per second can keep rising while users wait longer. System throughput and per-user speed pull against each other, and most levers below trade one for the other.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;vllm bench serve&lt;/code&gt; measures all of this: set &lt;code&gt;--request-rate&lt;/code&gt; and &lt;code&gt;--burstiness&lt;/code&gt; to your arrival pattern, cap load with &lt;code&gt;--max-concurrency&lt;/code&gt;, and pass your TTFT and TPOT limits in milliseconds to &lt;code&gt;--goodput&lt;/code&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[6]&lt;/a&gt;. Report p50 and p99 for each; averages hide the slow tail. Use your own prompt lengths, output lengths and shared prefixes. Published benchmarks show which levers are worth trying; your own traffic shows how much each one gives.&lt;/p&gt;
&lt;p&gt;In production, watch &lt;code&gt;vllm:num_requests_waiting&lt;/code&gt; for queueing, &lt;code&gt;vllm:kv_cache_usage_perc&lt;/code&gt; and &lt;code&gt;vllm:num_preemptions&lt;/code&gt; for memory pressure, and &lt;code&gt;vllm:prefix_cache_hits&lt;/code&gt; against &lt;code&gt;vllm:prefix_cache_queries&lt;/code&gt; for cache reuse &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[8]&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;get-the-batch-and-the-kv-cache-right&quot;&gt;Get the batch and the KV cache right&lt;/h2&gt;
&lt;p&gt;The first lever is the engine. Early servers ran a batch until its longest request finished, so new requests queued behind it. Orca introduced “iteration-level scheduling”: the scheduler decides after every decoding step, so finished requests leave and new ones join at once. On a 175B model the paper reports 36.9 times the throughput of NVIDIA FasterTransformer at the same latency &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[9]&lt;/a&gt;. Every current engine does this, under the names continuous batching or in-flight batching &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[10]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The next constraint is the KV cache, the stored keys and values for every token of every running request. For a 13B model the PagedAttention paper works it out at 800 KB per token in FP16, so a single 2,048-token request can hold 1.6 GB &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[11]&lt;/a&gt;. The general form is 2 × layers × KV heads × head dimension × bytes per element, per token. It grows linearly with context length and concurrency, so the KV cache sets the batch ceiling. Grouped-query attention, where several query heads share one key-value head, shrinks the KV heads term &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[12]&lt;/a&gt;; a model’s config lists its KV head count, so check it before you choose the model.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;GPU memory as 48 fixed-size KV cache blocks. Static batch with one slab per request: four requests each reserve a 12-block slab for their maximum length and fill part of it; B has finished but its slab stays held until the batch ends, so queued requests wait and 17 of 48 blocks hold tokens. Continuous batching with a paged cache: requests take small scattered blocks step by step; B&apos;s blocks free after step 3, new request E takes all of them at step 4, and by step 6 41 of 48 blocks hold tokens.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;A slab reserves a request’s maximum length up front and holds it until the request, or the batch, ends. A paged cache adds blocks as tokens arrive and returns them the moment a request finishes.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Before PagedAttention, only 20.4 to 38.2% of KV cache memory held actual token states. Storing the cache in small pages, as an operating system pages virtual memory, brought waste close to zero and gave 2 to 4 times the throughput of FasterTransformer and Orca at the same latency &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[11]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;With a current engine, the work is sizing. vLLM 0.30.0 sets &lt;code&gt;gpu_memory_utilization&lt;/code&gt; to 0.92 and picks &lt;code&gt;max_num_seqs&lt;/code&gt; and &lt;code&gt;max_num_batched_tokens&lt;/code&gt; by GPU &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[13]&lt;/a&gt;. When the cache runs out, requests are preempted and redo their prefill; the documented fixes include a higher &lt;code&gt;gpu_memory_utilization&lt;/code&gt;, lower &lt;code&gt;max_num_seqs&lt;/code&gt; or &lt;code&gt;max_num_batched_tokens&lt;/code&gt;, or more tensor or pipeline parallelism to free memory for the cache &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[14]&lt;/a&gt;. A rising &lt;code&gt;vllm:num_preemptions&lt;/code&gt; counter is the first thing we look at. On a new deployment we benchmark at the defaults first, then change one setting per run.&lt;/p&gt;
&lt;p&gt;Leave CUDA graphs on. A decode step is a series of tiny GPU jobs, and graph replay skips the CPU cost of launching each one &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[15]&lt;/a&gt;; vLLM captures graphs by default, at some cost in memory and startup time &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[16]&lt;/a&gt;. Keep &lt;code&gt;--enforce-eager&lt;/code&gt; for debugging &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[14]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The cost: a bigger batch raises tokens per second but also each user’s latency, the trade-off Sarathi-Serve is named for &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[1]&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;stop-recomputing-the-same-prompt&quot;&gt;Stop recomputing the same prompt&lt;/h2&gt;
&lt;p&gt;Many workloads resend the same long prefix: a system prompt, tool definitions, a document, earlier turns of a chat. Prefix caching keeps that prefix’s KV cache, and a new request that shares it skips that part of prefill &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[17]&lt;/a&gt;. SGLang’s RadixAttention keeps prefixes in a prefix tree (a radix tree) with least-recently-used eviction &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[18]&lt;/a&gt;. Both engines turn it on by default in their current releases &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[13]&lt;/a&gt;, &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[19]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;It “only reduces the time of processing the queries (the prefilling phase)” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[17]&lt;/a&gt;, so it pays most where prompts are long and answers short: &lt;a href=&quot;https://dravin.ai/services/#rag&quot;&gt;RAG&lt;/a&gt;, agents and document chat.&lt;/p&gt;
&lt;p&gt;A cache only hits on an identical token prefix. Put stable content first (system prompt, tools, documents), variable content last, and keep timestamps and request IDs out of the top. In SGLang, &lt;code&gt;--schedule-policy lpm&lt;/code&gt; reorders the queue by longest prefix match to raise the hit rate &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[20]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Routing that ignores the cache scatters requests that share a prefix across replicas, and one published benchmark shows the cost &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[21]&lt;/a&gt;. It ran 8 replicas of a 32B model on 16 H100s, with shared-prefix traffic whose cache demand was six times what one replica could hold. A router that tracked each replica’s actual cache contents produced 8,730 output tokens per second against 4,429 for random or load-only routing, and cut the 90th-percentile TTFT from over 90 seconds to about half a second; a router that estimated cache contents from its own routing history reached 6,944 tokens per second and 31 seconds &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[21]&lt;/a&gt;. SGLang’s router keeps such an estimate, an approximate copy of each worker’s prefix tree &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[22]&lt;/a&gt;, and NVIDIA Dynamo 1.5.0, an orchestration layer that runs on top of engines such as vLLM, SGLang and TensorRT-LLM, routes “based on worker load and KV cache overlap” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[23]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The cost: cache affinity fights load balance, so good routers blend a cache score with a load score, and the cache takes memory that could hold more running requests.&lt;/p&gt;
&lt;h2 id=&quot;keep-long-prompts-from-stalling-everyone&quot;&gt;Keep long prompts from stalling everyone&lt;/h2&gt;
&lt;p&gt;One 50-page document arriving mid-stream can freeze every other user’s output while its prefill runs. Sarathi-Serve’s fix is chunked prefill: cut the prompt into pieces and mix each piece into ordinary decode steps. Against the vLLM of early 2024, the paper reports 2.6 times the serving capacity for a 7B model on one A100 and up to 3.7 times for a 34B model on two A100s &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[1]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;vLLM has since made chunked prefill the default and schedules pending decodes before any prefill &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[14]&lt;/a&gt;. What remains is the dial. In vLLM’s docs, smaller &lt;code&gt;max_num_batched_tokens&lt;/code&gt; values such as 2,048 “achieve better ITL”, larger ones “achieve better time to first token”, and values above 8,192 suit throughput, “especially for smaller models on large GPUs” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[14]&lt;/a&gt;. vLLM 0.30.0’s server already defaults to 8,192 on an H100 and 2,048 on an A100 &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[13]&lt;/a&gt;. SGLang calls the same dial &lt;code&gt;--chunked-prefill-size&lt;/code&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[20]&lt;/a&gt;. Set it against your written TTFT and ITL targets, measured at p99.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;Timeline of GPU steps on one replica with five users streaming tokens while a 30,000-token prompt arrives. Without chunking, the prompt runs as one long prefill step and all five token streams stop until it ends, so inter-token latency spikes; the new user then gets a first token. With chunked prefill, the prompt is cut into eight equal slices, each riding along with a normal decode step; the five streams keep ticking at a slightly slower rate and the new user’s first token arrives a little later. A chunk-size scale shows the trade: small chunks such as 2,048 tokens keep streams smooth and delay the first token, large chunks such as 8,192 give an earlier first token and bumpier streams.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;Chunked prefill cuts a long prompt into slices and runs each one alongside ordinary decode steps, so other users keep receiving tokens while it loads. The two chunk sizes marked are the examples in vLLM’s docs [14].&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;spend-fewer-bytes-per-token&quot;&gt;Spend fewer bytes per token&lt;/h2&gt;
&lt;p&gt;Decode is limited by the bytes the GPU reads per token, so storing weights and cache in fewer bits gives more tokens per second.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;FP8 weights and activations&lt;/strong&gt; halve weight memory and need Ada (L40S), Hopper (H100, H200) or Blackwell (B200, GB200) GPUs &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[24]&lt;/a&gt;. One accuracy study ran more than 500,000 evaluations on an open model family at 8B, 70B and 405B, in vLLM on A6000, A100 and H100 GPUs. It found FP8 “effectively lossless across all model scales” and well-tuned INT8 within 1 to 3% &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[25]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;INT4 weight-only&lt;/strong&gt; (AWQ &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[26]&lt;/a&gt;, GPTQ &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[27]&lt;/a&gt;) cuts weight memory about four times against FP16. The same study found it the most cost-efficient when requests run one at a time, and 8-bit formats fastest under continuous batching &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[25]&lt;/a&gt;. Weight-only 4-bit speeds up memory-bound decode, and as the batch grows the GPU turns compute-bound: the authors of one 4-bit kernel report close to the full 4 times speed-up per layer up to batch 16 to 32, still significant but shrinking at 64 to 128, and up to 2.8 times end to end in vLLM &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[28]&lt;/a&gt;. For a busy server, that points to an 8-bit format first.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;NVFP4&lt;/strong&gt; stores 4-bit values in blocks of 16 that share an FP8 scale. NVIDIA reports about 3.5 times less memory than FP16 and 1.8 times less than FP8, and 1% or less accuracy loss moving a 671B mixture-of-experts (MoE) reasoning model from FP8 to NVFP4 &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[29]&lt;/a&gt;. FP4 arithmetic needs Blackwell, though older GPUs can still load NVFP4 weights: Nemotron 3.5 Lightning ships an NVFP4 checkpoint that runs on NVIDIA’s NVFP4 kernels across Blackwell, Hopper and Ampere &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[30]&lt;/a&gt;. On the older GPUs that saves memory and bandwidth, without the FP4 arithmetic.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The KV cache&lt;/strong&gt; can be quantised separately, which raises the number of requests that fit. vLLM supports an FP8 KV cache and recommends calibrated scales &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[31]&lt;/a&gt;. With Qwen3.8-27B on an RTX PRO 6000 Blackwell, the SGLang, Qwen and NVIDIA teams fitted 70 concurrent requests at 32K context with an NVFP4 KV cache, against 44 with FP8. At matched concurrency, peak-batch decode throughput rose 26 to 30%. SGLang still marks the support experimental &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[32]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The cost is accuracy and hardware lock-in. Formats are tied to GPU generations, uncalibrated KV scales are a risk, and upgrades can break checkpoints: vLLM 0.30.0 dropped GPTQ activation ordering (the “act-order” option), so check any GPTQ checkpoint before you upgrade &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[13]&lt;/a&gt;. We quantise before we add GPUs, and judge each format on the evaluation set of the product it serves.&lt;/p&gt;
&lt;h2 id=&quot;draft-tokens-while-there-is-spare-compute&quot;&gt;Draft tokens while there is spare compute&lt;/h2&gt;
&lt;p&gt;Speculative decoding lets a cheap drafter propose several tokens and the target model verify them in one pass, keeping the target model’s output distribution &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[3]&lt;/a&gt;. Leviathan and colleagues report 2 to 3 times on an 11B encoder-decoder model &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[3]&lt;/a&gt;. Each target pass yields (1 − α&lt;sup&gt;γ+1&lt;/sup&gt;) / (1 − α) tokens on average, where α is the acceptance rate and γ the number of drafted tokens: at α = 0.8 and γ = 4, about 3.4. The price is extra arithmetic: each step runs γ+1 positions through the target model in parallel, so “the number of concurrent arithmetic operations grows by a factor of γ+1”, and the method helps only when compute is spare &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;So the gain shrinks as the batch fills, and vLLM aims speculation at “medium-to-low QPS, memory-bound workloads” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[33]&lt;/a&gt;. EAGLE-3, a small draft head that reads the target model’s own hidden states, reports up to 6.5 times per-request speed-up on its best task and 1.38 times throughput at batch 64 in SGLang &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[34]&lt;/a&gt;. Plan a busy server around the second number.&lt;/p&gt;
&lt;figure&gt;&lt;p id=&quot;chart-per-user-speed-up-from-eagle-3-1-by-concurrency&quot;&gt;Per-user speed-up from EAGLE 3.1, by concurrency&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Concurrency 1: 2.03&lt;/li&gt;&lt;li&gt;Concurrency 4: 1.71&lt;/li&gt;&lt;li&gt;Concurrency 16: 1.66&lt;/li&gt;&lt;/ul&gt;&lt;figcaption&gt;Speed-up in each user’s output speed against no speculation, on a trillion-parameter-class MoE in NVFP4, tensor parallel across four GB200 GPUs, coding dataset [35].&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The chart shows the speed each user sees, the number speculation targets; even so, the gain narrows from 2.03 to 1.66 times between concurrency 1 and 16 &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[35]&lt;/a&gt;. P-EAGLE, a newer drafter, shows the same pattern against EAGLE-3: on a 20B open-weight MoE on one B200, its gain fell from 1.55 times at concurrency 1 to 1.05 times at 64 on one chat benchmark, and from 1.69 to 1.25 times on a coding benchmark &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[36]&lt;/a&gt;. The EAGLE 3.1 authors note that speculative decoding “often degrades under different chat templates, long-context inputs, or out-of-distribution system prompts”, the fragility 3.1 was built to reduce &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[35]&lt;/a&gt;, so track &lt;code&gt;vllm:spec_decode_num_accepted_tokens&lt;/code&gt; against &lt;code&gt;vllm:spec_decode_num_draft_tokens&lt;/code&gt; on your own prompts &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[33]&lt;/a&gt;. vLLM can also shrink the draft as load rises: dynamic speculative decoding sets the number of drafted tokens per batch-size range, and an opt-in adaptive verification mode (DSpark drafters only) sizes verification per request from the drafter’s confidence &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[33]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Multi-token prediction (MTP) trains the model to draft its own next tokens &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[37]&lt;/a&gt;. Nemotron 3.5 Lightning has it built in &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[30]&lt;/a&gt;, so there is no separate draft model to train or host; Gemma 4 ships ready-made MTP drafters that attach to the base model and share its embeddings, so there is nothing to train &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[38]&lt;/a&gt;. NVIDIA says “MTP is best suited for medium to high concurrency, with the optimal draft length decreasing as concurrency increases” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[30]&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;protect-the-latency-target&quot;&gt;Protect the latency target&lt;/h2&gt;
&lt;p&gt;Past a certain queue depth, one more request makes every request late. Cap concurrency where goodput peaks in your benchmark, reject or defer early with a retry signal (an HTTP 429 with a Retry-After header), and give interactive traffic priority: vLLM has &lt;code&gt;--scheduling-policy priority&lt;/code&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[39]&lt;/a&gt; and SGLang has priority scheduling &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[19]&lt;/a&gt;. Nightly jobs belong in their own lane, where “achieving a large batch size is the most important thing” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[20]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The cheapest request is one the big model never sees. We route classification and extraction to the smallest model that works, which takes that traffic off the serving fleet entirely; our post on &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/&quot;&gt;typed decision models&lt;/a&gt; covers that lever.&lt;/p&gt;
&lt;h2 id=&quot;scale-out-across-gpus&quot;&gt;Scale out across GPUs&lt;/h2&gt;
&lt;p&gt;Tensor parallelism splits every layer across GPUs and needs fast links between them; pipeline parallelism gives each GPU a run of layers; data parallelism runs whole copies of the model (replicas). vLLM advises tensor parallelism within a node, and pipeline parallelism across nodes or on GPUs without NVLink, NVIDIA’s fast GPU-to-GPU link &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[40]&lt;/a&gt;. SGLang says to “always favor data parallelism” for throughput when memory allows &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[20]&lt;/a&gt;. In practice: the smallest tensor-parallel group that holds the weights and a healthy KV cache, then more replicas.&lt;/p&gt;
&lt;p&gt;MoE models such as DeepSeek V4, Kimi K3, GLM-5.3 and Mistral Large 3 activate a small share of their weights per token. Expert parallelism places experts on separate GPUs and works best with data parallelism; since tokens spread unevenly, vLLM’s &lt;code&gt;--enable-eplb&lt;/code&gt; rebalances experts from live load statistics &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[41]&lt;/a&gt;. The SGLang team combined expert parallelism with prefill-decode disaggregation for a 671B MoE on H100s; in a decode-only test on 9 nodes of 8 GPUs, with prefill capacity assumed unlimited, it reached about 22,300 output tokens per second per node for 2,000-token inputs, about 5 times a 16-GPU tensor-parallel baseline &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[42]&lt;/a&gt;. That needs both techniques, several nodes and fast interconnect; on one node, tensor parallelism and replicas are simpler.&lt;/p&gt;
&lt;h2 id=&quot;split-prefill-and-decode-last&quot;&gt;Split prefill and decode last&lt;/h2&gt;
&lt;p&gt;The final lever runs prefill and decode on separate GPU pools. DistServe found that “adding a single prefill job to a batch of decoding requests significantly slows down both processes”, and reports 7.4 times more requests or 12.6 times tighter SLOs than colocated systems, with over 90% of requests within their latency limits &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[7]&lt;/a&gt;. Splitwise reports 1.4 times the throughput at 20% lower cost, or 2.35 times the throughput at the same cost and power &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[2]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Yet vLLM’s documentation says plainly: “Disaggregated prefill DOES NOT improve throughput” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[43]&lt;/a&gt;. Both claims hold. The papers measure goodput under latency targets and cost-normalised capacity; vLLM’s statement is best read as raw tokens per second. Disaggregation leaves the GPUs as fast as they were. What it changes is interference. Chunked prefill stops the worst stalls, but every mixed step still carries a slice of someone’s prompt, which adds to ITL, and colocated phases “share their resource and parallelism settings” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[7]&lt;/a&gt;. Separate pools remove both: decode steps carry no prompt work, and each pool gets its own parallelism and batch size, so more requests meet their targets and TTFT and ITL can be tuned separately &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[43]&lt;/a&gt;. The pools are sized separately too: longer prompts call for more prefill GPUs, longer answers for more decode GPUs.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;Prefill and decode on separate GPU pools. Prompts of varying length arrive at a router, which scores replicas by prefix overlap and load, and go to a prefill pool where each GPU runs a wide, compute-heavy prefill. Each finished prompt’s KV cache crosses an RDMA link as a stream of packets to a decode pool, where many requests stream tokens in one large, steady batch. A TTFT gauge and an ITL gauge stay inside their target bands. Longer prompts call for more prefill GPUs; longer answers call for more decode GPUs.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;The router scores replicas by prefix overlap and load. Prefill hands each prompt’s KV cache to the decode pool over RDMA, so a long prompt does not stall the running decode streams. Shaded bands mark the latency targets.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Conditions matter. On GB200 with a 671B reasoning MoE, TensorRT-LLM’s rate-matched curves, which assume ideal pool ratios and free KV transfer, show 1.4 to 1.8 times the output throughput per GPU at the same per-user speed, at 4,400 input and 1,200 output tokens, mostly at lower concurrency; on 8,192-input, 256-output traffic its real end-to-end runs fell 0 to 25% short of the rate-matched curve &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[44]&lt;/a&gt;. SGLang moves the KV cache between pools with RDMA (remote direct memory access, which copies memory between machines without going through the CPU) transfer engines, on InfiniBand or RoCE devices or AWS EFA &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[45]&lt;/a&gt;; vLLM labels the feature experimental &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[43]&lt;/a&gt; and TensorRT-LLM beta &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[10]&lt;/a&gt;. For a single model on a single GPU, Dynamo’s README says the engine alone “is probably sufficient” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[23]&lt;/a&gt;. Reach for disaggregation with long prompts, several nodes and fast interconnect. We treat it as the last step, after chunk size, routing and replicas have been tuned against goodput.&lt;/p&gt;
&lt;h2 id=&quot;checklist&quot;&gt;Checklist&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Write the SLO down: TTFT and ITL at p99, and the share of requests that must meet them.&lt;/li&gt;
&lt;li&gt;Benchmark on your own prompt lengths, output lengths, shared prefixes and arrival pattern, and report goodput.&lt;/li&gt;
&lt;li&gt;Dashboard queue depth, KV cache usage, preemptions and prefix-cache hit rate.&lt;/li&gt;
&lt;li&gt;Run a current engine with CUDA graphs on; size memory and batch settings until preemptions stop.&lt;/li&gt;
&lt;li&gt;Put stable prompt content first, confirm prefix caching is on, and route by prefix across replicas.&lt;/li&gt;
&lt;li&gt;Set the chunk size against your TTFT and ITL targets.&lt;/li&gt;
&lt;li&gt;Try FP8 on Hopper, NVFP4 on Blackwell, INT4 weight-only at low concurrency, and a quantised KV cache; check accuracy on your own evals.&lt;/li&gt;
&lt;li&gt;Add speculative decoding where concurrency is low to medium, try a model’s built-in MTP heads at higher concurrency with a shorter draft, and watch the acceptance rate.&lt;/li&gt;
&lt;li&gt;Cap concurrency where goodput peaks, prioritise interactive traffic and give batch jobs their own lane.&lt;/li&gt;
&lt;li&gt;Use the smallest tensor-parallel group that fits, then replicas; expert parallelism with load balancing for MoE.&lt;/li&gt;
&lt;li&gt;Split prefill and decode only with long prompts, several nodes and RDMA.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;sources&quot;&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2403.02310&quot;&gt;Agrawal et al., “Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve”, OSDI 2024&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2311.18677&quot;&gt;Patel et al., “Splitwise: Efficient generative LLM inference using phase splitting”, ISCA 2024&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2211.17192&quot;&gt;Leviathan, Kalman and Matias, “Fast Inference from Transformers via Speculative Decoding”, ICML 2023&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;PyPI release pages, checked 23 Sep 2026: &lt;a href=&quot;https://pypi.org/project/vllm/&quot;&gt;vLLM&lt;/a&gt;, &lt;a href=&quot;https://pypi.org/project/sglang/&quot;&gt;SGLang&lt;/a&gt;, &lt;a href=&quot;https://pypi.org/project/tensorrt-llm/&quot;&gt;TensorRT-LLM&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.nvidia.com/nim/benchmarking/llm/latest/metrics.html&quot;&gt;NVIDIA NIM docs, “LLM benchmarking metrics”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/benchmarking/cli.html&quot;&gt;vLLM docs, “Benchmark CLI”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2401.09670&quot;&gt;Zhong et al., “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving”, OSDI 2024&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/usage/metrics.html&quot;&gt;vLLM docs, “Production Metrics”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.usenix.org/conference/osdi22/presentation/yu&quot;&gt;Yu et al., “Orca: A Distributed Serving System for Transformer-Based Generative Models”, OSDI 2022&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://nvidia.github.io/TensorRT-LLM/overview.html&quot;&gt;TensorRT-LLM docs, “Overview” (updated 21 Sep 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2309.06180&quot;&gt;Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention”, SOSP 2023&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2305.13245&quot;&gt;Ainslie et al., “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”, 2023&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;vLLM v0.30.0 (22 Sep 2026): &lt;a href=&quot;https://github.com/vllm-project/vllm/releases/tag/v0.30.0&quot;&gt;release notes&lt;/a&gt;, tagged source &lt;a href=&quot;https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/config/cache.py&quot;&gt;vllm/config/cache.py&lt;/a&gt; and &lt;a href=&quot;https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/engine/arg_utils.py&quot;&gt;vllm/engine/arg_utils.py&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/configuration/optimization.html&quot;&gt;vLLM docs, “Optimization and Tuning”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://pytorch.org/blog/accelerating-pytorch-with-cuda-graphs/&quot;&gt;PyTorch blog, “Accelerating PyTorch with CUDA Graphs”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/design/cuda_graphs.html&quot;&gt;vLLM docs, “CUDA Graphs”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/features/automatic_prefix_caching.html&quot;&gt;vLLM docs, “Automatic Prefix Caching”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2312.07104&quot;&gt;Zheng et al., “SGLang: Efficient Execution of Structured Language Model Programs”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;SGLang v0.5.20 (18 Sep 2026): &lt;a href=&quot;https://github.com/sgl-project/sglang/releases/tag/v0.5.20&quot;&gt;release notes&lt;/a&gt; and &lt;a href=&quot;https://docs.sglang.io/advanced_features/server_arguments.html&quot;&gt;server arguments&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.sglang.io/advanced_features/hyperparameter_tuning.html&quot;&gt;SGLang docs, “Hyperparameter Tuning”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://llm-d.ai/blog/kvcache-wins-you-can-see&quot;&gt;Prefix-aware routing benchmark, “KV-Cache Wins You Can See” (24 Sep 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://lmsys.org/blog/2024-12-04-sglang-v0-4/&quot;&gt;LMSYS blog, “SGLang v0.4” (4 Dec 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/ai-dynamo/dynamo&quot;&gt;NVIDIA Dynamo README, v1.5.0 (21 Sep 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/features/quantization/llm_compressor/fp8/&quot;&gt;vLLM docs, “FP8 W8A8”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2411.02355&quot;&gt;Kurtic et al., “Give Me BF16 or Give Me Death? Accuracy-Performance Trade-Offs in LLM Quantization”, ACL 2025&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2306.00978&quot;&gt;Lin et al., “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration”, MLSys 2024&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2210.17323&quot;&gt;Frantar et al., “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers”, ICLR 2023&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2408.11743&quot;&gt;Frantar et al., mixed-precision 4-bit inference kernels for LLMs (Aug 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/&quot;&gt;NVIDIA blog, “Introducing NVFP4 for Efficient and Accurate Low-Precision Inference” (24 Jun 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/&quot;&gt;NVIDIA blog, Nemotron 3.5 Lightning (11 Aug 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache.html&quot;&gt;vLLM docs, “Quantized KV Cache”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.lmsys.org/blog/2026-09-16-nvfp4-kv-cache&quot;&gt;LMSYS blog, “Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache” (16 Sep 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;vLLM docs: &lt;a href=&quot;https://docs.vllm.ai/en/latest/features/speculative_decoding/&quot;&gt;Speculative Decoding&lt;/a&gt;, &lt;a href=&quot;https://docs.vllm.ai/en/latest/features/speculative_decoding/dynamic_speculative_decoding/&quot;&gt;Dynamic Speculative Decoding&lt;/a&gt; and &lt;a href=&quot;https://docs.vllm.ai/en/latest/features/speculative_decoding/adaptive_verification/&quot;&gt;Adaptive Verification&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2503.01840&quot;&gt;Li et al., “EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test”, NeurIPS 2025&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://vllm.ai/blog/2026-05-26-eagle-3-1&quot;&gt;vLLM blog, EAGLE 3.1 (26 May 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://vllm.ai/blog/2026-03-13-p-eagle&quot;&gt;vLLM blog, “P-EAGLE: Faster LLM inference with Parallel Speculative Decoding in vLLM” (13 Mar 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2404.19737&quot;&gt;Gloeckle et al., “Better &amp;amp; Faster Large Language Models via Multi-token Prediction” (Apr 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://ai.google.dev/gemma/docs/mtp/overview&quot;&gt;Google AI for Developers, Gemma MTP overview&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/configuration/engine_args.html&quot;&gt;vLLM docs, “Engine Arguments”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/serving/parallelism_scaling.html&quot;&gt;vLLM docs, “Parallelism and Scaling”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/serving/expert_parallel_deployment.html&quot;&gt;vLLM docs, “Expert Parallel Deployment”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://lmsys.org/blog/2025-05-05-large-scale-ep/&quot;&gt;SGLang team: large-scale expert parallelism on 96 H100s (May 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/features/disagg_prefill.html&quot;&gt;vLLM docs, “Disaggregated Prefilling (experimental)”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://nvidia.github.io/TensorRT-LLM/blogs/tech_blog/blog5_Disaggregated_Serving_in_TensorRT-LLM.html&quot;&gt;TensorRT-LLM tech blog, “Disaggregated Serving in TensorRT LLM” (updated 16 Mar 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.sglang.io/advanced_features/pd_disaggregation.html&quot;&gt;SGLang docs, “PD Disaggregation”&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This post first appeared on &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/&quot;&gt;Dravin AI&lt;/a&gt;.&lt;/p&gt;</content:encoded><category>LLM serving</category><category>Inference optimisation</category><category>Self-hosted LLMs</category><category>GPUs</category><author>info@dravin.ai (Dravin AI)</author></item><item><title>Typed decision models: what early tests show</title><link>https://dravin.ai/blog/jev-and-typed-decision-models/</link><guid isPermaLink="true">https://dravin.ai/blog/jev-and-typed-decision-models/</guid><description>Models that answer typed questions with a probability in one pass: what they are, what the first independent tests found, and how to test one.</description><pubDate>Tue, 22 Sep 2026 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Jev and Laya answer typed questions (a choice, a score or a yes/no) with a probability attached, instead of writing text.&lt;/li&gt;&lt;li&gt;Early independent tests: Jev was the fastest model measured, its calibration is mixed, and Laya is near chance until it is fine-tuned.&lt;/li&gt;&lt;li&gt;Before adopting one, test it on a held-out set from your own traffic, against a classifier fine-tuned on your own data.&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;On 15 September 2026 TypeSafe AI released Jev in early access. It calls Jev its first “System One Model”, a name taken from Daniel Kahneman’s fast, intuitive System 1 thinking &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt;. Jev does not write text. You give it a piece of state (a ticket, an email, a JSON object) and a set of typed questions, and it returns a choice, a score or a yes/no drawn from options you defined, with a probability attached. Three days later Convai Innovations published Laya, an open-weight model built on the same idea &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[13]&lt;/a&gt;. Since version 0.3.7 it also ships a server that accepts the same request shape as Jev &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Both are days old. The idea still deserves attention, because it describes what many production AI systems do on every request: make a small decision, such as routing a ticket, moderating a message or checking an agent’s tool call, and know when to leave it alone.&lt;/p&gt;
&lt;p&gt;The short version: Jev is fast, though the gap is large only against LLMs left in their slowest default modes &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[15]&lt;/a&gt;. On calibration, which is the point of the product, the independent tests are mixed &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[14]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[15]&lt;/a&gt;. Laya is near chance until it is &lt;a href=&quot;https://dravin.ai/services/#training&quot;&gt;fine-tuned&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;. The comparison that should decide adoption is against a classifier fine-tuned on your own data, and the only one we have found is Laya’s own, on a benchmark its checkpoint was fine-tuned on &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;what-a-typed-decision-model-is&quot;&gt;What a typed decision model is&lt;/h2&gt;
&lt;p&gt;TypeSafe describes Jev as “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt;. A request has two parts. The state is the input: a string, a JSON object or an array of text. Images, audio and video are not supported yet &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[2]&lt;/a&gt;. The questions use three primitives &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[2]&lt;/a&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Choice&lt;/strong&gt; picks one option from a list you supply, such as which team should handle a ticket.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Score&lt;/strong&gt; places the state on a rubric, for example customer frustration on a scale from 0 to 2.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Noul&lt;/strong&gt; is TypeSafe’s name for a yes/no question: it asks whether a statement is true and returns a probability between 0 and 1.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The Register gave an example of what a three-department routing question might return &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[17]&lt;/a&gt;:&lt;/p&gt;
&lt;pre&gt;&lt;code&gt;{&amp;quot;billing&amp;quot;: 0.08, &amp;quot;technical&amp;quot;: 0.85, &amp;quot;sales&amp;quot;: 0.07}&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;It paired that answer with a confidence of 0.82. The probabilities are per option. The confidence is a separate summary that TypeSafe derives from the shape of the whole distribution, and it is the number you set thresholds on; a Noul returns only its probability &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[5]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Jev’s internals are unpublished. TypeSafe says it built “a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD)” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt;. TechCrunch reports that the company has said little about that architecture, that outside observers suspect it sits on top of an open-weight LLM, and that the former OpenAI researcher who started the company says Jev is trained only on synthetic data &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[16]&lt;/a&gt;. So Jev can be judged only by what it returns, which is what the rest of this post does.&lt;/p&gt;
&lt;p&gt;Laya is the clearest public account of how such a model can work, because its model card spells it out &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;. The English checkpoint is ModernBERT-large, a bidirectional encoder (it reads the whole input at once), fully fine-tuned, with a small decision head on top: 421M parameters in total. A multilingual checkpoint built on mmBERT has 322M. Each option in a question is scored at its own &lt;code&gt;[MASK]&lt;/code&gt; token, the placeholder slot a BERT-style model is trained to fill, and the scores are normalised across that question’s options. The options arrive with the request, so a new schema needs no retraining, and every question in a call is answered in a single forward pass. Training rewards the model with strictly proper scoring rules, which pay the most only when the reported probabilities are honest.&lt;/p&gt;
&lt;p&gt;Laya, then, is a BERT-style classifier whose label set arrives with each request, trained with an objective that rewards honest probabilities. That places the category between two tools most teams already know. A fine-tuned classifier is fast and accurate, but it needs labelled data and knows only its fixed labels. A zero-shot classifier takes its labels at request time, and one served as the baseline in an independent test below &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[14]&lt;/a&gt;. What typed decision models add is calibration as a training objective, several typed questions answered in one call, and an interface built for software to branch on &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;how-this-differs-from-asking-an-llm-to-classify&quot;&gt;How this differs from asking an LLM to classify&lt;/h2&gt;
&lt;p&gt;Any LLM will produce a label from a prompt and a JSON schema. Three things are different, and the first is narrower than it looks.&lt;/p&gt;
&lt;p&gt;The format is guaranteed by construction. “By providing typed, structured values, Jev can avoid the parsing and validating that must be done to process text responses from LLMs,” as The Register put it &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[17]&lt;/a&gt;. An LLM gets most of that guarantee from constrained decoding, which limits its output to a JSON schema or a forced tool call. The independent benchmark below ran every LLM that way: seven of the eight returned malformed answers on at most 0.7% of decisions in any suite, one small model reached 6.2% on one suite, and Jev returned none &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[15]&lt;/a&gt;. The format advantage is real but small, so the differences that matter are the next two.&lt;/p&gt;
&lt;p&gt;The work happens in one step. An LLM writes its answer token by token; a typed decision model scores every option of every question at once. TypeSafe credits this parallel design for the speed &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[3]&lt;/a&gt;. The design pays most when one call asks several questions: Laya answers all of them in a single forward pass &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;, and TypeSafe says adding questions “barely changes the response time” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[9]&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;Two lanes answer the same ticket routing question. On the left an LLM writes a JSON answer one token at a time, route technical, and a parser then checks it. On the right a typed decision model reads the ticket once and fills a probability for every option: billing 0.08, technical 0.85, sales 0.07, then a confidence of 0.82. The right lane returns every probability in one step; the left lane writes its answer in order and then checks it.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;The same routing question, asked two ways. The LLM writes its answer one token at a time, then a parser checks it. The typed decision model scores every option in one pass and adds a confidence value. Values from The Register’s example [17].&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The probability is the product. A typed decision model is trained so that its per-option probabilities match outcomes, which is the aim of both TypeSafe’s RLCD and Laya’s proper scoring rules &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;. LLMs asked to state their confidence tend to be overconfident, which Xiong and colleagues found across several models &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[20]&lt;/a&gt;. The research is mixed: Tian and colleagues found verbalised confidence from RLHF-tuned models can be better calibrated than their token probabilities &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[21]&lt;/a&gt;. Where an API exposes token log-probabilities, an LLM’s probabilities over the labels can also be read directly, and Kadavath and colleagues found large models well calibrated on multiple-choice and true/false questions posed in the right format &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[22]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Type safety guarantees the shape of the answer and says nothing about whether the answer is right. TypeSafe says Jev “can’t hallucinate” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt;; The Register noted that structured output with probabilities “does not preclude the possibility of being incorrect” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[17]&lt;/a&gt;. And because the caller fixes the options, a Choice always picks one of them, even when none fits. TypeSafe’s own documentation says “the Choice is relative, settling &lt;em&gt;which&lt;/em&gt; option, while each Noul is absolute and can be low for all of them” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[4]&lt;/a&gt;. The practical fix is to add a “none of these” option, or to put a separate Noul in front of the Choice as a gate.&lt;/p&gt;

&lt;h2 id=&quot;why-calibrated-probabilities-matter-in-production&quot;&gt;Why calibrated probabilities matter in production&lt;/h2&gt;
&lt;p&gt;A model is calibrated when its probabilities match reality across many predictions. TypeSafe’s primer puts it plainly: “Outcomes assigned a probability of &lt;code&gt;0.2&lt;/code&gt; should occur about 20% of the time. Outcomes assigned a probability of &lt;code&gt;0.8&lt;/code&gt; should occur about 80% of the time” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[7]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Each small decision a production system makes carries two questions: what is the answer, and is it safe to act on. The CTO of Earendil, quoted by TechCrunch, described the trade: “The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it’s 95%, sure, then I can do something with it” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[16]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;TypeSafe’s documentation recommends three bands: act automatically when confidence is high; proceed with caution in the middle, for example by asking a person to confirm; and below that, route to a person, ask for clarification or fall back to another system. It adds that “a confidence threshold is not one number”, since actions with worse consequences deserve a stricter bar &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[5]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The operating number that follows is coverage at an error budget, an idea from selective classification, where a model is allowed to abstain &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[23]&lt;/a&gt;. Choose the error rate you can accept, say 5%. Set the threshold on labelled data so that errors among the items acted on stay within that rate, then measure what share of traffic the model handles alone. Everything else goes to a person or a larger model. One limit applies throughout: calibration is a property of groups of predictions, and in TypeSafe’s words it “does not guarantee that an individual answer is correct” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[2]&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;Tickets flow into a tagging gate, each with a confidence score. At 0.75 or above the item is acted on, from 0.50 it goes to confirm, and below 0.50 it goes to a person or an LLM. Four of six tickets clear the act threshold and are handled automatically. A stricter payment gate acts only at 0.97 or above, so a decision at 0.80, acted on at tagging, is held for confirmation before any money moves. Set your own thresholds on labelled data, one per action.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;A confidence gate turns a confidence value into an action. The bar rises with the cost of a mistake, so a decision at 0.80 is acted on at tagging and held before any money moves. Set your own thresholds on labelled data, one per action.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;what-the-first-tests-show&quot;&gt;What the first tests show&lt;/h2&gt;
&lt;p&gt;TypeSafe’s blog gives end-to-end response times of 70 to 500 ms. Its home page claims Jev is “193.6x faster”, a figure the blog says comes from its workflow evals, in which each LLM ran at its provider’s default reasoning setting, and TypeSafe expects these “are on the higher end of real world gains” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[10]&lt;/a&gt;. The same post says the evals were made by its own capability team, “so some bias could exist”, and that the reference answers are the average of GPT-6 Astra and Claude Fable 5.1 &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[10]&lt;/a&gt;. Accuracy in those evals therefore means agreement with two frontier LLMs, where most buyers would assume human labels. TypeSafe argues that this reference, if anything, favours the OpenAI and Anthropic models in the comparison &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Two independent benchmarks with published protocols followed within days. &lt;strong&gt;decision-model-benchmark&lt;/strong&gt; ran Jev against eight LLMs from five providers, each constrained to a JSON schema or a forced tool call, with thinking turned off where the provider allowed it. It declares no vendor sponsorship and publishes its raw logs &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[15]&lt;/a&gt;. Calibration error below is expected calibration error (ECE), the average gap between stated confidence and observed accuracy; lower is better and 0 is perfect.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Finding&lt;/th&gt;
&lt;th&gt;Jev&lt;/th&gt;
&lt;th&gt;The eight LLMs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy, 77-way banking intent&lt;/td&gt;
&lt;td&gt;76.3%&lt;/td&gt;
&lt;td&gt;70.9% to 81.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Accuracy, SMS spam&lt;/td&gt;
&lt;td&gt;93.0%&lt;/td&gt;
&lt;td&gt;66.1% to 94.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Median latency per request&lt;/td&gt;
&lt;td&gt;264 to 276 ms&lt;/td&gt;
&lt;td&gt;303 ms to 5.6 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Most options accepted in one question&lt;/td&gt;
&lt;td&gt;255&lt;/td&gt;
&lt;td&gt;512 (the most tested)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Low confidence (0.5 or below) when no option is correct&lt;/td&gt;
&lt;td&gt;49.7%&lt;/td&gt;
&lt;td&gt;97.3% to 100% for seven; 64.7% for one&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Answers changed by shuffling option order&lt;/td&gt;
&lt;td&gt;13%&lt;/td&gt;
&lt;td&gt;15% to 37%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Calibration error, banking intent&lt;/td&gt;
&lt;td&gt;0.083&lt;/td&gt;
&lt;td&gt;0.043 to 0.226&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Calibration error, SMS spam&lt;/td&gt;
&lt;td&gt;0.249&lt;/td&gt;
&lt;td&gt;0.042 to 0.197&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Calibration error, suite built to force uncertainty&lt;/td&gt;
&lt;td&gt;0.246&lt;/td&gt;
&lt;td&gt;0.039 to 0.122&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The authors’ summary is that “no class wins on quality”. Jev was the fastest model measured, about 1.2 times faster than the fastest LLM set-up (an open-weight model on specialised inference hardware) and 10 to 16 times faster than models in thinking mode, and they conclude that “the vendor’s speedup claim holds only against LLMs left in their slowest default mode” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[15]&lt;/a&gt;. Calibration is mixed: Jev sits inside the LLMs’ range on banking intent and above it on spam and on the forced-uncertainty suite. The 49.7% is the relative Choice described above at work.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;jev-benchmarks&lt;/strong&gt;, a pre-registered pilot on 300 held-out examples against an open zero-shot classifier, found Jev more accurate on two of three tasks, where it could handle 83% and 86% of items alone at a 5% error budget &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[14]&lt;/a&gt;. On the third, emotion labelling, Jev put zero probability on the true label for 16% of examples and could automate none of them at that budget; the open classifier managed 2%. The authors fitted those thresholds on the same small slice, so they describe this sample only &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[14]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Laya documents its own limits unusually well &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[12]&lt;/a&gt;. Its English base checkpoint scores 0.362 on a benchmark of four typed-decision workflows without fine-tuning: above the 0.318 random baseline, and below the 0.461 you would get by always guessing the most common answer. The headline 0.766 belongs to a checkpoint fine-tuned on that benchmark’s own training split. As shipped, the probabilities are overconfident. Fitting one temperature (a single number that sharpens or softens every probability) per question type and option count moves mean ECE from 0.466 to 0.081, per its model card, measured before a temperature change in a later release &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Held-out results are weaker than in-training ones. On held-out toxic chat, moderation accuracy is 0.530, barely above chance on a balanced split, and the authors put it bluntly: “hand-picked examples work, real traffic does not” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[12]&lt;/a&gt;. The English checkpoint scored 0.000 accuracy on Khmer, and its authors warn that it stays confident while wrong, which is why their router picks a checkpoint by script before the model runs &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[12]&lt;/a&gt;. The confidence figures from that language sweep were measured before the same temperature change; the maker has since re-run it, and accuracy held while the confidence figures moved &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[12]&lt;/a&gt;. A confidence gate protects you only on inputs like the ones the model was trained and calibrated on, so check the language and shape of inputs before the gate.&lt;/p&gt;
&lt;p&gt;Laya’s card also has a table that puts it ahead of Jev on most rows, while stating that the Jev figures are “third-party published, never measured here” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;. Its Jev calibration figure is the 0.246 from the forced-uncertainty suite above, and its own figure is measured after the temperature refit. Each side’s table suits its maker, and neither is a like-for-like comparison.&lt;/p&gt;
&lt;p&gt;Both products are young. Jev is eight days old at the time of writing. At launch, access was through a waitlist &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[1]&lt;/a&gt;, and TechCrunch reports that the company “briefly lost the ability to serve users from its API because demand was so high” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[16]&lt;/a&gt;. Its rate limits “can change without notice”, and the same weights serve every account, so it cannot be fine-tuned on your data. TypeSafe says “Jev is not trained on customer requests or responses”, and that English is where its accuracy is best &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Laya’s first release reached PyPI on 18 September 2026 &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[13]&lt;/a&gt;. Its weights and code are open under Apache 2.0, so you can inspect and run it yourself. It needs fine-tuning and a temperature fit on your task first. Its escalation head, &lt;code&gt;action.act_probability&lt;/code&gt;, “carries no usable signal yet”, and the card advises gating on confidence instead. Its bundled server listens on every network interface (&lt;code&gt;0.0.0.0&lt;/code&gt;) with no authentication unless &lt;code&gt;LAYA_API_KEY&lt;/code&gt; is set &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;, so set the key before the port is reachable.&lt;/p&gt;
&lt;p&gt;The evidence so far is vendor evals, one 300-example independent pilot and one independent benchmark of five small suites.&lt;/p&gt;
&lt;h2 id=&quot;choosing-between-an-llm-a-typed-model-and-a-fine-tuned-classifier&quot;&gt;Choosing between an LLM, a typed model and a fine-tuned classifier&lt;/h2&gt;
&lt;p&gt;TypeSafe’s list of failure modes for &lt;code&gt;jev-1.13&lt;/code&gt; covers literal reading, arithmetic, date comparisons, questions whose answer depends on first answering another question, and counting: “&lt;code&gt;jev-1.13&lt;/code&gt; does not count reliably” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[4]&lt;/a&gt;. Its advice is to keep arithmetic, dates and counting in code &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[4]&lt;/a&gt;, and to have each question ask one well-scoped thing, composing the answers in code &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[8]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;With that in mind, each of the three options has a place.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Option&lt;/th&gt;
&lt;th&gt;When it fits&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;An LLM&lt;/td&gt;
&lt;td&gt;The output is language, code or an explanation someone will read, or the decision needs several steps of reasoning. Jev “does not generate text, write code, or hold a conversation” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[6]&lt;/a&gt;. An LLM can also draft labels for training a classifier, with people checking a sample, and give a second opinion on a low-confidence case before it reaches a person.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A typed decision model&lt;/td&gt;
&lt;td&gt;The options are fixed per request, one call asks several questions, and you have no labelled data yet. Used zero-shot, Jev beat an open zero-shot classifier on two of three tasks in one pilot &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[14]&lt;/a&gt;. Laya fits only after fine-tuning, since its base checkpoints are near chance &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;, and its authors advise keeping Choice questions under about 20 options &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[12]&lt;/a&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A fine-tuned classifier&lt;/td&gt;
&lt;td&gt;You have labelled data and the label set is stable. A 2024 study by Bucher and Martini found small fine-tuned models consistently outperformed zero-shot prompted LLMs on text classification &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[24]&lt;/a&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&quot;how-we-evaluate-a-decision-model-before-trusting-it&quot;&gt;How we evaluate a decision model before trusting it&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Build a held-out set from your own traffic&lt;/strong&gt;, labelled by people and never used for prompt tuning or threshold fitting. Laya’s results show why: spam, which was in its training mix, scores 0.993, against the held-out moderation result above &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[12]&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Measure discrimination and probability quality separately.&lt;/strong&gt; Accuracy and macro-F1 (which counts rare labels as much as common ones) for the first. For the second, Brier score and log loss, which penalise confident wrong answers, and calibration error with a reliability diagram like the one below.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fit a temperature on a separate calibration split&lt;/strong&gt; and report calibration before and after. A hosted model’s weights cannot be tuned, but its output probabilities can still be recalibrated on your split. Temperature scaling has a single parameter and is often enough &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[18]&lt;/a&gt;. Laya found overconfidence on some tasks and underconfidence on others &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[12]&lt;/a&gt;, so check each question type.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Report coverage at your error budget for each action separately&lt;/strong&gt;, since each action has its own threshold &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[5]&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Probe the edges.&lt;/strong&gt; Shuffle option order, include items with no correct option, ask a question and its negation, and plant instructions inside the state. TypeSafe’s own docs show a question and its negation on the same ticket returning 0.72 and 0.47, which sum to 1.19 &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[4]&lt;/a&gt;. They also note that injected content “can move the answer” &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[4]&lt;/a&gt;, which matters most when the model is a guardrail.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Measure latency on the real path&lt;/strong&gt;, at the median and the 95th percentile. Laya’s 33 to 40 ms is one question on a T4 GPU, depending on the checkpoint; on CPU its card gives 193 to 464 ms &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[11]&lt;/a&gt;. Jev’s figures are hosted round trips.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Plan for drift.&lt;/strong&gt; Pin the model version, since TypeSafe warns that an alias moves when a new release ships &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[3]&lt;/a&gt;. Log every probability, and recheck calibration on fresh labelled samples on a schedule. Calibration fitted on past data degrades when the inputs shift &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[19]&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;A reliability diagram plots predicted probability against observed frequency, with the diagonal as perfect calibration. The bins start below the diagonal, meaning the model is overconfident. A single temperature dial turns and the bins settle close to the diagonal: for Laya&apos;s English checkpoint the model card reports mean expected calibration error falling from 0.466 to 0.081. The bars sketch the pattern; only the error figures are measured. In general, when the inputs shift, a fit made on past data drifts off the diagonal again, so calibration has to be rechecked on fresh labels.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;One temperature, fitted on a calibration split, pulls the bins close to the diagonal. A shift in the inputs pulls them off again, so the fit needs rechecking on fresh labels. The bars sketch the pattern; the error figures are from Laya’s model card [11].&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;where-this-fits-in-how-we-build&quot;&gt;Where this fits in how we build&lt;/h2&gt;
&lt;p&gt;We send classification to small trained classifiers and keep LLMs for work that needs language. Typed decision models share that view: a decision your software branches on should come from a model trained to return a calibrated label. Our view is that the most dependable path is still a small model fine-tuned on your own labelled data, for the reasons in the table above, and the gap between Laya’s zero-shot and fine-tuned results points the same way. We have not found a comparison of Jev with a classifier fine-tuned on a buyer’s own data, so that is the test worth running. Since Jev’s weights cannot be tuned per account &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/#sources&quot;&gt;[3]&lt;/a&gt;, the tuning moves into how you word the questions and criteria, and into the thresholds in your code.&lt;/p&gt;
&lt;p&gt;In practice we fine-tune encoder classifiers on a client’s data, calibrate and threshold their outputs, gate each action by confidence, and send the uncertain tail to an LLM or a person. Jev and Laya are worth testing against that baseline, on a held-out set, before either replaces it. If you have a classification step that is slow, expensive or hard to trust, &lt;a href=&quot;https://dravin.ai/contact/&quot;&gt;tell us about it&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;sources&quot;&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;TypeSafe AI, “Introducing System One Models &amp;amp; Jev”, 15 September 2026. &lt;a href=&quot;https://typesafe.ai/blog/introducing-system-one-models-and-jev&quot;&gt;typesafe.ai&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe docs, “System One”. &lt;a href=&quot;https://docs.typesafe.ai/concepts/system-one&quot;&gt;docs.typesafe.ai/concepts/system-one&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe docs, “Models” (jev-1.13.0, rate limits, fine-tuning, data handling, version pinning). &lt;a href=&quot;https://docs.typesafe.ai/models&quot;&gt;docs.typesafe.ai/models&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe docs, “Jev 1.13 jaggedness”, last reviewed 17 September 2026. &lt;a href=&quot;https://docs.typesafe.ai/model-jaggedness/jev-1.13&quot;&gt;docs.typesafe.ai/model-jaggedness/jev-1.13&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe docs, “Confidence”. &lt;a href=&quot;https://docs.typesafe.ai/confidence&quot;&gt;docs.typesafe.ai/confidence&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe docs, “Jev with coding agents”. &lt;a href=&quot;https://docs.typesafe.ai/introduction/coding-agents&quot;&gt;docs.typesafe.ai/introduction/coding-agents&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe docs, “AI primer”. &lt;a href=&quot;https://docs.typesafe.ai/introduction/machine-learning-primer&quot;&gt;docs.typesafe.ai/introduction/machine-learning-primer&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe docs, “Introduction”. &lt;a href=&quot;https://docs.typesafe.ai/introduction&quot;&gt;docs.typesafe.ai/introduction&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe docs, “Primitives”. &lt;a href=&quot;https://docs.typesafe.ai/primitives&quot;&gt;docs.typesafe.ai/primitives&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;TypeSafe, workflow evals. &lt;a href=&quot;https://evals.typesafe.ai/&quot;&gt;evals.typesafe.ai&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Convai Innovations, Laya model card (v0.3.7). &lt;a href=&quot;https://huggingface.co/convaiinnovations/laya&quot;&gt;huggingface.co/convaiinnovations/laya&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Laya, BENCHMARKS.md. &lt;a href=&quot;https://github.com/NandhaKishorM/laya/blob/main/BENCHMARKS.md&quot;&gt;github.com/NandhaKishorM/laya&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Laya on PyPI, first release 18 September 2026. &lt;a href=&quot;https://pypi.org/project/laya/&quot;&gt;pypi.org/project/laya&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;jev-benchmarks, pre-registered pilot, September 2026. &lt;a href=&quot;https://github.com/AbdelStark/jev-benchmarks&quot;&gt;github.com/AbdelStark/jev-benchmarks&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;decision-model-benchmark, README and v2 report of record, September 2026. &lt;a href=&quot;https://github.com/nibzard/decision-model-benchmark&quot;&gt;github.com/nibzard/decision-model-benchmark&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Tim Fernholz, “A new kind of AI model from a ChatGPT inventor is thrilling developers”, TechCrunch, 18 September 2026. &lt;a href=&quot;https://techcrunch.com/2026/09/18/a-new-kind-of-ai-model-from-a-chatgpt-inventor-is-thrilling-developers/&quot;&gt;techcrunch.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Thomas Claburn, “TypeSafe AI debuts model for machines that plays Doom”, The Register, 16 September 2026. &lt;a href=&quot;https://www.theregister.com/ai-and-ml/2026/09/16/typesafe-ai-debuts-model-for-machines-that-plays-doom/5296711&quot;&gt;theregister.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Guo, Pleiss, Sun and Weinberger, “On Calibration of Modern Neural Networks”, ICML 2017. &lt;a href=&quot;https://arxiv.org/abs/1706.04599&quot;&gt;arXiv:1706.04599&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Ovadia et al., “Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift”, NeurIPS 2019. &lt;a href=&quot;https://arxiv.org/abs/1906.02530&quot;&gt;arXiv:1906.02530&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Xiong et al., “Can LLMs Express Their Uncertainty?”, ICLR 2024. &lt;a href=&quot;https://arxiv.org/abs/2306.13063&quot;&gt;arXiv:2306.13063&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Tian et al., “Just Ask for Calibration”, 2023. &lt;a href=&quot;https://arxiv.org/abs/2305.14975&quot;&gt;arXiv:2305.14975&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Kadavath et al., “Language Models (Mostly) Know What They Know”, 2022. &lt;a href=&quot;https://arxiv.org/abs/2207.05221&quot;&gt;arXiv:2207.05221&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Geifman and El-Yaniv, “Selective Classification for Deep Neural Networks”, NeurIPS 2017. &lt;a href=&quot;https://arxiv.org/abs/1705.08500&quot;&gt;arXiv:1705.08500&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;Bucher and Martini, “Fine-Tuned ‘Small’ LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification”, 2024. &lt;a href=&quot;https://arxiv.org/abs/2406.08660&quot;&gt;arXiv:2406.08660&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This post first appeared on &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/&quot;&gt;Dravin AI&lt;/a&gt;.&lt;/p&gt;</content:encoded><category>Typed decision models</category><category>Classification</category><category>Calibration</category><category>Model evaluation</category><author>info@dravin.ai (Dravin AI)</author></item><item><title>The 90 firms that fix AI-generated code</title><link>https://dravin.ai/blog/who-cleans-up-ai-generated-code/</link><guid isPermaLink="true">https://dravin.ai/blog/who-cleans-up-ai-generated-code/</guid><description>What the 90 firms say breaks in AI-built apps, how they package the work, what none of them sells, and a checklist to use on any of them, including us.</description><pubDate>Mon, 21 Sep 2026 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;We found 90 firms that fix apps built with AI tools. Over half name security as the problem; only two name duplicated logic.&lt;/li&gt;&lt;li&gt;Not one of the 90 names runaway AI API spend as something they fix.&lt;/li&gt;&lt;li&gt;The market has settled on one sequence: read the code, stabilise what is dangerous, refactor or rebuild what will not hold, then watch it.&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;In September 2026 we went looking for companies that sell cleanup of AI-generated code: apps built with Lovable, Bolt, Cursor, Replit, v0, Claude Code and similar tools, by people who then got stuck. We found 90. This post is what those 90 say they fix, how they package the work, and what we think a buyer should demand from any of them.&lt;/p&gt;
&lt;h2 id=&quot;why-this-category-exists&quot;&gt;Why this category exists&lt;/h2&gt;
&lt;p&gt;In our reading of the 90 service descriptions, the typical caller in 2026 is a business owner, a product manager or a COO who built something real with an AI tool and then met real users, a payment that fails, or a security email from a stranger.&lt;/p&gt;
&lt;p&gt;The demand is visible in public. On r/ExperiencedDevs, the thread “Getting more calls to fix ai generated codebases than actual new builds lately” reached 405 upvotes and 102 comments; one comment: “Cleanup contracts are gonna be a whole market segment.” On r/SaaS, an MVP dev shop said it lost half its pipeline to Claude Code in 2025; of the prospects who tried it instead, about a third shipped, a third broke in production, and three came back for cleanups. On r/vibecoding, a post pointing to a vibe-code fix service, asking builders about getting their “mess” fixed, drew 1,086 upvotes and 125 comments.&lt;/p&gt;
&lt;p&gt;The failure rate is measurable. Veracode’s 2025 report found that AI-generated code introduced security flaws in 45% of its tests. How Bad Is It?, a public URL scanner for Lovable, Bolt and v0 apps, publishes a running counter: 4,271 apps reviewed, median score 38 out of 100.&lt;/p&gt;
&lt;figure&gt;&lt;ul&gt;&lt;li&gt;&lt;strong&gt;45%&lt;/strong&gt; of Veracode tests where AI-generated code introduced a security flaw. Source: Veracode, 2025&lt;/li&gt;&lt;li&gt;&lt;strong&gt;38&lt;/strong&gt; median score out of 100 across 4,271 Lovable, Bolt and v0 apps scanned. Source: How Bad Is It?&lt;/li&gt;&lt;/ul&gt;&lt;/figure&gt;
&lt;h2 id=&quot;who-the-90-are&quot;&gt;Who the 90 are&lt;/h2&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Type of firm&lt;/th&gt;
&lt;th&gt;Count&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Established dev agencies that added a cleanup service line&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Human audit specialists&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agencies formed for rescue work&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Automated scanners&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Solo fixers and freelance studios&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Marketplaces and expert networks&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Twenty are in the USA, 18 in Europe outside the UK, 7 in India, 6 in the UK, 4 in Canada, 11 elsewhere, and 24 give no location. Size is mostly undisclosed (53 of 90); of the rest, 6 have 250 or more people, 10 are mid-sized and 21 are small, micro or one-person.&lt;/p&gt;
&lt;p&gt;It is a young and thin market. We could open 75 of the 90 websites; 15 are known only from directory listings, and three domains (Vibe App Rescue, VibeCheckAudit, VibeFixLab) no longer resolve. Fifty-five name an AI builder in their service copy; Lovable leads (50), then Cursor (46), Bolt (44) and Replit (41).&lt;/p&gt;
&lt;h2 id=&quot;what-they-say-breaks&quot;&gt;What they say breaks&lt;/h2&gt;
&lt;p&gt;We counted which failure classes each firm’s service description names.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Failure class named in the service description&lt;/th&gt;
&lt;th&gt;Entries (of 90, scanners included)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Security in general (vulnerabilities, OWASP, hardening)&lt;/td&gt;
&lt;td&gt;48&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Missing or broken tests&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Performance and scalability&lt;/td&gt;
&lt;td&gt;25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No deploy pipeline, CI/CD or hosting setup&lt;/td&gt;
&lt;td&gt;23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authentication and access control (including Supabase RLS)&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data model or database problems&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Secrets and API keys in code&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monitoring and observability&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Documentation&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Payments and Stripe&lt;/td&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Duplicated logic and coupling&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hallucinated APIs&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runaway AI API spend&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Security is the sales pitch. Over half the market names it, and seven of the ten automated scanners are security scanners: Supabase row-level-security probes, leaked-key detection, public buckets. Beesoul, a US-listed audit shop (figures from its directory listing), reports that 10.3% of the Lovable apps it audited had critical RLS vulnerabilities, and that a typical MVP has 8 to 14 findings. Sidekick Interactive’s case study is typical of the genre: exposed API keys, database holes and broken payments in one app.&lt;/p&gt;
&lt;p&gt;Structure is under-sold. Only two of the 90 entries, one of them a scanner, mention duplicated logic or unsafe coupling by name, though that is what makes an AI-generated codebase slow and risky to change. A security patch is usually local; untangling several copies of the same auth check, each slightly different, is where the work goes.&lt;/p&gt;
&lt;p&gt;Nobody sells a spending cap. Not one of the 90 names uncontrolled AI API usage as something they fix; the two that mention “cost” at all mean something else. Yet on r/cursor, a thread about a Cursor bill that spiked within one hour, after a PM asked the agent to tag 87 tasks, reached 228 upvotes. An app that calls a model on every request with no budget, rate limit or per-user cap is a bill waiting to happen.&lt;/p&gt;
&lt;h2 id=&quot;how-they-package-it&quot;&gt;How they package it&lt;/h2&gt;
&lt;p&gt;The 80 human-service firms (the 90 minus the 10 scanners) use four patterns, often more than one at once.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Packaging pattern&lt;/th&gt;
&lt;th&gt;Firms (of 80)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Audit, assessment or review named as the first step&lt;/td&gt;
&lt;td&gt;61&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audit or assessment promised within 48 hours&lt;/td&gt;
&lt;td&gt;10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Work described in sprints&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rebuild or re-architecture offered as one option&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Explicitly against rebuilding&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Retainer, maintenance or monitoring mentioned&lt;/td&gt;
&lt;td&gt;16&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Audit first.&lt;/strong&gt; &lt;a href=&quot;https://vallettasoftware.com/vibe-coding-cleanup&quot;&gt;Valletta Software&lt;/a&gt; (Malta) is the cleanest version: a senior engineer reviews the repository across eight areas, delivers the audit in 48 hours with a debrief call, then scopes the cleanup, typically six to eight weeks. &lt;a href=&quot;https://www.metacto.com/solutions/product-development/vibe-rescue&quot;&gt;MetaCTO&lt;/a&gt; runs a 48-hour code audit, then promises production-ready in 30 days, with weekly demos.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fixed-scope rescue.&lt;/strong&gt; &lt;a href=&quot;https://axonbuild.com/&quot;&gt;AxonBuild&lt;/a&gt; sells a 10-working-day production-engineering sprint covering auth, secrets, payments, CI/CD and monitoring, ending in a verified handover with two weeks of defect cover and 30 days of async support. &lt;a href=&quot;https://relux.works/en/vibe-code-rescue/&quot;&gt;Relux Works&lt;/a&gt; (Armenia) runs a one-week audit, then a two-to-three-week stabilisation sprint, then what it calls “de-vibe-coding”: moving the app to a production architecture.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Rebuild.&lt;/strong&gt; Usually offered for the broken parts only, and five firms say outright that they avoid it. &lt;a href=&quot;https://bitnoise.pl/services/vibe-code-rescue&quot;&gt;Bitnoise&lt;/a&gt; (Poznań) shows the whole sequence: audit in 3 to 7 business days, stabilisation sprint 2 to 4 weeks, re-architecture 6 to 12 weeks.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Retainer.&lt;/strong&gt; &lt;a href=&quot;https://vibe-audit.com/&quot;&gt;Vibe-Audit&lt;/a&gt; (Barcelona, a one-person shop) runs a pre-launch security sweep delivered in 48 hours, done-for-you fixes, and then ongoing security monitoring.&lt;/p&gt;
&lt;p&gt;Side by side, the market has settled on one sequence: read the code, stabilise what is dangerous, refactor or rebuild what will not hold, then watch it.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;How cleanup is sold&lt;/p&gt;&lt;ol&gt;&lt;li&gt;Read the code: audit or review first&lt;/li&gt;&lt;li&gt;Stabilise: what is dangerous now&lt;/li&gt;&lt;li&gt;Refactor or rebuild: what will not hold&lt;/li&gt;&lt;li&gt;Watch it: retainer or monitoring&lt;/li&gt;&lt;/ol&gt;&lt;figcaption&gt;The sequence the 80 human-service firms describe, often combining more than one pattern.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;what-a-good-cleanup-engagement-includes&quot;&gt;What a good cleanup engagement includes&lt;/h2&gt;
&lt;p&gt;Use this on anyone you are considering, including us.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;A written review before any commitment, with findings ranked by severity and a verdict: ship, patch, refactor or rewrite. If the verdict is “not worth fixing”, they should say so.&lt;/li&gt;
&lt;li&gt;An experienced engineer who has actually read the code. A scanner finds leaked keys; it does not find the auth check that is bypassed on one route.&lt;/li&gt;
&lt;li&gt;Security tested against the running app, not only the repository: authentication, per-row authorisation, secrets rotated rather than just deleted from git, public storage buckets closed.&lt;/li&gt;
&lt;li&gt;Characterisation tests written before refactoring, so current behaviour is pinned and the cleanup cannot silently change it.&lt;/li&gt;
&lt;li&gt;Duplicated logic collapsed and the data model reviewed, because that decides whether the next feature takes a day or a month.&lt;/li&gt;
&lt;li&gt;A deploy pipeline and monitoring at the end: staging, CI, error tracking, backups. Working on one person’s laptop does not count as done.&lt;/li&gt;
&lt;li&gt;AI API spend bounded: budgets, rate limits, per-user caps and an alert. Nobody in the 90 sells this; ask for it anyway.&lt;/li&gt;
&lt;li&gt;Handover: documentation, and you holding the repository, the cloud accounts and the keys. You own the code from day one.&lt;/li&gt;
&lt;li&gt;A fixed scope with a definition of done and a defect-cover window afterwards.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;how-we-do-it&quot;&gt;How we do it&lt;/h2&gt;
&lt;p&gt;Dravin AI is an AI software company. We build AI systems and software, and we clean up &lt;a href=&quot;https://dravin.ai/services/#code-quality&quot;&gt;AI-built apps&lt;/a&gt;. Against the list above, this is what we commit to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;A written report first.&lt;/strong&gt; Read-only access to one repository, and within 48 hours a written report with findings ranked by severity: the file, the risk and the fix, and what to leave alone. If nothing needs fixing, it says so. There is no obligation to go further.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Secrets.&lt;/strong&gt; A secret we find in your repository is reported the same day, ahead of the report, and rotated within 24 hours.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Auth and data access.&lt;/strong&gt; Every route’s auth check, database access policies included, is traced in the code. Testing the running app is agreed on the call, because it needs access beyond the repository.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tests before refactoring.&lt;/strong&gt; Tests on the paths that make money, run on every pull request, and structure changed in small, tested steps.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Deploys.&lt;/strong&gt; From the repository, with a rollback that has been tried.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI API spend.&lt;/strong&gt; Timeouts, budgets and retry caps on every model call.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Ownership.&lt;/strong&gt; The code is yours from the first commit, we work on branches, never main, and our access is removed within seven days of handover, confirmed in writing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Defect cover.&lt;/strong&gt; Anything that misses the statement of work, reported in writing within 30 days of handover, is corrected under it.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;An engagement starts with a call about what you built, what it does in production and what worries you. If that is an app that got you to traction, &lt;a href=&quot;https://cal.com/dravin-ai/intro&quot;&gt;book one&lt;/a&gt; or &lt;a href=&quot;https://dravin.ai/review/&quot;&gt;request a code review&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;method&quot;&gt;Method&lt;/h2&gt;
&lt;p&gt;Every count comes from the cleanup list in our research workbook: 90 firms, each with a segment, a region, a size band, a description of what it does, its delivery model and its stated turnaround. We read the list on 22 September 2026. The workbook itself is not published; these are the counting rules.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Type, region, size&lt;/strong&gt;: counts of each firm’s segment, region and size band. “Small, micro or one-person” is Small (10 to 49) plus Micro (2 to 9) plus Solo.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reachability&lt;/strong&gt;: whether we could fetch the firm’s site or know it only from a directory listing, and whether its domain still resolves.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Tool mentions&lt;/strong&gt;: a case-insensitive search of each firm’s description for each builder’s name.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Failure classes&lt;/strong&gt;: a case-insensitive keyword search of the description only. Security: secur, vuln, OWASP, pentest. Tests: test, QA; Sherlock Forensics matches only through “pentest” and is excluded, so 25 rather than the mechanical 26. Performance: perf, scalab, bottleneck. Deploy: CI/CD, CI, deploy, pipeline, hosting, DevOps. Auth: auth, access control, RLS. Data model: data model, database, DB, schema, persistence, RLS. Secrets: secret, key(s), credential. Monitoring: monitor, observab, Sentry. Documentation: doc, docs, document or documentation as whole words. Payments: Stripe, payment. Duplication: duplic, coupling, dead code, untangle. Hallucination: hallucinat. AI spend: cost, spend, token, bill (the two hits, *instinctools’ “cost-benefit” and Vibe Code Rescue’s “perf &amp;amp; cost”, are not API spend). Scanner types: the descriptions of the 10 automated scanners.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Packaging&lt;/strong&gt;: over the 80 firms that are not automated scanners, searching the description, the delivery model and the stated turnaround unless a rule says otherwise. Audit-first: audit, assess, review, scan, diagnos or scorecard. 48 hours: the turnaround promises an audit, assessment, scan or check result in 48 hours or less; fix times and reply times excluded. The ten: MetaCTO, Pfaff Digital, Sonder, VibeAudits, Mitrix Technology, Valletta Software, ShipClarity, Vibe-Audit, VibeCodeBlue and one solo fixer. Sprints: “sprint” anywhere in those three fields. Rebuild offered: rebuild, re-architect, rewrite or full build (11 firms), minus the two of them that also say “without rebuild”, “no rewrites”, “not rebuild”, “diagnosis-over-rebuild” or “without starting from scratch”. Against rebuilding: one of those five phrases in those three fields, the firm’s hero line or the workbook’s notes on how each firm positions itself (5 firms). Retainer: retainer, maintenance, ongoing, monitoring, subscription, defect cover or post-rescue support (16 firms).&lt;/li&gt;
&lt;li&gt;Named offers, turnarounds and claims are quoted from what each firm publishes, unaudited.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Reddit and studies&lt;/strong&gt;: r/ExperiencedDevs (405 upvotes, 102 comments), r/SaaS (65 upvotes, 66 comments) and r/cursor (228 upvotes) from our AI-agency landscape research (48 searches, 2,837 posts, 17 September 2026); r/vibecoding (1,086 upvotes, 125 comments, 24 December 2025) from a separate Reddit sweep for the workbook; Veracode’s 45% from the same landscape research.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This post first appeared on &lt;a href=&quot;https://dravin.ai/blog/who-cleans-up-ai-generated-code/&quot;&gt;Dravin AI&lt;/a&gt;.&lt;/p&gt;</content:encoded><category>Code cleanup</category><category>Vibe coding</category><category>Security review</category><category>Market research</category><author>info@dravin.ai (Dravin AI)</author></item><item><title>What 843 AI agencies actually sell</title><link>https://dravin.ai/blog/what-843-ai-agencies-actually-sell/</link><guid isPermaLink="true">https://dravin.ai/blog/what-843-ai-agencies-actually-sell/</guid><description>We mapped 843 AI agencies across India, the US and Europe: what they sell, where they are, which offers are crowded, which are thin, and what to check.</description><pubDate>Mon, 21 Sep 2026 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Automation and custom builds account for 457 of the 843 firms we mapped. Only 31 sell audit, security or testing.&lt;/li&gt;&lt;li&gt;The thinnest part of the market is what a system needs once it is live: evaluation, monitoring and someone who answers when it breaks.&lt;/li&gt;&lt;li&gt;Size is unknown for 569 of the 843, so ask who will do the work, what happens after the build and how they will know it works.&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;Between 17 and 22 September 2026 we built a workbook of every AI agency, consultancy and AI-native delivery firm we could find: 843 firms across India, the US, Europe and a long tail of other regions. We wanted to see the market the way a buyer does. Every count below was computed from that workbook; the method is at the end.&lt;/p&gt;
&lt;p&gt;The short version: most of the market sells the same two things, automation and custom builds, which together account for 457 of the 843 firms. Only 31 sell audit, security or testing. And the work a team needs once something is in production, evaluation, monitoring and someone on the hook when it breaks, is the thinnest part of the map.&lt;/p&gt;
&lt;h2 id=&quot;what-we-counted-and-how&quot;&gt;What we counted and how&lt;/h2&gt;
&lt;p&gt;Eight research passes (Americas and ANZ, India, Europe and the Middle East, design-led studios, consulting and advisory, Asian, vertical and voice agencies, AI-code cleanup, freelancers), directory sweeps (Clutch, GoodFirms, DesignRush, Sortlist, and the n8n, Make and Zapier directories), and 211 firms carried over from our earlier landscape note. Duplicates were merged on domain. Of the 843 sites, 819 loaded and 619 homepages were read on the research date.&lt;/p&gt;
&lt;p&gt;Each firm has a region, one of 35 segments, a size band and a description of what it does. Kept separately, outside the 843: 90 firms that fix &lt;a href=&quot;https://dravin.ai/blog/who-cleans-up-ai-generated-code/&quot;&gt;AI-generated code&lt;/a&gt;, 103 independent consultants, and a 40-row catalogue of every offer type we saw, each rated by how many credible sellers it has.&lt;/p&gt;
&lt;p&gt;One caveat: size is known for only 274 of the 843. Most firms do not say how big they are, which matters later.&lt;/p&gt;
&lt;h2 id=&quot;the-regional-shape&quot;&gt;The regional shape&lt;/h2&gt;
&lt;figure&gt;&lt;p id=&quot;chart-firms-by-region-843-firms&quot;&gt;Firms by region · 843 firms&lt;/p&gt;&lt;ul&gt;&lt;li&gt;India: 209&lt;/li&gt;&lt;li&gt;USA: 162&lt;/li&gt;&lt;li&gt;Europe (as tagged): 146&lt;/li&gt;&lt;li&gt;UK: 55&lt;/li&gt;&lt;li&gt;Australia and NZ: 41&lt;/li&gt;&lt;li&gt;Asia (other): 39&lt;/li&gt;&lt;li&gt;Middle East: 35&lt;/li&gt;&lt;li&gt;Latin America: 32&lt;/li&gt;&lt;li&gt;Canada: 28&lt;/li&gt;&lt;li&gt;Africa: 14&lt;/li&gt;&lt;li&gt;Global, multi-country, other or unknown: 82&lt;/li&gt;&lt;/ul&gt;&lt;figcaption&gt;All 843 firms by the region each is tagged with. Europe is the tag as found and holds some UK-based firms; the Method breaks down the last row.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The regions sell different things. India’s 209 are build shops: 46 dev studios, 27 boutiques, 21 mid-sized IT firms, 29 automation agencies. The US mix is flatter: 21 vertical agencies, 21 automation agencies, 19 dev studios, 11 forward-deployed engineering firms and all 9 big-consultancy AI arms. Europe holds the incumbents: 31 IT majors and consultancies, 15 nearshore firms, 14 data engineering firms. The UK leans to advice: 6 of the 14 governance firms and 5 of the 11 training firms are British.&lt;/p&gt;
&lt;h2 id=&quot;what-firms-sell&quot;&gt;What firms sell&lt;/h2&gt;
&lt;p&gt;The 35 segments group into nine offer families. The grouping is ours; every firm lands in exactly one.&lt;/p&gt;
&lt;figure&gt;&lt;p id=&quot;chart-firms-by-offer-family-843-firms&quot;&gt;Firms by offer family · 843 firms&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Automation and agents as a service: 246&lt;/li&gt;&lt;li&gt;Custom build: 211&lt;/li&gt;&lt;li&gt;Embedded and enterprise delivery: 123&lt;/li&gt;&lt;li&gt;Advice and strategy: 117&lt;/li&gt;&lt;li&gt;Data engineering: 62&lt;/li&gt;&lt;li&gt;Audit, security and QA: 31&lt;/li&gt;&lt;li&gt;Platforms and tooling: 26&lt;/li&gt;&lt;li&gt;Training and community: 18&lt;/li&gt;&lt;li&gt;Other: 9&lt;/li&gt;&lt;/ul&gt;&lt;figcaption&gt;All 843 firms by offer family. Audit, security and QA, the rust bar, is 31 firms; automation and custom build together are 457. The table below lists the segments inside each family.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Offer family&lt;/th&gt;
&lt;th&gt;Firms&lt;/th&gt;
&lt;th&gt;Segments inside&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Automation and agents as a service&lt;/td&gt;
&lt;td&gt;246&lt;/td&gt;
&lt;td&gt;142 SMB automation agencies, 70 vertical agencies, 26 voice-AI agencies, 4 productised services, 4 automation-and-coaching hybrids&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom build&lt;/td&gt;
&lt;td&gt;211&lt;/td&gt;
&lt;td&gt;123 dev studios, 27 Indian boutiques, 26 design-led studios, 15 nearshore firms, 11 engineering firms with an AI practice, 9 AI-native delivery firms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Embedded and enterprise delivery&lt;/td&gt;
&lt;td&gt;123&lt;/td&gt;
&lt;td&gt;35 FDE firms, 31 IT majors based in Europe, 21 Indian mid-sized IT firms, 9 big-consultancy arms, 8 established Indian engineering firms, 7 platform-plus-services firms, 6 IT-major practices, 6 lab and hyperscaler arms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Advice and strategy&lt;/td&gt;
&lt;td&gt;117&lt;/td&gt;
&lt;td&gt;51 AI consultancies, 19 research and advisory, 17 strategy consultancies, 12 industry consultancies, 12 MSPs, 6 fractional-leadership firms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data engineering&lt;/td&gt;
&lt;td&gt;62&lt;/td&gt;
&lt;td&gt;50 data and AI engineering firms, 12 analytics consultancies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audit, security and QA&lt;/td&gt;
&lt;td&gt;31&lt;/td&gt;
&lt;td&gt;14 governance and compliance, 10 security and red-teaming, 7 QA and testing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platforms and tooling&lt;/td&gt;
&lt;td&gt;26&lt;/td&gt;
&lt;td&gt;15 AI-native startups, 11 developer-productivity and AI-impact platforms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training and community&lt;/td&gt;
&lt;td&gt;18&lt;/td&gt;
&lt;td&gt;11 training firms, 7 coaching communities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Other&lt;/td&gt;
&lt;td&gt;9&lt;/td&gt;
&lt;td&gt;Uncategorised&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Automation-as-a-service and custom build account for 457 of the 843; audit, security and QA are 31. That ratio is the most useful fact in the workbook. Almost everyone will build the thing. Very few will tell you whether the thing is safe, correct or still working a month later.&lt;/p&gt;
&lt;h2 id=&quot;what-is-crowded-and-what-is-thin&quot;&gt;What is crowded and what is thin&lt;/h2&gt;
&lt;p&gt;The catalogue rates each of the 40 offers by how many credible sellers we found: 22 crowded, 15 moderate, 3 thin.&lt;/p&gt;
&lt;p&gt;Crowded:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Custom builds&lt;/strong&gt; (agents, &lt;a href=&quot;https://dravin.ai/services/#rag&quot;&gt;RAG&lt;/a&gt; systems, &lt;a href=&quot;https://dravin.ai/services/#voice&quot;&gt;voice agents&lt;/a&gt;, internal tools, integrations) are the default offer of nearly every automation agency. The catalogue’s note is that capability is much the same everywhere; firms differ by vertical depth or proof.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Discovery and readiness audits&lt;/strong&gt; take the least capital to start, so almost every new operator leads with one.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Staff augmentation&lt;/strong&gt; is the default Indian services business, with hundreds of Indian shops selling the same thing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI support agents that resolve tickets&lt;/strong&gt; are sold by nine platforms, bundled into the inbox the buyer already owns.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data plumbing for AI&lt;/strong&gt; (pipelines, warehouses, vector databases) is offered by every data consultancy and is the easiest thing on the list to buy.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thin:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Evaluation harnesses, benchmarking and observability.&lt;/strong&gt; Dozens of funded tool vendors, very few independents doing the implementation, because it takes real ML judgement.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Data-quality and AI-data-readiness audits.&lt;/strong&gt; Crowded as software, thin as a service a person delivers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Managed agents:&lt;/strong&gt; monitoring, evals and a support SLA for AI systems already in production. Half a dozen vendors sell the dashboard; the human side, someone who answers when the agent breaks at 2am, is thin.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Codebase audit sits next to them, rated moderate: the automated side is crowded and well funded, and the human review is where the catalogue sees the gap. The cleanup list shows that gap starting to fill: 90 firms now sell cleanup of AI-generated code, 48 of them as a service line inside an existing dev agency and 13 as human audit specialists.&lt;/p&gt;
&lt;p&gt;The firm counts say the same thing. 457 firms build. 31 audit, secure or test. 11 are platforms in the developer-productivity and AI-impact segment, measuring what AI is doing to engineering teams. Meanwhile the landscape’s Reddit sweep found an r/ExperiencedDevs thread titled “Getting more calls to fix ai generated codebases than actual new builds lately” at 405 upvotes.&lt;/p&gt;
&lt;h2 id=&quot;what-to-check-before-hiring-anyone&quot;&gt;What to check before hiring anyone&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Ask who will be on the work.&lt;/strong&gt; Size is unknown for 569 of the 843, and a website rarely tells you whether you are talking to a two-person studio or a firm of 7,000. Ask how experienced the engineers are, who checks their code before it merges, and whether any of it is subcontracted.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Ask what happens after the build.&lt;/strong&gt; The market is crowded with builds and thin on evaluation, monitoring and support for systems already in production. Get a written answer on who maintains the thing, how a failure is detected and who answers when it breaks, before you sign for the thing.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Ask how they will know it works.&lt;/strong&gt; An agent or a RAG system without an evaluation harness is a demo. Few firms sell evaluation as a service, so ask to see the test set, the metrics and the regression checks for the system you are buying, and ask what happens when a model version changes underneath it.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;If you already have AI-generated code, get it read before you scale it.&lt;/strong&gt; 457 firms will happily build on top of it; 31 sell audit, security or testing, plus the 90 cleanup firms outside the main list. A careful read finds the auth check missing on one route, the key committed to the repository and the same logic copied into several places. Each is easier to fix before the next feature lands on top of it.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;what-we-do&quot;&gt;What we do&lt;/h2&gt;
&lt;p&gt;Dravin AI is an AI software company. We build AI systems and software. Coding agents draft in parallel, each in its own isolated environment, automated checks run on every pull request, and an engineer approves each change before it merges. We also audit code and clean up AI-built apps.&lt;/p&gt;
&lt;p&gt;What we build runs from agents, voice AI, fine-tuned models, classical ML and forecasting systems to web, mobile and desktop apps, and we deploy and run them in the cloud.&lt;/p&gt;
&lt;p&gt;We also work in the thin part of this map: evaluation harnesses, and AI-generated codebases made safe enough for a real team to own. An engagement starts with a call. If any of the above describes where you are, &lt;a href=&quot;https://cal.com/dravin-ai/intro&quot;&gt;book one&lt;/a&gt; and bring the repository.&lt;/p&gt;
&lt;h2 id=&quot;method&quot;&gt;Method&lt;/h2&gt;
&lt;p&gt;Every count comes from our research workbook, built between 17 and 22 September 2026 and counted on 22 September. The workbook itself is not published; these are the counting rules.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;843&lt;/strong&gt;: every firm on the main list; 211 of them came from our earlier landscape note.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Regions&lt;/strong&gt;: the region each firm is tagged with, exact match. The 82 is Global 33, Other 24, Unknown 23, Multi-country or remote 2. Europe is the tag as found; it holds some UK-based firms.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Segments and families&lt;/strong&gt;: each firm’s segment, exact match, filtered by region where one is named. The nine families are our grouping of the 35 segment values; the table is the mapping and it sums to 843. The 11 platforms are the segment “Dev-productivity / AI-impact platform”.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Size&lt;/strong&gt;: each firm’s size band: Solo 17, Micro 60, Small 97, Mid 76, Large 24, Unknown 569; 274 known.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;819 and 619&lt;/strong&gt;: sites that passed a link check, and homepages read on the research date.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;90, 48 and 13&lt;/strong&gt;: the firms on the cleanup list; those whose segment is a dev agency with a cleanup service line; those whose segment is a human audit of AI-generated code.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;103 consultants&lt;/strong&gt;: the firms on the list of independent consultants and freelancers.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;40 offers; 22, 15, 3&lt;/strong&gt;: the offer catalogue’s 40 entries, each rated crowded, moderate or thin by how many credible sellers we found. The crowded and thin notes paraphrase the catalogue’s notes on custom builds, discovery and readiness audits, staff augmentation, support agents, data plumbing, evaluation, data-quality audits, managed agents and codebase audit.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The r/ExperiencedDevs thread&lt;/strong&gt; (405 upvotes): from the Reddit sweep in our AI-agency landscape note.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This post first appeared on &lt;a href=&quot;https://dravin.ai/blog/what-843-ai-agencies-actually-sell/&quot;&gt;Dravin AI&lt;/a&gt;.&lt;/p&gt;</content:encoded><category>Market research</category><category>AI agencies</category><category>Code cleanup</category><author>info@dravin.ai (Dravin AI)</author></item></channel></rss>