<?xml version="1.0" encoding="UTF-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Dravin AI · Blog · LLM serving</title><description>Every post on LLM serving from the Dravin AI blog, newest first: long reads on building with current models and agents.</description><link>https://dravin.ai/blog/tags/llm-serving/</link><language>en-gb</language><image><url>https://dravin.ai/apple-touch-icon.png</url><title>Dravin AI · Blog · LLM serving</title><link>https://dravin.ai/blog/tags/llm-serving/</link></image><lastBuildDate>Tue, 22 Sep 2026 18:30:00 GMT</lastBuildDate><atom:link href="https://dravin.ai/blog/tags/llm-serving/rss.xml" rel="self" type="application/rss+xml"/><item><title>More users per GPU: a playbook for self-hosted LLM serving</title><link>https://dravin.ai/blog/raising-llm-serving-throughput/</link><guid isPermaLink="true">https://dravin.ai/blog/raising-llm-serving-throughput/</guid><description>What to measure first, then the levers that raise LLM serving throughput on your own GPUs, in the order we pull them and what each one trades away.</description><pubDate>Tue, 22 Sep 2026 18:30:00 GMT</pubDate><content:encoded>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Prefill keeps the GPU’s arithmetic busy while decode waits on memory, so most of the work is giving decode a bigger batch without letting prefill stall it.&lt;/li&gt;&lt;li&gt;Measure goodput first: the request rate each GPU serves while a target share of requests meet their latency limits.&lt;/li&gt;&lt;li&gt;Pull the levers in order of cost and risk: engine settings and prompt layout, then quantisation and speculative decoding, then parallelism and disaggregation.&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;A model that answers well in a notebook can be too slow and too expensive once hundreds of people use it at once. Before adding GPUs, we work through engine settings, prompt layout and a few serving techniques in a fixed order, and check each change against numbers we trust.&lt;/p&gt;
&lt;p&gt;One fact sits under everything here. Prefill, where the model reads the prompt, keeps the GPU’s arithmetic busy &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[1]&lt;/a&gt;, &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[2]&lt;/a&gt;. Decode, where it writes one token at a time per request, uses little of that arithmetic and mostly waits on memory bandwidth &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[3]&lt;/a&gt;. Most throughput work comes down to giving decode a bigger batch while stopping prefill from stalling it.&lt;/p&gt;
&lt;p&gt;We serve open-weight models privately, quantised, on vLLM. Versions are current as of 23 September 2026: vLLM 0.30.0, SGLang 0.5.20 and TensorRT-LLM 1.2.1, each the latest stable release &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[4]&lt;/a&gt;. Flags and metric names below are vLLM’s unless we name another engine.&lt;/p&gt;
&lt;p&gt;The order follows cost and risk. Batching, KV cache sizing, &lt;a href=&quot;https://dravin.ai/services/#inference&quot;&gt;prefix caching&lt;/a&gt; and chunk size are configuration on a current engine, and vLLM 0.30.0 turns most of them on by default, so the work is sizing and checking them. Quantisation and speculative decoding change the arithmetic the model runs, so they need your own evals and your own load. Parallelism and disaggregation change the topology, so they come last.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Lever&lt;/th&gt;
&lt;th&gt;What it buys&lt;/th&gt;
&lt;th&gt;What it costs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Batch and KV cache sizing&lt;/td&gt;
&lt;td&gt;More requests in flight&lt;/td&gt;
&lt;td&gt;Per-user latency rises with the batch&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prefix caching and routing&lt;/td&gt;
&lt;td&gt;Skipped prefill for shared prompts&lt;/td&gt;
&lt;td&gt;Cache memory, and uneven load across replicas&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chunk size&lt;/td&gt;
&lt;td&gt;Steady streams while long prompts load&lt;/td&gt;
&lt;td&gt;A later first token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantisation&lt;/td&gt;
&lt;td&gt;Fewer bytes per token, room for more requests&lt;/td&gt;
&lt;td&gt;Accuracy risk; formats tied to GPU generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Speculative decoding&lt;/td&gt;
&lt;td&gt;Faster streams at low load&lt;/td&gt;
&lt;td&gt;Spare compute; the gain fades as the batch fills&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Admission control&lt;/td&gt;
&lt;td&gt;Latency targets held under load&lt;/td&gt;
&lt;td&gt;Some requests deferred or rejected&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Replicas and expert parallelism&lt;/td&gt;
&lt;td&gt;Capacity beyond one GPU group&lt;/td&gt;
&lt;td&gt;More GPUs and fast interconnect&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Disaggregation&lt;/td&gt;
&lt;td&gt;Separate control of TTFT and ITL&lt;/td&gt;
&lt;td&gt;Two pools to size, RDMA, experimental support&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id=&quot;measure-before-you-tune&quot;&gt;Measure before you tune&lt;/h2&gt;
&lt;p&gt;Four numbers describe a serving system.&lt;/p&gt;
&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;What it measures&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;TTFT&lt;/td&gt;
&lt;td&gt;Time from sending a request to receiving the first token &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[5]&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;How long a user stares at a blank screen&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ITL and TPOT&lt;/td&gt;
&lt;td&gt;The gap between streamed outputs (ITL), or a request’s decode time spread over its output tokens (TPOT) &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[6]&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;How fast the answer streams&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output tokens per second&lt;/td&gt;
&lt;td&gt;Output tokens per second, summed over all concurrent requests &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[5]&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;What the hardware is worth to you&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Goodput&lt;/td&gt;
&lt;td&gt;The highest request rate per GPU served while a target share of requests meet their latency limits &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[7]&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;The number to optimise&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Tools disagree on whether ITL counts the first token, so NVIDIA’s guide says to “compare results only when definitions align” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[5]&lt;/a&gt;. Pick one tool and keep it.&lt;/p&gt;
&lt;p&gt;DistServe defines goodput as “the maximum request rate that can be served adhering to the SLO attainment goal (say, 90%) for each GPU provisioned” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[7]&lt;/a&gt;, where the SLO (service-level objective) is the latency limit you commit to. It leads the dashboard because raw tokens per second can keep rising while users wait longer. System throughput and per-user speed pull against each other, and most levers below trade one for the other.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;vllm bench serve&lt;/code&gt; measures all of this: set &lt;code&gt;--request-rate&lt;/code&gt; and &lt;code&gt;--burstiness&lt;/code&gt; to your arrival pattern, cap load with &lt;code&gt;--max-concurrency&lt;/code&gt;, and pass your TTFT and TPOT limits in milliseconds to &lt;code&gt;--goodput&lt;/code&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[6]&lt;/a&gt;. Report p50 and p99 for each; averages hide the slow tail. Use your own prompt lengths, output lengths and shared prefixes. Published benchmarks show which levers are worth trying; your own traffic shows how much each one gives.&lt;/p&gt;
&lt;p&gt;In production, watch &lt;code&gt;vllm:num_requests_waiting&lt;/code&gt; for queueing, &lt;code&gt;vllm:kv_cache_usage_perc&lt;/code&gt; and &lt;code&gt;vllm:num_preemptions&lt;/code&gt; for memory pressure, and &lt;code&gt;vllm:prefix_cache_hits&lt;/code&gt; against &lt;code&gt;vllm:prefix_cache_queries&lt;/code&gt; for cache reuse &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[8]&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;get-the-batch-and-the-kv-cache-right&quot;&gt;Get the batch and the KV cache right&lt;/h2&gt;
&lt;p&gt;The first lever is the engine. Early servers ran a batch until its longest request finished, so new requests queued behind it. Orca introduced “iteration-level scheduling”: the scheduler decides after every decoding step, so finished requests leave and new ones join at once. On a 175B model the paper reports 36.9 times the throughput of NVIDIA FasterTransformer at the same latency &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[9]&lt;/a&gt;. Every current engine does this, under the names continuous batching or in-flight batching &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[10]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The next constraint is the KV cache, the stored keys and values for every token of every running request. For a 13B model the PagedAttention paper works it out at 800 KB per token in FP16, so a single 2,048-token request can hold 1.6 GB &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[11]&lt;/a&gt;. The general form is 2 × layers × KV heads × head dimension × bytes per element, per token. It grows linearly with context length and concurrency, so the KV cache sets the batch ceiling. Grouped-query attention, where several query heads share one key-value head, shrinks the KV heads term &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[12]&lt;/a&gt;; a model’s config lists its KV head count, so check it before you choose the model.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;GPU memory as 48 fixed-size KV cache blocks. Static batch with one slab per request: four requests each reserve a 12-block slab for their maximum length and fill part of it; B has finished but its slab stays held until the batch ends, so queued requests wait and 17 of 48 blocks hold tokens. Continuous batching with a paged cache: requests take small scattered blocks step by step; B&apos;s blocks free after step 3, new request E takes all of them at step 4, and by step 6 41 of 48 blocks hold tokens.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;A slab reserves a request’s maximum length up front and holds it until the request, or the batch, ends. A paged cache adds blocks as tokens arrive and returns them the moment a request finishes.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Before PagedAttention, only 20.4 to 38.2% of KV cache memory held actual token states. Storing the cache in small pages, as an operating system pages virtual memory, brought waste close to zero and gave 2 to 4 times the throughput of FasterTransformer and Orca at the same latency &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[11]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;With a current engine, the work is sizing. vLLM 0.30.0 sets &lt;code&gt;gpu_memory_utilization&lt;/code&gt; to 0.92 and picks &lt;code&gt;max_num_seqs&lt;/code&gt; and &lt;code&gt;max_num_batched_tokens&lt;/code&gt; by GPU &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[13]&lt;/a&gt;. When the cache runs out, requests are preempted and redo their prefill; the documented fixes include a higher &lt;code&gt;gpu_memory_utilization&lt;/code&gt;, lower &lt;code&gt;max_num_seqs&lt;/code&gt; or &lt;code&gt;max_num_batched_tokens&lt;/code&gt;, or more tensor or pipeline parallelism to free memory for the cache &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[14]&lt;/a&gt;. A rising &lt;code&gt;vllm:num_preemptions&lt;/code&gt; counter is the first thing we look at. On a new deployment we benchmark at the defaults first, then change one setting per run.&lt;/p&gt;
&lt;p&gt;Leave CUDA graphs on. A decode step is a series of tiny GPU jobs, and graph replay skips the CPU cost of launching each one &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[15]&lt;/a&gt;; vLLM captures graphs by default, at some cost in memory and startup time &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[16]&lt;/a&gt;. Keep &lt;code&gt;--enforce-eager&lt;/code&gt; for debugging &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[14]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The cost: a bigger batch raises tokens per second but also each user’s latency, the trade-off Sarathi-Serve is named for &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[1]&lt;/a&gt;.&lt;/p&gt;

&lt;h2 id=&quot;stop-recomputing-the-same-prompt&quot;&gt;Stop recomputing the same prompt&lt;/h2&gt;
&lt;p&gt;Many workloads resend the same long prefix: a system prompt, tool definitions, a document, earlier turns of a chat. Prefix caching keeps that prefix’s KV cache, and a new request that shares it skips that part of prefill &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[17]&lt;/a&gt;. SGLang’s RadixAttention keeps prefixes in a prefix tree (a radix tree) with least-recently-used eviction &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[18]&lt;/a&gt;. Both engines turn it on by default in their current releases &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[13]&lt;/a&gt;, &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[19]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;It “only reduces the time of processing the queries (the prefilling phase)” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[17]&lt;/a&gt;, so it pays most where prompts are long and answers short: &lt;a href=&quot;https://dravin.ai/services/#rag&quot;&gt;RAG&lt;/a&gt;, agents and document chat.&lt;/p&gt;
&lt;p&gt;A cache only hits on an identical token prefix. Put stable content first (system prompt, tools, documents), variable content last, and keep timestamps and request IDs out of the top. In SGLang, &lt;code&gt;--schedule-policy lpm&lt;/code&gt; reorders the queue by longest prefix match to raise the hit rate &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[20]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Routing that ignores the cache scatters requests that share a prefix across replicas, and one published benchmark shows the cost &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[21]&lt;/a&gt;. It ran 8 replicas of a 32B model on 16 H100s, with shared-prefix traffic whose cache demand was six times what one replica could hold. A router that tracked each replica’s actual cache contents produced 8,730 output tokens per second against 4,429 for random or load-only routing, and cut the 90th-percentile TTFT from over 90 seconds to about half a second; a router that estimated cache contents from its own routing history reached 6,944 tokens per second and 31 seconds &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[21]&lt;/a&gt;. SGLang’s router keeps such an estimate, an approximate copy of each worker’s prefix tree &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[22]&lt;/a&gt;, and NVIDIA Dynamo 1.5.0, an orchestration layer that runs on top of engines such as vLLM, SGLang and TensorRT-LLM, routes “based on worker load and KV cache overlap” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[23]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The cost: cache affinity fights load balance, so good routers blend a cache score with a load score, and the cache takes memory that could hold more running requests.&lt;/p&gt;
&lt;h2 id=&quot;keep-long-prompts-from-stalling-everyone&quot;&gt;Keep long prompts from stalling everyone&lt;/h2&gt;
&lt;p&gt;One 50-page document arriving mid-stream can freeze every other user’s output while its prefill runs. Sarathi-Serve’s fix is chunked prefill: cut the prompt into pieces and mix each piece into ordinary decode steps. Against the vLLM of early 2024, the paper reports 2.6 times the serving capacity for a 7B model on one A100 and up to 3.7 times for a 34B model on two A100s &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[1]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;vLLM has since made chunked prefill the default and schedules pending decodes before any prefill &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[14]&lt;/a&gt;. What remains is the dial. In vLLM’s docs, smaller &lt;code&gt;max_num_batched_tokens&lt;/code&gt; values such as 2,048 “achieve better ITL”, larger ones “achieve better time to first token”, and values above 8,192 suit throughput, “especially for smaller models on large GPUs” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[14]&lt;/a&gt;. vLLM 0.30.0’s server already defaults to 8,192 on an H100 and 2,048 on an A100 &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[13]&lt;/a&gt;. SGLang calls the same dial &lt;code&gt;--chunked-prefill-size&lt;/code&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[20]&lt;/a&gt;. Set it against your written TTFT and ITL targets, measured at p99.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;Timeline of GPU steps on one replica with five users streaming tokens while a 30,000-token prompt arrives. Without chunking, the prompt runs as one long prefill step and all five token streams stop until it ends, so inter-token latency spikes; the new user then gets a first token. With chunked prefill, the prompt is cut into eight equal slices, each riding along with a normal decode step; the five streams keep ticking at a slightly slower rate and the new user’s first token arrives a little later. A chunk-size scale shows the trade: small chunks such as 2,048 tokens keep streams smooth and delay the first token, large chunks such as 8,192 give an earlier first token and bumpier streams.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;Chunked prefill cuts a long prompt into slices and runs each one alongside ordinary decode steps, so other users keep receiving tokens while it loads. The two chunk sizes marked are the examples in vLLM’s docs [14].&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&quot;spend-fewer-bytes-per-token&quot;&gt;Spend fewer bytes per token&lt;/h2&gt;
&lt;p&gt;Decode is limited by the bytes the GPU reads per token, so storing weights and cache in fewer bits gives more tokens per second.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;FP8 weights and activations&lt;/strong&gt; halve weight memory and need Ada (L40S), Hopper (H100, H200) or Blackwell (B200, GB200) GPUs &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[24]&lt;/a&gt;. One accuracy study ran more than 500,000 evaluations on an open model family at 8B, 70B and 405B, in vLLM on A6000, A100 and H100 GPUs. It found FP8 “effectively lossless across all model scales” and well-tuned INT8 within 1 to 3% &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[25]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;INT4 weight-only&lt;/strong&gt; (AWQ &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[26]&lt;/a&gt;, GPTQ &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[27]&lt;/a&gt;) cuts weight memory about four times against FP16. The same study found it the most cost-efficient when requests run one at a time, and 8-bit formats fastest under continuous batching &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[25]&lt;/a&gt;. Weight-only 4-bit speeds up memory-bound decode, and as the batch grows the GPU turns compute-bound: the authors of one 4-bit kernel report close to the full 4 times speed-up per layer up to batch 16 to 32, still significant but shrinking at 64 to 128, and up to 2.8 times end to end in vLLM &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[28]&lt;/a&gt;. For a busy server, that points to an 8-bit format first.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;NVFP4&lt;/strong&gt; stores 4-bit values in blocks of 16 that share an FP8 scale. NVIDIA reports about 3.5 times less memory than FP16 and 1.8 times less than FP8, and 1% or less accuracy loss moving a 671B mixture-of-experts (MoE) reasoning model from FP8 to NVFP4 &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[29]&lt;/a&gt;. FP4 arithmetic needs Blackwell, though older GPUs can still load NVFP4 weights: Nemotron 3.5 Lightning ships an NVFP4 checkpoint that runs on NVIDIA’s NVFP4 kernels across Blackwell, Hopper and Ampere &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[30]&lt;/a&gt;. On the older GPUs that saves memory and bandwidth, without the FP4 arithmetic.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The KV cache&lt;/strong&gt; can be quantised separately, which raises the number of requests that fit. vLLM supports an FP8 KV cache and recommends calibrated scales &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[31]&lt;/a&gt;. With Qwen3.8-27B on an RTX PRO 6000 Blackwell, the SGLang, Qwen and NVIDIA teams fitted 70 concurrent requests at 32K context with an NVFP4 KV cache, against 44 with FP8. At matched concurrency, peak-batch decode throughput rose 26 to 30%. SGLang still marks the support experimental &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[32]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The cost is accuracy and hardware lock-in. Formats are tied to GPU generations, uncalibrated KV scales are a risk, and upgrades can break checkpoints: vLLM 0.30.0 dropped GPTQ activation ordering (the “act-order” option), so check any GPTQ checkpoint before you upgrade &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[13]&lt;/a&gt;. We quantise before we add GPUs, and judge each format on the evaluation set of the product it serves.&lt;/p&gt;
&lt;h2 id=&quot;draft-tokens-while-there-is-spare-compute&quot;&gt;Draft tokens while there is spare compute&lt;/h2&gt;
&lt;p&gt;Speculative decoding lets a cheap drafter propose several tokens and the target model verify them in one pass, keeping the target model’s output distribution &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[3]&lt;/a&gt;. Leviathan and colleagues report 2 to 3 times on an 11B encoder-decoder model &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[3]&lt;/a&gt;. Each target pass yields (1 − α&lt;sup&gt;γ+1&lt;/sup&gt;) / (1 − α) tokens on average, where α is the acceptance rate and γ the number of drafted tokens: at α = 0.8 and γ = 4, about 3.4. The price is extra arithmetic: each step runs γ+1 positions through the target model in parallel, so “the number of concurrent arithmetic operations grows by a factor of γ+1”, and the method helps only when compute is spare &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[3]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;So the gain shrinks as the batch fills, and vLLM aims speculation at “medium-to-low QPS, memory-bound workloads” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[33]&lt;/a&gt;. EAGLE-3, a small draft head that reads the target model’s own hidden states, reports up to 6.5 times per-request speed-up on its best task and 1.38 times throughput at batch 64 in SGLang &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[34]&lt;/a&gt;. Plan a busy server around the second number.&lt;/p&gt;
&lt;figure&gt;&lt;p id=&quot;chart-per-user-speed-up-from-eagle-3-1-by-concurrency&quot;&gt;Per-user speed-up from EAGLE 3.1, by concurrency&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Concurrency 1: 2.03&lt;/li&gt;&lt;li&gt;Concurrency 4: 1.71&lt;/li&gt;&lt;li&gt;Concurrency 16: 1.66&lt;/li&gt;&lt;/ul&gt;&lt;figcaption&gt;Speed-up in each user’s output speed against no speculation, on a trillion-parameter-class MoE in NVFP4, tensor parallel across four GB200 GPUs, coding dataset [35].&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The chart shows the speed each user sees, the number speculation targets; even so, the gain narrows from 2.03 to 1.66 times between concurrency 1 and 16 &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[35]&lt;/a&gt;. P-EAGLE, a newer drafter, shows the same pattern against EAGLE-3: on a 20B open-weight MoE on one B200, its gain fell from 1.55 times at concurrency 1 to 1.05 times at 64 on one chat benchmark, and from 1.69 to 1.25 times on a coding benchmark &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[36]&lt;/a&gt;. The EAGLE 3.1 authors note that speculative decoding “often degrades under different chat templates, long-context inputs, or out-of-distribution system prompts”, the fragility 3.1 was built to reduce &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[35]&lt;/a&gt;, so track &lt;code&gt;vllm:spec_decode_num_accepted_tokens&lt;/code&gt; against &lt;code&gt;vllm:spec_decode_num_draft_tokens&lt;/code&gt; on your own prompts &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[33]&lt;/a&gt;. vLLM can also shrink the draft as load rises: dynamic speculative decoding sets the number of drafted tokens per batch-size range, and an opt-in adaptive verification mode (DSpark drafters only) sizes verification per request from the drafter’s confidence &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[33]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Multi-token prediction (MTP) trains the model to draft its own next tokens &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[37]&lt;/a&gt;. Nemotron 3.5 Lightning has it built in &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[30]&lt;/a&gt;, so there is no separate draft model to train or host; Gemma 4 ships ready-made MTP drafters that attach to the base model and share its embeddings, so there is nothing to train &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[38]&lt;/a&gt;. NVIDIA says “MTP is best suited for medium to high concurrency, with the optimal draft length decreasing as concurrency increases” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[30]&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&quot;protect-the-latency-target&quot;&gt;Protect the latency target&lt;/h2&gt;
&lt;p&gt;Past a certain queue depth, one more request makes every request late. Cap concurrency where goodput peaks in your benchmark, reject or defer early with a retry signal (an HTTP 429 with a Retry-After header), and give interactive traffic priority: vLLM has &lt;code&gt;--scheduling-policy priority&lt;/code&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[39]&lt;/a&gt; and SGLang has priority scheduling &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[19]&lt;/a&gt;. Nightly jobs belong in their own lane, where “achieving a large batch size is the most important thing” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[20]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The cheapest request is one the big model never sees. We route classification and extraction to the smallest model that works, which takes that traffic off the serving fleet entirely; our post on &lt;a href=&quot;https://dravin.ai/blog/jev-and-typed-decision-models/&quot;&gt;typed decision models&lt;/a&gt; covers that lever.&lt;/p&gt;
&lt;h2 id=&quot;scale-out-across-gpus&quot;&gt;Scale out across GPUs&lt;/h2&gt;
&lt;p&gt;Tensor parallelism splits every layer across GPUs and needs fast links between them; pipeline parallelism gives each GPU a run of layers; data parallelism runs whole copies of the model (replicas). vLLM advises tensor parallelism within a node, and pipeline parallelism across nodes or on GPUs without NVLink, NVIDIA’s fast GPU-to-GPU link &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[40]&lt;/a&gt;. SGLang says to “always favor data parallelism” for throughput when memory allows &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[20]&lt;/a&gt;. In practice: the smallest tensor-parallel group that holds the weights and a healthy KV cache, then more replicas.&lt;/p&gt;
&lt;p&gt;MoE models such as DeepSeek V4, Kimi K3, GLM-5.3 and Mistral Large 3 activate a small share of their weights per token. Expert parallelism places experts on separate GPUs and works best with data parallelism; since tokens spread unevenly, vLLM’s &lt;code&gt;--enable-eplb&lt;/code&gt; rebalances experts from live load statistics &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[41]&lt;/a&gt;. The SGLang team combined expert parallelism with prefill-decode disaggregation for a 671B MoE on H100s; in a decode-only test on 9 nodes of 8 GPUs, with prefill capacity assumed unlimited, it reached about 22,300 output tokens per second per node for 2,000-token inputs, about 5 times a 16-GPU tensor-parallel baseline &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[42]&lt;/a&gt;. That needs both techniques, several nodes and fast interconnect; on one node, tensor parallelism and replicas are simpler.&lt;/p&gt;
&lt;h2 id=&quot;split-prefill-and-decode-last&quot;&gt;Split prefill and decode last&lt;/h2&gt;
&lt;p&gt;The final lever runs prefill and decode on separate GPU pools. DistServe found that “adding a single prefill job to a batch of decoding requests significantly slows down both processes”, and reports 7.4 times more requests or 12.6 times tighter SLOs than colocated systems, with over 90% of requests within their latency limits &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[7]&lt;/a&gt;. Splitwise reports 1.4 times the throughput at 20% lower cost, or 2.35 times the throughput at the same cost and power &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[2]&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Yet vLLM’s documentation says plainly: “Disaggregated prefill DOES NOT improve throughput” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[43]&lt;/a&gt;. Both claims hold. The papers measure goodput under latency targets and cost-normalised capacity; vLLM’s statement is best read as raw tokens per second. Disaggregation leaves the GPUs as fast as they were. What it changes is interference. Chunked prefill stops the worst stalls, but every mixed step still carries a slice of someone’s prompt, which adds to ITL, and colocated phases “share their resource and parallelism settings” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[7]&lt;/a&gt;. Separate pools remove both: decode steps carry no prompt work, and each pool gets its own parallelism and batch size, so more requests meet their targets and TTFT and ITL can be tuned separately &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[43]&lt;/a&gt;. The pools are sized separately too: longer prompts call for more prefill GPUs, longer answers for more decode GPUs.&lt;/p&gt;
&lt;figure&gt;&lt;p&gt;&lt;em&gt;Prefill and decode on separate GPU pools. Prompts of varying length arrive at a router, which scores replicas by prefix overlap and load, and go to a prefill pool where each GPU runs a wide, compute-heavy prefill. Each finished prompt’s KV cache crosses an RDMA link as a stream of packets to a decode pool, where many requests stream tokens in one large, steady batch. A TTFT gauge and an ITL gauge stay inside their target bands. Longer prompts call for more prefill GPUs; longer answers call for more decode GPUs.&lt;/em&gt; &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/&quot;&gt;See the animated diagram in the post.&lt;/a&gt;&lt;/p&gt;&lt;figcaption&gt;The router scores replicas by prefix overlap and load. Prefill hands each prompt’s KV cache to the decode pool over RDMA, so a long prompt does not stall the running decode streams. Shaded bands mark the latency targets.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Conditions matter. On GB200 with a 671B reasoning MoE, TensorRT-LLM’s rate-matched curves, which assume ideal pool ratios and free KV transfer, show 1.4 to 1.8 times the output throughput per GPU at the same per-user speed, at 4,400 input and 1,200 output tokens, mostly at lower concurrency; on 8,192-input, 256-output traffic its real end-to-end runs fell 0 to 25% short of the rate-matched curve &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[44]&lt;/a&gt;. SGLang moves the KV cache between pools with RDMA (remote direct memory access, which copies memory between machines without going through the CPU) transfer engines, on InfiniBand or RoCE devices or AWS EFA &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[45]&lt;/a&gt;; vLLM labels the feature experimental &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[43]&lt;/a&gt; and TensorRT-LLM beta &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[10]&lt;/a&gt;. For a single model on a single GPU, Dynamo’s README says the engine alone “is probably sufficient” &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/#sources&quot;&gt;[23]&lt;/a&gt;. Reach for disaggregation with long prompts, several nodes and fast interconnect. We treat it as the last step, after chunk size, routing and replicas have been tuned against goodput.&lt;/p&gt;
&lt;h2 id=&quot;checklist&quot;&gt;Checklist&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Write the SLO down: TTFT and ITL at p99, and the share of requests that must meet them.&lt;/li&gt;
&lt;li&gt;Benchmark on your own prompt lengths, output lengths, shared prefixes and arrival pattern, and report goodput.&lt;/li&gt;
&lt;li&gt;Dashboard queue depth, KV cache usage, preemptions and prefix-cache hit rate.&lt;/li&gt;
&lt;li&gt;Run a current engine with CUDA graphs on; size memory and batch settings until preemptions stop.&lt;/li&gt;
&lt;li&gt;Put stable prompt content first, confirm prefix caching is on, and route by prefix across replicas.&lt;/li&gt;
&lt;li&gt;Set the chunk size against your TTFT and ITL targets.&lt;/li&gt;
&lt;li&gt;Try FP8 on Hopper, NVFP4 on Blackwell, INT4 weight-only at low concurrency, and a quantised KV cache; check accuracy on your own evals.&lt;/li&gt;
&lt;li&gt;Add speculative decoding where concurrency is low to medium, try a model’s built-in MTP heads at higher concurrency with a shorter draft, and watch the acceptance rate.&lt;/li&gt;
&lt;li&gt;Cap concurrency where goodput peaks, prioritise interactive traffic and give batch jobs their own lane.&lt;/li&gt;
&lt;li&gt;Use the smallest tensor-parallel group that fits, then replicas; expert parallelism with load balancing for MoE.&lt;/li&gt;
&lt;li&gt;Split prefill and decode only with long prompts, several nodes and RDMA.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id=&quot;sources&quot;&gt;Sources&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2403.02310&quot;&gt;Agrawal et al., “Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve”, OSDI 2024&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2311.18677&quot;&gt;Patel et al., “Splitwise: Efficient generative LLM inference using phase splitting”, ISCA 2024&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2211.17192&quot;&gt;Leviathan, Kalman and Matias, “Fast Inference from Transformers via Speculative Decoding”, ICML 2023&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;PyPI release pages, checked 23 Sep 2026: &lt;a href=&quot;https://pypi.org/project/vllm/&quot;&gt;vLLM&lt;/a&gt;, &lt;a href=&quot;https://pypi.org/project/sglang/&quot;&gt;SGLang&lt;/a&gt;, &lt;a href=&quot;https://pypi.org/project/tensorrt-llm/&quot;&gt;TensorRT-LLM&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.nvidia.com/nim/benchmarking/llm/latest/metrics.html&quot;&gt;NVIDIA NIM docs, “LLM benchmarking metrics”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/benchmarking/cli.html&quot;&gt;vLLM docs, “Benchmark CLI”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2401.09670&quot;&gt;Zhong et al., “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving”, OSDI 2024&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/usage/metrics.html&quot;&gt;vLLM docs, “Production Metrics”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.usenix.org/conference/osdi22/presentation/yu&quot;&gt;Yu et al., “Orca: A Distributed Serving System for Transformer-Based Generative Models”, OSDI 2022&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://nvidia.github.io/TensorRT-LLM/overview.html&quot;&gt;TensorRT-LLM docs, “Overview” (updated 21 Sep 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2309.06180&quot;&gt;Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention”, SOSP 2023&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2305.13245&quot;&gt;Ainslie et al., “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”, 2023&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;vLLM v0.30.0 (22 Sep 2026): &lt;a href=&quot;https://github.com/vllm-project/vllm/releases/tag/v0.30.0&quot;&gt;release notes&lt;/a&gt;, tagged source &lt;a href=&quot;https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/config/cache.py&quot;&gt;vllm/config/cache.py&lt;/a&gt; and &lt;a href=&quot;https://github.com/vllm-project/vllm/blob/v0.30.0/vllm/engine/arg_utils.py&quot;&gt;vllm/engine/arg_utils.py&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/configuration/optimization.html&quot;&gt;vLLM docs, “Optimization and Tuning”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://pytorch.org/blog/accelerating-pytorch-with-cuda-graphs/&quot;&gt;PyTorch blog, “Accelerating PyTorch with CUDA Graphs”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/design/cuda_graphs.html&quot;&gt;vLLM docs, “CUDA Graphs”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/features/automatic_prefix_caching.html&quot;&gt;vLLM docs, “Automatic Prefix Caching”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2312.07104&quot;&gt;Zheng et al., “SGLang: Efficient Execution of Structured Language Model Programs”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;SGLang v0.5.20 (18 Sep 2026): &lt;a href=&quot;https://github.com/sgl-project/sglang/releases/tag/v0.5.20&quot;&gt;release notes&lt;/a&gt; and &lt;a href=&quot;https://docs.sglang.io/advanced_features/server_arguments.html&quot;&gt;server arguments&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.sglang.io/advanced_features/hyperparameter_tuning.html&quot;&gt;SGLang docs, “Hyperparameter Tuning”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://llm-d.ai/blog/kvcache-wins-you-can-see&quot;&gt;Prefix-aware routing benchmark, “KV-Cache Wins You Can See” (24 Sep 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://lmsys.org/blog/2024-12-04-sglang-v0-4/&quot;&gt;LMSYS blog, “SGLang v0.4” (4 Dec 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://github.com/ai-dynamo/dynamo&quot;&gt;NVIDIA Dynamo README, v1.5.0 (21 Sep 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/features/quantization/llm_compressor/fp8/&quot;&gt;vLLM docs, “FP8 W8A8”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2411.02355&quot;&gt;Kurtic et al., “Give Me BF16 or Give Me Death? Accuracy-Performance Trade-Offs in LLM Quantization”, ACL 2025&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2306.00978&quot;&gt;Lin et al., “AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration”, MLSys 2024&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2210.17323&quot;&gt;Frantar et al., “GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers”, ICLR 2023&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2408.11743&quot;&gt;Frantar et al., mixed-precision 4-bit inference kernels for LLMs (Aug 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/&quot;&gt;NVIDIA blog, “Introducing NVFP4 for Efficient and Accurate Low-Precision Inference” (24 Jun 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://developer.nvidia.com/blog/nvidia-nemotron-3-5-lightning-delivers-fast-accurate-specialized-task-execution-for-long-running-agents/&quot;&gt;NVIDIA blog, Nemotron 3.5 Lightning (11 Aug 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/features/quantization/quantized_kvcache.html&quot;&gt;vLLM docs, “Quantized KV Cache”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://www.lmsys.org/blog/2026-09-16-nvfp4-kv-cache&quot;&gt;LMSYS blog, “Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache” (16 Sep 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;vLLM docs: &lt;a href=&quot;https://docs.vllm.ai/en/latest/features/speculative_decoding/&quot;&gt;Speculative Decoding&lt;/a&gt;, &lt;a href=&quot;https://docs.vllm.ai/en/latest/features/speculative_decoding/dynamic_speculative_decoding/&quot;&gt;Dynamic Speculative Decoding&lt;/a&gt; and &lt;a href=&quot;https://docs.vllm.ai/en/latest/features/speculative_decoding/adaptive_verification/&quot;&gt;Adaptive Verification&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2503.01840&quot;&gt;Li et al., “EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test”, NeurIPS 2025&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://vllm.ai/blog/2026-05-26-eagle-3-1&quot;&gt;vLLM blog, EAGLE 3.1 (26 May 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://vllm.ai/blog/2026-03-13-p-eagle&quot;&gt;vLLM blog, “P-EAGLE: Faster LLM inference with Parallel Speculative Decoding in vLLM” (13 Mar 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://arxiv.org/abs/2404.19737&quot;&gt;Gloeckle et al., “Better &amp;amp; Faster Large Language Models via Multi-token Prediction” (Apr 2024)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://ai.google.dev/gemma/docs/mtp/overview&quot;&gt;Google AI for Developers, Gemma MTP overview&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/configuration/engine_args.html&quot;&gt;vLLM docs, “Engine Arguments”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/serving/parallelism_scaling.html&quot;&gt;vLLM docs, “Parallelism and Scaling”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/serving/expert_parallel_deployment.html&quot;&gt;vLLM docs, “Expert Parallel Deployment”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://lmsys.org/blog/2025-05-05-large-scale-ep/&quot;&gt;SGLang team: large-scale expert parallelism on 96 H100s (May 2025)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.vllm.ai/en/latest/features/disagg_prefill.html&quot;&gt;vLLM docs, “Disaggregated Prefilling (experimental)”&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://nvidia.github.io/TensorRT-LLM/blogs/tech_blog/blog5_Disaggregated_Serving_in_TensorRT-LLM.html&quot;&gt;TensorRT-LLM tech blog, “Disaggregated Serving in TensorRT LLM” (updated 16 Mar 2026)&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&quot;https://docs.sglang.io/advanced_features/pd_disaggregation.html&quot;&gt;SGLang docs, “PD Disaggregation”&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This post first appeared on &lt;a href=&quot;https://dravin.ai/blog/raising-llm-serving-throughput/&quot;&gt;Dravin AI&lt;/a&gt;.&lt;/p&gt;</content:encoded><category>LLM serving</category><category>Inference optimisation</category><category>Self-hosted LLMs</category><category>GPUs</category><author>info@dravin.ai (Dravin AI)</author></item></channel></rss>