More users per GPU: a playbook for self-hosted LLM serving

What to measure first, then the levers that raise LLM serving throughput on your own GPUs, in the order we pull them and what each one trades away.

More posts