INFERENCE LAB / UPDATES
Publication updates
Substantive revisions and new publications from the notebook.
Follow substantive article revisions through RSS. Dates below reflect content changes, not automated freshness updates.
September 14, 2026
Added worked examples, diagnostic decision tables, a cache diagram and an inspectable memory estimator. Clarified the distinction between documentation-based guidance and measured results.
- Understanding FP8 KV Cache in vLLM — 2026-09-14
- Prefix Cache vs KV Cache in LLM Inference — 2026-09-14
- Qwen3.8-27B on 4×RTX 4090: Deployment Guide — 2026-09-14
- SGLang MTP / NEXTN Explained — 2026-09-14
- vLLM vs SGLang: Choosing an Inference Engine — 2026-09-14
- NVIDIA Xid 79: GPU Has Fallen Off the Bus — 2026-09-14
September 9, 2026
Updated canonical URLs, RSS and the sitemap to the publication's custom domain.
September 7, 2026
Published the first six reference guides.