INFERENCE LAB / INFERENCE

vLLM

Deploy, serve and optimize. Practical notes on the engines behind LLM inference.