← All blog posts

vLLM and LLM-compressor are here

August 19, 2024 · AI

Its very easy (and not so cheap) to use LLMs with your application. But only if you use a SaaS, like ChatGPT or Gemini. What happens if you want to scale your own infrastructure and serve your own models?

Not an easy task

If you want to serve your own models, it is becoming increasingly difficult, in correlation with the size of the models, thus it is easier to scale the llama 8B model than the 405B.

To be blunt, if you do not have a good amount of money to invest, there is no way to succeed, or even compete with the big players.

Accidentaly, I found vLLM, a company which focusing on creating solution to easily scale LLM models on services, without depending on external SaaS solutions.

vLLM

vLLM is a open-source library designed to do fast-inferencing. To do so, utilizes a new attention algorithm, named PageAttention. By implementing it, it achieves blazing speeds in inferencing, as depicted in the following figure:

It is also really, simple to use it, and embed it into your application. First, you have to install the dependency using pip.

pip install vllm

Then, write a Python program like the following, and include your favorite model from HuggingFace.

from vllm import LLM

prompts = ["Hello, my name is", "The capital of France is"]
llm = LLM(model="lmsys/vicuna-7b-v1.3")
outputs = llm.generate(prompts)

Perfect, right? It also has an embedded API service, which you can invoke from command line, like:

python -m vllm.entrypoints.openai.api_server --model lmsys/vicuna-7b-v1.3

Really, nice work. On top of that, they also created an LLM compression tool, which reduces the size of LLMs, making them easier to deploy them into production infrastructure.

LLM-compressor

So, this is a tool that optimizes LLM models for deployment. The tool includes various methods to do so, let us assume that we want to optimize the tiny-llama model (1.1B version) and compress it to deploy that into our production infrastructure.

You have to include a variant of the following script:

from llmcompressor.modifiers.quantization import GPTQModifier
from llmcompressor.modifiers.smoothquant import SmoothQuantModifier
from llmcompressor.transformers import oneshot

# * apply SmoothQuant to make the activations easier to quantize
# * quantize the weights to int8 with GPTQ (static per channel)
# * quantize the activations to int8 (dynamic per token)
recipe = [
SmoothQuantModifier(smoothing_strength=0.8),
GPTQModifier(scheme="W8A8", targets="Linear", ignore=["lm_head"]),
]

oneshot(
model="TinyLlama/TinyLlama-1.1B-Chat-v1.0",
dataset="open_platypus",
recipe=recipe,
output_dir="TinyLlama-1.1B-Chat-v1.0-INT8",
max_seq_length=2048,
num_calibration_samples=512,
)

The TinyLlama-1.1B-Chat-v1.0-INT8 model is created and ready to be included into your vLLM-powered application, like that:

from vllm import LLM

model = LLM("TinyLlama-1.1B-Chat-v1.0-INT8")
output = model.generate("My name is")

Sweet.

Epilogue

So far, I have not used these technologies in production, but I will surely use them in the near future and write detailed feedback!

So, stay tuned.