← All blog posts

BitNet is Here: Run you AI Models on CPUs

October 20, 2024 · AI

Image Generated by DALL-E

All the discussion about scaling and using LLMs is around GPUs. And everyone is focusing on how to built data centers around that purpose. Larry Elison recently said that in order to built a frontier model, you will need around 100bln dollars of investment. I think that rules out lots of people trying to do that (source: “the race of AI”).

So, training the models and calculating the weights is out of the question. Actually, you need GPUs. But the inference (aka prompt and get your answer) is another problem.

Llama.cpp and Focusing on CPUs

At first, I saw the llama.cpp. It was a pure C++ implementation that ran various models and also provided bindings for several languages, to use it effectively. This technology seemed ok, but nothing that it could work in great scale. And still we are not there.

What is the problem we are trying to solve? To break the dominance of the GPU, at least for the inference part, “for the training, nothing”.

So, llama.cpp was a good start, in order to begin doing some optimisations on the runtime part. But it had some limitations, on 13B or more models, it was impractical to use it. Only if you have the patience to wait and only for hobby tasks. So, it seemed that for the large models, we are stuck on services, like ChatGPT, Gemini or Groq.

BitNet

So, at 2023 I bumped into this paper.

BitNet: Scaling 1-bit Transformers for Large Language Models

At a glance, it says:

Practically it says, since the weights are 1-bit, the method of inference is simplified, which leads to less energy used when executing it.

So, almost one year has passed and bitnet.cpp is released on Github. An inference engine, coded in C++ that makes the AI models faster to execute on normal CPUs. It promises significant performance gains against llama.cpp.

The speedup is impressive and as you can see we are beginning to hav results for the 70B and 100B models, which were not feasible to efficiently ran them in the past.

Of course the model needs to be adapted to the 1-bit version of it, but I am sure that will happen on Hugging Face as usual.

Epilogue

In my view, only a handful of companies can sustain the money-intensive process of creating a frontier model or any model at all.

But the inference method is still open and a CPU-centric method could change a lot of things in the LLM world.

See also my previous article on GPU-related framework vLLM, which focus on model execution and deployment.