← All blog posts

Inlining a language model: four languages, three walls, and a memory bus

September 30, 2026 · AI

Four stages. Python kept the weights as a compressed blob in one file, not yet inline. Go compiled every weight as a literal: TinyStories at 3.7 million weights made a 22 MB binary at 600 tokens a second, but the compiler holds about 700 bytes per literal, so the 7B model was cut into 1,062 processes, 33 GB resident, at 1.5 tokens a second. C linked Qwen3-4B from 1,142 files into a 16 GB Mach-O, which codesign refuses, and an unsigned arm64 binary is killed before main. Java runs it: 17,067 classes, 20 GB of source, 33.5 seconds to load, 12 tokens a second. Below, the third wall, the memory bus: decode speed follows the bytes each token reads, 12 tokens a second for 16 GB, about 41 for bfloat16, about 100 for 4-bit, and about 30 for Apple's 3B on an iPhone, the last three second-hand. Twelve tokens of 16 GB is about 190 GB/s, against 215 for one hot matrix-vector product and 546 published.
The weekend in one picture. Two walls stopped the literals from becoming one program; the third, the memory bus, decides how fast the one that runs can go. Only the blue bars were measured here.

On Friday, 25 September, a small idea came up; download an open-weight model and write an inference engine with the weights compiled into the program. Not a checkpoint file next to the binary, not a blob glued at the end of it. Every weight an ordinary number in the source code, and a running program that never opens a file.

Why would anyone want that? The honest answer is curiosity. The slightly less honest one is the old hope that a compiler, told everything up front, will find something clever to do with it. When I was younger (in the late 90s) I was very optimistic about what compilers could offer, and it seems that some of that optimism never left. If the weights are constants, maybe the machine can run them faster than a generic engine that loads them at run time.

It took a weekend, four languages and a fair amount of disk to find out where that idea breaks. It breaks in a different place in each language, and none of those places is where I expected.

Small models are generous

The first version was Python, on one of Karpathy's tiny story models, the weights compressed at the bottom of a single file. It ran and it wrote a children's story that frayed after a few sentences. Then the engine moved to Go, on TinyStories-1M (which, counting every stored weight, is actually 3.7 million parameters). Every weight became a decimal float32 literal, the vocabulary became Go strings, and go build produced one binary of 22 MB.

It worked beautifully. The greedy output ("Once upon a time there was" continued into Lily and a shiny rock) matched an independent run of the same checkpoint, and it did about 600 tokens per second on my MacBook Pro (M4 Max, 64 GB). It compiled in 12 seconds.

The first lesson was hiding in there already, and I missed it. Most of the time went to one multiply, the last one, which projects onto a vocabulary of 50,257 words. The hidden state of the model is only 64 floats; the expensive part is the wide matrix at the end. The cost of a model is where its weights are, not where its logic is.

The Wall

So, the obvious next step; a real model. I asked for gpt-oss-20b first, which was ambitious (20.9 billion weights, about 300 GB of Go source if written the same way). Then a 7B, and I pushed the literal format until the compiler showed me where it stops.

The answer is about 50 million weights. The Go compiler holds a syntax tree node for every number, roughly 700 bytes of state for 4 bytes of payload. Fifteen million literals already needed 9.7 GB of RAM to compile; on a 64 GB machine the ceiling is somewhere around 50 million. A 7B model is 140 times past that. There was enough disk for the text, but not enough memory for the compiler to read it.

Ok, then cut the model into pieces.

For the small model, each layer became its own executable. A driver owns the vocabulary and the sampler, and for every token it pipes 64 floats to layer 0, takes the answer, pipes it to layer 1, and so on. The output stayed identical and the speed barely moved (the pipe is cheap, the vocabulary is not). For Qwen2.5-7B the same idea needed a finer knife; one layer alone is past the compiler, so every matrix was cut into strips of at most 8 million numbers, and every strip became a program. That is 1,062 processes, 97 GB of decimal source and 33 GB of binaries.

A sequence diagram. The driver embeds the token, then sends the hidden state to process 0 and gets the updated state back, then to process 1, and so on to process n; at the end the driver scores and samples. The processes never talk to each other.
The Go arrangement. The processes never talk to each other; the driver calls them one after another. On the small model a process is a whole layer, on the 7B it is one strip of a matrix, and the exchange is the same: a vector goes out, a slice of the result comes back.

It answered "Hello!" when asked to say hello. It did so at about 1.5 tokens per second, after 32 seconds of starting processes, with 33 GB of memory held before the first token. The alternative was to start each strip only when needed and let it exit afterwards; memory dropped to 0.43 GB and a single token took 63 seconds. Either you pay the compiler's limit in RAM, or you pay it in cold starts. You do not get to skip paying.

Anyone that typed a program from a computer magazine in the 80s remembers the pages of DATA statements at the end of the listing; the program was short and the data was everything. This was the same thing, forty years later, a thousand times over.

Signed, sealed, not delivered

On Saturday the target changed. I wanted to see if something the size of Apple's on-device model could be inlined, and since Apple does not hand out those weights, the nearest open model was Qwen3-4B-Instruct. Four billion parameters, 16 GB in float32. And this time the wish was the original one; a single executable.

Go could not do it without the same 500 strips. C looked more promising, because C has a linker; you still cannot compile four billion literals in one file, but you can compile 1,142 files and link them into one program. And that is where the walls started coming from the platform instead of the language.

The first wall was addressing. On Apple Silicon, code can only reach a global variable directly if it lives within 4 GB of the instructions, and a 16 GB image is well past that. The fix was a small table of 64-bit pointers sitting next to the code. The second wall was the file format; Apple's linker gives up once a segment passes a 4 GB offset, so the weights went into six segments and the link was done with LLVM's lld, which happily wrote a 16 GB Mach-O file.

It never ran. codesign refused it; the signature format does not cover a file that large. And on an Apple Silicon Mac an unsigned binary is killed before it reaches main. The tokenizer checked out, the engine was the real model, and the kernel would not let it start. Rust would produce the same kind of file and hit the same refusal. A float32 copy of this model cannot be both fully inlined and signed, at least on macOS.

That one hurt a bit. Every problem up to that point was engineering, and this one was policy (a very reasonable policy, to be fair, but a policy nonetheless).

Write once, run anywhere

This is where the word "inline" changed meaning, and I think it is worth saying so plainly rather than pretending the goal was met.

A single executable of literals was the original wish, and it is not possible here. What survived is weaker but still interesting; code specific to each layer. Layer0 is its own class with its own step method and its own named tensors, and the driver calls Layer0.step through Layer35.step by name. There is no loop over a table of layers and there is no weight file. The architecture of the Qwen3 block is fixed, so the code says it explicitly instead of discovering it from a config. The matrices stay arrays, and the values of those arrays live in the source.

The language that made this possible was Java, for a reason I would not have guessed on Friday; the signed executable is the JVM. The weights arrive by initialising classes, and nobody has to sign a 16 GB file.

Java has its own walls, of course, and they are much smaller than 4 GB. A method's bytecode is capped at 64 KB, a string constant at 65,535 bytes, and a class at 65,535 constant pool entries. So the big matrices are stored as base64 strings of float bytes, 32 strings per class, and a matrix is assembled by loader classes calling pieces by name. That is 17,067 generated files and about 20 GB of source. The tokenizer is Java too, so at run time there is literally nothing to open.

Loading has one nice trick. The base64 text is a second copy of the weights (about 21 GB of text next to 16 GB of floats), and a class that references it keeps it alive forever. So each weight class is loaded through a throwaway class loader, the floats are copied out, and the loader is dropped so the text can be collected. With twelve threads decoding pieces in parallel, bringing the whole model up takes 33.5 seconds.

Two rows of boxes. Weights come up once: check memory, Tok.init, LayerN.bind, Data.load, L0q.fill, pieces decode. Then the question: LayerN.prefill, Ops.prefill, LayerN.step, Ops.gemv, FinalNorm.apply, and Embed.logits then sample.
One question in the Java version. The top row runs once: LayerN means the same method on Layer0, then Layer1, through Layer35, and L0q is the query matrix of layer 0, filled by its pieces. The bottom row is the prompt and then every token.

And then it works. "What is 2 + 2?" comes back as 2 + 2 = 4., which became the sanity check for every change after that.

It's the bus

The first run was about 6 tokens per second. The matrix multiply then moved to Apple's Accelerate (cblas_sgemv and cblas_sgemm, called through Java's Panama interface without copying the arrays), decode was split across the 12 performance cores, and the prompt was processed as one big multiply per matrix. The final numbers, on a request to list twenty fruits; 63 tokens per second for the prompt and 12 tokens per second for the answer.

Why 12? Because every generated token has to read all 16 GB of weights, and 12 times 16 GB is about 190 GB/s. A hot matrix-vector product measured on the same machine moves about 210 to 220 GB/s. The chip is advertised at around 546 GB/s, but a token is many separate multiplies, not one long stream, and the loop never gets there. Running two matrices side by side on the cores made things slower. The bus was already full.

So here is the answer to the question I started with. Inlining the weights does not make them cheaper to read. The compiler cannot optimise a multiply by a number it only knows as data in a 16 GB array, and the CPU does not care whether those bytes came from a class file or a checkpoint. What decides the speed of a language model on a CPU is how many bytes you move per token, and nothing else comes close.

The neighbours make the point better than I can. I did not run them myself; these are other people's numbers, on an M4 Max for the same network, and Apple's own for its phone model.

Table 1. Decode speed against the bytes each generated token reads. The first row is measured here (61 tokens, greedy); the others are second-hand, and approximate.
EngineDecode (tok/s)Read per token
This program, float321216 GB
Same model, bfloat16 (MLX or llama.cpp)41about 8 GB
Same model, 4-bit100a few GB
Apple's on-device 3B, iPhone 15 Pro (Apple's figure)30under 2 GB

The ranking is almost exactly the ranking of the bytes. The engines that beat this program do not have better loops; they read less. A quick try with int8 weights was actually slower and a few percent off, so a smaller format has to earn its place, but four-bit is clearly the next cut to try.

Epilogue

Go will compile the literals, but it will not give you one program you can start. C will link them into one file, and macOS will refuse to sign it. Java starts, because the thing that is signed is the JVM, and it runs at the speed of the memory bus, like everything else.

There is also a smaller, more practical wall at the very end. GitHub will not take the tree; the working copy is around 28 GB, and a push is capped at 2 GB. What can be shared is the generator, a few handwritten classes and this article. The weights have to be regenerated on a machine that has the checkpoint, which means the thing I own, in the end, is the program that writes the program.

I wrote this down so that the next time I say "inline", I remember which wall I am standing in front of. Quantisation and a faster class load are the two doors that are still open. We will see where they lead.