Building PRISM: I spent months writing an LLM inference engine in C99
I built PRISM because I wanted to understand how LLM inference actually works. Not the high-level “call an API” kind of understanding. The “I can explain every instruction the CPU executes while running a forward pass” kind.
So I wrote a C99 inference engine. From scratch. No dependencies. It loads real GGUF models, runs on CPU, and serves text over HTTP. It’s called PRISM — Plugin-based Runtime for Inference and Model Serving.
It’s also about 5-8x slower than llama.cpp. But I’m getting ahead of myself.
The plugin architecture was the right call
I started with the most important decision: every component would be swappable at runtime. Kernels, models, tokenizers, samplers, IO layers — all loaded by name from a config file:
--boot.plugins "kernel.avx2,model.llama2c,tok.gguf,sampler.basic,loop.generate,io.http"
This forced me to write clean interfaces between components. Each one communicates through C function pointers — basically an explicit vtable struct:
/* Clean C99 plugin contract: dynamic dispatch without runtime overhead */
typedef struct prism_kernel_plugin {
const char *name;
int (*init)(prism_ctx_t *ctx);
void (*matmul)(float *out, const float *a, const float *b, int m, int k, int n);
void (*attention)(float *out, const float *q, const float *k, const float *v, int seq_len);
void (*deinit)(prism_ctx_t *ctx);
} prism_kernel_plugin_t;
Was it worth it? Absolutely. I could test each component in isolation, swap kernel implementations without touching model code, and actually debug things without tearing my hair out.
A comparison against baseline llama.cpp illustrates the trade-off:
| Engine | Architecture | Layer Matmul | CPU Overhead | Extensibility |
|---|---|---|---|---|
llama.cpp | Monolithic optimized engine | Handcrafted ASM / SIMD | Minimal / direct | Hard fork / deep refactor |
PRISM | Dynamic C99 plugin vtables | Function pointer dispatch | ~15–20% vtable overhead | Drop-in runtime plugins (.so) |
Getting the demo model running
First milestone: generate text. Any text. I built a synthetic model with random weights — 2 layers, 64 dimensions, 8 heads, 256 vocab.
It ran at 1900 tokens/second. Looked impressive until you realize the entire model fits in L2 cache and the output is complete garbage. The weights are random. The model has no knowledge. It just produces statistically plausible nonsense.
But the pipeline worked. Tokenize, embed, attention, FFN, sample, detokenize. That felt good.
The GGUF rabbit hole
Loading a real model was where things got interesting. GGUF is llama.cpp’s format — a binary container with metadata, tensor definitions, and tensor data.
I wrote the parser. It worked. I loaded the tensors. I ran the forward pass.
NaN. Every tensor produced NaN.
Turns out GGUF tensors are not contiguous in the file. My code assumed they were. The actual offset for layer 1’s attention weights was 3,764,736 bytes but I expected 117,504. I was reading garbage data from nowhere.
Fix: use the actual offset from the tensor definition. Simple. Obvious in hindsight. Cost me two days.
Then there was the dtype issue. The GGUF metadata said tensors were F32. They were actually Q8_0 (quantized). The type field lied. I trusted it. Another day gone.
Then the shared classifier problem. Some models tie the output projection to the input embedding — no separate output weight tensor. My code assumed it always existed. When it didn’t, I read uninitialized memory.
Three bugs. Three days. Welcome to parsing binary formats.
The performance reality
With SmolLM2-135M-Instruct-Q8_0 working, I benchmarked:
PRISM: 3.3 tok/s. llama.cpp: 30-50 tok/s. On the same machine. Same model.
A 10x gap. I spent weeks trying to close it.
I implemented cache-blocked matmul (1.5x speedup). AVX2 kernels (1.3x). Memory pool allocator (1.2x). Async pipeline (1.1x). Huge pages (1.05x). NUMA-aware allocation (1.1x).
Total: 4.5x improvement. Still 5-8x behind llama.cpp.
The problem is memory bandwidth. SmolLM2-135M is 135MB in Q8_0. CPU memory bandwidth is ~50 GB/s. Minimum time per token is 135MB / 50GB/s × 17 layers = 46ms. That’s a theoretical max of ~22 tok/s, and I’m not at theoretical max because memory access patterns are terrible and I have no SIMD in the hot path.
llama.cpp wins because they have hand-written AVX2/AVX-512 assembly (not compiler-generated), fused kernels that combine operations, K-quants that use 4.5 bytes per parameter instead of 8, and GPU support. Years of iteration. I can’t catch up in weeks.
The MoE trap
I wanted to support Mixture of Expert models. Built the router, got it working, saw correct expert probabilities. Then the classifier produced zeros.
The classifier is a matmul: logits = x @ classifier_weight. The weight is [32000, 2048] — 64M values. Manual matmul produces correct logits. My kernel produces zeros. I never found the root cause. It works for smaller matrices. Something about large dimensions breaks it.
MoE loading works. Router works. Classifier is broken. That’s where it sits.
What I actually learned
The technical stuff — cache locality matters more than algorithmic complexity, memory bandwidth is the ceiling on CPU, file formats lie — that’s all true but it’s also stuff you can read in a textbook.
The real lessons are different.
Debugging is the skill. Finding the NaN bug, the dtype bug, the shared classifier — that’s the actual work. Writing code is maybe 30% of building something. The rest is figuring out why it doesn’t work.
Scope creep is real. MoE support seemed simple. It wasn’t. Every “small feature” is a rabbit hole.
Knowing when to stop matters. I could spend another year chasing llama.cpp’s performance. I won’t. I learned what I set out to learn.
Building something from scratch teaches you more than reading about it. I understand transformer inference at a level I never would have gotten from calling an API.
Where PRISM stands
It works. Dense Llama-family models, GGUF loading, GQA attention, HTTP API with streaming, auth, rate limiting, TLS. The plugin architecture genuinely works — I can swap components at runtime.
It’s not fast. It’s not compatible with all models. It’s not a product. It’s a learning project that became a real engine.
I’m stopping here. Not because it’s done — it never will be. But because I’ve learned what I came to learn.
The code is published: github.com/wraithen0/prism. If you want to understand how inference works at the register level, read it. If you need speed, use llama.cpp. I’m not ashamed to say that.
I built something real. I understand every line of it. And honestly, that’s enough.