Building vek: A Zero-Dependency SIMD Vector Similarity Library in C
So I built vek. It’s a zero-dependency, hand-tuned SIMD vector similarity library in C that works on basically anything with a compiler.
I started it because I was annoyed. I was trying to build a small local RAG pipeline, nothing serious, and the library I found wanted to pull in half of PyTorch just to compute a dot product. A dot product. Twenty lines of C. I wrote those twenty lines and then I couldn’t stop.
Three weeks later I had a library.
What it does
vek computes dot product, L2 distance, and cosine similarity. It supports f32, f16, bf16, int8, uint8, and 1-bit binary vectors. Runtime CPU dispatch means one binary works on SSE2, AVX2, AVX-512, or ARM NEON. No compiler flags required from consumers, no recompilation.
The C ABI is stable. extern "C" everywhere. Call it from Rust, Zig, Go, Python, or whatever language needs it. I use it from Rust myself — the FFI binding is trivial.
The stuff that actually mattered
Neumaier summation. This is the kind of detail nobody thinks about until their results drift. Standard floating point addition loses precision when you add numbers of very different magnitudes because the smaller value gets rounded off. Imagine adding 1.0 to 0.0000001 a million times: you’d expect 1.0001, but you get 1.0 because the tiny additions keep getting swallowed.
Neumaier summation tracks those lost bits in a separate compensation variable and adds them back. For vector dot products, where you multiply large embedding dimensions and accumulate thousands of terms, this matters. I saw 0.1% differences in cosine similarity between naive summation and Neumaier. That doesn’t sound like much until you realize that’s the gap between a relevant top-1 vector search hit and irrelevant noise.
Masked loads on AVX-512. Most SIMD code handles tail elements (the trailing elements that don’t fill an entire vector lane) with scalar fallback loops. That means two code paths and twice the edge cases. AVX-512 masked loads let you process the last partial vector with the exact same SIMD instructions, just masking off the unused elements:
/* AVX-512 masked dot product tail handling */
__mmask16 mask = (1U << remainder) - 1;
__m512 va = _mm512_maskz_loadu_ps(mask, a + offset);
__m512 vb = _mm512_maskz_loadu_ps(mask, b + offset);
acc = _mm512_fmadd_ps(va, vb, acc);
One execution path. Cleaner, faster, and branch-free.
Zero heap allocations in the hot path. Stack only, cache friendly, zero malloc in the inner loop. Sounds obvious, but you’d be surprised how many embedding search libraries allocate per query call. When you’re processing millions of vectors, heap allocations destroy memory locality fast.
Performance comparison
Microbenchmarks on f32 vectors ($D = 1536$, OpenAI embedding dimension) across 100,000 runs on an Intel Core i7:
| Implementation | Kernel Mode | Latency (ns / vec) | Speedup vs Naive | Precision Loss |
|---|---|---|---|---|
| Naive C loop | Scalar (-O3) | ~1,420 ns | 1.0x (baseline) | High (drifts on large sums) |
vek | SSE2 | ~410 ns | 3.5x | Compensated (Neumaier) |
vek | AVX2 + FMA | ~135 ns | 10.5x | Compensated (Neumaier) |
vek | AVX-512 (masked) | ~72 ns | 19.7x | Compensated (Neumaier) |
What I messed up
My first attempt was a single generic kernel using template macros and void * casting. Classic C metaprogramming trap. It was unreadable, broke diagnostics, and the compiler couldn’t auto-vectorize safely.
The rewrite — separate, explicit kernels per precision and ISA — was tedious, but the code is actually debuggable now. I can read it at 2 AM without wanting to tear my hair out.
I also spent way too long on the build system. Cross-compilation for Linux, macOS, and Windows with proper CPU feature detection took almost as long as the kernels themselves. The Makefile is 200 lines of hardened detection logic.
Where it’s at
vek is in active development. The core kernels are solid, and we’re integrating them into our own local RAG pipelines.
Try it
git clone https://github.com/wraithen0/vek
cd vek
make
Dual licensed under MIT or Apache-2.0. If you build something with it, let us know — that’s the whole reason we open source our tools.