code-daemon-embed-v1
A 4-layer, 46.8M-parameter code-embedding model built for one job: embedding a whole repository fast enough to re-index it on every commit, as the dense half of a hybrid (vectors + BM25) code search. Short code units β signatures, docstrings, symbol names, one-line descriptions β and short keyword-shaped queries land in one 768-dim space.
What makes it different
- Speed first. 4β6Γ the throughput of any public 768-dim encoder we tested; a 1.1M-entity C++ repository embeds in about 80 seconds on a laptop GPU.
- A reranker's judgement baked into the vectors. A 4B cross-encoder's ranking was distilled into the bi-encoder at training time; in the system it was built for, a runtime reranker on top of it measured net-negative.
- Nothing to get wrong at the output. Mean pooling and L2 normalization are inside the graph:
the model returns unit-norm
[B, 768], ready for a dot product. - INT8 from quantization-aware training, so TensorRT and OpenVINO build an INT8 engine with no calibration pass.
Queries and documents are encoded the same way β no query: / passage: prefix. Hard cap of
128 tokens. English and code only.
Choosing it β the trade
Vector channel alone, every model at the same 128-token budget. Quality is nDCG@10 Γ 100; speed is texts/s through onnxruntime CUDA FP32, batch 64, at seq 64 (the mean serving text is 70 tokens). Laptop RTX 5060, 2026-09-11.
| model | layers | texts/s | cosqa | stack-qa | codeβcode | our queries | Russian hit@1 |
|---|---|---|---|---|---|---|---|
| code-daemon-embed-v1 | 4 | 3 457 | 28.5 | 49.9 | 42.9 | 60.1 | 0.06 |
| multilingual-e5-base | 12 | 854 | 29.8 | 79.4 | 57.7 | 60.7 | 0.24 |
| CodeRankEmbed | 12 | 723 | 35.5 | 73.3 | β | 64.6 | 0.00 |
| gte-modernbert-base (its teacher) | 22 | 606 | 36.6 | 81.4 | 82.5 | 70.5 | 0.15 |
- 4.0β5.7Γ the throughput for 1β10 points of nDCG on our queries. Pick it when indexing speed is the
constraint and a lexical channel runs beside it; pick
gte-modernbert-basewhen vector-only accuracy is. - Inside the hybrid search it was built for (80 real agent queries, file level, runtime reranker off): the right file is first 44 % of the time and in the top 5 69 %.
- Not multilingual. The vocabulary prune kept 942 of XLM-R's 31 671 Cyrillic pieces; for
non-English queries use
multilingual-e5-base. - Not for long text or general prose. On four CoIR tasks it averages 52.8 nDCG@10 β 89.8 on Python docstringβcode, 28.5 on cosqa.
Throughput in a real index
A full index through the Code-Daemon daemon, laptop RTX 5060, TensorRT INT8, pinned 114 W:
| workload | texts/s |
|---|---|
| C/C++ β mysql-server, 1.1M entities | 13 800 |
| TypeScript β vscode, 479k | 10 600 |
| Java β netty, 97k | 8 000 |
| C# β roslyn, 403k | 7 900 |
| small incremental updates (a few hundred to 15k texts) | 1 900 β 5 500 |
| no discrete GPU β Intel CPU + iGPU + NPU together | ~1 050 |
| Apple M4, Neural Engine | ~1 100 |
The spread between languages is text length: 71 % of C# texts fall in the longest bucket against
24 % for C/C++. Above ~820k texts the daemon adds the Intel iGPU beside the RTX card, so the
mysql-server row is a two-device number. The same daemon running multilingual-e5-base compiled
to the same INT8 engines is 2.3Γ slower on every row.
How to use it
Pick a length bucket
Every engine family ships four length buckets. A batch pays for its longest member, so route each text to the smallest bucket that fits it:
| bucket | batch | seq range (opt) | solo texts/s, RTX 5060 | typical content |
|---|---|---|---|---|
| s | 96 | 8 β 48 (48) | 19 800 at 48 | names, signatures, queries |
| m | 128 | 56 β 64 (56) | 13 100 at 64 | typical entity text β take this one if you load only one |
| l | 128 | 72 β 80 (72) | 9 900 at 80 | |
| xl | 256 | 88 β 128 (96) | 5 500 at 128 | long docstrings, doc chunks |
- Sort by length, then batch, and dispatch each batch at its own longest length rounded up to a multiple of 8. The TensorRT engines take a dynamic sequence range for exactly this β worth +10 % end to end.
- Batch size is not the lever: 96 Γ 48 already saturates the GPU, and a bigger batch was slower per text. Running the four engines concurrently does not add throughput either.
- All four TensorRT engines together take ~400 MB of VRAM.
- OpenVINO IRs are static shapes: batch 64 on CPU and iGPU, batch 16 on the NPU. The NPU has the lowest latency per batch (28β95 ms) and suits single interactive queries better than bulk indexing.
Shape the input
Documents β the model was trained on a compact, front-loaded text, and reproducing it is worth more than any inference tuning:
function acquireProjectLock in src/storage/multi_db.zig: zig
path: src storage multi db
[exported]
sig: (allocator: Allocator, project_hash: []const u8) -> !Lock
Acquires the exclusive SQLite lock for one project.
That is {type} {name} in {file}: {lang}, the path as words, flags, the signature, then the first
~200 characters of the doc comment and a one-sentence description; ~768 characters at most. Leave
the function body out β it belongs in the lexical channel.
Queries β raw, no prefix. Short keyword bags and behaviour descriptions are what it was tuned for. Longer than 128 tokens β windows of 128, stride 96; mean-pool the window vectors and re-normalize.
Run the ONNX
import onnxruntime as ort, sentencepiece as spm, numpy as np
sp = spm.SentencePieceProcessor(model_file="sentencepiece.bpe.model") # pad=0 unk=1 bos=2 eos=3
sess = ort.InferenceSession("model_int8qdt.onnx", providers=["CPUExecutionProvider"])
def embed(texts, max_len=128):
ids = [[2, *sp.encode(t)[: max_len - 2], 3] for t in texts] # bos β¦ eos
L = max(len(x) for x in ids)
inp = np.array([x + [0] * (L - len(x)) for x in ids], dtype=np.int64) # pad=0
mask = (inp != 0).astype(np.int64)
return sess.run(None, {"input_ids": inp, "attention_mask": mask})[0] # [B, 768], unit-norm
D = embed(["function acquireLock in src/db.zig: zig\npath: src db"])
Q = embed(["acquire database lock"])
print(Q @ D.T) # inner product = cosine
Tokenize with SentencePiece β the vocabulary is a unigram model, and a BPE merge loop segments it differently from training. Both inputs are int64.
Building your own engines
Feed model_int8qdt.onnx as it is. TensorRT 11 reads the precision from its Q/DQ nodes;
OpenVINO keeps the trained scales as FakeQuantize (run compress_quantize_weights_transformation
before saving, or the IR comes out twice the size). Never run a calibration / PTQ pass over
it: it replaces the trained scales with fitted ones, and the engine builds and loads cleanly
and quietly retrieves worse. Use the bucket profiles from the table above, for example:
trtexec --onnx=model_int8qdt.onnx --saveEngine=embed-m.engine --builderOptimizationLevel=5 \
--minShapes=input_ids:128x56,attention_mask:128x56 \
--optShapes=input_ids:128x56,attention_mask:128x56 \
--maxShapes=input_ids:128x64,attention_mask:128x64
How it was made
- Backbone:
intfloat/multilingual-e5-base(XLM-RoBERTa, 12 layers, 278M), cut in depth 12 β 8 β 6 β 4 with a healing pass after each cut. - Vocabulary pruned to code and English: 250k β 22.7k SentencePiece pieces, which takes the embedding table from 192M to 17.5M parameters β the largest single saving.
- Two dense teachers, mixed: the vector space is a distillation from
gte-modernbert-baseblended with the text tower ofgoogle/embeddinggemma-2(0.3 / 0.7 of the cosine, projected back to 768-d);Qwen3-Reranker-4Bremains the ranking teacher β the student learned the cross-encoder's score distribution over hard negatives, not only "positive above negative". - Trained on the serving text format itself: 1 546 752 entity texts and queries from 173 repositories, written the way the daemon writes them at index time.
- INT8 by quantization-aware training; the E5 prefixes, the pooler and MLM heads and the token-type input were dropped; pooling and L2 were fused into the export.
Files
| file | what it is |
|---|---|
model_int8qdt.onnx |
INT8 Q/DQ graph with the trained scales β the source of every engine |
model.onnx |
the FP32 twin, same weights, for runtimes that cannot read Q/DQ and for fine-tuning |
model.safetensors, config.json |
the float weights for transformers: AutoModel loads the encoder; mean-pool over the mask β proj β L2 yourself (config.json β ultracode) |
sentencepiece.bpe.model, tokenizer_config.json |
the tokenizer, raw SentencePiece ids |
β¦-{s,m,l,xl}_{win_x64,linux_x64}_trt11.3_sm_{75,80,86,89,120}.engine |
TensorRT, per bucket Γ OS Γ GPU; sm_90 (H100/H200) Linux only |
β¦-{s,m,l,xl}_ov2026.4_{cpu_int8,igpu_lnl_int8,npu_int4}_b*_s*.{xml,bin} |
OpenVINO 2026.4 for Intel CPU, iGPU, NPU |
β¦_{win_x64,linux_x64}_tvm0.25_vulkan.{dll,so} |
TVM Vulkan, for non-NVIDIA GPUs |
model_gpu_mlx0.32/ |
MLX, Apple GPU |
coreml_ane/embed.mlpackage/ |
Core ML for the Apple Neural Engine β one function per bucket shape; load with cpuAndNeuralEngine, fall back to MLX for any other shape |
A TensorRT engine loads only on the exact GPU architecture and TensorRT version that built it:
sm_75 Turing Β· sm_80 A100/A30 Β· sm_86 RTX 30xx Β· sm_89 RTX 40xx/L4 Β· sm_90 H100/H200 Β·
sm_120 RTX 50xx, all TensorRT 11.3. On sm_120 the m / l / xl engines run attention as one fused
fp16 kernel (+11β21 % per bucket); their vectors are interchangeable with the other engines'. If
your combination is missing, build it from the ONNX.
License
Released under the MIT license. The backbone (intfloat/multilingual-e5-base) is MIT; the
teachers (gte-modernbert-base, Qwen3-Reranker-4B) are Apache-2.0. As is standard practice for
distilled embedding models, the weights are released under MIT. Not legal advice.
Attribution
Backbone: intfloat/multilingual-e5-base (MIT). Dense teacher: Alibaba-NLP/gte-modernbert-base (Apache-2.0). Ranking teacher: Qwen/Qwen3-Reranker-4B (Apache-2.0).
- Downloads last month
- 401
Model tree for faxenoff/code-daemon-embed-v1
Base model
answerdotai/ModernBERT-base