code-daemon-embed-v1

A 4-layer, 46.8M-parameter code-embedding model built for one job: embedding a whole repository fast enough to re-index it on every commit, as the dense half of a hybrid (vectors + BM25) code search. Short code units β€” signatures, docstrings, symbol names, one-line descriptions β€” and short keyword-shaped queries land in one 768-dim space.

What makes it different

  • Speed first. 4–6Γ— the throughput of any public 768-dim encoder we tested; a 1.1M-entity C++ repository embeds in about 80 seconds on a laptop GPU.
  • A reranker's judgement baked into the vectors. A 4B cross-encoder's ranking was distilled into the bi-encoder at training time; in the system it was built for, a runtime reranker on top of it measured net-negative.
  • Nothing to get wrong at the output. Mean pooling and L2 normalization are inside the graph: the model returns unit-norm [B, 768], ready for a dot product.
  • INT8 from quantization-aware training, so TensorRT and OpenVINO build an INT8 engine with no calibration pass.

Queries and documents are encoded the same way β€” no query: / passage: prefix. Hard cap of 128 tokens. English and code only.

Choosing it β€” the trade

Vector channel alone, every model at the same 128-token budget. Quality is nDCG@10 Γ— 100; speed is texts/s through onnxruntime CUDA FP32, batch 64, at seq 64 (the mean serving text is 70 tokens). Laptop RTX 5060, 2026-09-11.

model layers texts/s cosqa stack-qa code→code our queries Russian hit@1
code-daemon-embed-v1 4 3 457 28.5 49.9 42.9 60.1 0.06
multilingual-e5-base 12 854 29.8 79.4 57.7 60.7 0.24
CodeRankEmbed 12 723 35.5 73.3 β€” 64.6 0.00
gte-modernbert-base (its teacher) 22 606 36.6 81.4 82.5 70.5 0.15
  • 4.0–5.7Γ— the throughput for 1–10 points of nDCG on our queries. Pick it when indexing speed is the constraint and a lexical channel runs beside it; pick gte-modernbert-base when vector-only accuracy is.
  • Inside the hybrid search it was built for (80 real agent queries, file level, runtime reranker off): the right file is first 44 % of the time and in the top 5 69 %.
  • Not multilingual. The vocabulary prune kept 942 of XLM-R's 31 671 Cyrillic pieces; for non-English queries use multilingual-e5-base.
  • Not for long text or general prose. On four CoIR tasks it averages 52.8 nDCG@10 β€” 89.8 on Python docstringβ†’code, 28.5 on cosqa.

Throughput in a real index

A full index through the Code-Daemon daemon, laptop RTX 5060, TensorRT INT8, pinned 114 W:

workload texts/s
C/C++ β€” mysql-server, 1.1M entities 13 800
TypeScript β€” vscode, 479k 10 600
Java β€” netty, 97k 8 000
C# β€” roslyn, 403k 7 900
small incremental updates (a few hundred to 15k texts) 1 900 – 5 500
no discrete GPU β€” Intel CPU + iGPU + NPU together ~1 050
Apple M4, Neural Engine ~1 100

The spread between languages is text length: 71 % of C# texts fall in the longest bucket against 24 % for C/C++. Above ~820k texts the daemon adds the Intel iGPU beside the RTX card, so the mysql-server row is a two-device number. The same daemon running multilingual-e5-base compiled to the same INT8 engines is 2.3Γ— slower on every row.

How to use it

Pick a length bucket

Every engine family ships four length buckets. A batch pays for its longest member, so route each text to the smallest bucket that fits it:

bucket batch seq range (opt) solo texts/s, RTX 5060 typical content
s 96 8 – 48 (48) 19 800 at 48 names, signatures, queries
m 128 56 – 64 (56) 13 100 at 64 typical entity text β€” take this one if you load only one
l 128 72 – 80 (72) 9 900 at 80
xl 256 88 – 128 (96) 5 500 at 128 long docstrings, doc chunks
  • Sort by length, then batch, and dispatch each batch at its own longest length rounded up to a multiple of 8. The TensorRT engines take a dynamic sequence range for exactly this β€” worth +10 % end to end.
  • Batch size is not the lever: 96 Γ— 48 already saturates the GPU, and a bigger batch was slower per text. Running the four engines concurrently does not add throughput either.
  • All four TensorRT engines together take ~400 MB of VRAM.
  • OpenVINO IRs are static shapes: batch 64 on CPU and iGPU, batch 16 on the NPU. The NPU has the lowest latency per batch (28–95 ms) and suits single interactive queries better than bulk indexing.

Shape the input

Documents β€” the model was trained on a compact, front-loaded text, and reproducing it is worth more than any inference tuning:

function acquireProjectLock in src/storage/multi_db.zig: zig
path: src storage multi db
[exported]
sig: (allocator: Allocator, project_hash: []const u8) -> !Lock
Acquires the exclusive SQLite lock for one project.

That is {type} {name} in {file}: {lang}, the path as words, flags, the signature, then the first ~200 characters of the doc comment and a one-sentence description; ~768 characters at most. Leave the function body out β€” it belongs in the lexical channel.

Queries β€” raw, no prefix. Short keyword bags and behaviour descriptions are what it was tuned for. Longer than 128 tokens β€” windows of 128, stride 96; mean-pool the window vectors and re-normalize.

Run the ONNX

import onnxruntime as ort, sentencepiece as spm, numpy as np

sp   = spm.SentencePieceProcessor(model_file="sentencepiece.bpe.model")  # pad=0 unk=1 bos=2 eos=3
sess = ort.InferenceSession("model_int8qdt.onnx", providers=["CPUExecutionProvider"])

def embed(texts, max_len=128):
    ids  = [[2, *sp.encode(t)[: max_len - 2], 3] for t in texts]            # bos … eos
    L    = max(len(x) for x in ids)
    inp  = np.array([x + [0] * (L - len(x)) for x in ids], dtype=np.int64)  # pad=0
    mask = (inp != 0).astype(np.int64)
    return sess.run(None, {"input_ids": inp, "attention_mask": mask})[0]    # [B, 768], unit-norm

D = embed(["function acquireLock in src/db.zig: zig\npath: src db"])
Q = embed(["acquire database lock"])
print(Q @ D.T)                                                              # inner product = cosine

Tokenize with SentencePiece β€” the vocabulary is a unigram model, and a BPE merge loop segments it differently from training. Both inputs are int64.

Building your own engines

Feed model_int8qdt.onnx as it is. TensorRT 11 reads the precision from its Q/DQ nodes; OpenVINO keeps the trained scales as FakeQuantize (run compress_quantize_weights_transformation before saving, or the IR comes out twice the size). Never run a calibration / PTQ pass over it: it replaces the trained scales with fitted ones, and the engine builds and loads cleanly and quietly retrieves worse. Use the bucket profiles from the table above, for example:

trtexec --onnx=model_int8qdt.onnx --saveEngine=embed-m.engine --builderOptimizationLevel=5 \
        --minShapes=input_ids:128x56,attention_mask:128x56 \
        --optShapes=input_ids:128x56,attention_mask:128x56 \
        --maxShapes=input_ids:128x64,attention_mask:128x64

How it was made

  • Backbone: intfloat/multilingual-e5-base (XLM-RoBERTa, 12 layers, 278M), cut in depth 12 β†’ 8 β†’ 6 β†’ 4 with a healing pass after each cut.
  • Vocabulary pruned to code and English: 250k β†’ 22.7k SentencePiece pieces, which takes the embedding table from 192M to 17.5M parameters β€” the largest single saving.
  • Two dense teachers, mixed: the vector space is a distillation from gte-modernbert-base blended with the text tower of google/embeddinggemma-2 (0.3 / 0.7 of the cosine, projected back to 768-d); Qwen3-Reranker-4B remains the ranking teacher β€” the student learned the cross-encoder's score distribution over hard negatives, not only "positive above negative".
  • Trained on the serving text format itself: 1 546 752 entity texts and queries from 173 repositories, written the way the daemon writes them at index time.
  • INT8 by quantization-aware training; the E5 prefixes, the pooler and MLM heads and the token-type input were dropped; pooling and L2 were fused into the export.

Files

file what it is
model_int8qdt.onnx INT8 Q/DQ graph with the trained scales β€” the source of every engine
model.onnx the FP32 twin, same weights, for runtimes that cannot read Q/DQ and for fine-tuning
model.safetensors, config.json the float weights for transformers: AutoModel loads the encoder; mean-pool over the mask β†’ proj β†’ L2 yourself (config.json β†’ ultracode)
sentencepiece.bpe.model, tokenizer_config.json the tokenizer, raw SentencePiece ids
…-{s,m,l,xl}_{win_x64,linux_x64}_trt11.3_sm_{75,80,86,89,120}.engine TensorRT, per bucket Γ— OS Γ— GPU; sm_90 (H100/H200) Linux only
…-{s,m,l,xl}_ov2026.4_{cpu_int8,igpu_lnl_int8,npu_int4}_b*_s*.{xml,bin} OpenVINO 2026.4 for Intel CPU, iGPU, NPU
…_{win_x64,linux_x64}_tvm0.25_vulkan.{dll,so} TVM Vulkan, for non-NVIDIA GPUs
model_gpu_mlx0.32/ MLX, Apple GPU
coreml_ane/embed.mlpackage/ Core ML for the Apple Neural Engine β€” one function per bucket shape; load with cpuAndNeuralEngine, fall back to MLX for any other shape

A TensorRT engine loads only on the exact GPU architecture and TensorRT version that built it: sm_75 Turing Β· sm_80 A100/A30 Β· sm_86 RTX 30xx Β· sm_89 RTX 40xx/L4 Β· sm_90 H100/H200 Β· sm_120 RTX 50xx, all TensorRT 11.3. On sm_120 the m / l / xl engines run attention as one fused fp16 kernel (+11–21 % per bucket); their vectors are interchangeable with the other engines'. If your combination is missing, build it from the ONNX.

License

Released under the MIT license. The backbone (intfloat/multilingual-e5-base) is MIT; the teachers (gte-modernbert-base, Qwen3-Reranker-4B) are Apache-2.0. As is standard practice for distilled embedding models, the weights are released under MIT. Not legal advice.

Attribution

Backbone: intfloat/multilingual-e5-base (MIT). Dense teacher: Alibaba-NLP/gte-modernbert-base (Apache-2.0). Ranking teacher: Qwen/Qwen3-Reranker-4B (Apache-2.0).

Downloads last month
401
Safetensors
Model size
46.8M params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for faxenoff/code-daemon-embed-v1

Quantized
(14)
this model