probe-inference weights

Trained activation probes in four architectures (linear, MLP, EFC (early-fusion covariance) and axial) for nine open-weight models. Each probe reads a model's residual-stream activations at six decoder layers and returns one score per transcript. Load them with the probe-inference package (AlignmentResearch/caught-in-the-act-probes):

from probe_inference import load_probe_from_hub

probe = load_probe_from_hub("qwen3.5-9b/efc")  # this repository at the package's pinned revision
score = probe.score(acts, probe.read_mask(prompt_mask, completion_mask, followup_start_position))

Layout

<model>/<arch>/, with <arch> in linear, mlp, efc, axial. Each directory is one trained probe. A linear or MLP probe is one small probe per layer (layer_<L>/config.json, layer_<L>/model.pt); an EFC or axial probe is one module that reads all its layers at once (config.json, model.pt). Every probe has probe_metadata.json: the model and revision, the architecture, the layers, the read window (obfuscate_over), the token aggregation (eval_sequence_aggregator) and, for linear and MLP, the layers whose sigmoids are averaged (layer_rule.used_layers, all six). model.pt files are plain float32 state dicts, with the input normaliser (input_scale, and input_mean for axial) the probe was trained with.

Directory Model (revision) Layers linear MiB MLP MiB EFC MiB axial MiB
qwen3.5-2b Qwen/Qwen3.5-2B (15852e8c) 7, 10, 13, 16, 19, 22 0.1 12.0 6.1 26.1
qwen3.5-9b Qwen/Qwen3.5-9B (c2022362) 10, 13, 18, 21, 26, 29 0.1 24.0 12.1 28.1
qwen3.5-27b Qwen/Qwen3.5-27B (fc05daec) 19, 27, 35, 43, 51, 58 0.1 30.0 15.1 29.2
qwen3.5-122b-a10b Qwen/Qwen3.5-122B-A10B (dc4d3484) 14, 20, 26, 32, 38, 43 0.1 18.0 9.1 27.1
qwen3.5-397b-a17b Qwen/Qwen3.5-397B-A17B (84726181) 18, 25, 33, 40, 48, 54 0.1 24.0 12.1 28.1
nemotron-3-nano-30b-a3b nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (bf77c317) 16, 22, 29, 35, 42, 47 0.1 15.8 8.0 26.7
nemotron-3-super-120b-a12b nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 (2dc98e2a) 26, 37, 48, 59, 70, 79 0.1 24.0 12.1 28.1
nemotron-3-ultra-550b-a55b nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 (77df655d) 32, 45, 59, 72, 86, 97 0.2 48.0 24.1 32.2
kimi-k3 moonshotai/Kimi-K3 (f831ab66) 28, 39, 51, 62, 74, 84 0.2 42.0 21.1 31.2

The layers are at depth fractions 0.3, 0.42, 0.55, 0.67, 0.8 and 0.9 of the model's decoder blocks, round(f * num_blocks). Linear and MLP scores are probabilities (the mean of per-layer sigmoids); EFC and axial scores are logits.

Training

Every probe was trained on FIBSv1 (AlignmentResearch/fibs-v1) at revision 65dccf12934620bfe01db665ea3d7d0e3c9f1355, on the model's own activations over FIBSv1's transcripts:

  • 152,980 training rows and 5,000 validation rows, the same rows for every model; seed 0;
  • linear and MLP: up to 6 epochs; EFC and axial: up to 4 epochs;
  • early stopping: training stops after 4 checkpoints (one epoch) in a row without a lower validation loss (6 epochs for linear/mlp). The published probe is the checkpoint with the lowest validation loss;
  • linear and MLP divide each layer's activations by input_scale, and axial subtracts input_mean and then divides by input_scale. These are computed per layer from up to 512 training rows and are stored in model.pt. EFC normalises each token by its own RMS and stores no statistic;
  • each published probe reproduces the trainer's validation loss and AUROC within 1e-4.

Validation AUROC on the 5,000 validation rows (ES: stopped early):

Directory linear MLP EFC axial
qwen3.5-2b 0.926 0.968 0.987 0.991 ES
qwen3.5-9b 0.965 0.986 0.996 0.999
qwen3.5-27b 0.972 0.988 0.998 0.999
qwen3.5-122b-a10b 0.938 0.972 0.998 ES 0.999
qwen3.5-397b-a17b 0.948 0.977 0.999 0.999
nemotron-3-nano-30b-a3b 0.951 0.980 0.993 0.996 ES
nemotron-3-super-120b-a12b 0.968 0.986 0.998 0.998 ES
nemotron-3-ultra-550b-a55b 0.982 0.993 0.999 0.999 ES
kimi-k3 0.948 0.959 0.999 0.999 ES

Activations the probes expect

  • Layer k is the output of decoder block k, i.e. Hugging Face hidden_states[k + 1], in bfloat16; the probes run in float32. All activations were captured with vLLM at the decoder-layer outputs.
  • Kimi K3's decoder layers use attention residuals, so a layer has no single residual stream. Its layer k is the attention-residual mixture that layer k + 1 reads, computed with the model's own attn_res op. Kimi K3 ran from its released MXFP4 checkpoint, not a bf16 one; its activations are bfloat16.
  • Linear and MLP read one token: the token before the final end-of-turn token. For Qwen and Nemotron-3 that is the answer's last token. For Kimi K3 it is the <|sep|> that closes <|close|>message, after the answer's <|close|>response<|sep|>. EFC and axial read every token from the start of the final user turn through the end-of-turn token.

Licences and attribution

The probe weights and this card are released by FAR AI, Inc. under the MIT licence (LICENSE).

They are derived from the models below. Their licence texts ship here unchanged, and their attribution notices are kept in NOTICE:

Model Licence Licence file
Qwen/Qwen3.5-2B, Qwen/Qwen3.5-9B, Qwen/Qwen3.5-27B, Qwen/Qwen3.5-122B-A10B, Qwen/Qwen3.5-397B-A17B Apache-2.0 LICENSE-QWEN-APACHE-2.0.txt
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16, nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16 NVIDIA Nemotron Open Model License (v. December 15, 2025) LICENSE-NVIDIA-NEMOTRON-OPEN-MODEL.txt
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 OpenMDW License Agreement, version 1.1 LICENSE-OPENMDW-1.1.txt
moonshotai/Kimi-K3 Kimi K3 License LICENSE-KIMI-K3.txt

Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AlignmentResearch/probe-inference-weights

Finetuned
(55)
this model

Dataset used to train AlignmentResearch/probe-inference-weights