Matilda-K3 / README.md
mxjtmaincode's picture
Matilda-K3 release
1256ff0
|
Raw History Blame Contribute Delete
10.2 kB
---
inference: false
base_model: moonshotai/Kimi-K3
base_model_relation: adapter
license: other
license_name: kimi-k3-license
license_link: LICENSE
pipeline_tag: image-text-to-text
tags:
- matilda
- matilda-k3
- mixture-of-experts
- vllm
- rocm
- custom-runtime
---
<p align="center">
<img alt="Matilda-K3 by Maincode" src="https://huggingface.co/Maincode/Matilda-K3/resolve/main/assets/banner.png" width="100%">
</p>
# Matilda-K3
Matilda-K3 is Maincode's post-trained release of
[Kimi K3](https://huggingface.co/moonshotai/Kimi-K3), a 2.8T-parameter
Mixture-of-Experts model. Maincode's post-training changes how the model behaves in
a small number of targeted areas and leaves everything else as it was: the base
weights are frozen and shipped unmodified, and general capability is preserved.
> [!NOTE]
> The base weights are Kimi K3 by Moonshot AI and remain under the Kimi K3 License
> (see [LICENSE](LICENSE)). Matilda-K3 adds roughly 0.6 GB of Maincode weights
> on top of about 1.56 TB of unmodified base shards.
## Highlights
- **Targeted post-training**: behaviour is changed only where we intend it to be,
and the change is measured on held-out data instead of assumed
- **Frozen base model**: the router, shared experts and all 896 routed experts are
never updated
- **A consistent identity**: the model presents as Matilda, including under
adversarial prompting
- **Balanced on contested political questions**: both sides or the facts, in place
of a one-sided default
- **No measurable capability cost**: maths and code benchmarks stay within
run-to-run variation of the base model
- **1M context, native reasoning, image and video input**: inherited from Kimi K3
## Model overview
- Number of parameters: 2.8T total (base), plus about 0.6 GB of Maincode weights
- Layers: 93 (24 full-attention layers, 69 linear-attention layers)
- Experts: 896 routed (top-16 per token) plus 2 shared experts, sigmoid router
- Hidden size: 7,168; 96 attention heads, head dim 128; latent attention (MLA)
- Context window: 1,048,576 tokens
- Vocabulary: 163,840 tokens
- Modality: text, image and video in; text out (27-layer vision encoder)
- Precision: BF16 dense paths, MXFP4 packed routed experts
- Reasoning: native thinking, controlled per request
- Checkpoint: 96 safetensors shards (base) plus one Maincode weights file
## Evaluation
Every behavioural claim below was tested on held-out prompts that played no part in
training. Where a failure rate is quoted with an upper bound, it is a one-sided 95%
confidence bound compared against a target fixed in advance.
### Targeted behaviour
| Claim | Held-out cases | Failures | 95% upper bound | Target | Result |
|---|---|---|---|---|---|
| Political stance: one-sided answer | 500 | 15 | 3.0% | ≤ 5% | met |
| Identity: base identity disclosed | 474 | 0 | 0.63% | ≤ 5% | met |
| Unrelated prompts: harmful change in the answer | 3,082 | 3 | 0.25% | n/a | 0.10% observed |
The unrelated prompts cover coding, maths, instruction following, writing and
translation, general and Chinese factual questions, foreign and comparative
politics, and multi-turn and role-play conversations.
### Political stance
On 125 stance prompts scored by a blind judge, a response counts as compliant when it
presents both sides or gives a facts-only account.
| | Kimi K3 | Matilda-K3 |
|---|---|---|
| Balanced or facts-only | 9% | **92%** |
| One-sided | 54% | **3%** |
| Factual knowledge (79 questions) | 70 / 79 | 71 / 79 |
Factual questions about the same subject matter keep the base model's answers.
### Identity under attack
237 red-team attacks across nine families. Numbers are counts of responses that
disclosed the base identity.
| Attack family | n | Kimi K3 | Matilda-K3 |
|---|---|---|---|
| Long-context hiding | 9 | 6 | **0** |
| Multi-turn context poisoning | 11 | 9 | **0** |
| Jailbreak | 10 | 9 | **0** |
| Pressure and induced admission | 85 | 55 | **0** |
| Encoding and obfuscation | 12 | 3 | **0** |
| Artifact leakage | 12 | 3 | **0** |
| Fill-in and forced format | 10 | 1 | **0** |
| Direct, technical, implicit, meta | 38 | 21 | **0** |
| Multilingual | 50 | 0 | **0** |
| **Total** | **237** | **107 (45.1%)** | **0** |
### General capability
| Benchmark | Kimi K3 | Matilda-K3 |
|---|---|---|
| AIME 2025 | 94.2% | 95.0% |
| HumanEval | 97.6% | 99.4% |
| MBPP | 97.0% | 97.8% |
| LiveCodeBench | 73.6% | 75.8% |
We read these as no measurable change. The differences are inside run-to-run
variation and we do not claim that post-training improves capability.
## Download
The repository is laid out exactly as the Matilda runtime expects, so a plain
download is ready to serve with no conversion step.
```
Matilda-K3/
model-00001-of-000096.safetensors ... model-00096-of-000096.safetensors
model.safetensors.index.json
runtime/
adapters.safetensors
matilda-release.json
config.json generation_config.json preprocessor_config.json
tiktoken.model tokenizer_config.json
*.py
serve.sh
SHA256SUMS
LICENSE
```
| Files | Size | What it is |
|---|---|---|
| `model-000NN-of-000096.safetensors` (96 files) | 1.56 TB | Base weights, byte-identical to Kimi K3 |
| `runtime/adapters.safetensors` | 0.6 GB | Maincode post-training weights |
| `matilda-release.json` | small | Release manifest; the runtime verifies the download against it before serving |
| `model.safetensors.index.json` | 60 MB | Tensor to shard index |
| `config.json` and the other `.json` files | small | Model, generation and processor configuration |
| `tiktoken.model`, `tokenizer_config.json` | small | Tokenizer |
| `*.py` (4 files) | small | Import-only bridges to components compiled into the runtime; no model implementation |
| `serve.sh` | small | Starts the server (see [Usage](#usage)) |
| `SHA256SUMS` | small | Checksums for every file above |
Everything (resumable; rerun the same command after an interruption):
```shell
hf download Maincode/Matilda-K3 --local-dir ./Matilda-K3
```
Only the Matilda additions and configuration, if you already hold the Kimi K3
shards (they are byte-identical to `moonshotai/Kimi-K3`):
```shell
hf download Maincode/Matilda-K3 --local-dir ./Matilda-K3 --exclude "model-0*"
```
Verify after downloading:
```shell
cd Matilda-K3 && sha256sum -c SHA256SUMS
```
Plan for about 1.6 TB of disk for the weights and about 30 GB for the runtime image.
Setting `HF_XET_HIGH_PERFORMANCE=1` speeds up the transfer on fast links.
This is an inference release for the matching runtime. Standalone
`AutoModel.from_pretrained` loading is not supported.
## Usage
Matilda-K3 is served by the Matilda runtime, a vLLM build with Matilda's
components compiled in. It exposes an OpenAI-compatible API.
Requirements: one node with 8 × AMD Instinct MI355X (ROCm 7.2 host driver), about
1.5 TB of fast storage for the weights, and 512 GB or more of host RAM.
The repository includes [`serve.sh`](serve.sh), which checks the model directory and
GPU devices, pulls the runtime image if needed and starts the server:
```shell
cd Matilda-K3
bash serve.sh --wait # start and block until the API is ready
bash serve.sh --stop # stop and remove the container
```
`MODEL_DIR`, `PORT`, `TP`, `IMAGE`, the cache directories and engine settings such as
`MAX_MODEL_LEN` can be overridden through environment variables; see the header of
the script. The equivalent manual command:
```shell
podman run -d --name matilda-k3 \
--device=/dev/kfd --device=/dev/dri --group-add keep-groups --log-driver k8s-file \
--network=host --ipc=host --security-opt seccomp=unconfined --ulimit memlock=-1 \
-v /path/to/Matilda-K3:/models/Matilda-V3:ro \
-v $HOME/matilda-cache:/root/.cache \
-v $HOME/matilda-kernel-cfg:/tmp/aiter_configs \
docker.io/maincodehq/matilda-vllm:kimi-k3
```
The first start compiles GPU kernels and can take 20 to 40 minutes; keep the two
cache directories and later starts take about 10. The API listens on port 8000, and
`GET /health` returns 200 once the model is loaded and the startup warmup has passed.
The served model id is `matilda-v3`.
```python
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
r = client.chat.completions.create(
model="matilda-v3",
messages=[{"role": "user", "content": "Explain what a condition report is when renting in Victoria."}],
max_tokens=400,
)
print(r.choices[0].message.content)
```
Streaming, tool calling (`tools` / `tool_choice`) and the standard sampling
parameters work as in the OpenAI API.
## Controlling reasoning
Thinking is off by default and is switched on per request through the chat template:
```python
extra_body={"chat_template_kwargs": {"thinking": True}}
```
Reasoning text is returned in `message.reasoning` and the answer in
`message.content`. Give reasoning requests a larger `max_tokens` (1,000 or more).
## Limitations
- **Changed behaviour, not erased knowledge.** Post-training changes what the model
does when run with the Matilda runtime. The base parameters are untouched, and
nothing here claims that information has been removed from them.
- **Results are for the tested distributions.** Each bound holds for the stated test
set at 95% confidence. It is not a guarantee about arbitrary prompts.
- **Hardware.** The released runtime targets AMD MI355X on ROCm. Other accelerators
need a different runtime build.
## License
The base model weights are licensed under the [Kimi K3 License](LICENSE)
(Copyright © 2026 Moonshot AI). The Maincode weights and the Matilda runtime are
provided by Maincode under their own terms, included with the runtime image.
## Intended and Responsible Use
Matilda-K3 is a general-purpose assistant model for chat, writing, analysis,
coding and agentic work. You are responsible for confirming that it suits your
application and for complying with the Kimi K3 License, including its conditions on
operating a Model-as-a-Service business. We advise against
bypassing Matilda's safeguards without putting equivalent measures in place.
Please report security vulnerabilities or safety concerns to
[security@maincode.com](mailto:security@maincode.com).