β‘ MiniTransformer-91M (Vivid86 Local SLM)
MiniTransformer-91M is an ultra-fast, local-first Small Language Model (SLM) trained completely from scratch on consumer hardware without third-party API dependencies or cloud GPU clusters.
Engineered specifically as a lightning-fast local co-processor for developer agent routing, intent classification, and structured code synthesis, it delivers sub-2ms latency and up to 671 QPS when compiled to NVIDIA TensorRT FP16.
"Sharp, confident, and relentless about writing clean code. Read before write, verify after edit, and never guess."
ποΈ Technical Architecture
MiniTransformer-91M follows a modern LLaMA-style autoregressive decoder architecture optimized for low-latency inference:
| Component | Specification | Architectural Purpose |
|---|---|---|
| Total Parameters | 91,245,312 (~91.2M) | Right-sized for microsecond response on consumer GPUs |
Layers (n_layers) |
12 Transformer Blocks | Balanced depth for coherent multi-step agent reasoning |
Hidden Dim (d_model) |
768 | Standard projection dimension matching GPT-2 scale |
| Attention Heads | 12 Query / 12 Key-Value | FlashAttention Scaled Dot-Product Attention (SDPA) |
Feed-Forward (d_ff) |
2,048 (SwiGLU) | Swish-Gated Linear Units (SiLU(W1(x)) * W3(x) -> W2) |
| Context Length | 1,024 tokens | Rotary Position Embeddings (RoPE) with strict KV-cache alignment |
| Vocabulary Size | 8,192 tokens | Custom Byte-Level BPE trained on technical & code corpora |
| Weight Tying | Enabled | Embedding table shared with output language modeling head ($V \times D$) |
| Normalization | RMSNorm ($\epsilon = 10^{-5}$) | Scaled root-mean-square normalization without mean-centering overhead |
| Precision | FP16 / BF16 | Native mixed-precision training and FP16 inference |
β‘ Workstation Inference Benchmarks
All benchmarks were measured on a single consumer workstation equipped with an NVIDIA GeForce RTX 5070 12GB (Blackwell sm_120) running CUDA 12.8 on Windows 11 / WSL2.
| Runtime / Engine | Execution Backend | Single-Token Latency | Throughput (QPS) | Inference VRAM |
|---|---|---|---|---|
| PyTorch 2.11 (Eager) | FP16 Autocast | 4.81 ms | 208 QPS | 420 MB |
PyTorch torch.compile |
Inductor + Triton | 3.10 ms | 322 QPS | 410 MB |
| ONNX Runtime | CUDA Execution Provider | 2.40 ms | 416 QPS | 360 MB |
| NVIDIA TensorRT 11.3 | FP16 Engine (sm_120) |
1.24 ms | 671 QPS | 290 MB |
Note: TensorRT throughput was measured using concurrent asynchronous execution streams over batch size 1 with continuous generation.
π¬ Training Innovation: CPU AdamW Offload (< 1.8GB Peak VRAM)
Pre-training and fine-tuning models on consumer GPUs typically fails due to optimizer memory overhead:
- Standard 32-bit AdamW stores two state vectors ($m_t$ and $v_t$) per parameter. For larger models, optimizer states alone consume 10+ GB of VRAM, crashing 12GB consumer cards.
- The Solution: AdamW optimizer states were offloaded directly to 32GB system DDR5 RAM over PCIe. The RTX 5070 GPU exclusively handles the forward and backward passes.
- The Result: Peak training VRAM stayed under 1.78 GB, leaving massive headroom for sequence lengths and batching on a standard desktop GPU.
Read the complete engineering deep-dive on the workbench:
π I Trained a 458M Model on an RTX 5070 with Under 2GB VRAM: Crazy or Frugal?
π Quickstart & Usage
1. Standard Hugging Face transformers (Universal 3 Lines)
Because this model exports directly to LlamaForCausalLM format, it can be loaded with any standard Hugging Face pipeline:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Vivid86/MiniTransformer-91M"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto"
)
# Format prompts using the standard Human / Assistant conversation format
prompt = "Human: Write a Python function to check if a word is a palindrome.\n\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=128,
temperature=0.2,
top_p=0.9,
repetition_penalty=1.1
)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
2. High-Throughput Serving with vLLM
You can deploy MiniTransformer-91M as an OpenAI-compatible HTTP API server using vLLM:
pip install vllm
vllm serve Vivid86/MiniTransformer-91M \
--port 8000 \
--max-model-len 1024 \
--dtype float16
Once running, query it with any OpenAI-compatible client (Cursor, Continue.dev, LangChain, or curl):
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Vivid86/MiniTransformer-91M",
"messages": [{"role": "user", "content": "Explain binary search simply."}],
"temperature": 0.2
}'
π― Intended Use & Honest Engineering Boundaries
β Where MiniTransformer-91M Excels:
- Local Agent Routing: Classifying user intent, selecting tools, and dispatching tasks in sub-2ms without round-trip network lag.
- Structured JSON Extraction: Enforcing rigid schemas and parsing noisy inputs on-device.
- Local Co-Processor Loops: Running continuous validation or verification passes in background agent swarms without burning cloud API budgets.
- Edge & Embedded Devices: Low memory footprint (< 300MB VRAM) allows it to run on entry-level GPUs, laptops, and mini-PCs.
β οΈ Honest Limitations (The 1,024 Token Ceiling):
- Context Ceiling: The model operates with a hard 1,024 token rotary context window. It is not designed to absorb 50 pages of documentation in a single prompt.
- Frontier Reasoning: It will not match frontier models (Claude 3.5 Sonnet, GPT-4o) on abstract multi-hop mathematical proofs or complex full-stack architectural design.
- Best Practice: Pair MiniTransformer as a local high-speed routing and triage layer in front of larger models or dedicated RAG pipelines.
π The Vivid Ecosystem
MiniTransformer-91M is part of the Vivid Local SLM Family developed for the Vivid Developer Agent OS:
- π οΈ Workbench & Technical Debates: Vivid's Tech Bench
- π€ Multi-Agent Runtime: Vivid Developer Agent GitHub
- π¦ Model Family:
Vivid86/MiniTransformer-91M(This model β 1.2ms latency, 671 QPS)Vivid86/MiniTransformer-220M(Balanced developer model β 369 QPS)Vivid86/MiniTransformer-458M(Deeper reasoning model trained with DDR5 offload)
π License & Citation
Released under the Apache 2.0 License. Free for personal, research, and commercial use.
@misc{vivid86_minitransformer_2026,
author = {Vivid86},
title = {MiniTransformer-91M: High-Throughput Local SLM for Consumer Silicon},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Vivid86/MiniTransformer-91M}}
}
- Downloads last month
- -