⚑ MiniTransformer-91M (Vivid86 Local SLM)

PyTorch CUDA TensorRT Hardware Parameters License Workbench

MiniTransformer-91M is an ultra-fast, local-first Small Language Model (SLM) trained completely from scratch on consumer hardware without third-party API dependencies or cloud GPU clusters.

Engineered specifically as a lightning-fast local co-processor for developer agent routing, intent classification, and structured code synthesis, it delivers sub-2ms latency and up to 671 QPS when compiled to NVIDIA TensorRT FP16.

"Sharp, confident, and relentless about writing clean code. Read before write, verify after edit, and never guess."


πŸ—οΈ Technical Architecture

MiniTransformer-91M follows a modern LLaMA-style autoregressive decoder architecture optimized for low-latency inference:

Component Specification Architectural Purpose
Total Parameters 91,245,312 (~91.2M) Right-sized for microsecond response on consumer GPUs
Layers (n_layers) 12 Transformer Blocks Balanced depth for coherent multi-step agent reasoning
Hidden Dim (d_model) 768 Standard projection dimension matching GPT-2 scale
Attention Heads 12 Query / 12 Key-Value FlashAttention Scaled Dot-Product Attention (SDPA)
Feed-Forward (d_ff) 2,048 (SwiGLU) Swish-Gated Linear Units (SiLU(W1(x)) * W3(x) -> W2)
Context Length 1,024 tokens Rotary Position Embeddings (RoPE) with strict KV-cache alignment
Vocabulary Size 8,192 tokens Custom Byte-Level BPE trained on technical & code corpora
Weight Tying Enabled Embedding table shared with output language modeling head ($V \times D$)
Normalization RMSNorm ($\epsilon = 10^{-5}$) Scaled root-mean-square normalization without mean-centering overhead
Precision FP16 / BF16 Native mixed-precision training and FP16 inference

⚑ Workstation Inference Benchmarks

All benchmarks were measured on a single consumer workstation equipped with an NVIDIA GeForce RTX 5070 12GB (Blackwell sm_120) running CUDA 12.8 on Windows 11 / WSL2.

Runtime / Engine Execution Backend Single-Token Latency Throughput (QPS) Inference VRAM
PyTorch 2.11 (Eager) FP16 Autocast 4.81 ms 208 QPS 420 MB
PyTorch torch.compile Inductor + Triton 3.10 ms 322 QPS 410 MB
ONNX Runtime CUDA Execution Provider 2.40 ms 416 QPS 360 MB
NVIDIA TensorRT 11.3 FP16 Engine (sm_120) 1.24 ms 671 QPS 290 MB

Note: TensorRT throughput was measured using concurrent asynchronous execution streams over batch size 1 with continuous generation.


πŸ”¬ Training Innovation: CPU AdamW Offload (< 1.8GB Peak VRAM)

Pre-training and fine-tuning models on consumer GPUs typically fails due to optimizer memory overhead:

  • Standard 32-bit AdamW stores two state vectors ($m_t$ and $v_t$) per parameter. For larger models, optimizer states alone consume 10+ GB of VRAM, crashing 12GB consumer cards.
  • The Solution: AdamW optimizer states were offloaded directly to 32GB system DDR5 RAM over PCIe. The RTX 5070 GPU exclusively handles the forward and backward passes.
  • The Result: Peak training VRAM stayed under 1.78 GB, leaving massive headroom for sequence lengths and batching on a standard desktop GPU.

Read the complete engineering deep-dive on the workbench:
πŸ‘‰ I Trained a 458M Model on an RTX 5070 with Under 2GB VRAM: Crazy or Frugal?


πŸš€ Quickstart & Usage

1. Standard Hugging Face transformers (Universal 3 Lines)

Because this model exports directly to LlamaForCausalLM format, it can be loaded with any standard Hugging Face pipeline:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Vivid86/MiniTransformer-91M"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, 
    torch_dtype=torch.float16, 
    device_map="auto"
)

# Format prompts using the standard Human / Assistant conversation format
prompt = "Human: Write a Python function to check if a word is a palindrome.\n\nAssistant:"

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
    **inputs, 
    max_new_tokens=128, 
    temperature=0.2, 
    top_p=0.9, 
    repetition_penalty=1.1
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

2. High-Throughput Serving with vLLM

You can deploy MiniTransformer-91M as an OpenAI-compatible HTTP API server using vLLM:

pip install vllm

vllm serve Vivid86/MiniTransformer-91M \
  --port 8000 \
  --max-model-len 1024 \
  --dtype float16

Once running, query it with any OpenAI-compatible client (Cursor, Continue.dev, LangChain, or curl):

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Vivid86/MiniTransformer-91M",
    "messages": [{"role": "user", "content": "Explain binary search simply."}],
    "temperature": 0.2
  }'

🎯 Intended Use & Honest Engineering Boundaries

βœ… Where MiniTransformer-91M Excels:

  1. Local Agent Routing: Classifying user intent, selecting tools, and dispatching tasks in sub-2ms without round-trip network lag.
  2. Structured JSON Extraction: Enforcing rigid schemas and parsing noisy inputs on-device.
  3. Local Co-Processor Loops: Running continuous validation or verification passes in background agent swarms without burning cloud API budgets.
  4. Edge & Embedded Devices: Low memory footprint (< 300MB VRAM) allows it to run on entry-level GPUs, laptops, and mini-PCs.

⚠️ Honest Limitations (The 1,024 Token Ceiling):

  • Context Ceiling: The model operates with a hard 1,024 token rotary context window. It is not designed to absorb 50 pages of documentation in a single prompt.
  • Frontier Reasoning: It will not match frontier models (Claude 3.5 Sonnet, GPT-4o) on abstract multi-hop mathematical proofs or complex full-stack architectural design.
  • Best Practice: Pair MiniTransformer as a local high-speed routing and triage layer in front of larger models or dedicated RAG pipelines.

🌐 The Vivid Ecosystem

MiniTransformer-91M is part of the Vivid Local SLM Family developed for the Vivid Developer Agent OS:

  • πŸ› οΈ Workbench & Technical Debates: Vivid's Tech Bench
  • πŸ€– Multi-Agent Runtime: Vivid Developer Agent GitHub
  • πŸ“¦ Model Family:
    • Vivid86/MiniTransformer-91M (This model β€” 1.2ms latency, 671 QPS)
    • Vivid86/MiniTransformer-220M (Balanced developer model β€” 369 QPS)
    • Vivid86/MiniTransformer-458M (Deeper reasoning model trained with DDR5 offload)

πŸ“œ License & Citation

Released under the Apache 2.0 License. Free for personal, research, and commercial use.

@misc{vivid86_minitransformer_2026,
  author = {Vivid86},
  title = {MiniTransformer-91M: High-Throughput Local SLM for Consumer Silicon},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Vivid86/MiniTransformer-91M}}
}
Downloads last month
-
Safetensors
Model size
97.5M params
Tensor type
F16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support