Wisp-5M

Wisp-5M is a tiny decoder-only language model from the Wisp family, trained from scratch on 3.00B tokens of English web text, code and math. It uses a plain LLaMA architecture, so it loads with transformers (LlamaForCausalLM) with no custom code.

It is a base model (not instruction-tuned). Its size makes it useful for research, education, speculative decoding drafts, on-device experiments and as a starting point for fine-tuning. Don't expect factual accuracy.

loss curve

The Wisp family

Model Params Non-embedding Hidden Layers Heads KV heads FFN Tokens
Wisp-5M 5.1M 3.5M 192 9 6 2 512 3B
Wisp-15M 15.5M 12.9M 320 12 10 2 864 3B
Wisp-60M (planned) 60.6M 56.4M 512 20 8 2 1408 —

All models share the same tokenizer and saw the same data in the same order: one fixed global shuffle of 2048-token windows. That makes the family directly comparable across sizes.

Architecture

  • LLaMA-style decoder: pre-norm RMSNorm, rotary position embeddings (θ = 10,000), SwiGLU MLP, grouped-query attention, no biases, tied input/output embeddings
  • Deep-and-thin shapes (following MobileLLM), which work better than wide-and-shallow at this scale
  • Context length: 2048 tokens
  • Tokenizer: Wisp byte-level BPE, 8,192 tokens, trained on the same data mix; digits are split individually (better for arithmetic); <|endoftext|> is used as BOS/EOS and document separator. Compression is ~3.8 characters/token on English web text and ~3.0 on code.

Training

Tokens 3.00B (11,444 steps × 262,144 tokens)
Sequence length 2048
Optimizer Muon (hidden matrices, Moonlight RMS-matched update) + AdamW (embeddings, norms)
Peak LR / weight decay 0.01 / 0.1
Schedule Warmup-Stable-Decay: 1% warmup, linear decay to 0 over the last 20%
Precision mixed precision (bf16/fp16 autocast, fp32 master weights), torch.compile
Hardware 1× NVIDIA Tesla T4 (Hugging Face Jobs), 340 min
  • Compute cost: ~US$2.28 on Hugging Face Jobs

Data mix

An 8B-token pool was pretokenized with the mix below and shuffled once as 2048-token windows. Each model trained on the first 3.00B tokens of that shuffle, which keeps the same proportions.

Source Subset / filter Share Tokens in pool
FineWeb-Edu sample-100BT 45% 3.6B
DCLM-Baseline global-shard_01 35% 2.8B
The Stack v3 40 popular languages, no vendor files, <=64KB 12% 0.96B
FineMath finemath-4plus 8% 0.64B

Evaluation

Held-out cross-entropy (nats/token, lower is better) on documents never seen in training:

FineWeb-Edu DCLM Stack v3 FineMath Average
3.228 3.456 1.768 2.575 2.757

The average is unweighted across the four sources. Code is much more predictable than prose, so it sits below the training loss, which follows the 45/35/12/8 training mix. Weighted by that mix, the held-out loss is 3.080.

Zero-shot benchmarks

lm-evaluation-harness, zero-shot; acc_norm for HellaSwag/ARC/PIQA/OpenBookQA/SciQ, acc for WinoGrande/LAMBADA. Pythia-70M (70M params, 300B tokens of the Pile) is shown as a reference from EleutherAI's published evaluations.

Task Wisp-5M Pythia-70M Random
HellaSwag 27.1 — 25.0
ARC-Easy 32.7 35.0 25.0
ARC-Challenge 22.7 22.1 25.0
PIQA 54.2 59.1 50.0
WinoGrande 48.6 52.8 50.0
OpenBookQA 24.6 — 25.0
SciQ 60.2 55.2 25.0
LAMBADA 16.8 18.5 0.0
LAMBADA ppl 266.9 142.4 —

Greedy samples right after training:

  • 'The meaning of life is the same as the meaning of life. The meaning of life is the same as the meaning of life.\nThe meaning of life is the same as the meaning of life. It means that we know what we are doing and how we can do'
  • 'Once upon a time, the world was not as much of an issue. The world had to be more than just a single day and a week.\nThe world was very different from what we have today. We are all in our own homes. It is a great'
  • 'def is_prime(n): 10\n# 2. The number of digits in the first digit is 3, and the last digit is 4.\n# 3. The number of digits in the second digit is 5, and the last digit'
  • 'The derivative of x^2 is the product of its value. The derivative of y^2 is the product of its value.\n\n## 1. What are the derivatives?\n\n### 3. How do you find the derivative of x^2?\n\n### '

Usage

from transformers import AutoTokenizer, AutoModelForCausalLM

repo = "DedeProGames/Wisp-5M"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo)

inputs = tok("The water cycle is", return_tensors="pt")
out = model.generate(**inputs, max_new_tokens=60, do_sample=True, temperature=0.8, top_p=0.95,
                     repetition_penalty=1.1)
print(tok.decode(out[0], skip_special_tokens=True))

The tokenizer prepends <|endoftext|> as BOS, which matches how documents were laid out during training.

Limitations

A model this small has very limited knowledge and reasoning. It will produce fluent-looking but often wrong or incoherent text, and it can reproduce biases present in web data. It is English-centric and was not aligned or safety-tuned. Don't use it for anything where correctness matters.

License

Apache-2.0 for the model weights. Training data licenses: FineWeb-Edu, FineMath and The Stack v3 are ODC-By; DCLM-Baseline is CC-BY-4.0. Code in The Stack v3 comes with its original licenses.

Downloads last month
8
Safetensors
Model size
5.12M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train DedeProGames/Wisp-5M

Spaces using DedeProGames/Wisp-5M 2