We are announcing the Supra3 family with four core SLM models: - Supra3 Flash Lite: 25M parameters, ~60B pretraining tokens - Supra3 Flash: 50M parameters, ~100B pretraining tokens - Supra3 Pro: 75M parameters, ~150B pretraining tokens - Supra3 Ultra: 100M parameters, ~200B pretraining tokens
For Supra3 Pro and Ultra, we search for sponsors who give us free access to compute like RTX 5090 32GB or so.
We estimate the total cost of the pro and ultra models at around $600.
For Supra3 Flash Lite and Flash, we do not need sponsors.
If anyone would apply for helping us, we would be really thankful and this person would get early access to new modele, insider information, credit and more!
I don't know if it was us or one of you guys or maybe all of us at once but lately we have seen a finetuning/pretraining explosion of models below 200m params and we can't be more happy about it keep coming tinkerers all of this is possible because of you!
Today, we are announcing a brand-new series of SupraLabs models: Supra2 This series will feature various models, including such as: - ๐ Supra2-Nano (0.4M) โ The smallest Supra2 model. - ๐ค Supra2-Small (1.4M) โ The tiny model that runs everywhere. - ๐ช Supra2-Medium (25M) โ Our medium class model in the Supra2 family. The powerful midsizer. - ๐ฅ Supra2-Pro (100M): base, instruct, reasoning, code, math and more! โ The most capable model yet! A real allrounder for all your everyday tasks. - ๐จ Supra2-IMG โ our generative text-to-image model ...and many more...
Current progress: - Nano (0.4M) and Small (1.4M): in training; almost done. Baseline set. - Medium (25M): coming soon... - Pro (100M): in training; finishes in 66 hours - Monday, 3rd August 2026, 12:00AM - IMG: coming soon...
You can support us with a like and follow if you want! Don't miss our next release! Stay tuned...
After a month of interacting with my AI Waifu, I noticed a few issues in the system; so I decided to spend this week revisiting the systems implemented in Phase 1.0, 1.5 and 2.0, and try to make them to be more like production-grade as much as possible:
1) Memory Degradation - recalled memories are not as good as in the beginning, causing AI Waifu to be more chaotic as she hallucinates over contaminated memories like a bad vicious cycle. So I transformed the original stateless sqlite-vec vector store to be a simple entity co-mention graph. And even make a studio to visualize the memories stored inside the vector db.
Just by looking at the graph, I saw a couple issues: a) After 1.5 months of interactions, there should be only one month of pinned memory (in green) over 1.5 months of active memory (in purple). How come pinned memory is in majority over active ones? I suppose the forgetting curve I had set too aggressive and memory half-life and shelf life too short, active memory got decayed way before monthly consolidation and got lost forever. b) I saw she memorized me into 3 different entities: my username, my nickname and my Github user ID (leaked into pinned memory, presumbly during nightly dreaming process). 3B small param LLM has hard time to correlation 3 different entities into single person, I may have to harden into one.
2) RAM burst during voice input - for some reason the tensor calculation of SileroVAD of the voice input uses PyTorch, and that's the only place in the whole codebase using torch after removing it from TTS synthesization. By switching to SileroVAD-onnx integrated in the ASR sherpa-onnx, the RAM usage drops at least 0.5GB (after shaving off ~1GB from TTS) by completely remove PyTorch dependencies.
3) Introduced a better Wake Word system using Livekit-Wake word instead of using ASR to do the wake word activation to save computation. Optional features like Speak Verification, Barge-in sensitivity, etc, need to find the optimum settings.