multi-agents orchestration is old news by now. Try multi-teams of agents ! that is a whole different nightmare ... This is the real signal about the sparks of AGI.
We trained a 10.9M byte-level recurrent Transformer on L3 and L6. (Loop 3 and Loop 6)
Yet L4/L5 improved too, L8 held up, and the L3โL6 gain grew during training.
Same weights. More compute. Better predictions.
This is a new architecture for effective compute after several steps beyond original training!
We mixed and matched components like time and mhc into an ouro-like byte-level language model and the result is BET, a byte-level step-elastic transformer that can run computation steps without significant degradation.
One of the coolest parts of this training was discovering how Gradient Descent decided to use the first layer as what we would consider a scratchpad! Totally destroyed for the decoder but somehow makes total sense for the next layer!
I believe looped-transformers are the future of edge computing and this is a first step towards it.
I made a 1.1M ModernBERT encoder play Doom in real time on a CPU Some time ago, VAGO Solutions released SauerkrautLM-Doom-MultiVec-1.3M, a tiny model trained to play Doom Defend the Center scenario from 31k human gameplay examples.
My first thought: cool! I love both Doom and Small Language Models.
Then another idea: I bet I can do better :-)
What I did? - evaluated the original model and found it's better than reported - changed a bit the architecture - generated SFT data with a scripted oracle - SFT + PPO refinement on consumer hardware
Got a smaller, faster and killer model Can even fit a floppy with int8 quantization ๐พ
I don't know what to feel about it. But would be great if huggingface gets something similar to Kaggle with free GPU hours (or even days) for training.