Activity Feed

AI & ML interests

None defined yet.

Recent Activity

alielfilali01ย 
posted an update 2 days ago
view post
Post
161
multi-agents orchestration is old news by now. Try multi-teams of agents ! that is a whole different nightmare ... This is the real signal about the sparks of AGI.
appvoidย 
posted an update 2 days ago
view post
Post
140
Much of the misunderstanding on loops and byte-level models comes from fear on inventing the wheel, not reinventing it
appvoidย 
posted an update 3 days ago
view post
Post
135
How much should OpenAI offer for a name like Dot Labs?
  • 1 reply
ยท
appvoidย 
posted an update 7 days ago
view post
Post
165
Reality has not limits if you know how to shape it, one step at a time.
  • 1 reply
ยท
appvoidย 
posted an update 9 days ago
view post
Post
84
How old were you when you discovered good AI agents favor the Strategy Design Pattern over any other?
  • 1 reply
ยท
appvoidย 
posted an update 19 days ago
view post
Post
131
LLMS for edge devices?
What if we go the reliable route instead of the speed route?
What if we make it run on sbcs with few megabytes available?

That's the idea for the next model.

Keep in tune.
appvoidย 
posted an update 23 days ago
view post
Post
141
We trained a 10.9M byte-level recurrent Transformer on L3 and L6. (Loop 3 and Loop 6)

Yet L4/L5 improved too, L8 held up, and the L3โ†’L6 gain grew during training.

Same weights. More compute. Better predictions.

This is a new architecture for effective compute after several steps beyond original training!

We mixed and matched components like time and mhc into an ouro-like byte-level language model and the result is BET, a byte-level step-elastic transformer that can run computation steps without significant degradation.

One of the coolest parts of this training was discovering how Gradient Descent decided to use the first layer as what we would consider a scratchpad! Totally destroyed for the decoder but somehow makes total sense for the next layer!

I believe looped-transformers are the future of edge computing and this is a first step towards it.

Blogpost: https://medium.com/@appvoidofficial/byte-level-elasticity-182fe2ed1d2f

appvoid/bet-10m
anakin87ย 
posted an update 25 days ago
view post
Post
3478
I made a 1.1M ModernBERT encoder play Doom in real time on a CPU

Some time ago, VAGO Solutions released SauerkrautLM-Doom-MultiVec-1.3M, a tiny model trained to play Doom Defend the Center scenario from 31k human gameplay examples.

My first thought: cool! I love both Doom and Small Language Models.

Then another idea: I bet I can do better :-)

What I did?
- evaluated the original model and found it's better than reported
- changed a bit the architecture
- generated SFT data with a scripted oracle
- SFT + PPO refinement on consumer hardware

Got a smaller, faster and killer model
Can even fit a floppy with int8 quantization ๐Ÿ’พ

Watch it play/read the article: anakin87/tiny-doom-defender
appvoidย 
posted an update 28 days ago
view post
Post
3916
We got gpt6 before gta6
  • 15 replies
ยท
appvoidย 
posted an update 29 days ago
view post
Post
132
Any thoughts on Nvidia acquiring this website?

I don't know what to feel about it. But would be great if huggingface gets something similar to Kaggle with free GPU hours (or even days) for training.
  • 3 replies
ยท