Lots of attentions are now on GPU residual value. I find car analogy to be useful.
* Both new and used cars can do the same job (you can run the latest model on 6-year-old A100) * New cars are more efficient (higher performance per power draw) * It's not the year, but mileage and maintenance
The last point is undeveloped. We need Carfax and KBB for GPUs.
Looking forward to the Qwen 3.8 model drop, and congratulations on joining the trillion parameters club.
But my eyes are on the promised 27B model. Small (<50B) models decline on OpenRouter, because they are being run on local devices. If people see the family trees here on HF they will understand.
DeepSeek plans to raise token prices. I don't think this is because they are bleeding, but they are overwhelmed. If your price is 1/10 of your affiliate vendors, you can't leverage their resources. Markup is the only way to diverge traffic away.
Sadly I haven't found discussions on differentiators enabling DS to balance cost at such low prices. All software solutions (that we know of) are accessible by other vendors. If you attribute it to electricity or hardware, you can't explain why GLM and Kimi charge so much for their APIs.
This is where our attention should be (but distracted by things above).
Many developers discovered that the native DeepSeek API has higher cache-hit rate than neocloud APIs hosting the same DS models.
My speculation is that DS aggressively kills its old models. It has released 18 models thus far, and only 2 are being served now (v4 pro and flash).
This is tough to customers who don't want to upgrade (migrate or leave), but effectively boost the serving capacity to the same model, i.e. more woods behind fewer arrows.
There has been a leaked memo (now struck down) from the founder of DeepSeek. I'm not here to circulate it, but comment on the minimum-effort evolutionary path he proposed.
This makes sense to me: even at the agent stage I learn world models much faster than when I learned LLM at the LLM stage.
But this means humans are still needed beyond the digital singularity, until robots can close their own loop: eval, manufacturing, self improvement, i.e. physical singularity.
The Nvidia paper came down to this: remove synchronization barriers. DeepSeek has already done that with DeepEP (which this paper cited) one layer above.
* Tacit knowledge won't be captured in words, much less accessible dataset * RL is not efficient at all, and labeling gets more and more expensive * AGI won't cover your specialized tasks
I am starting a new series on matrix. The idea came to me when I wrote about the Muon optimizer.
Matrix itself has lots of fascinating properties and is applied in STEM fields for many decades. Its application in ML is just the beginning. There are lots of low hanging fruits. At the very least, I hope this math perspective will give you a new lens.