arxiv:2602.06079
Liangyu Wang
ly4096
AI & ML interests
Efficient reinforcement learning (RL) for LLMs reasoning
Distributed training and inference of LLMs
Efficient algorithm and infrastructure design for LLMs
Recent Activity
upvoted a paper about 22 hours ago
Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight liked a model about 2 months ago
Qwen/Qwen3.8-2.4T-A95B-FP8 upvoted a paper 5 months ago
SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training