Benchmark data in "Beyond Ideal Instruction: A Comprehensive Framework for Evaluating LLMs in Realistic Interactions".
AI & ML interests
LLM reasoning
Recent Activity
View all activity
models 7
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch3
Text Generation • 9B • Updated • 297
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch2
Text Generation • 9B • Updated • 321
Miaow-Lab/Qwen3.5-9B-decision-subtrajectory-32k-epoch1
Text Generation • 9B • Updated • 325
Miaow-Lab/RLVR-Linearity-Checkpoints
Text Generation • Updated
Miaow-Lab/STT-Agent-RL
196k • Updated • 16 • 1
Miaow-Lab/STT-Agent-SFT
196k • Updated • 6 • 1
Miaow-Lab/SSAE-Checkpoints
Feature Extraction • Updated