Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
AI & ML interests
Factuality, reasoning, alignment, LLM applications
Recent Activity
View all activity
Papers
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning
Gaming the Judge: Unfaithful Chain-of-Thought Can Undermine Agent Evaluation
spaces 7
Running
LudoBench
🎲
Multimodal Game Reasoning Benchmark [ICLR 2026]
Sleeping
Agents
Answer Convergence Early Stopping
🛑
Demo for EMNLP Paper "Answer Convergence as a Signal..."
Sleeping
FactRBench
🏆
View and analyze long-form factuality leaderboard
Running
3
ExpertLongBench
🚀
Leaderboard for ExpertLongBench
Sleeping
1
ManyICLBench
🚀
Leaderboard for ManyICLBench
Running
MLRC-BENCH
📊
Display model performance rankings
models 15
launch/MET-D-Gemma3-4B-en-only
Text Generation • 4B • Updated • 305
launch/MET-D-Gemma3-4B
Text Generation • 4B • Updated • 312
launch/MET-D-Qwen3-8B-en-only
Text Generation • 8B • Updated • 299
launch/MET-D-Qwen3-8B
Text Generation • 8B • Updated • 319
launch/MET-D-Qwen3-4B-zh-only
Text Generation • 4B • Updated • 311
launch/MET-D-Qwen3-4B-ms-only
Text Generation • 4B • Updated • 311
launch/MET-D-Qwen3-4B-ko-only
Text Generation • 4B • Updated • 315
launch/MET-D-Qwen3-4B-hi-only
Text Generation • 4B • Updated • 309
launch/MET-D-Qwen3-4B-es-only
Text Generation • 4B • Updated • 313
launch/MET-D-Qwen3-4B-en-only
Text Generation • 4B • Updated • 314
datasets 14
launch/MCLASH
Viewer • Updated • 2.61k • 376
launch/CLASH
Viewer • Updated • 345 • 115 • 3
launch/thinkprm-1K-verification-cots
Viewer • Updated • 1k • 93 • 8
launch/LudoBench
Viewer • Updated • 638 • 39
launch/ExpertLongBench
Preview • Updated • 194 • 10
launch/ManyICLBench
Viewer • Updated • 66 • 478 • 1
launch/CMV
Viewer • Updated • 133 • 19
launch/FactRBench
Viewer • Updated • 1.06k • 42 • 2
launch/FactBench
Viewer • Updated • 1k • 97 • 3
launch/gov_report
Viewer • Updated • 58.4k • 516 • 14