Submitted by Remco Hendriks 32 MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes Continker 0 4