复现技术报告benchmark求助🙏🙏🙏

#6
by derek2073 - opened

各位作者好,

我们是一个做 looped / parallel-loop Transformer 代码模型复现的小组。
你们关于 PLT loop-count selection 的 gain–cost 分析是我们读过的对这个现象
讲得最清楚的一篇,非单调的 loop 曲线也正是我们想接着往下做的地方。所以我们
拿发布的 two-loop checkpoint 试着把 Table 2 复现了一遍。

能复现的部分我们是真复现出来了。用自己的 harness 跑出 BigCodeBench-Full
45.00(论文 46.1)、LiveCodeBench 35.64(35.4)、Mind2Web K=5 33.40(34.5)、
BFCL v3 38.31(40.1),都在 2 分以内;HumanEval+ 是 80.5,对应你们的 84.1。
另外我们还用同一条流水线跑了 Qwen2.5-Coder-7B-Instruct 做对照,它的 MultiPL-E
公开成绩我们也基本跑出来了(C++ 73.9 对 75.4,Java 64.8,JavaScript 74.3),
所以问题应该不在我们的生成和打分链路上。

复现不出来的部分,差距很大,我们不想靠猜:

  • MultiPL-E:18 语言口径 29.6、全 24 语言 23.8,你们是 73.9。我们的
    pass@20 oracle 上界只有 51.6,还低于你们的 pass@1,说明我们大概率
    评的不是同一个东西。
  • SWE-bench Verified:500 例分母 14.0%,有补丁的 305 例分母 23.0%,
    你们是 64.4。
  • SWE-bench Multilingual:整体 5.67%,agentic lane 13.1%,你们是 31.0。
  • Terminal-Bench:v1 官方 80 题子集 + 官方 terminus-1 得 28.2%,v2 得
    6.74%,你们是 34.2 和 21.0。

我们完全接受这些差距多半出在我们这边:scaffold、步数上限、分母口径都是我们
自己选的,不是你们写的。正因为如此才来请教——论文正文和参考文献里我们没找到
benchmark 的引用、版本号或子集定义,很可能我们猜错了协议。如果能指点下面几点,
可以省掉我们大量盲试:

  1. MultiPL-E 的 "multilingual avg." 具体是哪几个语言,聚合方式是不是各语言
    pass@1 的未加权平均?
  2. SWE-bench Verified 用的哪种 agent scaffold、步数上限和 patch 编辑契约?
    是否方便分享 harness 或 prompt 模板?
  3. SWE-M 这一列,Table 2 脚注写 SWE-bench Multilingual、摘要和 README 写
    Multi-SWE、§4.1 又写 SWE-bench-CC——31.0 这个数来自哪一个?
    SWE-bench-CC 是不是第四个独立测量?
  4. TB-v1 / TB-v2 分别对应哪个官方发布、哪个任务子集、多少步上限?
  5. Table 2 的数字取自 instruction-tuned 还是 Table 4 的 thinking 变体?
    采样参数(temperature、top-p、max tokens)是怎么设的?

如果能分享评测 harness 或配置文件——哪怕只是 prompt、分母和语言列表——我们会
非常感激。作为回报我们也乐意提供逐实例日志、复现产物和一份简短的书面小结,
并在所有报告中正确引用你们的工作。我们不是质疑结果,只是希望在往下做之前,
先把你们的配置忠实复现出来。

感谢你们开源 checkpoint,也感谢那部分逐 loop 的机制分析——正是它让我们觉得
这件事值得认真复现。

顺祝研安

如果能提供具体测评的代码库甚至是镜像就更好了🫰

Multilingual-Multimodal-NLP org

Thank you for your interest in LoopCoder-V2 and for asking about reproducing the scores reported in our paper. Due to ongoing upgrades to our evaluation infrastructure, fully reproducing the evaluation environment used for the paper may be challenging. Scores from reruns may therefore vary slightly from those reported in the paper.

To help with reproduction, we have rerun the evaluations and are sharing the updated scores and evaluation settings below.

Benchmark Score (pass rate)
SWE-bench Verified 60.24%
SWE-bench Multilingual 39.06%
Terminal-Bench 2.0 20.45%

All three evaluations used Harbor 0.20.0 with an OpenAI-compatible API, 32 concurrent trials, one evaluation attempt per task, and a timeout multiplier of 8. The framework allowed up to three retries, subject to the exclusions listed in the configuration below. Retries were not counted as additional independent samples when calculating scores.

Both SWE-bench evaluations used R2E-Gym 0.1.0 with the mopenhands scaffold, temperature=1, and tool_calling=false. The R2E-Gym implementation used here includes improvements developed internally by our team. The configured context length was 131,072 tokens, with a maximum output of 65,536 tokens per request. The standard step limit was 500, the absolute step limit was 520, and the maximum total agent runtime per task was 7,200 seconds. Proactive context compaction was triggered at 80,000 prompt tokens.

Terminal-Bench 2.0 used Terminus-2 2.0.0 with an XML parser, a maximum of 500 turns, max_input_tokens=131072, max_output_tokens=65536, and a proactive summarization threshold of 104,858 tokens. The configuration did not explicitly specify a sampling temperature for Terminal-Bench 2.0.

The full evaluation parameters are provided below. This YAML describes the evaluation settings; its format should be adapted to the configuration schema of the evaluation framework being used.

model:
  name: LoopCoder-V2
  api_format: openai
harness:
  name: harbor
  version: 0.20.0
evaluation:
  n_concurrent_trials: 32
  n_attempts: 1
  timeout_multiplier: 8
  retry:
    max_retries: 3
    exclude_exceptions:
    - AgentAuthenticationError
    - AgentSafetyRefusalError
    - AgentTimeoutError
    - ApiUsageLimitError
    - ModelNotFoundError
    - RewardFileEmptyError
    - RewardFileNotFoundError
    - VerifierOutputParseError
    - VerifierTimeoutError
    wait_multiplier: 1.0
    min_wait_sec: 1.0
    max_wait_sec: 60.0
benchmarks:
  swe_bench_verified:
    benchmark: SWE-bench Verified
    agent:
      name: r2e-gym
      version: 0.1.0
      kwargs:
        context_error_keep_last: 80
        llm_query_retries: 3
        llm_timeout_sec: 2400
        max_exec_time_sec: 1800
        max_observation_chars: 36000
        max_steps: 500
        max_steps_absolute: 520
        max_tokens: 65536
        max_total_time_sec: 7200
        model_max_context_tokens: 131072
        proactive_context_prompt_tokens: 80000
        request_format: openai
        scaffold: mopenhands
        summary_head_messages: 2
        summary_tail_messages: 60
        temperature: 1
        tool_calling: false
  swe_bench_multilingual:
    benchmark: SWE-bench Multilingual
    agent:
      name: r2e-gym
      version: 0.1.0
      kwargs:
        context_error_keep_last: 80
        llm_query_retries: 3
        llm_timeout_sec: 2400
        max_exec_time_sec: 1800
        max_observation_chars: 36000
        max_steps: 500
        max_steps_absolute: 520
        max_tokens: 65536
        max_total_time_sec: 7200
        model_max_context_tokens: 131072
        proactive_context_prompt_tokens: 80000
        request_format: openai
        scaffold: mopenhands
        summary_head_messages: 2
        summary_tail_messages: 60
        temperature: 1
        tool_calling: false
  terminal_bench:
    benchmark: Terminal-Bench 2.0
    agent:
      name: terminus-2
      version: 2.0.0
      kwargs:
        max_turns: 500
        parser_name: xml
        proactive_summarization_threshold: 104858
        model_info:
          max_input_tokens: 131072
          max_output_tokens: 65536

Sign up or log in to comment