Skip to content

[Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 #46

Description

@tt-a1i

摘要

本 Issue 记录 OpenPI 0.3.0、Bare Pi 与 OMP 在两个模型上的同条件对照证据:

  • 模型:seal/glm-5.3seal/gpt-5.6-luna
  • Harness:Bare Pi、Pi + OpenPI、OMP
  • 任务:3 个多轮 Terminal 任务
  • 重复:每题 3 次
  • 总计:2 models × 3 tasks × 3 repeats × 3 harnesses = 54 cells

主要观察:

  • GLM-5.3:三条 Harness 都通过 7/9;OpenPI 使用 331,851 tokens,OMP 使用 2,176,017 tokens,OpenPI 少 84.7%。
  • GPT-5.6 Luna:Bare Pi 与 OpenPI 都通过 7/9,OMP 通过 4/9;OpenPI 比 OMP 少 58.2% 的物理模型请求,总耗时低 36.2%。
  • Git 与 SQLite 两类任务中,Bare Pi 和 OpenPI 在两个模型上均为 12/12。

这组结果没有证明 OpenPI 相对 Bare Pi 存在通用质量提升。它支持的更窄结论是:在本次普通 Terminal 任务里,OpenPI 保持了 Pi 的成功数;相对 OMP,本轮观察到了更低的模型调用开销。

结果边界

这是在 Linux ARM64 VM 上运行的 Terminal-Bench 2.1 源构建派生诊断,不是 Terminal-Bench 官方成绩或排行榜提交

任务数只有 3,每个 cell 重复 3 次。cancel-async-tasks 对 SIGINT/process-group 行为较敏感,方差明显,因此不能把小幅胜负外推成 Harness 的稳定总体能力。

GPT-5.6 Luna 路线没有返回可用 token usage;该模型只比较 verifier、物理请求数与 wall time,不作 token 或成本结论。

冻结条件

项目
Pi@earendil-works/pi-coding-agent 0.84.2
OpenPI@tt-a1i/openpi 0.3.0
OMP@oh-my-pi/pi-coding-agent 17.2.12
Bun1.3.14
Harbor0.20.0
Container runtimePodman server 6.0.2,Linux ARM64 VM
Task sourceTerminal-Bench 2.1 ARM64 source-build-derived,commit d1f1920f2d817a831f466d0ff363ef795a9a3b00
Tasksgit-leak-recovery, sqlite-db-truncate, cancel-async-tasks
Repeats每个模型、任务、Harness 3 次
Schedulestrict serial Latin-square
Cell deadline1,800 秒
Schedule SHA-256974fdc67aad36c4c890a83d6705e7a687610a0ecbef00a0c0ca4e7efac369a62
Credential boundaryhost forwarding proxy + 每格短命 bearer;candidate 不持有真实 provider key

模型参数:

模型Thinking
seal/glm-5.3high
seal/gpt-5.6-lunahigh

GLM-5.3

结果格式为 pass / fail / indeterminate。Wall 为 9 个 cell 的总耗时;indeterminate 的等待时间不隐藏。

Harness结果WallProvider tokensPhysical POST
Bare Pi7 / 1 / 12,647.724s249,62594
OpenPI7 / 2 / 01,107.098s331,85187
OMP7 / 2 / 01,081.746s2,176,017106

逐任务:

TaskBare PiOpenPIOMP
Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
SQLite truncate3 / 0 / 03 / 0 / 03 / 0 / 0
Cancel async tasks1 / 1 / 11 / 2 / 01 / 2 / 0

效率观察:

  • OpenPI 与 OMP 都是 7 pass;OpenPI tokens 为 OMP 的 15.3%,少 84.7%。
  • OMP tokens 为 OpenPI 的 6.56 倍。
  • OpenPI 相对 Bare Pi 多使用 32.9% tokens;本轮不能声称 OpenPI 比原生 Pi 更省 token。
  • Bare Pi 的 cancel 总时长包含一个约 30 分钟的 indeterminate,因此不能用 GLM 总 wall 宣称 OpenPI 相对 Pi 有稳定提速。

GPT-5.6 Luna

Harness结果WallPhysical POSTLogical attempts额外物理 POST
Bare Pi7 / 2 / 01,129.885s43430
OpenPI7 / 2 / 01,151.213s46460
OMP4 / 5 / 01,803.368s1109218

逐任务:

TaskBare PiOpenPIOMP
Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
SQLite truncate3 / 0 / 03 / 0 / 01 / 2 / 0
Cancel async tasks1 / 2 / 01 / 2 / 00 / 3 / 0

效率观察:

  • OpenPI 与 Bare Pi 都是 7/9;OpenPI wall 比 Bare Pi 高 1.9%,没有提速证据。
  • OpenPI 相对 OMP 多 3 pass,物理模型请求少 58.2%,总 wall 低 36.2%。
  • OMP 的 110 个物理 POST 对应 92 个已完成 logical attempts;18 个额外请求按 provider 重试保留,没有从开销中隐藏。
  • Seal 未返回可用 usage,因此不能根据该轮声称 OpenPI 节省 Luna token 或成本。

完整性与安全收据

两组运行均满足:

  • 27/27 Pi/OpenPI/OMP cells 落盘且身份唯一;
  • 最终冻结源校验通过;
  • retained artifact credential scan 通过:0 missing roots、0 credential leak、0 scan failure;
  • 真实 provider key 未注入 candidate;
  • 全局并发为 1,所有 cell 按冻结 Latin-square 严格串行执行。

运行身份:

模型Lock fingerprintReceipt
GLM-5.3sha256:037179456f3223457e428b9d5543911144b045bcea7ee7127728b27dd735780ecompleted_with_indeterminate
GPT-5.6 Lunasha256:00b9db64606832521df7f853d63d5e50d1e6754920e0ac51e1b32c2aacc4866ccompleted_with_indeterminate

证据 controller 的 POST/attempt reconciliation 另外覆盖了 provider 物理重试:physical POST < logical attempts 时 fail closed;非负差值记录为 unattributedProviderPostRequests

验证:

  • controller 定向 Node tests:23 pass,1 个环境用例 skip;
  • 项目全量 Node tests:719/719;
  • Vitest:29/29;
  • bun run check:format、lint、typecheck 通过;typecheck 仅有既存 Effect advisory warnings。

当前可支持的结论

  1. 在这两个完整模型批次中,OpenPI 与 Bare Pi 都取得 7/9;Git 与 SQLite 合计均为 12/12。
  2. 在本轮两个完整模型批次中,OpenPI 保持了与原生 Pi 相同的通过数;不同任务的耗时方向不一致,因此暂不对相对 Pi 的速度优势作结论。
  3. 相对 OMP,OpenPI 在 GLM-5.3 上以相同 pass 数使用少 84.7% tokens;在 Luna 上取得更多 pass,并使用少 58.2% 的物理请求和少 36.2% 的 wall time。
  4. 这些数据支持继续保持 Pi-native、普通回合 zero-resident、按需披露能力的方向;尚不足以证明 Subagent/Workflow 的收益,因为这三道题没有专门要求编排能力。
  5. 下一轮应扩大任务集,并预注册会自然触发 Subagent/Workflow 的多文件任务,将能力采用率、质量与成本分开报告。

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
       blocks
      (function() {
      function addCopyButtons() {
      document.querySelectorAll('pre code').forEach(function(codeBlock) {
      if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
      codeBlock.parentElement.setAttribute('data-copy-added', 'true');
      var btn = document.createElement('button');
      btn.textContent = 'Copy';
      btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
      btn.onmouseover = function() { this.style.opacity = '1'; };
      btn.onmouseout = function() { this.style.opacity = '0.7'; };
      btn.onclick = function() {
      navigator.clipboard.writeText(codeBlock.textContent).then(function() {
      btn.textContent = 'Copied!';
      setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
      });
      };
      codeBlock.parentElement.style.position = 'relative';
      codeBlock.parentElement.appendChild(btn);
      });
      }
      addCopyButtons();
      // Re-run on dynamic content
      var observer = new MutationObserver(addCopyButtons);
      observer.observe(document.body, { childList: true, subtree: true });
      })();
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 · Issue #46 · openpi-dev/openpi · GitHub
      Skip to content

      [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 #46

      Description

      @tt-a1i

      摘要

      本 Issue 记录 OpenPI 0.3.0、Bare Pi 与 OMP 在两个模型上的同条件对照证据:

      • 模型:seal/glm-5.3seal/gpt-5.6-luna
      • Harness:Bare Pi、Pi + OpenPI、OMP
      • 任务:3 个多轮 Terminal 任务
      • 重复:每题 3 次
      • 总计:2 models × 3 tasks × 3 repeats × 3 harnesses = 54 cells

      主要观察:

      • GLM-5.3:三条 Harness 都通过 7/9;OpenPI 使用 331,851 tokens,OMP 使用 2,176,017 tokens,OpenPI 少 84.7%。
      • GPT-5.6 Luna:Bare Pi 与 OpenPI 都通过 7/9,OMP 通过 4/9;OpenPI 比 OMP 少 58.2% 的物理模型请求,总耗时低 36.2%。
      • Git 与 SQLite 两类任务中,Bare Pi 和 OpenPI 在两个模型上均为 12/12。

      这组结果没有证明 OpenPI 相对 Bare Pi 存在通用质量提升。它支持的更窄结论是:在本次普通 Terminal 任务里,OpenPI 保持了 Pi 的成功数;相对 OMP,本轮观察到了更低的模型调用开销。

      结果边界

      这是在 Linux ARM64 VM 上运行的 Terminal-Bench 2.1 源构建派生诊断,不是 Terminal-Bench 官方成绩或排行榜提交

      任务数只有 3,每个 cell 重复 3 次。cancel-async-tasks 对 SIGINT/process-group 行为较敏感,方差明显,因此不能把小幅胜负外推成 Harness 的稳定总体能力。

      GPT-5.6 Luna 路线没有返回可用 token usage;该模型只比较 verifier、物理请求数与 wall time,不作 token 或成本结论。

      冻结条件

      项目
      Pi@earendil-works/pi-coding-agent 0.84.2
      OpenPI@tt-a1i/openpi 0.3.0
      OMP@oh-my-pi/pi-coding-agent 17.2.12
      Bun1.3.14
      Harbor0.20.0
      Container runtimePodman server 6.0.2,Linux ARM64 VM
      Task sourceTerminal-Bench 2.1 ARM64 source-build-derived,commit d1f1920f2d817a831f466d0ff363ef795a9a3b00
      Tasksgit-leak-recovery, sqlite-db-truncate, cancel-async-tasks
      Repeats每个模型、任务、Harness 3 次
      Schedulestrict serial Latin-square
      Cell deadline1,800 秒
      Schedule SHA-256974fdc67aad36c4c890a83d6705e7a687610a0ecbef00a0c0ca4e7efac369a62
      Credential boundaryhost forwarding proxy + 每格短命 bearer;candidate 不持有真实 provider key

      模型参数:

      模型Thinking
      seal/glm-5.3high
      seal/gpt-5.6-lunahigh

      GLM-5.3

      结果格式为 pass / fail / indeterminate。Wall 为 9 个 cell 的总耗时;indeterminate 的等待时间不隐藏。

      Harness结果WallProvider tokensPhysical POST
      Bare Pi7 / 1 / 12,647.724s249,62594
      OpenPI7 / 2 / 01,107.098s331,85187
      OMP7 / 2 / 01,081.746s2,176,017106

      逐任务:

      TaskBare PiOpenPIOMP
      Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
      SQLite truncate3 / 0 / 03 / 0 / 03 / 0 / 0
      Cancel async tasks1 / 1 / 11 / 2 / 01 / 2 / 0

      效率观察:

      • OpenPI 与 OMP 都是 7 pass;OpenPI tokens 为 OMP 的 15.3%,少 84.7%。
      • OMP tokens 为 OpenPI 的 6.56 倍。
      • OpenPI 相对 Bare Pi 多使用 32.9% tokens;本轮不能声称 OpenPI 比原生 Pi 更省 token。
      • Bare Pi 的 cancel 总时长包含一个约 30 分钟的 indeterminate,因此不能用 GLM 总 wall 宣称 OpenPI 相对 Pi 有稳定提速。

      GPT-5.6 Luna

      Harness结果WallPhysical POSTLogical attempts额外物理 POST
      Bare Pi7 / 2 / 01,129.885s43430
      OpenPI7 / 2 / 01,151.213s46460
      OMP4 / 5 / 01,803.368s1109218

      逐任务:

      TaskBare PiOpenPIOMP
      Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
      SQLite truncate3 / 0 / 03 / 0 / 01 / 2 / 0
      Cancel async tasks1 / 2 / 01 / 2 / 00 / 3 / 0

      效率观察:

      • OpenPI 与 Bare Pi 都是 7/9;OpenPI wall 比 Bare Pi 高 1.9%,没有提速证据。
      • OpenPI 相对 OMP 多 3 pass,物理模型请求少 58.2%,总 wall 低 36.2%。
      • OMP 的 110 个物理 POST 对应 92 个已完成 logical attempts;18 个额外请求按 provider 重试保留,没有从开销中隐藏。
      • Seal 未返回可用 usage,因此不能根据该轮声称 OpenPI 节省 Luna token 或成本。

      完整性与安全收据

      两组运行均满足:

      • 27/27 Pi/OpenPI/OMP cells 落盘且身份唯一;
      • 最终冻结源校验通过;
      • retained artifact credential scan 通过:0 missing roots、0 credential leak、0 scan failure;
      • 真实 provider key 未注入 candidate;
      • 全局并发为 1,所有 cell 按冻结 Latin-square 严格串行执行。

      运行身份:

      模型Lock fingerprintReceipt
      GLM-5.3sha256:037179456f3223457e428b9d5543911144b045bcea7ee7127728b27dd735780ecompleted_with_indeterminate
      GPT-5.6 Lunasha256:00b9db64606832521df7f853d63d5e50d1e6754920e0ac51e1b32c2aacc4866ccompleted_with_indeterminate

      证据 controller 的 POST/attempt reconciliation 另外覆盖了 provider 物理重试:physical POST < logical attempts 时 fail closed;非负差值记录为 unattributedProviderPostRequests

      验证:

      • controller 定向 Node tests:23 pass,1 个环境用例 skip;
      • 项目全量 Node tests:719/719;
      • Vitest:29/29;
      • bun run check:format、lint、typecheck 通过;typecheck 仅有既存 Effect advisory warnings。

      当前可支持的结论

      1. 在这两个完整模型批次中,OpenPI 与 Bare Pi 都取得 7/9;Git 与 SQLite 合计均为 12/12。
      2. 在本轮两个完整模型批次中,OpenPI 保持了与原生 Pi 相同的通过数;不同任务的耗时方向不一致,因此暂不对相对 Pi 的速度优势作结论。
      3. 相对 OMP,OpenPI 在 GLM-5.3 上以相同 pass 数使用少 84.7% tokens;在 Luna 上取得更多 pass,并使用少 58.2% 的物理请求和少 36.2% 的 wall time。
      4. 这些数据支持继续保持 Pi-native、普通回合 zero-resident、按需披露能力的方向;尚不足以证明 Subagent/Workflow 的收益,因为这三道题没有专门要求编排能力。
      5. 下一轮应扩大任务集,并预注册会自然触发 Subagent/Workflow 的多文件任务,将能力采用率、质量与成本分开报告。

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        documentationImprovements or additions to documentation

        Type

        No type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 · Issue #46 · openpi-dev/openpi · GitHub
          Skip to content

          [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 #46

          Description

          @tt-a1i

          摘要

          本 Issue 记录 OpenPI 0.3.0、Bare Pi 与 OMP 在两个模型上的同条件对照证据:

          • 模型:seal/glm-5.3seal/gpt-5.6-luna
          • Harness:Bare Pi、Pi + OpenPI、OMP
          • 任务:3 个多轮 Terminal 任务
          • 重复:每题 3 次
          • 总计:2 models × 3 tasks × 3 repeats × 3 harnesses = 54 cells

          主要观察:

          • GLM-5.3:三条 Harness 都通过 7/9;OpenPI 使用 331,851 tokens,OMP 使用 2,176,017 tokens,OpenPI 少 84.7%。
          • GPT-5.6 Luna:Bare Pi 与 OpenPI 都通过 7/9,OMP 通过 4/9;OpenPI 比 OMP 少 58.2% 的物理模型请求,总耗时低 36.2%。
          • Git 与 SQLite 两类任务中,Bare Pi 和 OpenPI 在两个模型上均为 12/12。

          这组结果没有证明 OpenPI 相对 Bare Pi 存在通用质量提升。它支持的更窄结论是:在本次普通 Terminal 任务里,OpenPI 保持了 Pi 的成功数;相对 OMP,本轮观察到了更低的模型调用开销。

          结果边界

          这是在 Linux ARM64 VM 上运行的 Terminal-Bench 2.1 源构建派生诊断,不是 Terminal-Bench 官方成绩或排行榜提交

          任务数只有 3,每个 cell 重复 3 次。cancel-async-tasks 对 SIGINT/process-group 行为较敏感,方差明显,因此不能把小幅胜负外推成 Harness 的稳定总体能力。

          GPT-5.6 Luna 路线没有返回可用 token usage;该模型只比较 verifier、物理请求数与 wall time,不作 token 或成本结论。

          冻结条件

          项目
          Pi@earendil-works/pi-coding-agent 0.84.2
          OpenPI@tt-a1i/openpi 0.3.0
          OMP@oh-my-pi/pi-coding-agent 17.2.12
          Bun1.3.14
          Harbor0.20.0
          Container runtimePodman server 6.0.2,Linux ARM64 VM
          Task sourceTerminal-Bench 2.1 ARM64 source-build-derived,commit d1f1920f2d817a831f466d0ff363ef795a9a3b00
          Tasksgit-leak-recovery, sqlite-db-truncate, cancel-async-tasks
          Repeats每个模型、任务、Harness 3 次
          Schedulestrict serial Latin-square
          Cell deadline1,800 秒
          Schedule SHA-256974fdc67aad36c4c890a83d6705e7a687610a0ecbef00a0c0ca4e7efac369a62
          Credential boundaryhost forwarding proxy + 每格短命 bearer;candidate 不持有真实 provider key

          模型参数:

          模型Thinking
          seal/glm-5.3high
          seal/gpt-5.6-lunahigh

          GLM-5.3

          结果格式为 pass / fail / indeterminate。Wall 为 9 个 cell 的总耗时;indeterminate 的等待时间不隐藏。

          Harness结果WallProvider tokensPhysical POST
          Bare Pi7 / 1 / 12,647.724s249,62594
          OpenPI7 / 2 / 01,107.098s331,85187
          OMP7 / 2 / 01,081.746s2,176,017106

          逐任务:

          TaskBare PiOpenPIOMP
          Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
          SQLite truncate3 / 0 / 03 / 0 / 03 / 0 / 0
          Cancel async tasks1 / 1 / 11 / 2 / 01 / 2 / 0

          效率观察:

          • OpenPI 与 OMP 都是 7 pass;OpenPI tokens 为 OMP 的 15.3%,少 84.7%。
          • OMP tokens 为 OpenPI 的 6.56 倍。
          • OpenPI 相对 Bare Pi 多使用 32.9% tokens;本轮不能声称 OpenPI 比原生 Pi 更省 token。
          • Bare Pi 的 cancel 总时长包含一个约 30 分钟的 indeterminate,因此不能用 GLM 总 wall 宣称 OpenPI 相对 Pi 有稳定提速。

          GPT-5.6 Luna

          Harness结果WallPhysical POSTLogical attempts额外物理 POST
          Bare Pi7 / 2 / 01,129.885s43430
          OpenPI7 / 2 / 01,151.213s46460
          OMP4 / 5 / 01,803.368s1109218

          逐任务:

          TaskBare PiOpenPIOMP
          Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
          SQLite truncate3 / 0 / 03 / 0 / 01 / 2 / 0
          Cancel async tasks1 / 2 / 01 / 2 / 00 / 3 / 0

          效率观察:

          • OpenPI 与 Bare Pi 都是 7/9;OpenPI wall 比 Bare Pi 高 1.9%,没有提速证据。
          • OpenPI 相对 OMP 多 3 pass,物理模型请求少 58.2%,总 wall 低 36.2%。
          • OMP 的 110 个物理 POST 对应 92 个已完成 logical attempts;18 个额外请求按 provider 重试保留,没有从开销中隐藏。
          • Seal 未返回可用 usage,因此不能根据该轮声称 OpenPI 节省 Luna token 或成本。

          完整性与安全收据

          两组运行均满足:

          • 27/27 Pi/OpenPI/OMP cells 落盘且身份唯一;
          • 最终冻结源校验通过;
          • retained artifact credential scan 通过:0 missing roots、0 credential leak、0 scan failure;
          • 真实 provider key 未注入 candidate;
          • 全局并发为 1,所有 cell 按冻结 Latin-square 严格串行执行。

          运行身份:

          模型Lock fingerprintReceipt
          GLM-5.3sha256:037179456f3223457e428b9d5543911144b045bcea7ee7127728b27dd735780ecompleted_with_indeterminate
          GPT-5.6 Lunasha256:00b9db64606832521df7f853d63d5e50d1e6754920e0ac51e1b32c2aacc4866ccompleted_with_indeterminate

          证据 controller 的 POST/attempt reconciliation 另外覆盖了 provider 物理重试:physical POST < logical attempts 时 fail closed;非负差值记录为 unattributedProviderPostRequests

          验证:

          • controller 定向 Node tests:23 pass,1 个环境用例 skip;
          • 项目全量 Node tests:719/719;
          • Vitest:29/29;
          • bun run check:format、lint、typecheck 通过;typecheck 仅有既存 Effect advisory warnings。

          当前可支持的结论

          1. 在这两个完整模型批次中,OpenPI 与 Bare Pi 都取得 7/9;Git 与 SQLite 合计均为 12/12。
          2. 在本轮两个完整模型批次中,OpenPI 保持了与原生 Pi 相同的通过数;不同任务的耗时方向不一致,因此暂不对相对 Pi 的速度优势作结论。
          3. 相对 OMP,OpenPI 在 GLM-5.3 上以相同 pass 数使用少 84.7% tokens;在 Luna 上取得更多 pass,并使用少 58.2% 的物理请求和少 36.2% 的 wall time。
          4. 这些数据支持继续保持 Pi-native、普通回合 zero-resident、按需披露能力的方向;尚不足以证明 Subagent/Workflow 的收益,因为这三道题没有专门要求编排能力。
          5. 下一轮应扩大任务集,并预注册会自然触发 Subagent/Workflow 的多文件任务,将能力采用率、质量与成本分开报告。

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            documentationImprovements or additions to documentation

            Type

            No type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 · Issue #46 · openpi-dev/openpi · GitHub
              Skip to content

              [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 #46

              Description

              @tt-a1i

              摘要

              本 Issue 记录 OpenPI 0.3.0、Bare Pi 与 OMP 在两个模型上的同条件对照证据:

              • 模型:seal/glm-5.3seal/gpt-5.6-luna
              • Harness:Bare Pi、Pi + OpenPI、OMP
              • 任务:3 个多轮 Terminal 任务
              • 重复:每题 3 次
              • 总计:2 models × 3 tasks × 3 repeats × 3 harnesses = 54 cells

              主要观察:

              • GLM-5.3:三条 Harness 都通过 7/9;OpenPI 使用 331,851 tokens,OMP 使用 2,176,017 tokens,OpenPI 少 84.7%。
              • GPT-5.6 Luna:Bare Pi 与 OpenPI 都通过 7/9,OMP 通过 4/9;OpenPI 比 OMP 少 58.2% 的物理模型请求,总耗时低 36.2%。
              • Git 与 SQLite 两类任务中,Bare Pi 和 OpenPI 在两个模型上均为 12/12。

              这组结果没有证明 OpenPI 相对 Bare Pi 存在通用质量提升。它支持的更窄结论是:在本次普通 Terminal 任务里,OpenPI 保持了 Pi 的成功数;相对 OMP,本轮观察到了更低的模型调用开销。

              结果边界

              这是在 Linux ARM64 VM 上运行的 Terminal-Bench 2.1 源构建派生诊断,不是 Terminal-Bench 官方成绩或排行榜提交

              任务数只有 3,每个 cell 重复 3 次。cancel-async-tasks 对 SIGINT/process-group 行为较敏感,方差明显,因此不能把小幅胜负外推成 Harness 的稳定总体能力。

              GPT-5.6 Luna 路线没有返回可用 token usage;该模型只比较 verifier、物理请求数与 wall time,不作 token 或成本结论。

              冻结条件

              项目
              Pi@earendil-works/pi-coding-agent 0.84.2
              OpenPI@tt-a1i/openpi 0.3.0
              OMP@oh-my-pi/pi-coding-agent 17.2.12
              Bun1.3.14
              Harbor0.20.0
              Container runtimePodman server 6.0.2,Linux ARM64 VM
              Task sourceTerminal-Bench 2.1 ARM64 source-build-derived,commit d1f1920f2d817a831f466d0ff363ef795a9a3b00
              Tasksgit-leak-recovery, sqlite-db-truncate, cancel-async-tasks
              Repeats每个模型、任务、Harness 3 次
              Schedulestrict serial Latin-square
              Cell deadline1,800 秒
              Schedule SHA-256974fdc67aad36c4c890a83d6705e7a687610a0ecbef00a0c0ca4e7efac369a62
              Credential boundaryhost forwarding proxy + 每格短命 bearer;candidate 不持有真实 provider key

              模型参数:

              模型Thinking
              seal/glm-5.3high
              seal/gpt-5.6-lunahigh

              GLM-5.3

              结果格式为 pass / fail / indeterminate。Wall 为 9 个 cell 的总耗时;indeterminate 的等待时间不隐藏。

              Harness结果WallProvider tokensPhysical POST
              Bare Pi7 / 1 / 12,647.724s249,62594
              OpenPI7 / 2 / 01,107.098s331,85187
              OMP7 / 2 / 01,081.746s2,176,017106

              逐任务:

              TaskBare PiOpenPIOMP
              Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
              SQLite truncate3 / 0 / 03 / 0 / 03 / 0 / 0
              Cancel async tasks1 / 1 / 11 / 2 / 01 / 2 / 0

              效率观察:

              • OpenPI 与 OMP 都是 7 pass;OpenPI tokens 为 OMP 的 15.3%,少 84.7%。
              • OMP tokens 为 OpenPI 的 6.56 倍。
              • OpenPI 相对 Bare Pi 多使用 32.9% tokens;本轮不能声称 OpenPI 比原生 Pi 更省 token。
              • Bare Pi 的 cancel 总时长包含一个约 30 分钟的 indeterminate,因此不能用 GLM 总 wall 宣称 OpenPI 相对 Pi 有稳定提速。

              GPT-5.6 Luna

              Harness结果WallPhysical POSTLogical attempts额外物理 POST
              Bare Pi7 / 2 / 01,129.885s43430
              OpenPI7 / 2 / 01,151.213s46460
              OMP4 / 5 / 01,803.368s1109218

              逐任务:

              TaskBare PiOpenPIOMP
              Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
              SQLite truncate3 / 0 / 03 / 0 / 01 / 2 / 0
              Cancel async tasks1 / 2 / 01 / 2 / 00 / 3 / 0

              效率观察:

              • OpenPI 与 Bare Pi 都是 7/9;OpenPI wall 比 Bare Pi 高 1.9%,没有提速证据。
              • OpenPI 相对 OMP 多 3 pass,物理模型请求少 58.2%,总 wall 低 36.2%。
              • OMP 的 110 个物理 POST 对应 92 个已完成 logical attempts;18 个额外请求按 provider 重试保留,没有从开销中隐藏。
              • Seal 未返回可用 usage,因此不能根据该轮声称 OpenPI 节省 Luna token 或成本。

              完整性与安全收据

              两组运行均满足:

              • 27/27 Pi/OpenPI/OMP cells 落盘且身份唯一;
              • 最终冻结源校验通过;
              • retained artifact credential scan 通过:0 missing roots、0 credential leak、0 scan failure;
              • 真实 provider key 未注入 candidate;
              • 全局并发为 1,所有 cell 按冻结 Latin-square 严格串行执行。

              运行身份:

              模型Lock fingerprintReceipt
              GLM-5.3sha256:037179456f3223457e428b9d5543911144b045bcea7ee7127728b27dd735780ecompleted_with_indeterminate
              GPT-5.6 Lunasha256:00b9db64606832521df7f853d63d5e50d1e6754920e0ac51e1b32c2aacc4866ccompleted_with_indeterminate

              证据 controller 的 POST/attempt reconciliation 另外覆盖了 provider 物理重试:physical POST < logical attempts 时 fail closed;非负差值记录为 unattributedProviderPostRequests

              验证:

              • controller 定向 Node tests:23 pass,1 个环境用例 skip;
              • 项目全量 Node tests:719/719;
              • Vitest:29/29;
              • bun run check:format、lint、typecheck 通过;typecheck 仅有既存 Effect advisory warnings。

              当前可支持的结论

              1. 在这两个完整模型批次中,OpenPI 与 Bare Pi 都取得 7/9;Git 与 SQLite 合计均为 12/12。
              2. 在本轮两个完整模型批次中,OpenPI 保持了与原生 Pi 相同的通过数;不同任务的耗时方向不一致,因此暂不对相对 Pi 的速度优势作结论。
              3. 相对 OMP,OpenPI 在 GLM-5.3 上以相同 pass 数使用少 84.7% tokens;在 Luna 上取得更多 pass,并使用少 58.2% 的物理请求和少 36.2% 的 wall time。
              4. 这些数据支持继续保持 Pi-native、普通回合 zero-resident、按需披露能力的方向;尚不足以证明 Subagent/Workflow 的收益,因为这三道题没有专门要求编排能力。
              5. 下一轮应扩大任务集,并预注册会自然触发 Subagent/Workflow 的多文件任务,将能力采用率、质量与成本分开报告。

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                documentationImprovements or additions to documentation

                Type

                No type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + ' [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 · Issue #46 · openpi-dev/openpi · GitHub
                  Skip to content

                  [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 #46

                  Description

                  @tt-a1i

                  摘要

                  本 Issue 记录 OpenPI 0.3.0、Bare Pi 与 OMP 在两个模型上的同条件对照证据:

                  • 模型:seal/glm-5.3seal/gpt-5.6-luna
                  • Harness:Bare Pi、Pi + OpenPI、OMP
                  • 任务:3 个多轮 Terminal 任务
                  • 重复:每题 3 次
                  • 总计:2 models × 3 tasks × 3 repeats × 3 harnesses = 54 cells

                  主要观察:

                  • GLM-5.3:三条 Harness 都通过 7/9;OpenPI 使用 331,851 tokens,OMP 使用 2,176,017 tokens,OpenPI 少 84.7%。
                  • GPT-5.6 Luna:Bare Pi 与 OpenPI 都通过 7/9,OMP 通过 4/9;OpenPI 比 OMP 少 58.2% 的物理模型请求,总耗时低 36.2%。
                  • Git 与 SQLite 两类任务中,Bare Pi 和 OpenPI 在两个模型上均为 12/12。

                  这组结果没有证明 OpenPI 相对 Bare Pi 存在通用质量提升。它支持的更窄结论是:在本次普通 Terminal 任务里,OpenPI 保持了 Pi 的成功数;相对 OMP,本轮观察到了更低的模型调用开销。

                  结果边界

                  这是在 Linux ARM64 VM 上运行的 Terminal-Bench 2.1 源构建派生诊断,不是 Terminal-Bench 官方成绩或排行榜提交

                  任务数只有 3,每个 cell 重复 3 次。cancel-async-tasks 对 SIGINT/process-group 行为较敏感,方差明显,因此不能把小幅胜负外推成 Harness 的稳定总体能力。

                  GPT-5.6 Luna 路线没有返回可用 token usage;该模型只比较 verifier、物理请求数与 wall time,不作 token 或成本结论。

                  冻结条件

                  项目
                  Pi@earendil-works/pi-coding-agent 0.84.2
                  OpenPI@tt-a1i/openpi 0.3.0
                  OMP@oh-my-pi/pi-coding-agent 17.2.12
                  Bun1.3.14
                  Harbor0.20.0
                  Container runtimePodman server 6.0.2,Linux ARM64 VM
                  Task sourceTerminal-Bench 2.1 ARM64 source-build-derived,commit d1f1920f2d817a831f466d0ff363ef795a9a3b00
                  Tasksgit-leak-recovery, sqlite-db-truncate, cancel-async-tasks
                  Repeats每个模型、任务、Harness 3 次
                  Schedulestrict serial Latin-square
                  Cell deadline1,800 秒
                  Schedule SHA-256974fdc67aad36c4c890a83d6705e7a687610a0ecbef00a0c0ca4e7efac369a62
                  Credential boundaryhost forwarding proxy + 每格短命 bearer;candidate 不持有真实 provider key

                  模型参数:

                  模型Thinking
                  seal/glm-5.3high
                  seal/gpt-5.6-lunahigh

                  GLM-5.3

                  结果格式为 pass / fail / indeterminate。Wall 为 9 个 cell 的总耗时;indeterminate 的等待时间不隐藏。

                  Harness结果WallProvider tokensPhysical POST
                  Bare Pi7 / 1 / 12,647.724s249,62594
                  OpenPI7 / 2 / 01,107.098s331,85187
                  OMP7 / 2 / 01,081.746s2,176,017106

                  逐任务:

                  TaskBare PiOpenPIOMP
                  Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
                  SQLite truncate3 / 0 / 03 / 0 / 03 / 0 / 0
                  Cancel async tasks1 / 1 / 11 / 2 / 01 / 2 / 0

                  效率观察:

                  • OpenPI 与 OMP 都是 7 pass;OpenPI tokens 为 OMP 的 15.3%,少 84.7%。
                  • OMP tokens 为 OpenPI 的 6.56 倍。
                  • OpenPI 相对 Bare Pi 多使用 32.9% tokens;本轮不能声称 OpenPI 比原生 Pi 更省 token。
                  • Bare Pi 的 cancel 总时长包含一个约 30 分钟的 indeterminate,因此不能用 GLM 总 wall 宣称 OpenPI 相对 Pi 有稳定提速。

                  GPT-5.6 Luna

                  Harness结果WallPhysical POSTLogical attempts额外物理 POST
                  Bare Pi7 / 2 / 01,129.885s43430
                  OpenPI7 / 2 / 01,151.213s46460
                  OMP4 / 5 / 01,803.368s1109218

                  逐任务:

                  TaskBare PiOpenPIOMP
                  Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
                  SQLite truncate3 / 0 / 03 / 0 / 01 / 2 / 0
                  Cancel async tasks1 / 2 / 01 / 2 / 00 / 3 / 0

                  效率观察:

                  • OpenPI 与 Bare Pi 都是 7/9;OpenPI wall 比 Bare Pi 高 1.9%,没有提速证据。
                  • OpenPI 相对 OMP 多 3 pass,物理模型请求少 58.2%,总 wall 低 36.2%。
                  • OMP 的 110 个物理 POST 对应 92 个已完成 logical attempts;18 个额外请求按 provider 重试保留,没有从开销中隐藏。
                  • Seal 未返回可用 usage,因此不能根据该轮声称 OpenPI 节省 Luna token 或成本。

                  完整性与安全收据

                  两组运行均满足:

                  • 27/27 Pi/OpenPI/OMP cells 落盘且身份唯一;
                  • 最终冻结源校验通过;
                  • retained artifact credential scan 通过:0 missing roots、0 credential leak、0 scan failure;
                  • 真实 provider key 未注入 candidate;
                  • 全局并发为 1,所有 cell 按冻结 Latin-square 严格串行执行。

                  运行身份:

                  模型Lock fingerprintReceipt
                  GLM-5.3sha256:037179456f3223457e428b9d5543911144b045bcea7ee7127728b27dd735780ecompleted_with_indeterminate
                  GPT-5.6 Lunasha256:00b9db64606832521df7f853d63d5e50d1e6754920e0ac51e1b32c2aacc4866ccompleted_with_indeterminate

                  证据 controller 的 POST/attempt reconciliation 另外覆盖了 provider 物理重试:physical POST < logical attempts 时 fail closed;非负差值记录为 unattributedProviderPostRequests

                  验证:

                  • controller 定向 Node tests:23 pass,1 个环境用例 skip;
                  • 项目全量 Node tests:719/719;
                  • Vitest:29/29;
                  • bun run check:format、lint、typecheck 通过;typecheck 仅有既存 Effect advisory warnings。

                  当前可支持的结论

                  1. 在这两个完整模型批次中,OpenPI 与 Bare Pi 都取得 7/9;Git 与 SQLite 合计均为 12/12。
                  2. 在本轮两个完整模型批次中,OpenPI 保持了与原生 Pi 相同的通过数;不同任务的耗时方向不一致,因此暂不对相对 Pi 的速度优势作结论。
                  3. 相对 OMP,OpenPI 在 GLM-5.3 上以相同 pass 数使用少 84.7% tokens;在 Luna 上取得更多 pass,并使用少 58.2% 的物理请求和少 36.2% 的 wall time。
                  4. 这些数据支持继续保持 Pi-native、普通回合 zero-resident、按需披露能力的方向;尚不足以证明 Subagent/Workflow 的收益,因为这三道题没有专门要求编排能力。
                  5. 下一轮应扩大任务集,并预注册会自然触发 Subagent/Workflow 的多文件任务,将能力采用率、质量与成本分开报告。

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    documentationImprovements or additions to documentation

                    Type

                    No type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 · Issue #46 · openpi-dev/openpi · GitHub
                      Skip to content

                      [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 #46

                      Description

                      @tt-a1i

                      摘要

                      本 Issue 记录 OpenPI 0.3.0、Bare Pi 与 OMP 在两个模型上的同条件对照证据:

                      • 模型:seal/glm-5.3seal/gpt-5.6-luna
                      • Harness:Bare Pi、Pi + OpenPI、OMP
                      • 任务:3 个多轮 Terminal 任务
                      • 重复:每题 3 次
                      • 总计:2 models × 3 tasks × 3 repeats × 3 harnesses = 54 cells

                      主要观察:

                      • GLM-5.3:三条 Harness 都通过 7/9;OpenPI 使用 331,851 tokens,OMP 使用 2,176,017 tokens,OpenPI 少 84.7%。
                      • GPT-5.6 Luna:Bare Pi 与 OpenPI 都通过 7/9,OMP 通过 4/9;OpenPI 比 OMP 少 58.2% 的物理模型请求,总耗时低 36.2%。
                      • Git 与 SQLite 两类任务中,Bare Pi 和 OpenPI 在两个模型上均为 12/12。

                      这组结果没有证明 OpenPI 相对 Bare Pi 存在通用质量提升。它支持的更窄结论是:在本次普通 Terminal 任务里,OpenPI 保持了 Pi 的成功数;相对 OMP,本轮观察到了更低的模型调用开销。

                      结果边界

                      这是在 Linux ARM64 VM 上运行的 Terminal-Bench 2.1 源构建派生诊断,不是 Terminal-Bench 官方成绩或排行榜提交

                      任务数只有 3,每个 cell 重复 3 次。cancel-async-tasks 对 SIGINT/process-group 行为较敏感,方差明显,因此不能把小幅胜负外推成 Harness 的稳定总体能力。

                      GPT-5.6 Luna 路线没有返回可用 token usage;该模型只比较 verifier、物理请求数与 wall time,不作 token 或成本结论。

                      冻结条件

                      项目
                      Pi@earendil-works/pi-coding-agent 0.84.2
                      OpenPI@tt-a1i/openpi 0.3.0
                      OMP@oh-my-pi/pi-coding-agent 17.2.12
                      Bun1.3.14
                      Harbor0.20.0
                      Container runtimePodman server 6.0.2,Linux ARM64 VM
                      Task sourceTerminal-Bench 2.1 ARM64 source-build-derived,commit d1f1920f2d817a831f466d0ff363ef795a9a3b00
                      Tasksgit-leak-recovery, sqlite-db-truncate, cancel-async-tasks
                      Repeats每个模型、任务、Harness 3 次
                      Schedulestrict serial Latin-square
                      Cell deadline1,800 秒
                      Schedule SHA-256974fdc67aad36c4c890a83d6705e7a687610a0ecbef00a0c0ca4e7efac369a62
                      Credential boundaryhost forwarding proxy + 每格短命 bearer;candidate 不持有真实 provider key

                      模型参数:

                      模型Thinking
                      seal/glm-5.3high
                      seal/gpt-5.6-lunahigh

                      GLM-5.3

                      结果格式为 pass / fail / indeterminate。Wall 为 9 个 cell 的总耗时;indeterminate 的等待时间不隐藏。

                      Harness结果WallProvider tokensPhysical POST
                      Bare Pi7 / 1 / 12,647.724s249,62594
                      OpenPI7 / 2 / 01,107.098s331,85187
                      OMP7 / 2 / 01,081.746s2,176,017106

                      逐任务:

                      TaskBare PiOpenPIOMP
                      Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
                      SQLite truncate3 / 0 / 03 / 0 / 03 / 0 / 0
                      Cancel async tasks1 / 1 / 11 / 2 / 01 / 2 / 0

                      效率观察:

                      • OpenPI 与 OMP 都是 7 pass;OpenPI tokens 为 OMP 的 15.3%,少 84.7%。
                      • OMP tokens 为 OpenPI 的 6.56 倍。
                      • OpenPI 相对 Bare Pi 多使用 32.9% tokens;本轮不能声称 OpenPI 比原生 Pi 更省 token。
                      • Bare Pi 的 cancel 总时长包含一个约 30 分钟的 indeterminate,因此不能用 GLM 总 wall 宣称 OpenPI 相对 Pi 有稳定提速。

                      GPT-5.6 Luna

                      Harness结果WallPhysical POSTLogical attempts额外物理 POST
                      Bare Pi7 / 2 / 01,129.885s43430
                      OpenPI7 / 2 / 01,151.213s46460
                      OMP4 / 5 / 01,803.368s1109218

                      逐任务:

                      TaskBare PiOpenPIOMP
                      Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
                      SQLite truncate3 / 0 / 03 / 0 / 01 / 2 / 0
                      Cancel async tasks1 / 2 / 01 / 2 / 00 / 3 / 0

                      效率观察:

                      • OpenPI 与 Bare Pi 都是 7/9;OpenPI wall 比 Bare Pi 高 1.9%,没有提速证据。
                      • OpenPI 相对 OMP 多 3 pass,物理模型请求少 58.2%,总 wall 低 36.2%。
                      • OMP 的 110 个物理 POST 对应 92 个已完成 logical attempts;18 个额外请求按 provider 重试保留,没有从开销中隐藏。
                      • Seal 未返回可用 usage,因此不能根据该轮声称 OpenPI 节省 Luna token 或成本。

                      完整性与安全收据

                      两组运行均满足:

                      • 27/27 Pi/OpenPI/OMP cells 落盘且身份唯一;
                      • 最终冻结源校验通过;
                      • retained artifact credential scan 通过:0 missing roots、0 credential leak、0 scan failure;
                      • 真实 provider key 未注入 candidate;
                      • 全局并发为 1,所有 cell 按冻结 Latin-square 严格串行执行。

                      运行身份:

                      模型Lock fingerprintReceipt
                      GLM-5.3sha256:037179456f3223457e428b9d5543911144b045bcea7ee7127728b27dd735780ecompleted_with_indeterminate
                      GPT-5.6 Lunasha256:00b9db64606832521df7f853d63d5e50d1e6754920e0ac51e1b32c2aacc4866ccompleted_with_indeterminate

                      证据 controller 的 POST/attempt reconciliation 另外覆盖了 provider 物理重试:physical POST < logical attempts 时 fail closed;非负差值记录为 unattributedProviderPostRequests

                      验证:

                      • controller 定向 Node tests:23 pass,1 个环境用例 skip;
                      • 项目全量 Node tests:719/719;
                      • Vitest:29/29;
                      • bun run check:format、lint、typecheck 通过;typecheck 仅有既存 Effect advisory warnings。

                      当前可支持的结论

                      1. 在这两个完整模型批次中,OpenPI 与 Bare Pi 都取得 7/9;Git 与 SQLite 合计均为 12/12。
                      2. 在本轮两个完整模型批次中,OpenPI 保持了与原生 Pi 相同的通过数;不同任务的耗时方向不一致,因此暂不对相对 Pi 的速度优势作结论。
                      3. 相对 OMP,OpenPI 在 GLM-5.3 上以相同 pass 数使用少 84.7% tokens;在 Luna 上取得更多 pass,并使用少 58.2% 的物理请求和少 36.2% 的 wall time。
                      4. 这些数据支持继续保持 Pi-native、普通回合 zero-resident、按需披露能力的方向;尚不足以证明 Subagent/Workflow 的收益,因为这三道题没有专门要求编排能力。
                      5. 下一轮应扩大任务集,并预注册会自然触发 Subagent/Workflow 的多文件任务,将能力采用率、质量与成本分开报告。

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        documentationImprovements or additions to documentation

                        Type

                        No type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + ' [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 · Issue #46 · openpi-dev/openpi · GitHub
                          Skip to content

                          [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 #46

                          Description

                          @tt-a1i

                          摘要

                          本 Issue 记录 OpenPI 0.3.0、Bare Pi 与 OMP 在两个模型上的同条件对照证据:

                          • 模型:seal/glm-5.3seal/gpt-5.6-luna
                          • Harness:Bare Pi、Pi + OpenPI、OMP
                          • 任务:3 个多轮 Terminal 任务
                          • 重复:每题 3 次
                          • 总计:2 models × 3 tasks × 3 repeats × 3 harnesses = 54 cells

                          主要观察:

                          • GLM-5.3:三条 Harness 都通过 7/9;OpenPI 使用 331,851 tokens,OMP 使用 2,176,017 tokens,OpenPI 少 84.7%。
                          • GPT-5.6 Luna:Bare Pi 与 OpenPI 都通过 7/9,OMP 通过 4/9;OpenPI 比 OMP 少 58.2% 的物理模型请求,总耗时低 36.2%。
                          • Git 与 SQLite 两类任务中,Bare Pi 和 OpenPI 在两个模型上均为 12/12。

                          这组结果没有证明 OpenPI 相对 Bare Pi 存在通用质量提升。它支持的更窄结论是:在本次普通 Terminal 任务里,OpenPI 保持了 Pi 的成功数;相对 OMP,本轮观察到了更低的模型调用开销。

                          结果边界

                          这是在 Linux ARM64 VM 上运行的 Terminal-Bench 2.1 源构建派生诊断,不是 Terminal-Bench 官方成绩或排行榜提交

                          任务数只有 3,每个 cell 重复 3 次。cancel-async-tasks 对 SIGINT/process-group 行为较敏感,方差明显,因此不能把小幅胜负外推成 Harness 的稳定总体能力。

                          GPT-5.6 Luna 路线没有返回可用 token usage;该模型只比较 verifier、物理请求数与 wall time,不作 token 或成本结论。

                          冻结条件

                          项目
                          Pi@earendil-works/pi-coding-agent 0.84.2
                          OpenPI@tt-a1i/openpi 0.3.0
                          OMP@oh-my-pi/pi-coding-agent 17.2.12
                          Bun1.3.14
                          Harbor0.20.0
                          Container runtimePodman server 6.0.2,Linux ARM64 VM
                          Task sourceTerminal-Bench 2.1 ARM64 source-build-derived,commit d1f1920f2d817a831f466d0ff363ef795a9a3b00
                          Tasksgit-leak-recovery, sqlite-db-truncate, cancel-async-tasks
                          Repeats每个模型、任务、Harness 3 次
                          Schedulestrict serial Latin-square
                          Cell deadline1,800 秒
                          Schedule SHA-256974fdc67aad36c4c890a83d6705e7a687610a0ecbef00a0c0ca4e7efac369a62
                          Credential boundaryhost forwarding proxy + 每格短命 bearer;candidate 不持有真实 provider key

                          模型参数:

                          模型Thinking
                          seal/glm-5.3high
                          seal/gpt-5.6-lunahigh

                          GLM-5.3

                          结果格式为 pass / fail / indeterminate。Wall 为 9 个 cell 的总耗时;indeterminate 的等待时间不隐藏。

                          Harness结果WallProvider tokensPhysical POST
                          Bare Pi7 / 1 / 12,647.724s249,62594
                          OpenPI7 / 2 / 01,107.098s331,85187
                          OMP7 / 2 / 01,081.746s2,176,017106

                          逐任务:

                          TaskBare PiOpenPIOMP
                          Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
                          SQLite truncate3 / 0 / 03 / 0 / 03 / 0 / 0
                          Cancel async tasks1 / 1 / 11 / 2 / 01 / 2 / 0

                          效率观察:

                          • OpenPI 与 OMP 都是 7 pass;OpenPI tokens 为 OMP 的 15.3%,少 84.7%。
                          • OMP tokens 为 OpenPI 的 6.56 倍。
                          • OpenPI 相对 Bare Pi 多使用 32.9% tokens;本轮不能声称 OpenPI 比原生 Pi 更省 token。
                          • Bare Pi 的 cancel 总时长包含一个约 30 分钟的 indeterminate,因此不能用 GLM 总 wall 宣称 OpenPI 相对 Pi 有稳定提速。

                          GPT-5.6 Luna

                          Harness结果WallPhysical POSTLogical attempts额外物理 POST
                          Bare Pi7 / 2 / 01,129.885s43430
                          OpenPI7 / 2 / 01,151.213s46460
                          OMP4 / 5 / 01,803.368s1109218

                          逐任务:

                          TaskBare PiOpenPIOMP
                          Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
                          SQLite truncate3 / 0 / 03 / 0 / 01 / 2 / 0
                          Cancel async tasks1 / 2 / 01 / 2 / 00 / 3 / 0

                          效率观察:

                          • OpenPI 与 Bare Pi 都是 7/9;OpenPI wall 比 Bare Pi 高 1.9%,没有提速证据。
                          • OpenPI 相对 OMP 多 3 pass,物理模型请求少 58.2%,总 wall 低 36.2%。
                          • OMP 的 110 个物理 POST 对应 92 个已完成 logical attempts;18 个额外请求按 provider 重试保留,没有从开销中隐藏。
                          • Seal 未返回可用 usage,因此不能根据该轮声称 OpenPI 节省 Luna token 或成本。

                          完整性与安全收据

                          两组运行均满足:

                          • 27/27 Pi/OpenPI/OMP cells 落盘且身份唯一;
                          • 最终冻结源校验通过;
                          • retained artifact credential scan 通过:0 missing roots、0 credential leak、0 scan failure;
                          • 真实 provider key 未注入 candidate;
                          • 全局并发为 1,所有 cell 按冻结 Latin-square 严格串行执行。

                          运行身份:

                          模型Lock fingerprintReceipt
                          GLM-5.3sha256:037179456f3223457e428b9d5543911144b045bcea7ee7127728b27dd735780ecompleted_with_indeterminate
                          GPT-5.6 Lunasha256:00b9db64606832521df7f853d63d5e50d1e6754920e0ac51e1b32c2aacc4866ccompleted_with_indeterminate

                          证据 controller 的 POST/attempt reconciliation 另外覆盖了 provider 物理重试:physical POST < logical attempts 时 fail closed;非负差值记录为 unattributedProviderPostRequests

                          验证:

                          • controller 定向 Node tests:23 pass,1 个环境用例 skip;
                          • 项目全量 Node tests:719/719;
                          • Vitest:29/29;
                          • bun run check:format、lint、typecheck 通过;typecheck 仅有既存 Effect advisory warnings。

                          当前可支持的结论

                          1. 在这两个完整模型批次中,OpenPI 与 Bare Pi 都取得 7/9;Git 与 SQLite 合计均为 12/12。
                          2. 在本轮两个完整模型批次中,OpenPI 保持了与原生 Pi 相同的通过数;不同任务的耗时方向不一致,因此暂不对相对 Pi 的速度优势作结论。
                          3. 相对 OMP,OpenPI 在 GLM-5.3 上以相同 pass 数使用少 84.7% tokens;在 Luna 上取得更多 pass,并使用少 58.2% 的物理请求和少 36.2% 的 wall time。
                          4. 这些数据支持继续保持 Pi-native、普通回合 zero-resident、按需披露能力的方向;尚不足以证明 Subagent/Workflow 的收益,因为这三道题没有专门要求编排能力。
                          5. 下一轮应扩大任务集,并预注册会自然触发 Subagent/Workflow 的多文件任务,将能力采用率、质量与成本分开报告。

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            documentationImprovements or additions to documentation

                            Type

                            No type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })(); [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 · Issue #46 · openpi-dev/openpi · GitHub
                              Skip to content

                              [Benchmark] OpenPI 0.3.0 / Pi / OMP:GLM-5.3 与 GPT-5.6 Luna 54-cell ARM64 派生诊断 #46

                              Description

                              @tt-a1i

                              摘要

                              本 Issue 记录 OpenPI 0.3.0、Bare Pi 与 OMP 在两个模型上的同条件对照证据:

                              • 模型:seal/glm-5.3seal/gpt-5.6-luna
                              • Harness:Bare Pi、Pi + OpenPI、OMP
                              • 任务:3 个多轮 Terminal 任务
                              • 重复:每题 3 次
                              • 总计:2 models × 3 tasks × 3 repeats × 3 harnesses = 54 cells

                              主要观察:

                              • GLM-5.3:三条 Harness 都通过 7/9;OpenPI 使用 331,851 tokens,OMP 使用 2,176,017 tokens,OpenPI 少 84.7%。
                              • GPT-5.6 Luna:Bare Pi 与 OpenPI 都通过 7/9,OMP 通过 4/9;OpenPI 比 OMP 少 58.2% 的物理模型请求,总耗时低 36.2%。
                              • Git 与 SQLite 两类任务中,Bare Pi 和 OpenPI 在两个模型上均为 12/12。

                              这组结果没有证明 OpenPI 相对 Bare Pi 存在通用质量提升。它支持的更窄结论是:在本次普通 Terminal 任务里,OpenPI 保持了 Pi 的成功数;相对 OMP,本轮观察到了更低的模型调用开销。

                              结果边界

                              这是在 Linux ARM64 VM 上运行的 Terminal-Bench 2.1 源构建派生诊断,不是 Terminal-Bench 官方成绩或排行榜提交

                              任务数只有 3,每个 cell 重复 3 次。cancel-async-tasks 对 SIGINT/process-group 行为较敏感,方差明显,因此不能把小幅胜负外推成 Harness 的稳定总体能力。

                              GPT-5.6 Luna 路线没有返回可用 token usage;该模型只比较 verifier、物理请求数与 wall time,不作 token 或成本结论。

                              冻结条件

                              项目
                              Pi@earendil-works/pi-coding-agent 0.84.2
                              OpenPI@tt-a1i/openpi 0.3.0
                              OMP@oh-my-pi/pi-coding-agent 17.2.12
                              Bun1.3.14
                              Harbor0.20.0
                              Container runtimePodman server 6.0.2,Linux ARM64 VM
                              Task sourceTerminal-Bench 2.1 ARM64 source-build-derived,commit d1f1920f2d817a831f466d0ff363ef795a9a3b00
                              Tasksgit-leak-recovery, sqlite-db-truncate, cancel-async-tasks
                              Repeats每个模型、任务、Harness 3 次
                              Schedulestrict serial Latin-square
                              Cell deadline1,800 秒
                              Schedule SHA-256974fdc67aad36c4c890a83d6705e7a687610a0ecbef00a0c0ca4e7efac369a62
                              Credential boundaryhost forwarding proxy + 每格短命 bearer;candidate 不持有真实 provider key

                              模型参数:

                              模型Thinking
                              seal/glm-5.3high
                              seal/gpt-5.6-lunahigh

                              GLM-5.3

                              结果格式为 pass / fail / indeterminate。Wall 为 9 个 cell 的总耗时;indeterminate 的等待时间不隐藏。

                              Harness结果WallProvider tokensPhysical POST
                              Bare Pi7 / 1 / 12,647.724s249,62594
                              OpenPI7 / 2 / 01,107.098s331,85187
                              OMP7 / 2 / 01,081.746s2,176,017106

                              逐任务:

                              TaskBare PiOpenPIOMP
                              Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
                              SQLite truncate3 / 0 / 03 / 0 / 03 / 0 / 0
                              Cancel async tasks1 / 1 / 11 / 2 / 01 / 2 / 0

                              效率观察:

                              • OpenPI 与 OMP 都是 7 pass;OpenPI tokens 为 OMP 的 15.3%,少 84.7%。
                              • OMP tokens 为 OpenPI 的 6.56 倍。
                              • OpenPI 相对 Bare Pi 多使用 32.9% tokens;本轮不能声称 OpenPI 比原生 Pi 更省 token。
                              • Bare Pi 的 cancel 总时长包含一个约 30 分钟的 indeterminate,因此不能用 GLM 总 wall 宣称 OpenPI 相对 Pi 有稳定提速。

                              GPT-5.6 Luna

                              Harness结果WallPhysical POSTLogical attempts额外物理 POST
                              Bare Pi7 / 2 / 01,129.885s43430
                              OpenPI7 / 2 / 01,151.213s46460
                              OMP4 / 5 / 01,803.368s1109218

                              逐任务:

                              TaskBare PiOpenPIOMP
                              Git leak recovery3 / 0 / 03 / 0 / 03 / 0 / 0
                              SQLite truncate3 / 0 / 03 / 0 / 01 / 2 / 0
                              Cancel async tasks1 / 2 / 01 / 2 / 00 / 3 / 0

                              效率观察:

                              • OpenPI 与 Bare Pi 都是 7/9;OpenPI wall 比 Bare Pi 高 1.9%,没有提速证据。
                              • OpenPI 相对 OMP 多 3 pass,物理模型请求少 58.2%,总 wall 低 36.2%。
                              • OMP 的 110 个物理 POST 对应 92 个已完成 logical attempts;18 个额外请求按 provider 重试保留,没有从开销中隐藏。
                              • Seal 未返回可用 usage,因此不能根据该轮声称 OpenPI 节省 Luna token 或成本。

                              完整性与安全收据

                              两组运行均满足:

                              • 27/27 Pi/OpenPI/OMP cells 落盘且身份唯一;
                              • 最终冻结源校验通过;
                              • retained artifact credential scan 通过:0 missing roots、0 credential leak、0 scan failure;
                              • 真实 provider key 未注入 candidate;
                              • 全局并发为 1,所有 cell 按冻结 Latin-square 严格串行执行。

                              运行身份:

                              模型Lock fingerprintReceipt
                              GLM-5.3sha256:037179456f3223457e428b9d5543911144b045bcea7ee7127728b27dd735780ecompleted_with_indeterminate
                              GPT-5.6 Lunasha256:00b9db64606832521df7f853d63d5e50d1e6754920e0ac51e1b32c2aacc4866ccompleted_with_indeterminate

                              证据 controller 的 POST/attempt reconciliation 另外覆盖了 provider 物理重试:physical POST < logical attempts 时 fail closed;非负差值记录为 unattributedProviderPostRequests

                              验证:

                              • controller 定向 Node tests:23 pass,1 个环境用例 skip;
                              • 项目全量 Node tests:719/719;
                              • Vitest:29/29;
                              • bun run check:format、lint、typecheck 通过;typecheck 仅有既存 Effect advisory warnings。

                              当前可支持的结论

                              1. 在这两个完整模型批次中,OpenPI 与 Bare Pi 都取得 7/9;Git 与 SQLite 合计均为 12/12。
                              2. 在本轮两个完整模型批次中,OpenPI 保持了与原生 Pi 相同的通过数;不同任务的耗时方向不一致,因此暂不对相对 Pi 的速度优势作结论。
                              3. 相对 OMP,OpenPI 在 GLM-5.3 上以相同 pass 数使用少 84.7% tokens;在 Luna 上取得更多 pass,并使用少 58.2% 的物理请求和少 36.2% 的 wall time。
                              4. 这些数据支持继续保持 Pi-native、普通回合 zero-resident、按需披露能力的方向;尚不足以证明 Subagent/Workflow 的收益,因为这三道题没有专门要求编排能力。
                              5. 下一轮应扩大任务集,并预注册会自然触发 Subagent/Workflow 的多文件任务,将能力采用率、质量与成本分开报告。

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                documentationImprovements or additions to documentation

                                Type

                                No type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions