revert: Token estimation values should use existing auto-calculation - #29

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values
Draft

revert: Token estimation values should use existing auto-calculation#29
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

After review, the proposed changes to extract hardcoded token estimation values to configuration have been reverted.

Reason

The rate limiting concurrency is already automatically calculated using existing environment variables:

  • LLM_REQUESTS_PER_MINUTE (default: 800)
  • LLM_TOKENS_PER_MINUTE (default: 40000)

The rate_limiter.py module contains an internal avg_tokens_map that provides token estimation values for each service (LLM: 3500, Embedding: 20000, Rerank: 500, DS_OCR: 3500). These internal values are used to auto-calculate the appropriate concurrency level.

Conclusion

No additional configuration is needed. The existing LLM_REQUESTS_PER_MINUTE and LLM_TOKENS_PER_MINUTE environment variables, combined with the internal token estimation in rate_limiter.py, already provide the correct behavior for rate limiting.

Original prompt

This section details on the original issue you should resolve

<issue_title>Token estimation values hardcoded multiple places - inaccurate rate limiting</issue_title>
<issue_description>## 问题描述
Token 估算值硬编码在多个文件中,导致速率限制不准确。

受影响的文件和位置

  • multi_tenant.py 行 113:estimated_output = 3000(LLM)
  • multi_tenant.py 行 313:estimated_output = 500(VLM)
  • multi_tenant.py 行 339:max_tokens = 500
  • deepseek_ocr_client.py 行 285, 352

问题

  • 无法按不同模型调整估算值
  • 同一个值 "3000" 在多处重复,维护困难
  • 使用小模型时估算过高,并发受限;使用大模型时估算不足

解决方案

应该提取到 .env 配置:

  • LLM_ESTIMATED_OUTPUT_TOKENS
  • VLM_ESTIMATED_OUTPUT_TOKENS 等</issue_description>

Comments on the Issue (you are @copilot in this section)


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

- 在 LLMConfig 中添加 estimated_output_tokens, vlm_estimated_output_tokens, vlm_max_tokens, vlm_image_tokens_estimate
- 在 DeepSeekOCRConfig 中添加 image_tokens_estimate
- 更新 multi_tenant.py 使用配置值替代硬编码
- 更新 deepseek_ocr_client.py 使用配置值替代硬编码
- 更新 env.example 添加新的环境变量说明
Closes #issue
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Fix hardcoded token estimation values for accurate rate limitingfeat: Extract hardcoded token estimation values to configurationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:47

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)

这两个环境变量来算

Comment threadenv.example Outdated
Comment on lines +34 to +41
# --- LLM Token 估算配置(用于速率限制) ---
# 估算输出 tokens 数量,用于速率限制计算
# 如果估算过高,并发会受限;如果估算不足,可能触发 429 错误
# LLM_ESTIMATED_OUTPUT_TOKENS=3000 # LLM 输出估算(实体提取约 3000 tokens,默认 3000)
# LLM_VLM_ESTIMATED_OUTPUT_TOKENS=500 # VLM 输出估算(图片描述较短,默认 500)
# LLM_VLM_MAX_TOKENS=500 # VLM API 最大输出 tokens(默认 500)
# LLM_VLM_IMAGE_TOKENS_ESTIMATE=200 # VLM 图片输入估算 tokens(默认 200)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个变量值和LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)
是不是有冲突?明明是自动计算,为什么还要这些估算值? @copilot@claude

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确实有冲突。已在 e2cb6e7 中撤销这些配置项。

rate_limiter.pyavg_tokens_map 已内置各服务的 token 估算值,通过 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 自动计算并发数,不需要额外暴露这些内部值。

Comment threadsrc/multi_tenant.py Outdated
Comment on lines +67 to +72
# Token 估算配置
self.llm_estimated_output_tokens = config.llm.estimated_output_tokens
self.vlm_estimated_output_tokens = config.llm.vlm_estimated_output_tokens
self.vlm_max_tokens = config.llm.vlm_max_tokens
self.vlm_image_tokens_estimate = config.llm.vlm_image_tokens_estimate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题同.env.example

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 e2cb6e7 中撤销相关修改。

根据用户反馈,这些估算值应该通过 LLM_REQUESTS_PER_MINUTE 和 LLM_TOKENS_PER_MINUTE 自动计算,不需要额外的配置项。
rate_limiter.py 中已有 avg_tokens_map 用于自动计算并发数。
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)...

已在 e2cb6e7 中撤销所有修改。

速率限制的并发数确实应该通过 LLM_REQUESTS_PER_MINUTELLM_TOKENS_PER_MINUTE 自动计算,rate_limiter.py 中的 avg_tokens_map 已包含各服务的内部 token 估算值用于计算。

CopilotAI changed the title feat: Extract hardcoded token estimation values to configurationrevert: Token estimation values should use existing auto-calculationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 10:05
@BukeLy

Copy link
Copy Markdown
Owner

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

建议改进方案:将分散的硬编码值统一到 rate_limiter.pyavg_tokens_map 中复用。

当前问题:

  • multi_tenant.py:113 - estimated_output = 3000
  • multi_tenant.py:312-313 - image_tokens = 200, estimated_output = 500
  • deepseek_ocr_client.py:285,352 - 1000 + self.config.max_tokens

改进方案:

  1. rate_limiter.py 导出 avg_tokens_map 或创建 get_avg_tokens(service) 函数
  2. multi_tenant.pydeepseek_ocr_client.py 调用此函数获取估算值
  3. 这样只需维护 rate_limiter.py 一处,无需添加新环境变量

这样可保持与 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 的自动计算逻辑一致,同时解决硬编码分散的问题。

是否需要我实现这个方案?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token estimation values hardcoded multiple places - inaccurate rate limiting

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

revert: Token estimation values should use existing auto-calculation - #29

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values
Draft

revert: Token estimation values should use existing auto-calculation#29
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

After review, the proposed changes to extract hardcoded token estimation values to configuration have been reverted.

Reason

The rate limiting concurrency is already automatically calculated using existing environment variables:

  • LLM_REQUESTS_PER_MINUTE (default: 800)
  • LLM_TOKENS_PER_MINUTE (default: 40000)

The rate_limiter.py module contains an internal avg_tokens_map that provides token estimation values for each service (LLM: 3500, Embedding: 20000, Rerank: 500, DS_OCR: 3500). These internal values are used to auto-calculate the appropriate concurrency level.

Conclusion

No additional configuration is needed. The existing LLM_REQUESTS_PER_MINUTE and LLM_TOKENS_PER_MINUTE environment variables, combined with the internal token estimation in rate_limiter.py, already provide the correct behavior for rate limiting.

Original prompt

This section details on the original issue you should resolve

<issue_title>Token estimation values hardcoded multiple places - inaccurate rate limiting</issue_title>
<issue_description>## 问题描述
Token 估算值硬编码在多个文件中,导致速率限制不准确。

受影响的文件和位置

  • multi_tenant.py 行 113:estimated_output = 3000(LLM)
  • multi_tenant.py 行 313:estimated_output = 500(VLM)
  • multi_tenant.py 行 339:max_tokens = 500
  • deepseek_ocr_client.py 行 285, 352

问题

  • 无法按不同模型调整估算值
  • 同一个值 "3000" 在多处重复,维护困难
  • 使用小模型时估算过高,并发受限;使用大模型时估算不足

解决方案

应该提取到 .env 配置:

  • LLM_ESTIMATED_OUTPUT_TOKENS
  • VLM_ESTIMATED_OUTPUT_TOKENS 等</issue_description>

Comments on the Issue (you are @copilot in this section)


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

- 在 LLMConfig 中添加 estimated_output_tokens, vlm_estimated_output_tokens, vlm_max_tokens, vlm_image_tokens_estimate
- 在 DeepSeekOCRConfig 中添加 image_tokens_estimate
- 更新 multi_tenant.py 使用配置值替代硬编码
- 更新 deepseek_ocr_client.py 使用配置值替代硬编码
- 更新 env.example 添加新的环境变量说明
Closes #issue
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Fix hardcoded token estimation values for accurate rate limitingfeat: Extract hardcoded token estimation values to configurationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:47

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)

这两个环境变量来算

Comment threadenv.example Outdated
Comment on lines +34 to +41
# --- LLM Token 估算配置(用于速率限制) ---
# 估算输出 tokens 数量,用于速率限制计算
# 如果估算过高,并发会受限;如果估算不足,可能触发 429 错误
# LLM_ESTIMATED_OUTPUT_TOKENS=3000 # LLM 输出估算(实体提取约 3000 tokens,默认 3000)
# LLM_VLM_ESTIMATED_OUTPUT_TOKENS=500 # VLM 输出估算(图片描述较短,默认 500)
# LLM_VLM_MAX_TOKENS=500 # VLM API 最大输出 tokens(默认 500)
# LLM_VLM_IMAGE_TOKENS_ESTIMATE=200 # VLM 图片输入估算 tokens(默认 200)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个变量值和LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)
是不是有冲突?明明是自动计算,为什么还要这些估算值? @copilot@claude

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确实有冲突。已在 e2cb6e7 中撤销这些配置项。

rate_limiter.pyavg_tokens_map 已内置各服务的 token 估算值,通过 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 自动计算并发数,不需要额外暴露这些内部值。

Comment threadsrc/multi_tenant.py Outdated
Comment on lines +67 to +72
# Token 估算配置
self.llm_estimated_output_tokens = config.llm.estimated_output_tokens
self.vlm_estimated_output_tokens = config.llm.vlm_estimated_output_tokens
self.vlm_max_tokens = config.llm.vlm_max_tokens
self.vlm_image_tokens_estimate = config.llm.vlm_image_tokens_estimate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题同.env.example

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 e2cb6e7 中撤销相关修改。

根据用户反馈,这些估算值应该通过 LLM_REQUESTS_PER_MINUTE 和 LLM_TOKENS_PER_MINUTE 自动计算,不需要额外的配置项。
rate_limiter.py 中已有 avg_tokens_map 用于自动计算并发数。
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)...

已在 e2cb6e7 中撤销所有修改。

速率限制的并发数确实应该通过 LLM_REQUESTS_PER_MINUTELLM_TOKENS_PER_MINUTE 自动计算,rate_limiter.py 中的 avg_tokens_map 已包含各服务的内部 token 估算值用于计算。

CopilotAI changed the title feat: Extract hardcoded token estimation values to configurationrevert: Token estimation values should use existing auto-calculationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 10:05
@BukeLy

Copy link
Copy Markdown
Owner

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

建议改进方案:将分散的硬编码值统一到 rate_limiter.pyavg_tokens_map 中复用。

当前问题:

  • multi_tenant.py:113 - estimated_output = 3000
  • multi_tenant.py:312-313 - image_tokens = 200, estimated_output = 500
  • deepseek_ocr_client.py:285,352 - 1000 + self.config.max_tokens

改进方案:

  1. rate_limiter.py 导出 avg_tokens_map 或创建 get_avg_tokens(service) 函数
  2. multi_tenant.pydeepseek_ocr_client.py 调用此函数获取估算值
  3. 这样只需维护 rate_limiter.py 一处,无需添加新环境变量

这样可保持与 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 的自动计算逻辑一致,同时解决硬编码分散的问题。

是否需要我实现这个方案?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token estimation values hardcoded multiple places - inaccurate rate limiting

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

revert: Token estimation values should use existing auto-calculation - #29

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values
Draft

revert: Token estimation values should use existing auto-calculation#29
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

After review, the proposed changes to extract hardcoded token estimation values to configuration have been reverted.

Reason

The rate limiting concurrency is already automatically calculated using existing environment variables:

  • LLM_REQUESTS_PER_MINUTE (default: 800)
  • LLM_TOKENS_PER_MINUTE (default: 40000)

The rate_limiter.py module contains an internal avg_tokens_map that provides token estimation values for each service (LLM: 3500, Embedding: 20000, Rerank: 500, DS_OCR: 3500). These internal values are used to auto-calculate the appropriate concurrency level.

Conclusion

No additional configuration is needed. The existing LLM_REQUESTS_PER_MINUTE and LLM_TOKENS_PER_MINUTE environment variables, combined with the internal token estimation in rate_limiter.py, already provide the correct behavior for rate limiting.

Original prompt

This section details on the original issue you should resolve

<issue_title>Token estimation values hardcoded multiple places - inaccurate rate limiting</issue_title>
<issue_description>## 问题描述
Token 估算值硬编码在多个文件中,导致速率限制不准确。

受影响的文件和位置

  • multi_tenant.py 行 113:estimated_output = 3000(LLM)
  • multi_tenant.py 行 313:estimated_output = 500(VLM)
  • multi_tenant.py 行 339:max_tokens = 500
  • deepseek_ocr_client.py 行 285, 352

问题

  • 无法按不同模型调整估算值
  • 同一个值 "3000" 在多处重复,维护困难
  • 使用小模型时估算过高,并发受限;使用大模型时估算不足

解决方案

应该提取到 .env 配置:

  • LLM_ESTIMATED_OUTPUT_TOKENS
  • VLM_ESTIMATED_OUTPUT_TOKENS 等</issue_description>

Comments on the Issue (you are @copilot in this section)


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

- 在 LLMConfig 中添加 estimated_output_tokens, vlm_estimated_output_tokens, vlm_max_tokens, vlm_image_tokens_estimate
- 在 DeepSeekOCRConfig 中添加 image_tokens_estimate
- 更新 multi_tenant.py 使用配置值替代硬编码
- 更新 deepseek_ocr_client.py 使用配置值替代硬编码
- 更新 env.example 添加新的环境变量说明
Closes #issue
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Fix hardcoded token estimation values for accurate rate limitingfeat: Extract hardcoded token estimation values to configurationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:47

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)

这两个环境变量来算

Comment threadenv.example Outdated
Comment on lines +34 to +41
# --- LLM Token 估算配置(用于速率限制) ---
# 估算输出 tokens 数量,用于速率限制计算
# 如果估算过高,并发会受限;如果估算不足,可能触发 429 错误
# LLM_ESTIMATED_OUTPUT_TOKENS=3000 # LLM 输出估算(实体提取约 3000 tokens,默认 3000)
# LLM_VLM_ESTIMATED_OUTPUT_TOKENS=500 # VLM 输出估算(图片描述较短,默认 500)
# LLM_VLM_MAX_TOKENS=500 # VLM API 最大输出 tokens(默认 500)
# LLM_VLM_IMAGE_TOKENS_ESTIMATE=200 # VLM 图片输入估算 tokens(默认 200)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个变量值和LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)
是不是有冲突?明明是自动计算,为什么还要这些估算值? @copilot@claude

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确实有冲突。已在 e2cb6e7 中撤销这些配置项。

rate_limiter.pyavg_tokens_map 已内置各服务的 token 估算值,通过 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 自动计算并发数,不需要额外暴露这些内部值。

Comment threadsrc/multi_tenant.py Outdated
Comment on lines +67 to +72
# Token 估算配置
self.llm_estimated_output_tokens = config.llm.estimated_output_tokens
self.vlm_estimated_output_tokens = config.llm.vlm_estimated_output_tokens
self.vlm_max_tokens = config.llm.vlm_max_tokens
self.vlm_image_tokens_estimate = config.llm.vlm_image_tokens_estimate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题同.env.example

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 e2cb6e7 中撤销相关修改。

根据用户反馈,这些估算值应该通过 LLM_REQUESTS_PER_MINUTE 和 LLM_TOKENS_PER_MINUTE 自动计算,不需要额外的配置项。
rate_limiter.py 中已有 avg_tokens_map 用于自动计算并发数。
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)...

已在 e2cb6e7 中撤销所有修改。

速率限制的并发数确实应该通过 LLM_REQUESTS_PER_MINUTELLM_TOKENS_PER_MINUTE 自动计算,rate_limiter.py 中的 avg_tokens_map 已包含各服务的内部 token 估算值用于计算。

CopilotAI changed the title feat: Extract hardcoded token estimation values to configurationrevert: Token estimation values should use existing auto-calculationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 10:05
@BukeLy

Copy link
Copy Markdown
Owner

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

建议改进方案:将分散的硬编码值统一到 rate_limiter.pyavg_tokens_map 中复用。

当前问题:

  • multi_tenant.py:113 - estimated_output = 3000
  • multi_tenant.py:312-313 - image_tokens = 200, estimated_output = 500
  • deepseek_ocr_client.py:285,352 - 1000 + self.config.max_tokens

改进方案:

  1. rate_limiter.py 导出 avg_tokens_map 或创建 get_avg_tokens(service) 函数
  2. multi_tenant.pydeepseek_ocr_client.py 调用此函数获取估算值
  3. 这样只需维护 rate_limiter.py 一处,无需添加新环境变量

这样可保持与 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 的自动计算逻辑一致,同时解决硬编码分散的问题。

是否需要我实现这个方案?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token estimation values hardcoded multiple places - inaccurate rate limiting

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

revert: Token estimation values should use existing auto-calculation - #29

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values
Draft

revert: Token estimation values should use existing auto-calculation#29
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

After review, the proposed changes to extract hardcoded token estimation values to configuration have been reverted.

Reason

The rate limiting concurrency is already automatically calculated using existing environment variables:

  • LLM_REQUESTS_PER_MINUTE (default: 800)
  • LLM_TOKENS_PER_MINUTE (default: 40000)

The rate_limiter.py module contains an internal avg_tokens_map that provides token estimation values for each service (LLM: 3500, Embedding: 20000, Rerank: 500, DS_OCR: 3500). These internal values are used to auto-calculate the appropriate concurrency level.

Conclusion

No additional configuration is needed. The existing LLM_REQUESTS_PER_MINUTE and LLM_TOKENS_PER_MINUTE environment variables, combined with the internal token estimation in rate_limiter.py, already provide the correct behavior for rate limiting.

Original prompt

This section details on the original issue you should resolve

<issue_title>Token estimation values hardcoded multiple places - inaccurate rate limiting</issue_title>
<issue_description>## 问题描述
Token 估算值硬编码在多个文件中,导致速率限制不准确。

受影响的文件和位置

  • multi_tenant.py 行 113:estimated_output = 3000(LLM)
  • multi_tenant.py 行 313:estimated_output = 500(VLM)
  • multi_tenant.py 行 339:max_tokens = 500
  • deepseek_ocr_client.py 行 285, 352

问题

  • 无法按不同模型调整估算值
  • 同一个值 "3000" 在多处重复,维护困难
  • 使用小模型时估算过高,并发受限;使用大模型时估算不足

解决方案

应该提取到 .env 配置:

  • LLM_ESTIMATED_OUTPUT_TOKENS
  • VLM_ESTIMATED_OUTPUT_TOKENS 等</issue_description>

Comments on the Issue (you are @copilot in this section)


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

- 在 LLMConfig 中添加 estimated_output_tokens, vlm_estimated_output_tokens, vlm_max_tokens, vlm_image_tokens_estimate
- 在 DeepSeekOCRConfig 中添加 image_tokens_estimate
- 更新 multi_tenant.py 使用配置值替代硬编码
- 更新 deepseek_ocr_client.py 使用配置值替代硬编码
- 更新 env.example 添加新的环境变量说明
Closes #issue
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Fix hardcoded token estimation values for accurate rate limitingfeat: Extract hardcoded token estimation values to configurationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:47

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)

这两个环境变量来算

Comment threadenv.example Outdated
Comment on lines +34 to +41
# --- LLM Token 估算配置(用于速率限制) ---
# 估算输出 tokens 数量,用于速率限制计算
# 如果估算过高,并发会受限;如果估算不足,可能触发 429 错误
# LLM_ESTIMATED_OUTPUT_TOKENS=3000 # LLM 输出估算(实体提取约 3000 tokens,默认 3000)
# LLM_VLM_ESTIMATED_OUTPUT_TOKENS=500 # VLM 输出估算(图片描述较短,默认 500)
# LLM_VLM_MAX_TOKENS=500 # VLM API 最大输出 tokens(默认 500)
# LLM_VLM_IMAGE_TOKENS_ESTIMATE=200 # VLM 图片输入估算 tokens(默认 200)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个变量值和LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)
是不是有冲突?明明是自动计算,为什么还要这些估算值? @copilot@claude

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确实有冲突。已在 e2cb6e7 中撤销这些配置项。

rate_limiter.pyavg_tokens_map 已内置各服务的 token 估算值,通过 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 自动计算并发数,不需要额外暴露这些内部值。

Comment threadsrc/multi_tenant.py Outdated
Comment on lines +67 to +72
# Token 估算配置
self.llm_estimated_output_tokens = config.llm.estimated_output_tokens
self.vlm_estimated_output_tokens = config.llm.vlm_estimated_output_tokens
self.vlm_max_tokens = config.llm.vlm_max_tokens
self.vlm_image_tokens_estimate = config.llm.vlm_image_tokens_estimate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题同.env.example

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 e2cb6e7 中撤销相关修改。

根据用户反馈,这些估算值应该通过 LLM_REQUESTS_PER_MINUTE 和 LLM_TOKENS_PER_MINUTE 自动计算,不需要额外的配置项。
rate_limiter.py 中已有 avg_tokens_map 用于自动计算并发数。
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)...

已在 e2cb6e7 中撤销所有修改。

速率限制的并发数确实应该通过 LLM_REQUESTS_PER_MINUTELLM_TOKENS_PER_MINUTE 自动计算,rate_limiter.py 中的 avg_tokens_map 已包含各服务的内部 token 估算值用于计算。

CopilotAI changed the title feat: Extract hardcoded token estimation values to configurationrevert: Token estimation values should use existing auto-calculationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 10:05
@BukeLy

Copy link
Copy Markdown
Owner

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

建议改进方案:将分散的硬编码值统一到 rate_limiter.pyavg_tokens_map 中复用。

当前问题:

  • multi_tenant.py:113 - estimated_output = 3000
  • multi_tenant.py:312-313 - image_tokens = 200, estimated_output = 500
  • deepseek_ocr_client.py:285,352 - 1000 + self.config.max_tokens

改进方案:

  1. rate_limiter.py 导出 avg_tokens_map 或创建 get_avg_tokens(service) 函数
  2. multi_tenant.pydeepseek_ocr_client.py 调用此函数获取估算值
  3. 这样只需维护 rate_limiter.py 一处,无需添加新环境变量

这样可保持与 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 的自动计算逻辑一致,同时解决硬编码分散的问题。

是否需要我实现这个方案?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token estimation values hardcoded multiple places - inaccurate rate limiting

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

revert: Token estimation values should use existing auto-calculation - #29

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values
Draft

revert: Token estimation values should use existing auto-calculation#29
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

After review, the proposed changes to extract hardcoded token estimation values to configuration have been reverted.

Reason

The rate limiting concurrency is already automatically calculated using existing environment variables:

  • LLM_REQUESTS_PER_MINUTE (default: 800)
  • LLM_TOKENS_PER_MINUTE (default: 40000)

The rate_limiter.py module contains an internal avg_tokens_map that provides token estimation values for each service (LLM: 3500, Embedding: 20000, Rerank: 500, DS_OCR: 3500). These internal values are used to auto-calculate the appropriate concurrency level.

Conclusion

No additional configuration is needed. The existing LLM_REQUESTS_PER_MINUTE and LLM_TOKENS_PER_MINUTE environment variables, combined with the internal token estimation in rate_limiter.py, already provide the correct behavior for rate limiting.

Original prompt

This section details on the original issue you should resolve

<issue_title>Token estimation values hardcoded multiple places - inaccurate rate limiting</issue_title>
<issue_description>## 问题描述
Token 估算值硬编码在多个文件中,导致速率限制不准确。

受影响的文件和位置

  • multi_tenant.py 行 113:estimated_output = 3000(LLM)
  • multi_tenant.py 行 313:estimated_output = 500(VLM)
  • multi_tenant.py 行 339:max_tokens = 500
  • deepseek_ocr_client.py 行 285, 352

问题

  • 无法按不同模型调整估算值
  • 同一个值 "3000" 在多处重复,维护困难
  • 使用小模型时估算过高,并发受限;使用大模型时估算不足

解决方案

应该提取到 .env 配置:

  • LLM_ESTIMATED_OUTPUT_TOKENS
  • VLM_ESTIMATED_OUTPUT_TOKENS 等</issue_description>

Comments on the Issue (you are @copilot in this section)


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

- 在 LLMConfig 中添加 estimated_output_tokens, vlm_estimated_output_tokens, vlm_max_tokens, vlm_image_tokens_estimate
- 在 DeepSeekOCRConfig 中添加 image_tokens_estimate
- 更新 multi_tenant.py 使用配置值替代硬编码
- 更新 deepseek_ocr_client.py 使用配置值替代硬编码
- 更新 env.example 添加新的环境变量说明
Closes #issue
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Fix hardcoded token estimation values for accurate rate limitingfeat: Extract hardcoded token estimation values to configurationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:47

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)

这两个环境变量来算

Comment threadenv.example Outdated
Comment on lines +34 to +41
# --- LLM Token 估算配置(用于速率限制) ---
# 估算输出 tokens 数量,用于速率限制计算
# 如果估算过高,并发会受限;如果估算不足,可能触发 429 错误
# LLM_ESTIMATED_OUTPUT_TOKENS=3000 # LLM 输出估算(实体提取约 3000 tokens,默认 3000)
# LLM_VLM_ESTIMATED_OUTPUT_TOKENS=500 # VLM 输出估算(图片描述较短,默认 500)
# LLM_VLM_MAX_TOKENS=500 # VLM API 最大输出 tokens(默认 500)
# LLM_VLM_IMAGE_TOKENS_ESTIMATE=200 # VLM 图片输入估算 tokens(默认 200)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个变量值和LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)
是不是有冲突?明明是自动计算,为什么还要这些估算值? @copilot@claude

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确实有冲突。已在 e2cb6e7 中撤销这些配置项。

rate_limiter.pyavg_tokens_map 已内置各服务的 token 估算值,通过 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 自动计算并发数,不需要额外暴露这些内部值。

Comment threadsrc/multi_tenant.py Outdated
Comment on lines +67 to +72
# Token 估算配置
self.llm_estimated_output_tokens = config.llm.estimated_output_tokens
self.vlm_estimated_output_tokens = config.llm.vlm_estimated_output_tokens
self.vlm_max_tokens = config.llm.vlm_max_tokens
self.vlm_image_tokens_estimate = config.llm.vlm_image_tokens_estimate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题同.env.example

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 e2cb6e7 中撤销相关修改。

根据用户反馈,这些估算值应该通过 LLM_REQUESTS_PER_MINUTE 和 LLM_TOKENS_PER_MINUTE 自动计算,不需要额外的配置项。
rate_limiter.py 中已有 avg_tokens_map 用于自动计算并发数。
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)...

已在 e2cb6e7 中撤销所有修改。

速率限制的并发数确实应该通过 LLM_REQUESTS_PER_MINUTELLM_TOKENS_PER_MINUTE 自动计算,rate_limiter.py 中的 avg_tokens_map 已包含各服务的内部 token 估算值用于计算。

CopilotAI changed the title feat: Extract hardcoded token estimation values to configurationrevert: Token estimation values should use existing auto-calculationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 10:05
@BukeLy

Copy link
Copy Markdown
Owner

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

建议改进方案:将分散的硬编码值统一到 rate_limiter.pyavg_tokens_map 中复用。

当前问题:

  • multi_tenant.py:113 - estimated_output = 3000
  • multi_tenant.py:312-313 - image_tokens = 200, estimated_output = 500
  • deepseek_ocr_client.py:285,352 - 1000 + self.config.max_tokens

改进方案:

  1. rate_limiter.py 导出 avg_tokens_map 或创建 get_avg_tokens(service) 函数
  2. multi_tenant.pydeepseek_ocr_client.py 调用此函数获取估算值
  3. 这样只需维护 rate_limiter.py 一处,无需添加新环境变量

这样可保持与 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 的自动计算逻辑一致,同时解决硬编码分散的问题。

是否需要我实现这个方案?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token estimation values hardcoded multiple places - inaccurate rate limiting

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

revert: Token estimation values should use existing auto-calculation - #29

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values
Draft

revert: Token estimation values should use existing auto-calculation#29
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

After review, the proposed changes to extract hardcoded token estimation values to configuration have been reverted.

Reason

The rate limiting concurrency is already automatically calculated using existing environment variables:

  • LLM_REQUESTS_PER_MINUTE (default: 800)
  • LLM_TOKENS_PER_MINUTE (default: 40000)

The rate_limiter.py module contains an internal avg_tokens_map that provides token estimation values for each service (LLM: 3500, Embedding: 20000, Rerank: 500, DS_OCR: 3500). These internal values are used to auto-calculate the appropriate concurrency level.

Conclusion

No additional configuration is needed. The existing LLM_REQUESTS_PER_MINUTE and LLM_TOKENS_PER_MINUTE environment variables, combined with the internal token estimation in rate_limiter.py, already provide the correct behavior for rate limiting.

Original prompt

This section details on the original issue you should resolve

<issue_title>Token estimation values hardcoded multiple places - inaccurate rate limiting</issue_title>
<issue_description>## 问题描述
Token 估算值硬编码在多个文件中,导致速率限制不准确。

受影响的文件和位置

  • multi_tenant.py 行 113:estimated_output = 3000(LLM)
  • multi_tenant.py 行 313:estimated_output = 500(VLM)
  • multi_tenant.py 行 339:max_tokens = 500
  • deepseek_ocr_client.py 行 285, 352

问题

  • 无法按不同模型调整估算值
  • 同一个值 "3000" 在多处重复,维护困难
  • 使用小模型时估算过高,并发受限;使用大模型时估算不足

解决方案

应该提取到 .env 配置:

  • LLM_ESTIMATED_OUTPUT_TOKENS
  • VLM_ESTIMATED_OUTPUT_TOKENS 等</issue_description>

Comments on the Issue (you are @copilot in this section)


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

- 在 LLMConfig 中添加 estimated_output_tokens, vlm_estimated_output_tokens, vlm_max_tokens, vlm_image_tokens_estimate
- 在 DeepSeekOCRConfig 中添加 image_tokens_estimate
- 更新 multi_tenant.py 使用配置值替代硬编码
- 更新 deepseek_ocr_client.py 使用配置值替代硬编码
- 更新 env.example 添加新的环境变量说明
Closes #issue
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Fix hardcoded token estimation values for accurate rate limitingfeat: Extract hardcoded token estimation values to configurationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:47

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)

这两个环境变量来算

Comment threadenv.example Outdated
Comment on lines +34 to +41
# --- LLM Token 估算配置(用于速率限制) ---
# 估算输出 tokens 数量,用于速率限制计算
# 如果估算过高,并发会受限;如果估算不足,可能触发 429 错误
# LLM_ESTIMATED_OUTPUT_TOKENS=3000 # LLM 输出估算(实体提取约 3000 tokens,默认 3000)
# LLM_VLM_ESTIMATED_OUTPUT_TOKENS=500 # VLM 输出估算(图片描述较短,默认 500)
# LLM_VLM_MAX_TOKENS=500 # VLM API 最大输出 tokens(默认 500)
# LLM_VLM_IMAGE_TOKENS_ESTIMATE=200 # VLM 图片输入估算 tokens(默认 200)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个变量值和LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)
是不是有冲突?明明是自动计算,为什么还要这些估算值? @copilot@claude

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确实有冲突。已在 e2cb6e7 中撤销这些配置项。

rate_limiter.pyavg_tokens_map 已内置各服务的 token 估算值,通过 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 自动计算并发数,不需要额外暴露这些内部值。

Comment threadsrc/multi_tenant.py Outdated
Comment on lines +67 to +72
# Token 估算配置
self.llm_estimated_output_tokens = config.llm.estimated_output_tokens
self.vlm_estimated_output_tokens = config.llm.vlm_estimated_output_tokens
self.vlm_max_tokens = config.llm.vlm_max_tokens
self.vlm_image_tokens_estimate = config.llm.vlm_image_tokens_estimate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题同.env.example

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 e2cb6e7 中撤销相关修改。

根据用户反馈,这些估算值应该通过 LLM_REQUESTS_PER_MINUTE 和 LLM_TOKENS_PER_MINUTE 自动计算,不需要额外的配置项。
rate_limiter.py 中已有 avg_tokens_map 用于自动计算并发数。
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)...

已在 e2cb6e7 中撤销所有修改。

速率限制的并发数确实应该通过 LLM_REQUESTS_PER_MINUTELLM_TOKENS_PER_MINUTE 自动计算,rate_limiter.py 中的 avg_tokens_map 已包含各服务的内部 token 估算值用于计算。

CopilotAI changed the title feat: Extract hardcoded token estimation values to configurationrevert: Token estimation values should use existing auto-calculationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 10:05
@BukeLy

Copy link
Copy Markdown
Owner

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

建议改进方案:将分散的硬编码值统一到 rate_limiter.pyavg_tokens_map 中复用。

当前问题:

  • multi_tenant.py:113 - estimated_output = 3000
  • multi_tenant.py:312-313 - image_tokens = 200, estimated_output = 500
  • deepseek_ocr_client.py:285,352 - 1000 + self.config.max_tokens

改进方案:

  1. rate_limiter.py 导出 avg_tokens_map 或创建 get_avg_tokens(service) 函数
  2. multi_tenant.pydeepseek_ocr_client.py 调用此函数获取估算值
  3. 这样只需维护 rate_limiter.py 一处,无需添加新环境变量

这样可保持与 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 的自动计算逻辑一致,同时解决硬编码分散的问题。

是否需要我实现这个方案?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token estimation values hardcoded multiple places - inaccurate rate limiting

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

revert: Token estimation values should use existing auto-calculation - #29

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values
Draft

revert: Token estimation values should use existing auto-calculation#29
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

After review, the proposed changes to extract hardcoded token estimation values to configuration have been reverted.

Reason

The rate limiting concurrency is already automatically calculated using existing environment variables:

  • LLM_REQUESTS_PER_MINUTE (default: 800)
  • LLM_TOKENS_PER_MINUTE (default: 40000)

The rate_limiter.py module contains an internal avg_tokens_map that provides token estimation values for each service (LLM: 3500, Embedding: 20000, Rerank: 500, DS_OCR: 3500). These internal values are used to auto-calculate the appropriate concurrency level.

Conclusion

No additional configuration is needed. The existing LLM_REQUESTS_PER_MINUTE and LLM_TOKENS_PER_MINUTE environment variables, combined with the internal token estimation in rate_limiter.py, already provide the correct behavior for rate limiting.

Original prompt

This section details on the original issue you should resolve

<issue_title>Token estimation values hardcoded multiple places - inaccurate rate limiting</issue_title>
<issue_description>## 问题描述
Token 估算值硬编码在多个文件中,导致速率限制不准确。

受影响的文件和位置

  • multi_tenant.py 行 113:estimated_output = 3000(LLM)
  • multi_tenant.py 行 313:estimated_output = 500(VLM)
  • multi_tenant.py 行 339:max_tokens = 500
  • deepseek_ocr_client.py 行 285, 352

问题

  • 无法按不同模型调整估算值
  • 同一个值 "3000" 在多处重复,维护困难
  • 使用小模型时估算过高,并发受限;使用大模型时估算不足

解决方案

应该提取到 .env 配置:

  • LLM_ESTIMATED_OUTPUT_TOKENS
  • VLM_ESTIMATED_OUTPUT_TOKENS 等</issue_description>

Comments on the Issue (you are @copilot in this section)


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

- 在 LLMConfig 中添加 estimated_output_tokens, vlm_estimated_output_tokens, vlm_max_tokens, vlm_image_tokens_estimate
- 在 DeepSeekOCRConfig 中添加 image_tokens_estimate
- 更新 multi_tenant.py 使用配置值替代硬编码
- 更新 deepseek_ocr_client.py 使用配置值替代硬编码
- 更新 env.example 添加新的环境变量说明
Closes #issue
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Fix hardcoded token estimation values for accurate rate limitingfeat: Extract hardcoded token estimation values to configurationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:47

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)

这两个环境变量来算

Comment threadenv.example Outdated
Comment on lines +34 to +41
# --- LLM Token 估算配置(用于速率限制) ---
# 估算输出 tokens 数量,用于速率限制计算
# 如果估算过高,并发会受限;如果估算不足,可能触发 429 错误
# LLM_ESTIMATED_OUTPUT_TOKENS=3000 # LLM 输出估算(实体提取约 3000 tokens,默认 3000)
# LLM_VLM_ESTIMATED_OUTPUT_TOKENS=500 # VLM 输出估算(图片描述较短,默认 500)
# LLM_VLM_MAX_TOKENS=500 # VLM API 最大输出 tokens(默认 500)
# LLM_VLM_IMAGE_TOKENS_ESTIMATE=200 # VLM 图片输入估算 tokens(默认 200)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个变量值和LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)
是不是有冲突?明明是自动计算,为什么还要这些估算值? @copilot@claude

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确实有冲突。已在 e2cb6e7 中撤销这些配置项。

rate_limiter.pyavg_tokens_map 已内置各服务的 token 估算值,通过 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 自动计算并发数,不需要额外暴露这些内部值。

Comment threadsrc/multi_tenant.py Outdated
Comment on lines +67 to +72
# Token 估算配置
self.llm_estimated_output_tokens = config.llm.estimated_output_tokens
self.vlm_estimated_output_tokens = config.llm.vlm_estimated_output_tokens
self.vlm_max_tokens = config.llm.vlm_max_tokens
self.vlm_image_tokens_estimate = config.llm.vlm_image_tokens_estimate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题同.env.example

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 e2cb6e7 中撤销相关修改。

根据用户反馈,这些估算值应该通过 LLM_REQUESTS_PER_MINUTE 和 LLM_TOKENS_PER_MINUTE 自动计算,不需要额外的配置项。
rate_limiter.py 中已有 avg_tokens_map 用于自动计算并发数。
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)...

已在 e2cb6e7 中撤销所有修改。

速率限制的并发数确实应该通过 LLM_REQUESTS_PER_MINUTELLM_TOKENS_PER_MINUTE 自动计算,rate_limiter.py 中的 avg_tokens_map 已包含各服务的内部 token 估算值用于计算。

CopilotAI changed the title feat: Extract hardcoded token estimation values to configurationrevert: Token estimation values should use existing auto-calculationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 10:05
@BukeLy

Copy link
Copy Markdown
Owner

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

建议改进方案:将分散的硬编码值统一到 rate_limiter.pyavg_tokens_map 中复用。

当前问题:

  • multi_tenant.py:113 - estimated_output = 3000
  • multi_tenant.py:312-313 - image_tokens = 200, estimated_output = 500
  • deepseek_ocr_client.py:285,352 - 1000 + self.config.max_tokens

改进方案:

  1. rate_limiter.py 导出 avg_tokens_map 或创建 get_avg_tokens(service) 函数
  2. multi_tenant.pydeepseek_ocr_client.py 调用此函数获取估算值
  3. 这样只需维护 rate_limiter.py 一处,无需添加新环境变量

这样可保持与 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 的自动计算逻辑一致,同时解决硬编码分散的问题。

是否需要我实现这个方案?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token estimation values hardcoded multiple places - inaccurate rate limiting

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

revert: Token estimation values should use existing auto-calculation - #29

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values
Draft

revert: Token estimation values should use existing auto-calculation#29
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/fix-token-estimation-values

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

After review, the proposed changes to extract hardcoded token estimation values to configuration have been reverted.

Reason

The rate limiting concurrency is already automatically calculated using existing environment variables:

  • LLM_REQUESTS_PER_MINUTE (default: 800)
  • LLM_TOKENS_PER_MINUTE (default: 40000)

The rate_limiter.py module contains an internal avg_tokens_map that provides token estimation values for each service (LLM: 3500, Embedding: 20000, Rerank: 500, DS_OCR: 3500). These internal values are used to auto-calculate the appropriate concurrency level.

Conclusion

No additional configuration is needed. The existing LLM_REQUESTS_PER_MINUTE and LLM_TOKENS_PER_MINUTE environment variables, combined with the internal token estimation in rate_limiter.py, already provide the correct behavior for rate limiting.

Original prompt

This section details on the original issue you should resolve

<issue_title>Token estimation values hardcoded multiple places - inaccurate rate limiting</issue_title>
<issue_description>## 问题描述
Token 估算值硬编码在多个文件中,导致速率限制不准确。

受影响的文件和位置

  • multi_tenant.py 行 113:estimated_output = 3000(LLM)
  • multi_tenant.py 行 313:estimated_output = 500(VLM)
  • multi_tenant.py 行 339:max_tokens = 500
  • deepseek_ocr_client.py 行 285, 352

问题

  • 无法按不同模型调整估算值
  • 同一个值 "3000" 在多处重复,维护困难
  • 使用小模型时估算过高,并发受限;使用大模型时估算不足

解决方案

应该提取到 .env 配置:

  • LLM_ESTIMATED_OUTPUT_TOKENS
  • VLM_ESTIMATED_OUTPUT_TOKENS 等</issue_description>

Comments on the Issue (you are @copilot in this section)


✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.

- 在 LLMConfig 中添加 estimated_output_tokens, vlm_estimated_output_tokens, vlm_max_tokens, vlm_image_tokens_estimate
- 在 DeepSeekOCRConfig 中添加 image_tokens_estimate
- 更新 multi_tenant.py 使用配置值替代硬编码
- 更新 deepseek_ocr_client.py 使用配置值替代硬编码
- 更新 env.example 添加新的环境变量说明
Closes #issue
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Fix hardcoded token estimation values for accurate rate limitingfeat: Extract hardcoded token estimation values to configurationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:47

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)

这两个环境变量来算

Comment threadenv.example Outdated
Comment on lines +34 to +41
# --- LLM Token 估算配置(用于速率限制) ---
# 估算输出 tokens 数量,用于速率限制计算
# 如果估算过高,并发会受限;如果估算不足,可能触发 429 错误
# LLM_ESTIMATED_OUTPUT_TOKENS=3000 # LLM 输出估算(实体提取约 3000 tokens,默认 3000)
# LLM_VLM_ESTIMATED_OUTPUT_TOKENS=500 # VLM 输出估算(图片描述较短,默认 500)
# LLM_VLM_MAX_TOKENS=500 # VLM API 最大输出 tokens(默认 500)
# LLM_VLM_IMAGE_TOKENS_ESTIMATE=200 # VLM 图片输入估算 tokens(默认 200)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这个变量值和LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)
LLM_TOKENS_PER_MINUTE=40000 # 每分钟最大令牌数(包含输入+输出,默认 40000)
是不是有冲突?明明是自动计算,为什么还要这些估算值? @copilot@claude

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

确实有冲突。已在 e2cb6e7 中撤销这些配置项。

rate_limiter.pyavg_tokens_map 已内置各服务的 token 估算值,通过 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 自动计算并发数,不需要额外暴露这些内部值。

Comment threadsrc/multi_tenant.py Outdated
Comment on lines +67 to +72
# Token 估算配置
self.llm_estimated_output_tokens = config.llm.estimated_output_tokens
self.vlm_estimated_output_tokens = config.llm.vlm_estimated_output_tokens
self.vlm_max_tokens = config.llm.vlm_max_tokens
self.vlm_image_tokens_estimate = config.llm.vlm_image_tokens_estimate

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

问题同.env.example

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已在 e2cb6e7 中撤销相关修改。

根据用户反馈,这些估算值应该通过 LLM_REQUESTS_PER_MINUTE 和 LLM_TOKENS_PER_MINUTE 自动计算,不需要额外的配置项。
rate_limiter.py 中已有 avg_tokens_map 用于自动计算并发数。
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

整个估算的配置就不应该这么加,按照设计目标来说应该是通过

LLM_REQUESTS_PER_MINUTE=800 # 每分钟最大请求数(默认 800)...

已在 e2cb6e7 中撤销所有修改。

速率限制的并发数确实应该通过 LLM_REQUESTS_PER_MINUTELLM_TOKENS_PER_MINUTE 自动计算,rate_limiter.py 中的 avg_tokens_map 已包含各服务的内部 token 估算值用于计算。

CopilotAI changed the title feat: Extract hardcoded token estimation values to configurationrevert: Token estimation values should use existing auto-calculationDec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 10:05
@BukeLy

Copy link
Copy Markdown
Owner

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
ContributorAuthor

@copilot 后续如何改进呢?issue是存在的,请问怎么解决?

建议改进方案:将分散的硬编码值统一到 rate_limiter.pyavg_tokens_map 中复用。

当前问题:

  • multi_tenant.py:113 - estimated_output = 3000
  • multi_tenant.py:312-313 - image_tokens = 200, estimated_output = 500
  • deepseek_ocr_client.py:285,352 - 1000 + self.config.max_tokens

改进方案:

  1. rate_limiter.py 导出 avg_tokens_map 或创建 get_avg_tokens(service) 函数
  2. multi_tenant.pydeepseek_ocr_client.py 调用此函数获取估算值
  3. 这样只需维护 rate_limiter.py 一处,无需添加新环境变量

这样可保持与 LLM_REQUESTS_PER_MINUTE / LLM_TOKENS_PER_MINUTE 的自动计算逻辑一致,同时解决硬编码分散的问题。

是否需要我实现这个方案?

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token estimation values hardcoded multiple places - inaccurate rate limiting

2 participants

@BukeLy