feat: 支持按租户/服务配置平均 token 数以优化并发计算 - #27

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration
Draft

feat: 支持按租户/服务配置平均 token 数以优化并发计算#27
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

平均 Token 数硬编码在 rate_limiter.py 导致并发计算不准确。Insert 和 Query 场景的 token 消耗差异大,固定保守值限制了系统吞吐量。

Changes

  • src/config.py: 各服务配置类新增 avg_tokens_per_request 字段

    • LLM/DS_OCR: 3500 (default)
    • Embedding: 20000
    • Rerank: 500
  • src/rate_limiter.py:

    • get_rate_limiter() 新增 avg_tokens_per_request 参数
    • 提取默认配置到 SERVICE_DEFAULTS 字典
  • src/tenant_config.py: 各服务合并方法传递 avg_tokens_per_request

  • src/multi_tenant.py / src/deepseek_ocr_client.py: 更新调用传参

  • env.example: 添加环境变量说明

配置优先级

  1. 租户配置 (API)
  2. 环境变量 (LLM_AVG_TOKENS_PER_REQUEST)
  3. 代码默认值

使用示例

# Insert 密集场景,降低平均 token 提升并发
LLM_AVG_TOKENS_PER_REQUEST=2500
// 租户 APIPUT /tenants/{tenant_id}/config
{
"llm_config": {
"avg_tokens_per_request": 4000
}
}
Original prompt

This section details on the original issue you should resolve

<issue_title>Average tokens per request hardcoded - inaccurate concurrency calculation</issue_title>
<issue_description>## 问题描述
平均 Token 数硬编码导致并发计算不准确。

受影响的文件

  • src/rate_limiter.py 行 420-423

硬编码值

  • llm: 3500
  • embedding: 500
  • rerank: 500
  • ds_ocr: 3500

问题

这些值因场景而异(Query vs Insert),但被固定为保守值,导致不必要的速率限制,降低系统吞吐量。

解决方案

应该支持按租户/服务配置平均 token 数。</issue_description>

Comments on the Issue (you are @copilot in this section)


💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

CopilotAIand others added 2 commits December 15, 2025 09:43
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Update average tokens per request for better concurrency calculationfeat: 支持按租户/服务配置平均 token 数以优化并发计算Dec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:48

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢?

Comment on lines -83 to +84
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute)
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute),
avg_tokens_per_request=getattr(self.config, 'avg_tokens_per_request', config.ds_ocr.avg_tokens_per_request)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢? @copilot

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这两个参数的作用不同:

  1. tokens_per_minute - API 服务商的 TPM 限制(如硅基流动 40000),用于 RateLimiter 控制每分钟最大 token 消耗,防止 429 错误。

  2. avg_tokens_per_request - 用于启动时预估安全并发数的参数。在 calculate_optimal_concurrent() 中计算:concurrent = TPM / avg_tokens / max_in_flight

虽然 count_tokens() 用 tiktoken 精确计算每个请求的实际 token(运行时扣减 TPM 配额),但 avg_tokens_per_request初始化时预估应允许多少并发请求。这是两个不同用途:

  • 运行时:count_tokens() 精确计算实际消耗
  • 初始化:avg_tokens_per_request 预估并发上限

不同场景(Insert vs Query)的平均 token 差异大,现在可配置这个预估值来优化并发。

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Average tokens per request hardcoded - inaccurate concurrency calculation

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Add copy buttons to all
 blocks\n(function() {\n function addCopyButtons() {\n document.querySelectorAll('pre code').forEach(function(codeBlock) {\n if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;\n codeBlock.parentElement.setAttribute('data-copy-added', 'true');\n \n var btn = document.createElement('button');\n btn.textContent = 'Copy';\n btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';\n btn.onmouseover = function() { this.style.opacity = '1'; };\n btn.onmouseout = function() { this.style.opacity = '0.7'; };\n btn.onclick = function() {\n navigator.clipboard.writeText(codeBlock.textContent).then(function() {\n btn.textContent = 'Copied!';\n setTimeout(function() { btn.textContent = 'Copy'; }, 1500);\n });\n };\n codeBlock.parentElement.style.position = 'relative';\n codeBlock.parentElement.appendChild(btn);\n });\n }\n \n addCopyButtons();\n \n // Re-run on dynamic content\n var observer = new MutationObserver(addCopyButtons);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Add Copy Buttons to Code Blocks");
}
} catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
})();
(function(){
try {
var __m = "github.com";
var __re = new RegExp('^' + "github\\.com" + '
Skip to content

feat: 支持按租户/服务配置平均 token 数以优化并发计算 - #27

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration
Draft

feat: 支持按租户/服务配置平均 token 数以优化并发计算#27
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

平均 Token 数硬编码在 rate_limiter.py 导致并发计算不准确。Insert 和 Query 场景的 token 消耗差异大,固定保守值限制了系统吞吐量。

Changes

  • src/config.py: 各服务配置类新增 avg_tokens_per_request 字段

    • LLM/DS_OCR: 3500 (default)
    • Embedding: 20000
    • Rerank: 500
  • src/rate_limiter.py:

    • get_rate_limiter() 新增 avg_tokens_per_request 参数
    • 提取默认配置到 SERVICE_DEFAULTS 字典
  • src/tenant_config.py: 各服务合并方法传递 avg_tokens_per_request

  • src/multi_tenant.py / src/deepseek_ocr_client.py: 更新调用传参

  • env.example: 添加环境变量说明

配置优先级

  1. 租户配置 (API)
  2. 环境变量 (LLM_AVG_TOKENS_PER_REQUEST)
  3. 代码默认值

使用示例

# Insert 密集场景,降低平均 token 提升并发
LLM_AVG_TOKENS_PER_REQUEST=2500
// 租户 APIPUT /tenants/{tenant_id}/config
{
"llm_config": {
"avg_tokens_per_request": 4000
}
}
Original prompt

This section details on the original issue you should resolve

<issue_title>Average tokens per request hardcoded - inaccurate concurrency calculation</issue_title>
<issue_description>## 问题描述
平均 Token 数硬编码导致并发计算不准确。

受影响的文件

  • src/rate_limiter.py 行 420-423

硬编码值

  • llm: 3500
  • embedding: 500
  • rerank: 500
  • ds_ocr: 3500

问题

这些值因场景而异(Query vs Insert),但被固定为保守值,导致不必要的速率限制,降低系统吞吐量。

解决方案

应该支持按租户/服务配置平均 token 数。</issue_description>

Comments on the Issue (you are @copilot in this section)


💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

CopilotAIand others added 2 commits December 15, 2025 09:43
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Update average tokens per request for better concurrency calculationfeat: 支持按租户/服务配置平均 token 数以优化并发计算Dec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:48

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢?

Comment on lines -83 to +84
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute)
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute),
avg_tokens_per_request=getattr(self.config, 'avg_tokens_per_request', config.ds_ocr.avg_tokens_per_request)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢? @copilot

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这两个参数的作用不同:

  1. tokens_per_minute - API 服务商的 TPM 限制(如硅基流动 40000),用于 RateLimiter 控制每分钟最大 token 消耗,防止 429 错误。

  2. avg_tokens_per_request - 用于启动时预估安全并发数的参数。在 calculate_optimal_concurrent() 中计算:concurrent = TPM / avg_tokens / max_in_flight

虽然 count_tokens() 用 tiktoken 精确计算每个请求的实际 token(运行时扣减 TPM 配额),但 avg_tokens_per_request初始化时预估应允许多少并发请求。这是两个不同用途:

  • 运行时:count_tokens() 精确计算实际消耗
  • 初始化:avg_tokens_per_request 预估并发上限

不同场景(Insert vs Query)的平均 token 差异大,现在可配置这个预估值来优化并发。

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Average tokens per request hardcoded - inaccurate concurrency calculation

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Force GitHub README to respect dark mode\n(function() {\n var style = document.createElement('style');\n style.textContent = '\n .markdown-body {\n color-scheme: dark light;\n }\n .markdown-body pre { background: #161b22 !important; }\n .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; }\n .markdown-body table th, .markdown-body table td { border-color: #30363d !important; }\n .markdown-body img { background: #0d1117; }\n .markdown-body blockquote { border-left-color: #8b949e; }\n .markdown-body hr { border-color: #30363d; }\n ';\n document.head.appendChild(style);\n})();", "GitHub Dark Mode README Fix"); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: 支持按租户/服务配置平均 token 数以优化并发计算 - #27

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration
Draft

feat: 支持按租户/服务配置平均 token 数以优化并发计算#27
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

平均 Token 数硬编码在 rate_limiter.py 导致并发计算不准确。Insert 和 Query 场景的 token 消耗差异大,固定保守值限制了系统吞吐量。

Changes

  • src/config.py: 各服务配置类新增 avg_tokens_per_request 字段

    • LLM/DS_OCR: 3500 (default)
    • Embedding: 20000
    • Rerank: 500
  • src/rate_limiter.py:

    • get_rate_limiter() 新增 avg_tokens_per_request 参数
    • 提取默认配置到 SERVICE_DEFAULTS 字典
  • src/tenant_config.py: 各服务合并方法传递 avg_tokens_per_request

  • src/multi_tenant.py / src/deepseek_ocr_client.py: 更新调用传参

  • env.example: 添加环境变量说明

配置优先级

  1. 租户配置 (API)
  2. 环境变量 (LLM_AVG_TOKENS_PER_REQUEST)
  3. 代码默认值

使用示例

# Insert 密集场景,降低平均 token 提升并发
LLM_AVG_TOKENS_PER_REQUEST=2500
// 租户 APIPUT /tenants/{tenant_id}/config
{
"llm_config": {
"avg_tokens_per_request": 4000
}
}
Original prompt

This section details on the original issue you should resolve

<issue_title>Average tokens per request hardcoded - inaccurate concurrency calculation</issue_title>
<issue_description>## 问题描述
平均 Token 数硬编码导致并发计算不准确。

受影响的文件

  • src/rate_limiter.py 行 420-423

硬编码值

  • llm: 3500
  • embedding: 500
  • rerank: 500
  • ds_ocr: 3500

问题

这些值因场景而异(Query vs Insert),但被固定为保守值,导致不必要的速率限制,降低系统吞吐量。

解决方案

应该支持按租户/服务配置平均 token 数。</issue_description>

Comments on the Issue (you are @copilot in this section)


💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

CopilotAIand others added 2 commits December 15, 2025 09:43
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Update average tokens per request for better concurrency calculationfeat: 支持按租户/服务配置平均 token 数以优化并发计算Dec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:48

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢?

Comment on lines -83 to +84
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute)
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute),
avg_tokens_per_request=getattr(self.config, 'avg_tokens_per_request', config.ds_ocr.avg_tokens_per_request)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢? @copilot

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这两个参数的作用不同:

  1. tokens_per_minute - API 服务商的 TPM 限制(如硅基流动 40000),用于 RateLimiter 控制每分钟最大 token 消耗,防止 429 错误。

  2. avg_tokens_per_request - 用于启动时预估安全并发数的参数。在 calculate_optimal_concurrent() 中计算:concurrent = TPM / avg_tokens / max_in_flight

虽然 count_tokens() 用 tiktoken 精确计算每个请求的实际 token(运行时扣减 TPM 配额),但 avg_tokens_per_request初始化时预估应允许多少并发请求。这是两个不同用途:

  • 运行时:count_tokens() 精确计算实际消耗
  • 初始化:avg_tokens_per_request 预估并发上限

不同场景(Insert vs Query)的平均 token 差异大,现在可配置这个预估值来优化并发。

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Average tokens per request hardcoded - inaccurate concurrency calculation

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Highlight search terms from Google/DuckDuckGo/Bing referrer\n(function() {\n var ref = document.referrer;\n var terms = [];\n \n if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) {\n var url = new URL(ref);\n var q = url.searchParams.get('q') || url.searchParams.get('p');\n if (q) {\n terms = q.split(/\\s+/).filter(function(t) { return t.length > 2; });\n }\n }\n \n if (terms.length === 0) return;\n \n var style = document.createElement('style');\n style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }';\n document.head.appendChild(style);\n \n function highlight(node) {\n if (node.nodeType === 3) { // text node\n var text = node.textContent;\n var found = false;\n terms.forEach(function(term) {\n var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\\]\\\\]/g, '\\\\') + ')', 'gi');\n if (regex.test(text)) {\n found = true;\n var frag = document.createDocumentFragment();\n var parts = text.split(regex);\n parts.forEach(function(part, i) {\n if (i % 2 === 0) {\n frag.appendChild(document.createTextNode(part));\n } else {\n var span = document.createElement('span');\n span.className = 'userscript-highlight';\n span.textContent = part;\n frag.appendChild(span);\n }\n });\n node.parentNode.replaceChild(frag, node);\n }\n });\n } else if (node.nodeType === 1 && node.childNodes) { // element\n var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT'];\n if (!skipTags.includes(node.tagName)) {\n Array.from(node.childNodes).forEach(highlight);\n }\n }\n }\n \n highlight(document.body);\n \n // Re-highlight on dynamic content\n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1 || node.nodeType === 3) highlight(node);\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Highlight Search Terms"); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: 支持按租户/服务配置平均 token 数以优化并发计算 - #27

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration
Draft

feat: 支持按租户/服务配置平均 token 数以优化并发计算#27
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

平均 Token 数硬编码在 rate_limiter.py 导致并发计算不准确。Insert 和 Query 场景的 token 消耗差异大,固定保守值限制了系统吞吐量。

Changes

  • src/config.py: 各服务配置类新增 avg_tokens_per_request 字段

    • LLM/DS_OCR: 3500 (default)
    • Embedding: 20000
    • Rerank: 500
  • src/rate_limiter.py:

    • get_rate_limiter() 新增 avg_tokens_per_request 参数
    • 提取默认配置到 SERVICE_DEFAULTS 字典
  • src/tenant_config.py: 各服务合并方法传递 avg_tokens_per_request

  • src/multi_tenant.py / src/deepseek_ocr_client.py: 更新调用传参

  • env.example: 添加环境变量说明

配置优先级

  1. 租户配置 (API)
  2. 环境变量 (LLM_AVG_TOKENS_PER_REQUEST)
  3. 代码默认值

使用示例

# Insert 密集场景,降低平均 token 提升并发
LLM_AVG_TOKENS_PER_REQUEST=2500
// 租户 APIPUT /tenants/{tenant_id}/config
{
"llm_config": {
"avg_tokens_per_request": 4000
}
}
Original prompt

This section details on the original issue you should resolve

<issue_title>Average tokens per request hardcoded - inaccurate concurrency calculation</issue_title>
<issue_description>## 问题描述
平均 Token 数硬编码导致并发计算不准确。

受影响的文件

  • src/rate_limiter.py 行 420-423

硬编码值

  • llm: 3500
  • embedding: 500
  • rerank: 500
  • ds_ocr: 3500

问题

这些值因场景而异(Query vs Insert),但被固定为保守值,导致不必要的速率限制,降低系统吞吐量。

解决方案

应该支持按租户/服务配置平均 token 数。</issue_description>

Comments on the Issue (you are @copilot in this section)


💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

CopilotAIand others added 2 commits December 15, 2025 09:43
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Update average tokens per request for better concurrency calculationfeat: 支持按租户/服务配置平均 token 数以优化并发计算Dec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:48

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢?

Comment on lines -83 to +84
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute)
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute),
avg_tokens_per_request=getattr(self.config, 'avg_tokens_per_request', config.ds_ocr.avg_tokens_per_request)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢? @copilot

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这两个参数的作用不同:

  1. tokens_per_minute - API 服务商的 TPM 限制(如硅基流动 40000),用于 RateLimiter 控制每分钟最大 token 消耗,防止 429 错误。

  2. avg_tokens_per_request - 用于启动时预估安全并发数的参数。在 calculate_optimal_concurrent() 中计算:concurrent = TPM / avg_tokens / max_in_flight

虽然 count_tokens() 用 tiktoken 精确计算每个请求的实际 token(运行时扣减 TPM 配额),但 avg_tokens_per_request初始化时预估应允许多少并发请求。这是两个不同用途:

  • 运行时:count_tokens() 精确计算实际消耗
  • 初始化:avg_tokens_per_request 预估并发上限

不同场景(Insert vs Query)的平均 token 差异大,现在可配置这个预估值来优化并发。

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Average tokens per request hardcoded - inaccurate concurrency calculation

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Strip utm_, fbclid, gclid, etc. from all links on page\n(function() {\n var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content',\n 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid',\n 'ref', 'ref_src', 'source', 'medium', 'campaign'];\n \n function cleanUrl(url) {\n try {\n var u = new URL(url, window.location.origin);\n var changed = false;\n trackingParams.forEach(function(p) {\n if (u.searchParams.has(p)) {\n u.searchParams.delete(p);\n changed = true;\n }\n });\n return changed ? u.toString() : url;\n } catch (e) {\n return url;\n }\n }\n \n function cleanLinks() {\n document.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n \n cleanLinks();\n \n var observer = new MutationObserver(function(mutations) {\n mutations.forEach(function(m) {\n m.addedNodes.forEach(function(node) {\n if (node.nodeType === 1) {\n if (node.tagName === 'A') cleanLinks();\n node.querySelectorAll('a[href]').forEach(function(a) {\n var clean = cleanUrl(a.href);\n if (clean !== a.href) a.href = clean;\n });\n }\n });\n });\n });\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "Remove Tracking Parameters from Links"); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
Skip to content

feat: 支持按租户/服务配置平均 token 数以优化并发计算 - #27

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration
Draft

feat: 支持按租户/服务配置平均 token 数以优化并发计算#27
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

平均 Token 数硬编码在 rate_limiter.py 导致并发计算不准确。Insert 和 Query 场景的 token 消耗差异大,固定保守值限制了系统吞吐量。

Changes

  • src/config.py: 各服务配置类新增 avg_tokens_per_request 字段

    • LLM/DS_OCR: 3500 (default)
    • Embedding: 20000
    • Rerank: 500
  • src/rate_limiter.py:

    • get_rate_limiter() 新增 avg_tokens_per_request 参数
    • 提取默认配置到 SERVICE_DEFAULTS 字典
  • src/tenant_config.py: 各服务合并方法传递 avg_tokens_per_request

  • src/multi_tenant.py / src/deepseek_ocr_client.py: 更新调用传参

  • env.example: 添加环境变量说明

配置优先级

  1. 租户配置 (API)
  2. 环境变量 (LLM_AVG_TOKENS_PER_REQUEST)
  3. 代码默认值

使用示例

# Insert 密集场景,降低平均 token 提升并发
LLM_AVG_TOKENS_PER_REQUEST=2500
// 租户 APIPUT /tenants/{tenant_id}/config
{
"llm_config": {
"avg_tokens_per_request": 4000
}
}
Original prompt

This section details on the original issue you should resolve

<issue_title>Average tokens per request hardcoded - inaccurate concurrency calculation</issue_title>
<issue_description>## 问题描述
平均 Token 数硬编码导致并发计算不准确。

受影响的文件

  • src/rate_limiter.py 行 420-423

硬编码值

  • llm: 3500
  • embedding: 500
  • rerank: 500
  • ds_ocr: 3500

问题

这些值因场景而异(Query vs Insert),但被固定为保守值,导致不必要的速率限制,降低系统吞吐量。

解决方案

应该支持按租户/服务配置平均 token 数。</issue_description>

Comments on the Issue (you are @copilot in this section)


💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

CopilotAIand others added 2 commits December 15, 2025 09:43
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Update average tokens per request for better concurrency calculationfeat: 支持按租户/服务配置平均 token 数以优化并发计算Dec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:48

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢?

Comment on lines -83 to +84
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute)
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute),
avg_tokens_per_request=getattr(self.config, 'avg_tokens_per_request', config.ds_ocr.avg_tokens_per_request)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢? @copilot

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这两个参数的作用不同:

  1. tokens_per_minute - API 服务商的 TPM 限制(如硅基流动 40000),用于 RateLimiter 控制每分钟最大 token 消耗,防止 429 错误。

  2. avg_tokens_per_request - 用于启动时预估安全并发数的参数。在 calculate_optimal_concurrent() 中计算:concurrent = TPM / avg_tokens / max_in_flight

虽然 count_tokens() 用 tiktoken 精确计算每个请求的实际 token(运行时扣减 TPM 配额),但 avg_tokens_per_request初始化时预估应允许多少并发请求。这是两个不同用途:

  • 运行时:count_tokens() 精确计算实际消耗
  • 初始化:avg_tokens_per_request 预估并发上限

不同场景(Insert vs Query)的平均 token 差异大,现在可配置这个预估值来优化并发。

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Average tokens per request hardcoded - inaccurate concurrency calculation

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Auto-enable theater mode on YouTube\n(function() {\n function tryTheater() {\n var btn = document.querySelector('button[aria-label=\"Theater mode\"], ytd-player #player button[title=\"Theater mode\"]');\n if (btn && !btn.classList.contains('activated')) {\n btn.click();\n }\n }\n \n // Try immediately\n tryTheater();\n \n // Try after navigation (SPA)\n var lastUrl = location.href;\n setInterval(function() {\n if (location.href !== lastUrl) {\n lastUrl = location.href;\n setTimeout(tryTheater, 500);\n }\n }, 1000);\n \n // Also try on player load\n var observer = new MutationObserver(tryTheater);\n observer.observe(document.body, { childList: true, subtree: true });\n})();", "YouTube Theater Mode Default"); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: 支持按租户/服务配置平均 token 数以优化并发计算 - #27

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration
Draft

feat: 支持按租户/服务配置平均 token 数以优化并发计算#27
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

平均 Token 数硬编码在 rate_limiter.py 导致并发计算不准确。Insert 和 Query 场景的 token 消耗差异大,固定保守值限制了系统吞吐量。

Changes

  • src/config.py: 各服务配置类新增 avg_tokens_per_request 字段

    • LLM/DS_OCR: 3500 (default)
    • Embedding: 20000
    • Rerank: 500
  • src/rate_limiter.py:

    • get_rate_limiter() 新增 avg_tokens_per_request 参数
    • 提取默认配置到 SERVICE_DEFAULTS 字典
  • src/tenant_config.py: 各服务合并方法传递 avg_tokens_per_request

  • src/multi_tenant.py / src/deepseek_ocr_client.py: 更新调用传参

  • env.example: 添加环境变量说明

配置优先级

  1. 租户配置 (API)
  2. 环境变量 (LLM_AVG_TOKENS_PER_REQUEST)
  3. 代码默认值

使用示例

# Insert 密集场景,降低平均 token 提升并发
LLM_AVG_TOKENS_PER_REQUEST=2500
// 租户 APIPUT /tenants/{tenant_id}/config
{
"llm_config": {
"avg_tokens_per_request": 4000
}
}
Original prompt

This section details on the original issue you should resolve

<issue_title>Average tokens per request hardcoded - inaccurate concurrency calculation</issue_title>
<issue_description>## 问题描述
平均 Token 数硬编码导致并发计算不准确。

受影响的文件

  • src/rate_limiter.py 行 420-423

硬编码值

  • llm: 3500
  • embedding: 500
  • rerank: 500
  • ds_ocr: 3500

问题

这些值因场景而异(Query vs Insert),但被固定为保守值,导致不必要的速率限制,降低系统吞吐量。

解决方案

应该支持按租户/服务配置平均 token 数。</issue_description>

Comments on the Issue (you are @copilot in this section)


💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

CopilotAIand others added 2 commits December 15, 2025 09:43
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Update average tokens per request for better concurrency calculationfeat: 支持按租户/服务配置平均 token 数以优化并发计算Dec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:48

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢?

Comment on lines -83 to +84
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute)
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute),
avg_tokens_per_request=getattr(self.config, 'avg_tokens_per_request', config.ds_ocr.avg_tokens_per_request)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢? @copilot

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这两个参数的作用不同:

  1. tokens_per_minute - API 服务商的 TPM 限制(如硅基流动 40000),用于 RateLimiter 控制每分钟最大 token 消耗,防止 429 错误。

  2. avg_tokens_per_request - 用于启动时预估安全并发数的参数。在 calculate_optimal_concurrent() 中计算:concurrent = TPM / avg_tokens / max_in_flight

虽然 count_tokens() 用 tiktoken 精确计算每个请求的实际 token(运行时扣减 TPM 配额),但 avg_tokens_per_request初始化时预估应允许多少并发请求。这是两个不同用途:

  • 运行时:count_tokens() 精确计算实际消耗
  • 初始化:avg_tokens_per_request 预估并发上限

不同场景(Insert vs Query)的平均 token 差异大,现在可配置这个预估值来优化并发。

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Average tokens per request hardcoded - inaccurate concurrency calculation

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Remove or un-stick sticky/fixed headers that block content\n(function() {\n function unstick() {\n document.querySelectorAll('header, nav, [role=\"banner\"], .header, .navbar, .sticky, .fixed-top, [style*=\"position: fixed\"], [style*=\"position:sticky\"]').forEach(function(el) {\n if (el.style.position === 'fixed' || el.style.position === 'sticky' || \n getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') {\n el.style.position = 'static';\n el.style.top = 'auto';\n el.style.zIndex = 'auto';\n }\n });\n }\n \n unstick();\n \n var observer = new MutationObserver(unstick);\n observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] });\n})();", "Kill Sticky Headers"); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
Skip to content

feat: 支持按租户/服务配置平均 token 数以优化并发计算 - #27

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration
Draft

feat: 支持按租户/服务配置平均 token 数以优化并发计算#27
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

平均 Token 数硬编码在 rate_limiter.py 导致并发计算不准确。Insert 和 Query 场景的 token 消耗差异大,固定保守值限制了系统吞吐量。

Changes

  • src/config.py: 各服务配置类新增 avg_tokens_per_request 字段

    • LLM/DS_OCR: 3500 (default)
    • Embedding: 20000
    • Rerank: 500
  • src/rate_limiter.py:

    • get_rate_limiter() 新增 avg_tokens_per_request 参数
    • 提取默认配置到 SERVICE_DEFAULTS 字典
  • src/tenant_config.py: 各服务合并方法传递 avg_tokens_per_request

  • src/multi_tenant.py / src/deepseek_ocr_client.py: 更新调用传参

  • env.example: 添加环境变量说明

配置优先级

  1. 租户配置 (API)
  2. 环境变量 (LLM_AVG_TOKENS_PER_REQUEST)
  3. 代码默认值

使用示例

# Insert 密集场景,降低平均 token 提升并发
LLM_AVG_TOKENS_PER_REQUEST=2500
// 租户 APIPUT /tenants/{tenant_id}/config
{
"llm_config": {
"avg_tokens_per_request": 4000
}
}
Original prompt

This section details on the original issue you should resolve

<issue_title>Average tokens per request hardcoded - inaccurate concurrency calculation</issue_title>
<issue_description>## 问题描述
平均 Token 数硬编码导致并发计算不准确。

受影响的文件

  • src/rate_limiter.py 行 420-423

硬编码值

  • llm: 3500
  • embedding: 500
  • rerank: 500
  • ds_ocr: 3500

问题

这些值因场景而异(Query vs Insert),但被固定为保守值,导致不必要的速率限制,降低系统吞吐量。

解决方案

应该支持按租户/服务配置平均 token 数。</issue_description>

Comments on the Issue (you are @copilot in this section)


💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

CopilotAIand others added 2 commits December 15, 2025 09:43
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Update average tokens per request for better concurrency calculationfeat: 支持按租户/服务配置平均 token 数以优化并发计算Dec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:48

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢?

Comment on lines -83 to +84
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute)
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute),
avg_tokens_per_request=getattr(self.config, 'avg_tokens_per_request', config.ds_ocr.avg_tokens_per_request)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢? @copilot

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这两个参数的作用不同:

  1. tokens_per_minute - API 服务商的 TPM 限制(如硅基流动 40000),用于 RateLimiter 控制每分钟最大 token 消耗,防止 429 错误。

  2. avg_tokens_per_request - 用于启动时预估安全并发数的参数。在 calculate_optimal_concurrent() 中计算:concurrent = TPM / avg_tokens / max_in_flight

虽然 count_tokens() 用 tiktoken 精确计算每个请求的实际 token(运行时扣减 TPM 配额),但 avg_tokens_per_request初始化时预估应允许多少并发请求。这是两个不同用途:

  • 运行时:count_tokens() 精确计算实际消耗
  • 初始化:avg_tokens_per_request 预估并发上限

不同场景(Insert vs Query)的平均 token 差异大,现在可配置这个预估值来优化并发。

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Average tokens per request hardcoded - inaccurate concurrency calculation

2 participants

@BukeLy
, 'i'); if (__m === '*' || __re.test(location.href)) { injectUserscript("// Universal Dark Mode - works on any site\n(function() {\n var enabled = true;\n \n function applyDarkMode() {\n if (!enabled) return;\n \n // Create style element if it doesn't exist\n var style = document.getElementById('universal-dark-mode-style');\n if (!style) {\n style = document.createElement('style');\n style.id = 'universal-dark-mode-style';\n document.head.appendChild(style);\n }\n \n // Dark mode CSS - inverts colors but preserves images/video\n style.textContent = '\n /* Invert everything except media */\n html {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #1a1a2e !important;\n }\n \n /* Restore images, videos, iframes, canvas */\n img, video, iframe, canvas, svg, picture, [style*=\"background-image\"] {\n filter: invert(1) hue-rotate(180deg) !important;\n }\n \n /* Preserve specific elements that should not be inverted */\n .no-dark-mode, .no-dark-mode *,\n [data-theme=\"light\"], [data-theme=\"light\"],\n .ace_editor, .ace_editor *,\n .CodeMirror, .CodeMirror *,\n .monaco-editor, .monaco-editor *,\n .markdown-body pre, .markdown-body pre *,\n .highlight, .highlight *,\n pre code, pre code * {\n filter: none !important;\n }\n \n /* Fix common UI elements */\n .modal, .popup, .dropdown-menu, .tooltip, .popover {\n filter: invert(1) hue-rotate(180deg) !important;\n background: #2d2d44 !important;\n border-color: #444 !important;\n }\n \n /* Scrollbars */\n ::-webkit-scrollbar { background: #1a1a2e !important; }\n ::-webkit-scrollbar-thumb { background: #444 !important; }\n ::-webkit-scrollbar-thumb:hover { background: #555 !important; }\n \n /* Selection */\n ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; }\n ';\n }\n \n function removeDarkMode() {\n var style = document.getElementById('universal-dark-mode-style');\n if (style) style.remove();\n }\n \n // Toggle with Alt+Shift+D\n document.addEventListener('keydown', function(e) {\n if (e.altKey && e.shiftKey && e.key === 'D') {\n e.preventDefault();\n enabled = !enabled;\n if (enabled) {\n applyDarkMode();\n console.log('[Universal Dark Mode] Enabled');\n } else {\n removeDarkMode();\n console.log('[Universal Dark Mode] Disabled');\n }\n }\n });\n \n // Apply on load\n applyDarkMode();\n \n // Re-apply on dynamic content\n var observer = new MutationObserver(function(mutations) {\n if (enabled && !document.getElementById('universal-dark-mode-style')) {\n applyDarkMode();\n }\n });\n observer.observe(document.head, { childList: true });\n \n console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle');\n})();", "Universal Dark Mode"); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
Skip to content

feat: 支持按租户/服务配置平均 token 数以优化并发计算 - #27

Draft
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration
Draft

feat: 支持按租户/服务配置平均 token 数以优化并发计算#27
BukeLy with Copilot wants to merge 3 commits into
mainfrom
copilot/update-average-tokens-configuration

Conversation

CopilotAI commented Dec 15, 2025

Copy link
Copy Markdown
Contributor

平均 Token 数硬编码在 rate_limiter.py 导致并发计算不准确。Insert 和 Query 场景的 token 消耗差异大,固定保守值限制了系统吞吐量。

Changes

  • src/config.py: 各服务配置类新增 avg_tokens_per_request 字段

    • LLM/DS_OCR: 3500 (default)
    • Embedding: 20000
    • Rerank: 500
  • src/rate_limiter.py:

    • get_rate_limiter() 新增 avg_tokens_per_request 参数
    • 提取默认配置到 SERVICE_DEFAULTS 字典
  • src/tenant_config.py: 各服务合并方法传递 avg_tokens_per_request

  • src/multi_tenant.py / src/deepseek_ocr_client.py: 更新调用传参

  • env.example: 添加环境变量说明

配置优先级

  1. 租户配置 (API)
  2. 环境变量 (LLM_AVG_TOKENS_PER_REQUEST)
  3. 代码默认值

使用示例

# Insert 密集场景,降低平均 token 提升并发
LLM_AVG_TOKENS_PER_REQUEST=2500
// 租户 APIPUT /tenants/{tenant_id}/config
{
"llm_config": {
"avg_tokens_per_request": 4000
}
}
Original prompt

This section details on the original issue you should resolve

<issue_title>Average tokens per request hardcoded - inaccurate concurrency calculation</issue_title>
<issue_description>## 问题描述
平均 Token 数硬编码导致并发计算不准确。

受影响的文件

  • src/rate_limiter.py 行 420-423

硬编码值

  • llm: 3500
  • embedding: 500
  • rerank: 500
  • ds_ocr: 3500

问题

这些值因场景而异(Query vs Insert),但被固定为保守值,导致不必要的速率限制,降低系统吞吐量。

解决方案

应该支持按租户/服务配置平均 token 数。</issue_description>

Comments on the Issue (you are @copilot in this section)


💬 We'd love your input! Share your thoughts on Copilot coding agent in our 2 minute survey.

CopilotAIand others added 2 commits December 15, 2025 09:43
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
Co-authored-by: BukeLy <19304666+BukeLy@users.noreply.github.com>
CopilotAI changed the title [WIP] Update average tokens per request for better concurrency calculationfeat: 支持按租户/服务配置平均 token 数以优化并发计算Dec 15, 2025
CopilotAI requested a review from BukeLyDecember 15, 2025 09:48

@BukeLyBukeLy left a comment

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢?

Comment on lines -83 to +84
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute)
tokens_per_minute=getattr(self.config, 'tokens_per_minute', config.ds_ocr.tokens_per_minute),
avg_tokens_per_request=getattr(self.config, 'avg_tokens_per_request', config.ds_ocr.avg_tokens_per_request)

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ds_ocr.avg_tokens_per_request
ds_ocr.tokens_per_minute
这两个参数的目的是什么?在没有设置TPM/RPM情况下估算吗?还有就是我记得关于Token我引入了Transform的tokenizer来计算,为什么还需要这两个参数呢? @copilot

Copy link
Copy Markdown
ContributorAuthor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这两个参数的作用不同:

  1. tokens_per_minute - API 服务商的 TPM 限制(如硅基流动 40000),用于 RateLimiter 控制每分钟最大 token 消耗,防止 429 错误。

  2. avg_tokens_per_request - 用于启动时预估安全并发数的参数。在 calculate_optimal_concurrent() 中计算:concurrent = TPM / avg_tokens / max_in_flight

虽然 count_tokens() 用 tiktoken 精确计算每个请求的实际 token(运行时扣减 TPM 配额),但 avg_tokens_per_request初始化时预估应允许多少并发请求。这是两个不同用途:

  • 运行时:count_tokens() 精确计算实际消耗
  • 初始化:avg_tokens_per_request 预估并发上限

不同场景(Insert vs Query)的平均 token 差异大,现在可配置这个预估值来优化并发。

Sign up for freeto join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Average tokens per request hardcoded - inaccurate concurrency calculation

2 participants

@BukeLy