[!Note]
本仓库包含后训练模型的权重与配置文件,格式为 Hugging Face Transformers 兼容格式。这些产物兼容 Hugging Face Transformers、vLLM、SGLang、TokenSpeed 等框架。
[!Tip]
对于需要托管式、可弹性扩展推理服务、且无需自行维护基础设施的用户,官方 Qwen API 服务由 Qwen Cloud 提供。 其中,Qwen3.8-27B 将提供托管版本,内置更多生产级功能,例如默认支持 1M 上下文长度、官方内置工具等。更多信息请参阅 Qwen3.8-27B 概览。该服务即将上线,敬请期待。
继 Qwen3.5 和 Qwen3.6 系列获得社区的广泛采用之后,我们欣然推出 Qwen3.8——迄今为止 Qwen 开源模型家族中能力最强的一代。
Qwen3.8 以 Qwen3.5 的架构为基础,在代码编写、专业工作、科研探索以及长时程智能体任务等方面均实现了显著提升。Qwen3.8-27B 将上述进步凝聚于一个紧凑、易于部署的稠密模型之中:它是一款原生视觉语言模型,能够理解图像与视频,支持灵活的思维控制,并专为可靠地完成复杂多步任务而设计。
Qwen3.8-27B 具备以下增强特性:
reasoning_effort 进行调节,历史消息中的推理上下文可通过 preserve_thinking 予以保留。| Qwen3.8-27B | Qwen3.6-27B | Qwen3.7-Plus | Muse Glimmer-30B | Opus4.6 Max | |
|---|---|---|---|---|---|
| 编程能力 | |||||
智能体终端编程 Terminal Bench 2.1 (Terminus) | 73.0 | 63.4 | 64.0 | 51.7 | 78.2 |
智能体编程 SWE-bench Pro | 61.7 | 53.5 | 57.6 | 51.2 | 53.4 |
仓库级代码生成 NL2Repo-Bench | 42.3 | 36.2 | 41.1 | -- | 47.6 |
智能体编程 DeepSWE 1.1 | 42.2 | 13.3 | 14.2 | -- | -- |
软件工程 QwenSWEBench | 79.0 | 49.3 | 59.2 | -- | 63.8 |
| 智能体能力 | |||||
长周期办公任务 CoWorkBench | 70.7 | 61.0 | 65.1 | -- | 68.2 |
专业岗位任务 JobBench | 33.4 | 21.8 | 27.6 | -- | -- |
前沿智能体任务 Agents' Last Exam | Pass@1 20.4 得分 42.9 | Pass@1 10.6 得分 27.3 | Pass@1 13.2 得分 33.6 | -- | -- |
| 通用能力 | |||||
指令遵循 IFBench | 79.5 | 69.1 | 79.1 | 77.0 | 62.5 |
科学推理 GPQA Diamond | 89.2 | 87.8 | 90.3 | 83.5 | 91.3 |
多学科推理 HLE | 30.8 | 24.0 | 34.7 | 22.0 | 40.0 |
竞赛级编程 LiveCodeBench v6 | 90.3 | 83.9 | 89.6 | -- | 88.8 |
| Qwen3.8-27B | Qwen3.6-27B | Qwen3.7-Plus | Muse Glimmer-30B | Opus4.6 Max | |
|---|---|---|---|---|---|
| 智能体多模态能力 | |||||
计算机操作 OSWorld-Verified | 84.3 | 63.9 | 73.3 | 65.9 | 72.7 |
浏览器操作 WebArena-Verified | 64.8 | 48.8 | 55.3 | -- | -- |
移动设备操作 AndroidWorld | 81.9 | 70.3 | 81.0 | -- | 62.0 |
应用还原 RecreationBench | 47.1 | 29.8 | 30.2 | -- | -- |
多模态工具调用 ClawEval-MM | Pass@3 57.4 平均 56.9 | Pass@3 42.6 平均 50.4 | Pass@3 57.4 平均 60.1 | -- | Pass@3 52.5 平均 54.7 |
多模态软件工程 SWE-MM | 38.6 | 25.7 | 30.0 | -- | 27.1 |
可视化网页开发 Vision2Web | 62.9 | 45.0 | 42.1 | -- | -- |
| 通用多模态能力 | |||||
视觉数学问题求解 MathVision | 无思维链 90.0 有思维链 94.6 | 无思维链 85.1 | 无思维链 90.3 | -- | 无思维链 65.5 |
通用视觉推理 BabyVision | 无思维链 65.7 有思维链 85.6 | 无思维链 28.9 | 无思维链 64.7 有思维链 70.4 | -- | 无思维链 12.6 |
科学图表分析 CharXiv (RQ) | 无CI 83.7 有CI 90.2 | 无CI 78.4 | 无CI 85.8 有CI 85.9 | 78.8 | 无CI 66.0 |
文档智能 OmniDocBench 1.5 | 91.1 | 89.4 | 91.4 | 75.8 | 86.6 |
真实世界感知 RealWorldQA | 85.9 | 84.1 | 86.9 | -- | 73.9 |
具身智能 ERQA | 65.5 | 62.5 | 69.8 | -- | 40.8 |
\boxed{} 中。”对于其余模型,我们报告两种提示变体中得分较高的结果——一种包含 \boxed{} 格式要求,另一种不包含。gpt-5.4-2026-03-05 进行评判。为简化集成流程,我们推荐通过API使用Qwen3.8。
[!重要] 不同推理框架的推理效率和吞吐量差异显著。 建议使用最新版本的框架,以确保最佳性能和兼容性。 对于生产环境负载或高吞吐场景,推荐使用专用推理引擎,如SGLang、vLLM或TokenSpeed。
Qwen3.8可通过主流推理框架进行部署,例如:
[!重要] Qwen3.8模型默认以思考模式运行,在生成最终回复之前,会先输出以
<think>\n...\n</think>\n\n标记的思考内容。 如需禁用思考内容并直接获取回复,请参阅此处的示例。
[!提示] 我们推荐使用以下采样参数组合进行生成:
- 思考模式:
temperature=1.0,top_p=0.95,top_k=20,min_p=0.0,presence_penalty=0.0,repetition_penalty=1.0- 指令(或非思考)模式:
temperature=0.7,top_p=0.80,top_k=20,min_p=0.0,presence_penalty=1.5,repetition_penalty=1.0请注意,不同推理框架对采样参数的支持可能有所差异。
Qwen3.8官方支持reasoning_effort参数,可用于调节推理深度并控制成本:
xhigh(默认):适用于需要深入分析的复杂任务
medium:在准确性和速度之间取得平衡low:高效推理,优先考虑速度和成本此外,为提供最佳的开箱即用体验,所有工作负载默认启用preserve_thinking。如需禁用保留思考功能,请参阅此处的示例。
[!提示] 在多轮智能体任务中,较低的推理强度并不总能减少整体任务完成时间。虽然每轮响应可能更快,但也可能导致分析不充分、失败率上升和重复重试,从而增加总延迟和Token消耗。
聊天补全API可配合大多数推理框架使用,也可用于Qwen Cloud。 在开始之前,请确保已安装OpenAI Python SDK,并配置好API密钥和API基础URL,例如:
pip install -U openai
# Set the following accordingly
export OPENAI_BASE_URL='your-base-url'
export OPENAI_API_KEY='your-api-key'from openai import OpenAI
# Configured by environment variables
client = OpenAI()
messages = [{"role": "user", "content": "Write a Python function to merge two sorted linked lists."}]
completion = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=messages,
extra_body={
"chat_template_kwargs": {
"enable_thinking": True, # on by default
"preserve_thinking": True, # on by default
},
},
reasoning_effort="xhigh", # xhigh by default; supported levels are xhigh, medium, and low
stream=True,
stream_options={"include_usage": True},
)
reasoning_content = ""
answer_content = ""
is_answering = False
print("\n" + "=" * 20 + "Reasoning" + "=" * 20 + "\n")
for chunk in completion:
if not chunk.choices:
print("\nUsage:")
print(chunk.usage)
continue
delta = chunk.choices[0].delta
if hasattr(delta, "reasoning_content") and delta.reasoning_content is not None:
if not is_answering:
print(delta.reasoning_content, end="", flush=True)
reasoning_content += delta.reasoning_content
elif hasattr(delta, "reasoning") and delta.reasoning is not None:
if not is_answering:
print(delta.reasoning, end="", flush=True)
reasoning_content += delta.reasoning
if hasattr(delta, "content") and delta.content:
if not is_answering:
print("\n" + "=" * 20 + "Answer" + "=" * 20 + "\n")
is_answering = True
print(delta.content, end="", flush=True)
answer_content += delta.content
messages.append({
"role": "assistant",
"content": answer_content,
"reasoning_content": reasoning_content,
"reasoning": reasoning_content,
})from openai import OpenAI
# Configured by environment variables
client = OpenAI()
messages = [
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/CI_Demo/mathv-1327.jpg"
}
},
{
"type": "text",
"text": "The centres of the four illustrated circles are in the corners of the square. The two big circles touch each other and also the two little circles. With which factor do you have to multiply the radii of the little circles to obtain the radius of the big circles?\nChoices:\n(A) $\\frac{2}{9}$\n(B) $\\sqrt{5}$\n(C) $0.8 \\cdot \\pi$\n(D) 2.5\n(E) $1+\\sqrt{2}$"
}
]
}
]
chat_response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=messages,
)
print("Chat response:", chat_response)from openai import OpenAI
# Configured by environment variables
client = OpenAI()
messages = [
{
"role": "user",
"content": [
{
"type": "video_url",
"video_url": {
"url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/video/N1cdUjctpG8.mp4"
}
},
{
"type": "text",
"text": "How many porcelain jars were discovered in the niches located in the primary chamber of the tomb?"
}
]
}
]
chat_response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=messages,
)
# When vLLM is launched with `--media-io-kwargs '{"video": {"num_frames": -1}}'`,
# video frame sampling can be configured via `extra_body` (e.g., by setting `fps`).
# This feature is currently supported only in vLLM.
#
# By default, `fps=2` and `do_sample_frames=True`.
# With `do_sample_frames=True`, you can customize the `fps` value to set your desired video sampling rate.
# chat_response = client.chat.completions.create(
# model="Qwen/Qwen3.8-27B",
# messages=messages,
# extra_body={
# "mm_processor_kwargs": {"fps": 2, "do_sample_frames": True},
# },
# )
print("Chat response:", chat_response)Qwen3.8-27B 默认会在响应前进行思考。 您可以通过配置 API 参数,直接获取模型的响应而无需其进行思考。 例如,
from openai import OpenAI
# Configured by environment variables
client = OpenAI()
messages = [
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": "https://qianwen-res.oss-accelerate.aliyuncs.com/Qwen3.5/demo/RealWorld/RealWorld-04.png"
}
},
{
"type": "text",
"text": "Where is this?"
}
]
}
]
chat_response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=messages,
temperature=0.7,
top_p=0.8,
presence_penalty=1.5,
extra_body={
"top_k": 20,
"chat_template_kwargs": {"enable_thinking": False},
},
)
print("Chat response:", chat_response)[!Note]
如果您使用的是 Qwen 云的 API,除了更改model外,请使用"enable_thinking": False,而不是"chat_template_kwargs": {"enable_thinking": False}。
默认情况下,Qwen3.8 会保留所有历史消息中的思考块,从而在整段对话中维持完整的推理轨迹。这种被称为“保留思考”的行为,确保了上下文的全方位连贯性,尤其对于智能体场景而言至关重要——在这些场景中,决策的一致性和减少冗余推理尤为关键。此外,它还能优化 KV 缓存利用率,在思考模式和非思考模式下均能提升推理效率。
如果您希望仅保留最新用户消息中的思考块,可以通过将 preserve_thinking 设置为 False 来禁用此功能:
from openai import OpenAI
# Configured by environment variables
client = OpenAI()
messages = [...]
chat_response = client.chat.completions.create(
model="Qwen/Qwen3.8-27B",
messages=messages,
extra_body={
"chat_template_kwargs": {"preserve_thinking": False},
},
)
print("Chat response:", chat_response)[!Note]
如果您使用的是来自 Qwen 云的 API,除了更改model之外,请直接使用"preserve_thinking": False,而无需将其包装在chat_template_kwargs中。
为获得最佳性能,我们建议采用以下设置:
采样参数:建议使用以下各组采样参数:
temperature=1.0,top_p=0.95,top_k=20,min_p=0.0,presence_penalty=0.0,repetition_penalty=1.0temperature=0.7,top_p=0.80,top_k=20,min_p=0.0,presence_penalty=1.5,repetition_penalty=1.0对于受支持的框架,您可以将 presence_penalty 参数调整至 0 到 2 之间,以减少无休止的重复。不过,使用较高的数值有时可能会导致语言混杂,并略微降低模型性能。
充足的输出长度:为在智能体任务上优化性能,建议分配足够的输出长度,使模型能够生成详尽且全面的回答。对于支持分别设置内部推理和最终输出 token 上限的框架,建议在 1M 上下文长度内采用以下配置:
这些设置为复杂推理提供了必要的能力,同时确保为高质量最终输出保留充足空间。
处理超长文本:Qwen3.8-27B 原生支持最长 262,144 个 token 的上下文长度。对于总长度(包括输入和输出)超过此上限的长时程任务,建议使用 RoPE 缩放技术有效处理长文本,例如 YaRN。
目前 YaRN 已获得多种推理框架的支持,例如 vLLM、SGLang 和 TokenSpeed。
通常,为受支持框架启用 YaRN 有以下两种方式:
修改模型配置文件:
在 config.json 文件的 text_config 中,将 rope_parameters 字段修改为:
{
"mrope_interleaved": true,
"mrope_section": [
11,
11,
10
],
"rope_type": "yarn",
"rope_theta": 10000000,
"partial_rotary_factor": 0.25,
"factor": 4.0,
"original_max_position_embeddings": 262144
} 传递命令行参数:
对于 vLLM,您可以使用
VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1000000 对于 SGLang,您可以使用
SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... --json-model-override-args '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --context-length 1000000 对于 TokenSpeed,您可以使用
TOKENSPEED_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 tokenspeed serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1000000 [!NOTE]
所有知名的开源框架均实现了静态 YaRN,这意味着缩放因子始终保持不变,与输入长度无关,这可能会对较短文本的性能产生影响。
我们建议仅在需要处理长上下文时修改rope_parameters配置。
同时建议根据实际需求修改factor。例如,如果您的应用通常处理 524,288 个 token 的上下文长度,最好将factor设置为 2.0。
长视频理解:为优化纯文本和图像的推理效率,发布的 video_preprocessor_config.json 中的 size 参数配置较为保守。建议将 video_preprocessor_config 文件中的 longest_edge 参数设置为 469,762,048(对应 224k 视频 token),以便对小时级视频实现更高帧率的采样,从而获得更优性能。例如:
{"longest_edge": 469762048, "shortest_edge": 4096} 如果您觉得我们的工作有所帮助,欢迎引用我们。
@misc{qwen38,
title = {{Qwen3.8-Max}: A New Bar for Coding and Cowork},
url = {https://qwen.ai/blog?id=qwen3.8},
author = {{Qwen Team}},
month = {August},
year = {2026}
}