tencent_hunyuan/ContextPilot-8B
模型介绍
文件和版本
Pull Requests
讨论
分析

ContextPilot-8B

ContextPilot:通过细粒度强化学习训练智能体进行主动上下文管理

GitHub 仓库 ContextPilot 在线演示 论文 Hugging Face 模型

ContextPilot-8B 是 ContextPilot 基于 Qwen3-8B 的模型检查点。ContextPilot 是一个面向长程语言模型智能体的主动上下文管理框架,能够教会智能体在持续推理与调用工具的过程中进行规划、维护长期记忆,并卸载价值较低的上下文。更多细节请参阅我们的论文和代码仓库。

ContextPilot 概览

概述

ContextPilot 融合了三大主要组件:

  • 一套扩展的上下文管理工具集,涵盖规划、结构化记忆、检索和软上下文卸载;
  • 一种上下文感知的部分 rollout,将探索聚焦于敏感的上下文编辑决策;以及
  • 一种细粒度信用分配机制,利用下游分支的结果训练中间快照。

训练后的智能体在长上下文问答与深度搜索任务上进行了评估;详见评测说明。

加载中

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "tencent/ContextPilot-8B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

Note that loading the checkpoint alone does not execute context-management tools; the tool definitions, agent runtime, and evaluation pipeline are provided in the ContextPilot repository. See the inference guide for the full setup.

Intended Use

This checkpoint is intended for research on proactive context management, long-horizon agents, long-context QA, and deep search.

License

LICENSE.

Citation

@inproceedings{pan-etal-2026-contextpilot,
    title = "ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL",
    author = "Pan, Zhuoshi  and
      Pei, Qizhi  and
      Lu, Junru  and
      Lin, Honglin  and
      Zhao, H. Vicky  and
      Yin, Di  and
      Sun, Xing",
    booktitle = "Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2026",
    address = "Budapest, Hungary",
    publisher = "Association for Computational Linguistics",
    abstract = "Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing working context. Recent proactive context management methods allow models to edit their own working context with specialized tools, yet they still face three key limitations: (1) a limited toolset restricted to search, deletion, and summarization, with no support for global planning, long-term memory, and adaptive compression; (2) inefficient exploration that treats context management actions uniformly despite their heterogeneous impacts on final outcomes; and (3) coarse-grained credit assignment that assigns the final trajectory-level reward to all intermediate context editing actions during RL. To bridge these gaps, we introduce ContextPilot, a proactive context management framework for long-horizon agentic reasoning. Our approach systematically augments the toolset with planning, long-term memory, and soft context offloading tools. We further propose an RL method tailored for context management, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action. Experiments on long-context QA and deep search tasks show that ContextPilot achieves stronger performance with a more compact working context, consistently outperforming existing baselines across various base models and benchmarks. Code is available at \url{https://github.com/Tencent/ContextPilot}."
}