ContextPilot:通过细粒度强化学习,训练智能体进行主动上下文管理
ContextPilot-E4B 是 ContextPilot(一个面向长程语言模型智能体的主动上下文管理框架)的 Gemma4-E4B 检查点。它教会智能体在持续推理和调用工具的同时,进行规划、维护长期记忆,并将效用较低的上下文卸载出去。更多详情,请参阅我们的论文和代码仓库。

ContextPilot 融合了三个核心组件:
所得智能体在长上下文问答和深度搜索任务上进行评估;详见评估指南。
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "tencent/ContextPilot-E4B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)请注意,仅加载检查点并不会执行上下文管理工具;工具定义、智能体运行时以及评估流水线已提供在 ContextPilot 仓库 中。完整的部署配置请参见 推理指南。
该检查点适用于主动上下文管理、长程智能体、长上下文问答以及深度搜索等领域的研究。
@inproceedings{pan-etal-2026-contextpilot,
title = "ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL",
author = "Pan, Zhuoshi and
Pei, Qizhi and
Lu, Junru and
Lin, Honglin and
Zhao, H. Vicky and
Yin, Di and
Sun, Xing",
booktitle = "Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2026",
address = "Budapest, Hungary",
publisher = "Association for Computational Linguistics",
abstract = "Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing working context. Recent proactive context management methods allow models to edit their own working context with specialized tools, yet they still face three key limitations: (1) a limited toolset restricted to search, deletion, and summarization, with no support for global planning, long-term memory, and adaptive compression; (2) inefficient exploration that treats context management actions uniformly despite their heterogeneous impacts on final outcomes; and (3) coarse-grained credit assignment that assigns the final trajectory-level reward to all intermediate context editing actions during RL. To bridge these gaps, we introduce ContextPilot, a proactive context management framework for long-horizon agentic reasoning. Our approach systematically augments the toolset with planning, long-term memory, and soft context offloading tools. We further propose an RL method tailored for context management, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action. Experiments on long-context QA and deep search tasks show that ContextPilot achieves stronger performance with a more compact working context, consistently outperforming existing baselines across various base models and benchmarks. Code is available at \url{https://github.com/Tencent/ContextPilot}."
}