类别
标签
显示 46 篇内容
Paper

ReTopK:在排序前先召回,用历史 Top-K 决策加速长上下文注意力

Wenshuai Yao · Wenyong Zhou · Hanyong Shao · Yizhe Chen · Zhiyuan Ning · Yuannuo Feng · Ru Huang · Kechao Tang

School of Integrated Circuits, Peking University, Beijing, China · Department of Electrical and Computer Engineering, The University of Hong Kong, Hong Kong SAR, China · School of Integrated Circuit Science and Engineering, Beihang University, Beijing, China · arXiv preprint arXiv:2607.27692

解读 ReTopK 如何通过相似 Query 召回历史 Top-K 支持集合,并在紧凑候选集上精确重排,从而降低长上下文稀疏注意力的索引发现成本。

Paper

FlexiCache:利用注意力头的时间稳定性管理分层 KV Cache

Nazmul Takbir · HamidReza Alikhani Koshkak · Nikil Dutt · Sangeetha Abdu Jyothi

Department of Computer Science, University of California, Irvine · Proceedings of Machine Learning and Systems (MLSys 2026)

FlexiCache 根据注意力头的时间稳定性差异,在 GPU 与主机内存间分层放置 KV 页面,降低显存占用、重排开销与在线推理延迟。

Paper

HitKV:哪些 Token 真正重要?

今夜白 · Collaborators

Chongqing University · Proceedings of the AAAI Conference on Artificial Intelligence

从激活频率出发,理解 KV Cache 重要性评估的一种思路。