Skip to content

2026-05-10 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM ServingZhiqing Zhong, Zhijing Ye, Jian Zhang, Weijian Zheng, Bolun Sun, Xiaodong Yu2026-05-10下载Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request lengths differ, EOS ev...
Emerging 2D Materials for Beyond von Neumann Computing: A PerspectiveYaser Banad2026-05-10下载The end of conventional Dennard scaling and the widening gap between memory bandwidth and arithmetic throughput have made the von Neumann partition a structural bottleneck rather than a transient one.
Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM ReasoningAojie Yuan, Tianqi Shen, Dajun Zhang2026-05-10下载Reasoning LLMs produce thousands of chain-of-thought tokens whose KV cache must reside in scarce GPU HBM. The dominant response -- permanently evicting low-importance tokens -- is catastrophic for rea...
31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-DecodingPingcheng Dong, Yonghao Tan, Xuejiao Liu, Peng Luo, Yu Liu, Di Pang, Songchen Ma, Xijie Huang, Shih-Yang Liu, Dong Zhang, Zhichao Lu, Luhong Liang, Chi-Ying Tsui, Fengbin Tu, Liang Zhao, Kwang-Ting Cheng2026-05-10下载This work presents a 55nm speculative decoding-based LLM accelerator with bumping-based face-to-face ReRAM-on-logic stacking technology. It features a local rotation unit for outlier-free low-bit quan...
Scaling Qubit Mapping and Routing With Position Graph Abstraction and MemoizationBrent Russon, Bao Bach, Ed Younis, Ilya Safro2026-05-10下载Scalable qubit mapping and routing remain major bottlenecks in quantum compilation, especially for Trapped-Ion Quantum Charge-Coupled device (TI-QCCD) architectures, where qubit interactions require p...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Optimizing Server Placement for Vertical Federated Learning in Dynamic Edge/Fog NetworksSu Wang, Mung Chiang, H. Vincent Poor2026-05-10下载We investigate the control and optimization of vertical federated learning (VFL), a class of distributed machine learning (ML) methods in which edge/fog devices contain separate data features, in dyna...
Multi-Tier Labeling and Physics-Informed Learning for Orbital Anomaly Detection at ScaleYong Fu2026-05-10下载Detecting orbital anomalies, such as maneuvers, atmospheric decay, and attitude upsets, across the rapidly growing population of low-Earth-orbit (LEO) satellites is a prerequisite for collision avoida...
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case StudyShimul Debnath, William Hart, Lori Pollock, Donald Lien, Wei Wang2026-05-10下载Cloud performance fluctuates due to factors such as resource contention and workload changes. These factors can be short-term, seasonal, or long-term.
Learning from Acceptance: Cumulative Regret in the Game of CodingHanzaleh Akbari Nodehi, Parsa Moradi, Mohammad Ali Maddah-Ali2026-05-10下载Classical coding-theoretic guarantees often rely on trust assumptions, such as requiring sufficiently many honest nodes compared with adversarial ones.
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM ServingZhiqing Zhong, Zhijing Ye, Jian Zhang, Weijian Zheng, Bolun Sun, Xiaodong Yu2026-05-10下载Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request lengths differ, EOS ev...
Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple SiliconVíctor Gallego2026-05-10下载We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in nn-body problems, multi-field Boltzmann, neighbor...
A Scalable and Unified Framework to Weighted Rank AggregationAmir Carmel, Debarati Das, Tien-Long Nguyen2026-05-10下载The rank aggregation problem seeks to combine multiple rank orderings of the same set of candidates into a single consensus ordering. Such problems arise in diverse domains, including web search, empl...
Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud ContinuumAkuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis, Praveen Kumar Donta2026-05-10下载In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approaches to DNN partitioning and offloading across the edge-cloud c...
Categorical Message Passing Language (CaMPL) for programmersDaniel Kiyoshi Hashimoto, Alexanna Little Berg, Priyaa Varshinee Srinivasan2026-05-10下载Categorical Message Passing Language (CaMPL) is a functional-style concurrent programming language whose semantics is in category theory, more specifically, linear actegories.
PoHAR: Understanding Hyperlocal Human Activities with Pollution Sensor NetworksPrasenjit Karmakar, Karthik Reddy, Sandip Chakraborty2026-05-10下载Low-cost air quality sensors are becoming ubiquitous in our daily lives as public awareness of air pollution continues to grow, and people take measures to monitor and improve the air they breathe ind...
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural NetworksPranjal Naman, Yogesh Simmhan2026-05-10下载Graph Neural Network (GNN) inference on billion-scale graphs is critical for domains like fintech and recommendation systems. Full-graph inference on these large graphs can be challenging due to high ...
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUsDaemyung Kang, Eunjin Hwang, Hanjeong Lee, HyeokJin Kim, Hyunhoi Koo, Jeongkyu Shin, Jeongseok Kang, Jihyun Kang, Joongi Kim, Junbum Lee, Jungseung Yang, Kyujin Cho, Youngsook Song2026-05-10下载Large-scale AI training is now fundamentally a distributed systems problem, and hardware failures have become routine operating conditions rather than rare exceptions.
Split CNN Inference on Networked MicrocontrollersJunyu Lu, Shashwath Suresh, Hao Liu, Qi Hong, Qing Wang2026-05-10下载Running deep neural networks on microcontroller units (MCUs) is severely constrained by limited memory resources. While TinyML techniques reduce model size and computation, they often fail in practice...
Enforcing Attestable Workflows across Untrusted NetworksHung Dang, Tue Nguyen2026-05-10下载Confidential high-performance computing orchestrates workloads across federated domains, yet existing frameworks rely on high-overhead user-space library operating systems or assume single-host execut...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Optimizing Server Placement for Vertical Federated Learning in Dynamic Edge/Fog NetworksSu Wang, Mung Chiang, H. Vincent Poor2026-05-10下载We investigate the control and optimization of vertical federated learning (VFL), a class of distributed machine learning (ML) methods in which edge/fog devices contain separate data features, in dyna...
Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud ContinuumAkuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis, Praveen Kumar Donta2026-05-10下载In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approaches to DNN partitioning and offloading across the edge-cloud c...
TSNBench: Benchmarking LLM Proficiency in Time-Sensitive NetworkingRubi Debnath, Daniel Bujosa Mateu, Luxi Zhao, Silviu S. Craciunas, Paul Pop, Sebastian Steinhorst2026-05-10下载We present TSNBench, the first benchmark for evaluating large language model (LLM) proficiency in Time-Sensitive Networking (TSN), a suite of IEEE 802.
PolicyCache-SDN: Hierarchical Intra-Path Learning for Adaptive SDN Traffic ControlWenyang Jia, Jingjing Wang, Ziwei Yan, Tanren Liu, Yakun Ren, Kai Lei2026-05-10下载Software defined networks offer global visibility, yet centralized control loops are too slow for transient congestion and bursty traffic dynamics.
The Carrier Pigeon Internet Protocol: An Algorithmic (and Lighthearted) PerspectiveMatthias Bentert, Shay Kutten, Darya Melnyk, Tijana Milentijevic, Stefan Schmid2026-05-10下载The theoretical model behind the pigeon post as a link layer in a communication network was introduced by Shannon (under the guise of studying One-Time Pads for cryptography).
Function-Space ADMM for Decentralized Federated Learning: A Control Theoretic PerspectiveAkihito Taya, Yuuki Nishiyama, Kaoru Sezaki2026-05-10下载Decentralized federated learning (FL) is a promising approach for training machine learning models on sensor networks, Internet of Things (IoT) devices, and other edge systems where no central server ...
CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian SplattingDaheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, Fang Dong2026-05-10下载Volumetric video (VV) streaming enables real-time, immersive access to remote 3D environments, powering telepresence, ecological monitoring, and robotic teleoperation.
Chain-of-Thought Reasoning Enhances In-Context Learning for LLM-Based Mobile Traffic PredictionMohammadMahdi Ghadaksaz, Mohammad Farzanullah, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci2026-05-10下载Accurate short-term mobile traffic prediction is important for proactive resource allocation and low-latency network management in fifth generation (5G) and sixth generation (6G).

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM ServingZhiqing Zhong, Zhijing Ye, Jian Zhang, Weijian Zheng, Bolun Sun, Xiaodong Yu2026-05-10下载Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request lengths differ, EOS ev...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Cloud Performance Decomposition for Long-Term Performance Engineering: A Case StudyShimul Debnath, William Hart, Lori Pollock, Donald Lien, Wei Wang2026-05-10下载Cloud performance fluctuates due to factors such as resource contention and workload changes. These factors can be short-term, seasonal, or long-term.
Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud ContinuumAkuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis, Praveen Kumar Donta2026-05-10下载In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approaches to DNN partitioning and offloading across the edge-cloud c...

基于 VitePress 构建 · 使用本地搜索查找论文