2026-05-10
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving | Zhiqing Zhong, Zhijing Ye, Jian Zhang, Weijian Zheng, Bolun Sun, Xiaodong Yu | 2026-05-10 | 下载 | Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request lengths differ, EOS ev... |
| Emerging 2D Materials for Beyond von Neumann Computing: A Perspective | Yaser Banad | 2026-05-10 | 下载 | The end of conventional Dennard scaling and the widening gap between memory bandwidth and arithmetic throughput have made the von Neumann partition a structural bottleneck rather than a transient one. |
| Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning | Aojie Yuan, Tianqi Shen, Dajun Zhang | 2026-05-10 | 下载 | Reasoning LLMs produce thousands of chain-of-thought tokens whose KV cache must reside in scarce GPU HBM. The dominant response -- permanently evicting low-importance tokens -- is catastrophic for rea... |
| 31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding | Pingcheng Dong, Yonghao Tan, Xuejiao Liu, Peng Luo, Yu Liu, Di Pang, Songchen Ma, Xijie Huang, Shih-Yang Liu, Dong Zhang, Zhichao Lu, Luhong Liang, Chi-Ying Tsui, Fengbin Tu, Liang Zhao, Kwang-Ting Cheng | 2026-05-10 | 下载 | This work presents a 55nm speculative decoding-based LLM accelerator with bumping-based face-to-face ReRAM-on-logic stacking technology. It features a local rotation unit for outlier-free low-bit quan... |
| Scaling Qubit Mapping and Routing With Position Graph Abstraction and Memoization | Brent Russon, Bao Bach, Ed Younis, Ilya Safro | 2026-05-10 | 下载 | Scalable qubit mapping and routing remain major bottlenecks in quantum compilation, especially for Trapped-Ion Quantum Charge-Coupled device (TI-QCCD) architectures, where qubit interactions require p... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Optimizing Server Placement for Vertical Federated Learning in Dynamic Edge/Fog Networks | Su Wang, Mung Chiang, H. Vincent Poor | 2026-05-10 | 下载 | We investigate the control and optimization of vertical federated learning (VFL), a class of distributed machine learning (ML) methods in which edge/fog devices contain separate data features, in dyna... |
| Multi-Tier Labeling and Physics-Informed Learning for Orbital Anomaly Detection at Scale | Yong Fu | 2026-05-10 | 下载 | Detecting orbital anomalies, such as maneuvers, atmospheric decay, and attitude upsets, across the rapidly growing population of low-Earth-orbit (LEO) satellites is a prerequisite for collision avoida... |
| Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study | Shimul Debnath, William Hart, Lori Pollock, Donald Lien, Wei Wang | 2026-05-10 | 下载 | Cloud performance fluctuates due to factors such as resource contention and workload changes. These factors can be short-term, seasonal, or long-term. |
| Learning from Acceptance: Cumulative Regret in the Game of Coding | Hanzaleh Akbari Nodehi, Parsa Moradi, Mohammad Ali Maddah-Ali | 2026-05-10 | 下载 | Classical coding-theoretic guarantees often rely on trust assumptions, such as requiring sufficiently many honest nodes compared with adversarial ones. |
| KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving | Zhiqing Zhong, Zhijing Ye, Jian Zhang, Weijian Zheng, Bolun Sun, Xiaodong Yu | 2026-05-10 | 下载 | Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request lengths differ, EOS ev... |
| Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon | Víctor Gallego | 2026-05-10 | 下载 | We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, all-pairs in -body problems, multi-field Boltzmann, neighbor... |
| A Scalable and Unified Framework to Weighted Rank Aggregation | Amir Carmel, Debarati Das, Tien-Long Nguyen | 2026-05-10 | 下载 | The rank aggregation problem seeks to combine multiple rank orderings of the same set of candidates into a single consensus ordering. Such problems arise in diverse domains, including web search, empl... |
| Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud Continuum | Akuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis, Praveen Kumar Donta | 2026-05-10 | 下载 | In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approaches to DNN partitioning and offloading across the edge-cloud c... |
| Categorical Message Passing Language (CaMPL) for programmers | Daniel Kiyoshi Hashimoto, Alexanna Little Berg, Priyaa Varshinee Srinivasan | 2026-05-10 | 下载 | Categorical Message Passing Language (CaMPL) is a functional-style concurrent programming language whose semantics is in category theory, more specifically, linear actegories. |
| PoHAR: Understanding Hyperlocal Human Activities with Pollution Sensor Networks | Prasenjit Karmakar, Karthik Reddy, Sandip Chakraborty | 2026-05-10 | 下载 | Low-cost air quality sensors are becoming ubiquitous in our daily lives as public awareness of air pollution continues to grow, and people take measures to monitor and improve the air they breathe ind... |
| ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks | Pranjal Naman, Yogesh Simmhan | 2026-05-10 | 下载 | Graph Neural Network (GNN) inference on billion-scale graphs is critical for domains like fintech and recommendation systems. Full-graph inference on these large graphs can be challenging due to high ... |
| From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs | Daemyung Kang, Eunjin Hwang, Hanjeong Lee, HyeokJin Kim, Hyunhoi Koo, Jeongkyu Shin, Jeongseok Kang, Jihyun Kang, Joongi Kim, Junbum Lee, Jungseung Yang, Kyujin Cho, Youngsook Song | 2026-05-10 | 下载 | Large-scale AI training is now fundamentally a distributed systems problem, and hardware failures have become routine operating conditions rather than rare exceptions. |
| Split CNN Inference on Networked Microcontrollers | Junyu Lu, Shashwath Suresh, Hao Liu, Qi Hong, Qing Wang | 2026-05-10 | 下载 | Running deep neural networks on microcontroller units (MCUs) is severely constrained by limited memory resources. While TinyML techniques reduce model size and computation, they often fail in practice... |
| Enforcing Attestable Workflows across Untrusted Networks | Hung Dang, Tue Nguyen | 2026-05-10 | 下载 | Confidential high-performance computing orchestrates workloads across federated domains, yet existing frameworks rely on high-overhead user-space library operating systems or assume single-host execut... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Optimizing Server Placement for Vertical Federated Learning in Dynamic Edge/Fog Networks | Su Wang, Mung Chiang, H. Vincent Poor | 2026-05-10 | 下载 | We investigate the control and optimization of vertical federated learning (VFL), a class of distributed machine learning (ML) methods in which edge/fog devices contain separate data features, in dyna... |
| Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud Continuum | Akuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis, Praveen Kumar Donta | 2026-05-10 | 下载 | In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approaches to DNN partitioning and offloading across the edge-cloud c... |
| TSNBench: Benchmarking LLM Proficiency in Time-Sensitive Networking | Rubi Debnath, Daniel Bujosa Mateu, Luxi Zhao, Silviu S. Craciunas, Paul Pop, Sebastian Steinhorst | 2026-05-10 | 下载 | We present TSNBench, the first benchmark for evaluating large language model (LLM) proficiency in Time-Sensitive Networking (TSN), a suite of IEEE 802. |
| PolicyCache-SDN: Hierarchical Intra-Path Learning for Adaptive SDN Traffic Control | Wenyang Jia, Jingjing Wang, Ziwei Yan, Tanren Liu, Yakun Ren, Kai Lei | 2026-05-10 | 下载 | Software defined networks offer global visibility, yet centralized control loops are too slow for transient congestion and bursty traffic dynamics. |
| The Carrier Pigeon Internet Protocol: An Algorithmic (and Lighthearted) Perspective | Matthias Bentert, Shay Kutten, Darya Melnyk, Tijana Milentijevic, Stefan Schmid | 2026-05-10 | 下载 | The theoretical model behind the pigeon post as a link layer in a communication network was introduced by Shannon (under the guise of studying One-Time Pads for cryptography). |
| Function-Space ADMM for Decentralized Federated Learning: A Control Theoretic Perspective | Akihito Taya, Yuuki Nishiyama, Kaoru Sezaki | 2026-05-10 | 下载 | Decentralized federated learning (FL) is a promising approach for training machine learning models on sensor networks, Internet of Things (IoT) devices, and other edge systems where no central server ... |
| CAGS: Color-Adaptive Volumetric Video Streaming with Dynamic 3D Gaussian Splatting | Daheng Yin, Yili Jin, Jianxin Shi, Isaac Ding, Miao Zhang, Fangxin Wang, Zhaowu Huang, Cong Zhang, Jiangchuan Liu, Fang Dong | 2026-05-10 | 下载 | Volumetric video (VV) streaming enables real-time, immersive access to remote 3D environments, powering telepresence, ecological monitoring, and robotic teleoperation. |
| Chain-of-Thought Reasoning Enhances In-Context Learning for LLM-Based Mobile Traffic Prediction | MohammadMahdi Ghadaksaz, Mohammad Farzanullah, Akram Bin Sediq, Ali Afana, Melike Erol-Kantarci | 2026-05-10 | 下载 | Accurate short-term mobile traffic prediction is important for proactive resource allocation and low-latency network management in fifth generation (5G) and sixth generation (6G). |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving | Zhiqing Zhong, Zhijing Ye, Jian Zhang, Weijian Zheng, Bolun Sun, Xiaodong Yu | 2026-05-10 | 下载 | Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highly irregular KV-cache behavior: request lengths differ, EOS ev... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Cloud Performance Decomposition for Long-Term Performance Engineering: A Case Study | Shimul Debnath, William Hart, Lori Pollock, Donald Lien, Wei Wang | 2026-05-10 | 下载 | Cloud performance fluctuates due to factors such as resource contention and workload changes. These factors can be short-term, seasonal, or long-term. |
| Adaptive DNN Partitioning and Offloading in Heterogeneous Edge-Cloud Continuum | Akuen Akoi Deng, Eimantas Butkus, Alfreds Lapkovskis, Praveen Kumar Donta | 2026-05-10 | 下载 | In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approaches to DNN partitioning and offloading across the edge-cloud c... |