2026-09-27
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| QBX: A Compiler for 2-local Qubit Hamiltonian Simulation on Quantum Chiplets | Zikun Li, Zhuoming Chen, Zhihao Jia | 2026-09-27 | 下载 | 2-local qubit Hamiltonian simulation, a fundamental task in quantum computing, is widely applied in various applications. This paper presents QBX, the first quantum compiler designed for 2-local qubit... |
| S-ALSA: Co-Design of Adiabatic Logic-based Sensing and Balanced Bit-Cells for Secure and Energy-Efficient MRAM | Wu Yang, Amit Degada, Himanshu Thapliyal | 2026-09-27 | 下载 | Magnetoresistive Random Access Memory (MRAM) technologies such as Spin-Transfer Torque (STT-MRAM) and Spin-Orbit Torque assisted (SOT-STT-MRAM) offer nonvolatility and low leakage, making them attract... |
| MorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision Transformers | Rachmad Vidya Wicaksana Putra, Amirhesam Jafari Rad, Muhammad Shafique | 2026-09-27 | 下载 | Spiking Vision Transformers (SViTs) are developed as an energy-efficient alternative to conventional ViTs for computer vision tasks at the edge. |
| Resource-Efficient Speculative Decoding for Long-Context LLM Serving | Fei Li, Song Liu, Shiqiang Nie, Jinyu Wang, Weiguo Wu | 2026-09-27 | 下载 | Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under constrained GPU memory. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence Modelling | Matei-Ioan Stan, Oliver Rhodes | 2026-09-27 | 下载 | A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities tha... |
| Validating Memory-Optimal Transformer Kernels on Real Hardware: From Formal Derivation to Measured Performance Across Two HPC Clusters | Lenore M. Mullin, Gaetan Hains | 2026-09-27 | 下载 | We validate memory-optimal cost functions for transformer kernels derived via the Mathematics of Arrays (MoA). Companion Papers I-IV formally derive kernels for attention forward, backward, fused forw... |
| QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents | Weiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, Jiemin Jiang, Wentao Yao, Chujie Zheng, JianWei Zhang | 2026-09-27 | 下载 | Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens ... |
| Performance vs Portability in Heterogeneous HPC Environments: Why Pre-execution Benchmarking is Required | Mindaugas Macernis | 2026-09-27 | 下载 | Cloud computing and high-performance computing (HPC) typically follow different paradigms: cloud services are often orchestrated using Kubernetes, whereas HPC workloads are managed through batch sched... |
| EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents? | Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li, Chaofan Tao, Yanli Wang, Huanxin Lin, Kwang-Ting Cheng, Chi Ying Tsui, Haoli Bai | 2026-09-27 | 下载 | LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when ... |
| Resource-Aware Parameter-Efficient Model Adaptation for Onboard High-Dimensional Data | Qiyang Zhang, Xinhao Li, Lei Shi, Zheng Lin, Jinfeng Wen, Ao Zhou, Shangguang Wang | 2026-09-27 | 下载 | Onboard satellite models often require frequent updates, but the weights adapted to earlier data distributions can quickly become outdated. However, updating large-scale model parameters in orbit pres... |
| Adaptive Client Clustering and Coordination for Federated Learning Workflow Management in Edge Networks | Jieping Luo, Qiyue Li, Yuxuan Chen, Hang Qi, Jiaying Yin, Jingjin Wu, Qian Wang | 2026-09-27 | 下载 | Federated learning (FL) is increasingly deployed as a managed learning service rather than as a set of isolated training jobs. In networked edge environments, dependent FL service flows must coordinat... |
| Toward System-of-Systems Integration for Composable Cloud-HPC-Edge AI Platforms | Sumit Rakesh, Rajkumar Saini | 2026-09-27 | 下载 | Modern AI platforms increasingly combine infrastructure stacks and operating models designed around different assumptions, including cloud-style service platforms, HPC workload-management systems, clo... |
| FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic Backward | Sriman Achanta | 2026-09-27 | 下载 | Autoregressive decode repeatedly streams a growing KV cache, making attention a major cost at long context. Existing high-performance kernels use online softmax, which discovers a row's normalization ... |
| OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading | Jingyuan Xiao, Jiayue Wang, Yitao Hu, Xinning Wang, Shi Chen, Ziqi Gong, Zhengchao Wang, Guotao Yang, Sheng Chen, Keqiu Li | 2026-09-27 | 下载 | Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a larg... |
| AgentLoop: Runtime Control of Slot-closed Execution Loops for Tool-augmented LLM Agents | Wanyi Zheng, Minxian Xu, Kan Hu, Kejiang Ye, Chengzhong Xu | 2026-09-27 | 下载 | Tool-augmented large language model (LLM) agents are becoming an important execution unit in service computing, but existing agent loops still lack explicit runtime signals for assessing task completi... |
| When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions | Dipankar Sarkar | 2026-09-27 | 下载 | Moving ML-mediated decision making onto privacy-preserving clients decentralises the economic decision along with the inference. Shared budget constraints then depend on information that cannot be glo... |
| CascadeEP: Asynchronous Expert Execution for MoE Prefill under Attention Imbalance | Jin Qin, Tiancheng Hu, Shiyan Wang, Junhao Hu, Zexin Jian, Yuzheng Wang, Haoyu Li, Chunwei Xia, Ying Liu, Pixian Zhan, Di Wang, Zhongzhe Hu, Huimin Cui, Tao Xie, Chenxi Wang | 2026-09-27 | 下载 | Mixture-of-experts (MoE) serving commonly deploys data and expert parallelism (DEP): attention replicas run distinct request batches while routed experts are sharded across an expert-parallel (EP) gro... |
| PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale | Zhiyuan Tan, Dejiang Zhu, Jingzhe Jiang, Yihao Zheng, Yang Tian, Tao Wang, Minchen Yu | 2026-09-27 | 下载 | Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared... |
| MpFA: Hardware-Efficient Train-Free QK4V8 FlashAttention Kernels on Blackwell GPUs | Chencheng Deng, Jianbin Fang, Dezun Dong | 2026-09-27 | 下载 | Long-context LLM inference pushes modern GPU serving stacks into an attention-bound regime, where both compute and memory are dominated by the softmax-GEMM pipeline. |
| Splitting Prompt Prefill from Response Replay for Context-Parallel Long-Context LLM Post-Training | Yubing Bao, Zhihui Lu, Qiang Duan, Yuedong Xu, Sen Liu, Pan Zhou | 2026-09-27 | 下载 | Training long-context LLM policies with RL requires re-evaluating groups of sampled responses under the updated policy, an update-stage attention workload that differs sharply from pre-training: each ... |
| ParallelPilot: Supporting Coordination and Monitoring in Parallel AI Coding | Tao Long, Weili Shi, Hussein Mozannar, Maya Murad, Rafah Hosn | 2026-09-27 | 下载 | As coding assistants become increasingly autonomous, developers run multiple sessions in parallel, shifting the challenge from code generation alone to coordinating and monitoring concurrent agent wor... |
| Hierarchical Secure Distributed Linearly Separable Computation with Arbitrary Heterogeneous Data Assignment | Ziting Zhang, Chenyi Sun, Kai Wan, Xiang Zhang | 2026-09-27 | 下载 | This paper studies secure distributed linearly separable computation over a three-layer hierarchical network, where clustered users communicate with a central server through relays. |
| SketchSSM: Write to the Full State, Read from a Compact Sketch | Omin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer, Sehoon Kim, Jae W. Lee | 2026-09-27 | 下载 | Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent state access becomes a major bottleneck... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Type-Safe Decision Frameworks for Agentic 5G Control: A Theory-Driven Testbed Characterization of Where They Can Be Applied | Michail-Alexandros Kourtis, George Xilouris | 2026-09-27 | 下载 | This paper presents a theory-driven characterization of type-safe decision frameworks for the agentic control of 5G networks, where every decision must be an element of a declared option set rather th... |
| Resource-Aware Parameter-Efficient Model Adaptation for Onboard High-Dimensional Data | Qiyang Zhang, Xinhao Li, Lei Shi, Zheng Lin, Jinfeng Wen, Ao Zhou, Shangguang Wang | 2026-09-27 | 下载 | Onboard satellite models often require frequent updates, but the weights adapted to earlier data distributions can quickly become outdated. However, updating large-scale model parameters in orbit pres... |
| A Novel Approach for the SDIR Epidemic Model on Online Social Networks | Nguyen Hong Phuc, Duong Khanh Ly, Hoang Phi Dung | 2026-09-27 | 下载 | Information diffusion can be controlled by restricting or removing links (edges) in online social networks, as well as in real-world networks. |
| StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC Networks | Mustafa Bora Çelik, Ceren Çelik, Orhan Gazi | 2026-09-27 | 下载 | In Integrated Sensing and Communications (ISAC), radar sensing must operate under chirp subsampling with up to 90% missing data. An attention-based baseline, limited to a 52~ms buffer, collapses towa... |
| SafePar: Monitoring Asynchrony in Microservices | Karuna Grewal, P. Brighten Godfrey, Justin Hsu, Umang Mathur | 2026-09-27 | 下载 | Modern cloud applications are built from loosely-coupled microservices that coordinate through well-defined APIs to service user requests. A single API request often triggers multiple downstream API c... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Where Activation Sparsity and KV-Cache Sparsity Cross in LLM Decoding | Jungseob Lee, Seungyoon Lee, Seongtae Hong, Sugyeong Eo, Heuiseok Lim | 2026-09-27 | 下载 | At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context. |
| JET: Justification Evaluation in Transformer | Shenghao Ding | 2026-09-27 | 下载 | JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across c... |
| Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs | Yirui Liu, Ruoling Qi, Xuaner Wu, Yuxin Jin, Jian Chen, Penghang Liu, Yafei Huang, Jiawei Shao, Xuelong Li | 2026-09-27 | 下载 | Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching. |