Skip to content

2026-09-27 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
QBX: A Compiler for 2-local Qubit Hamiltonian Simulation on Quantum ChipletsZikun Li, Zhuoming Chen, Zhihao Jia2026-09-27下载2-local qubit Hamiltonian simulation, a fundamental task in quantum computing, is widely applied in various applications. This paper presents QBX, the first quantum compiler designed for 2-local qubit...
S-ALSA: Co-Design of Adiabatic Logic-based Sensing and Balanced Bit-Cells for Secure and Energy-Efficient MRAMWu Yang, Amit Degada, Himanshu Thapliyal2026-09-27下载Magnetoresistive Random Access Memory (MRAM) technologies such as Spin-Transfer Torque (STT-MRAM) and Spin-Orbit Torque assisted (SOT-STT-MRAM) offer nonvolatility and low leakage, making them attract...
MorphAtt: A Neuromorphic Accelerator for Efficient Multi-Head Attention Processing in Spiking Vision TransformersRachmad Vidya Wicaksana Putra, Amirhesam Jafari Rad, Muhammad Shafique2026-09-27下载Spiking Vision Transformers (SViTs) are developed as an energy-efficient alternative to conventional ViTs for computer vision tasks at the edge.
Resource-Efficient Speculative Decoding for Long-Context LLM ServingFei Li, Song Liu, Shiqiang Nie, Jinyu Wang, Weiguo Wu2026-09-27下载Speculative decoding reduces sequential Target model calls by verifying multiple tokens from the Draft model in parallel. Yet KV Cache growth limits long-context serving under constrained GPU memory.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
ADPTNet: Adaptive with Prescriptive Timescales Non-Linear SSM for Sequence ModellingMatei-Ioan Stan, Oliver Rhodes2026-09-27下载A central aim of neuromorphic computing is to provide a viable alternative to highly energy-intensive Transformer-based AI. However, efficient alternatives struggle to capture the set of qualities tha...
Validating Memory-Optimal Transformer Kernels on Real Hardware: From Formal Derivation to Measured Performance Across Two HPC ClustersLenore M. Mullin, Gaetan Hains2026-09-27下载We validate memory-optimal cost functions for transformer kernels derived via the Mathematics of Arrays (MoA). Companion Papers I-IV formally derive kernels for attention forward, backward, fused forw...
QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon AgentsWeiqi Wang, Yuxin Zhou, Mouxiang Chen, Siyuan Zhang, Yi Zhang, Yuyan Luo, Zhiyu Yin, Chencan Wu, Jiemin Jiang, Wentao Yao, Chujie Zheng, JianWei Zhang2026-09-27下载Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens ...
Performance vs Portability in Heterogeneous HPC Environments: Why Pre-execution Benchmarking is RequiredMindaugas Macernis2026-09-27下载Cloud computing and high-performance computing (HPC) typically follow different paradigms: cloud services are often orchestrated using Kubernetes, whereas HPC workloads are managed through batch sched...
EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?Kunming Shao, Jierun Chen, Jiangnan Yu, Xiao-Hui Li, Chaofan Tao, Yanli Wang, Huanxin Lin, Kwang-Ting Cheng, Chi Ying Tsui, Haoli Bai2026-09-27下载LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when ...
Resource-Aware Parameter-Efficient Model Adaptation for Onboard High-Dimensional DataQiyang Zhang, Xinhao Li, Lei Shi, Zheng Lin, Jinfeng Wen, Ao Zhou, Shangguang Wang2026-09-27下载Onboard satellite models often require frequent updates, but the weights adapted to earlier data distributions can quickly become outdated. However, updating large-scale model parameters in orbit pres...
Adaptive Client Clustering and Coordination for Federated Learning Workflow Management in Edge NetworksJieping Luo, Qiyue Li, Yuxuan Chen, Hang Qi, Jiaying Yin, Jingjin Wu, Qian Wang2026-09-27下载Federated learning (FL) is increasingly deployed as a managed learning service rather than as a set of isolated training jobs. In networked edge environments, dependent FL service flows must coordinat...
Toward System-of-Systems Integration for Composable Cloud-HPC-Edge AI PlatformsSumit Rakesh, Rajkumar Saini2026-09-27下载Modern AI platforms increasingly combine infrastructure stacks and operating models designed around different assumptions, including cloud-style service platforms, HPC workload-management systems, clo...
FoldAttention: Declared-Reference Softmax for Fast Decode and Deterministic BackwardSriman Achanta2026-09-27下载Autoregressive decode repeatedly streams a growing KV cache, making attention a major cost at long context. Existing high-performance kernels use online softmax, which discovers a row's normalization ...
OLED-MoE: Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert OffloadingJingyuan Xiao, Jiayue Wang, Yitao Hu, Xinning Wang, Shi Chen, Ziqi Gong, Zhengchao Wang, Guotao Yang, Sheng Chen, Keqiu Li2026-09-27下载Semi-autoregressive diffusion large language models (dLLMs) improve decoding parallelism through iterative block-wise denoising, but scaling them with mixture-of-experts (MoE) layers introduces a larg...
AgentLoop: Runtime Control of Slot-closed Execution Loops for Tool-augmented LLM AgentsWanyi Zheng, Minxian Xu, Kan Hu, Kejiang Ye, Chengzhong Xu2026-09-27下载Tool-augmented large language model (LLM) agents are becoming an important execution unit in service computing, but existing agent loops still lack explicit runtime signals for assessing task completi...
When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in AuctionsDipankar Sarkar2026-09-27下载Moving ML-mediated decision making onto privacy-preserving clients decentralises the economic decision along with the inference. Shared budget constraints then depend on information that cannot be glo...
CascadeEP: Asynchronous Expert Execution for MoE Prefill under Attention ImbalanceJin Qin, Tiancheng Hu, Shiyan Wang, Junhao Hu, Zexin Jian, Yuzheng Wang, Haoyu Li, Chunwei Xia, Ying Liu, Pixian Zhan, Di Wang, Zhongzhe Hu, Huimin Cui, Tao Xie, Chenxi Wang2026-09-27下载Mixture-of-experts (MoE) serving commonly deploys data and expert parallelism (DEP): attention replicas run distinct request batches while routed experts are sharded across an expert-parallel (EP) gro...
PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at ScaleZhiyuan Tan, Dejiang Zhu, Jingzhe Jiang, Yihao Zheng, Yang Tian, Tao Wang, Minchen Yu2026-09-27下载Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared...
MpFA: Hardware-Efficient Train-Free QK4V8 FlashAttention Kernels on Blackwell GPUsChencheng Deng, Jianbin Fang, Dezun Dong2026-09-27下载Long-context LLM inference pushes modern GPU serving stacks into an attention-bound regime, where both compute and memory are dominated by the softmax-GEMM pipeline.
Splitting Prompt Prefill from Response Replay for Context-Parallel Long-Context LLM Post-TrainingYubing Bao, Zhihui Lu, Qiang Duan, Yuedong Xu, Sen Liu, Pan Zhou2026-09-27下载Training long-context LLM policies with RL requires re-evaluating groups of sampled responses under the updated policy, an update-stage attention workload that differs sharply from pre-training: each ...
ParallelPilot: Supporting Coordination and Monitoring in Parallel AI CodingTao Long, Weili Shi, Hussein Mozannar, Maya Murad, Rafah Hosn2026-09-27下载As coding assistants become increasingly autonomous, developers run multiple sessions in parallel, shifting the challenge from code generation alone to coordinating and monitoring concurrent agent wor...
Hierarchical Secure Distributed Linearly Separable Computation with Arbitrary Heterogeneous Data AssignmentZiting Zhang, Chenyi Sun, Kai Wan, Xiang Zhang2026-09-27下载This paper studies secure distributed linearly separable computation over a three-layer hierarchical network, where clustered users communicate with a central server through relays.
SketchSSM: Write to the Full State, Read from a Compact SketchOmin Kwon, JoongWon Shin, Minseo Kim, Kurt Keutzer, Sehoon Kim, Jae W. Lee2026-09-27下载Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent state access becomes a major bottleneck...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Type-Safe Decision Frameworks for Agentic 5G Control: A Theory-Driven Testbed Characterization of Where They Can Be AppliedMichail-Alexandros Kourtis, George Xilouris2026-09-27下载This paper presents a theory-driven characterization of type-safe decision frameworks for the agentic control of 5G networks, where every decision must be an element of a declared option set rather th...
Resource-Aware Parameter-Efficient Model Adaptation for Onboard High-Dimensional DataQiyang Zhang, Xinhao Li, Lei Shi, Zheng Lin, Jinfeng Wen, Ao Zhou, Shangguang Wang2026-09-27下载Onboard satellite models often require frequent updates, but the weights adapted to earlier data distributions can quickly become outdated. However, updating large-scale model parameters in orbit pres...
A Novel Approach for the SDIR Epidemic Model on Online Social NetworksNguyen Hong Phuc, Duong Khanh Ly, Hoang Phi Dung2026-09-27下载Information diffusion can be controlled by restricting or removing links (edges) in online social networks, as well as in real-world networks.
StarBOA: Real-Time Mamba State-Space Unrolling for Sparse Radar Micro-Doppler in ISAC NetworksMustafa Bora Çelik, Ceren Çelik, Orhan Gazi2026-09-27下载In Integrated Sensing and Communications (ISAC), radar sensing must operate under chirp subsampling with up to 90% missing data. An attention-based baseline, limited to a 52~ms buffer, collapses towa...
SafePar: Monitoring Asynchrony in MicroservicesKaruna Grewal, P. Brighten Godfrey, Justin Hsu, Umang Mathur2026-09-27下载Modern cloud applications are built from loosely-coupled microservices that coordinate through well-defined APIs to service user requests. A single API request often triggers multiple downstream API c...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Where Activation Sparsity and KV-Cache Sparsity Cross in LLM DecodingJungseob Lee, Seungyoon Lee, Seongtae Hong, Sugyeong Eo, Heuiseok Lim2026-09-27下载At each step, decoding one sequence with a large language model rereads the projection weights, whose traffic is fixed, and the key-value (KV) cache, whose traffic grows with context.
JET: Justification Evaluation in TransformerShenghao Ding2026-09-27下载JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across c...
Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMsYirui Liu, Ruoling Qi, Xuaner Wu, Yuxin Jin, Jian Chen, Penghang Liu, Yafei Huang, Jiawei Shao, Xuelong Li2026-09-27下载Hybrid LLMs interleave full-attention layers with linear-attention layers to reduce long-context inference cost, but this structure complicates prefix caching.

基于 VitePress 构建 · 使用本地搜索查找论文