2026-09-29
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Structure-augmented LLMs for High-Level Synthesis Pragma Optimization | Haocheng Xu, Ye Qiao, Phyo Pyae Moe Aung, Alok Mishra, Pavana Prakash, Rolando Pablo Hong Enriquez, Adam Han Wu, Zhiheng Chen, Dejan Milojicic, Sitao Huang | 2026-09-29 | 下载 | Pragma insertion drives the quality of high-level synthesis (HLS) designs. Choosing the right directives demands expert knowledge and reasoning about loop nesting, data dependences, and memory layout. |
| NDS: Programmer-Free Offload of High-Performance Near Data Strands | Shreyas Singh, Pratyush Nandi, Lin Jia, Shankar Balachandran, Rajeev Balasubramonian | 2026-09-29 | 下载 | Near Data Processing (NDP) has the potential to significantly improve system performance and energy by alleviating data movement bottlenecks. However, most NDP proposals pose heavy requirements for th... |
| Zephyr: An Efficient Audio Denoising System Using Spiking Neural Networks Enabled With A Sparsity-Aware Flexible FPGA PE Array | Cheng-En Chang, Chi-Wei Kao, Chung-Lun Yang, Yan-Lin Jiang, Yi-Chen Huang, Sebastian Fieldhouse, Kea-Tiong Tang | 2026-09-29 | 下载 | In this work we look to neuromorphic computing to solve the power consumption problem that audio denoising neural networks face on edge devices like smartphones, wireless headphones and hearing aids. |
| MEDEM: Multi-Engine DL Accelerator Design Methodology | Fareed Qararyah, Mohammad Ali Maleki, Pedro Trancoso | 2026-09-29 | 下载 | Multi-engine deep learning (DL) accelerators are becoming increasingly prevalent as they address the heterogeneity and growing complexity of modern DL workloads. |
| Mixed-Precision Computing for Scientific Discovery: Formats, Co-Design, and Responsible Approximation | Emmanuel Agullo, Hartwig Anzt, Daniel Bauer, David Bindel, Alfredo Buttari, Alexandru Calotoiu, Erin Claire Carson, Pasqua D'Ambra, Ieva Daužickaitė, James W. Demmel, Jack Dongarra, Iain Duff, Massimiliano Fasi, Dominik Göddeke, Stef Graillat, Laslo Hunhold, Roman Iakymchuk, Fabienne Jézéquel, Nils Kohl, Harald Köstler, Jakub Kružík, Julien Langou, Xiaoye Sherry Li, Hatem Ltaief, Piotr Luszczek, Yuxin Ma, Theo Mary, Mantas Mikaitis, Hiroyuki Ootomo, Daniel Osei-Kuffuor, Enrique S. Quintana-Ortí, Ulrich Rüde, Jennifer Scott, John Shalf, Linda Stals, Rasmus Tamstorf, Stefan Turek, Petr Vacek, Bastien Vieublé, Rio Yokota | 2026-09-29 | 下载 | Reduced and mixed precision have moved from a niche optimization to a central design axis in scientific computing and engineering, driven by energy constraints, heterogeneous accelerators, and the con... |
| Low-level optimizations in high-level HDLs: Is there a benefit? | Oliver Keszocze, Tjark Petersen, Arved Friedemann, Matthias Bo Stuart | 2026-09-29 | 下载 | This paper explores the applicability of functional programming to the design of Application-specific Integrated Circuits (ASICs). We investigate the impact of designing ASICs using high-level, abstra... |
| cktFormer: Transformer-Based Approach for Automated Analog Circuit Design | Pasindu Dodampegama, Praveen Wijesinghe, Naveen Basnayake, Keshawa Jayasundara, Tharindu Bandaragoda | 2026-09-29 | 下载 | Circuit design is a complex and iterative process that requires expertise in electronic engineering. It involves selecting components while meeting performance constraints, such as power efficiency, c... |
| Efficient Linkage-Based Compartmentalization on CHERI | Dapeng Gao, John Baldwin, Jessica Clarke, Nicholas C. Connolly, Brooks Davis, Franz A. Fuchs, Alfredo Mazzinghi, Daniel Moghimi, Peter Rugg, Domagoj Stolfa, Konrad Witaszczyk, Simon W. Moore, Robert N. M. Watson | 2026-09-29 | 下载 | We present an efficient linkage-based model for in-process compartmentalization built on CHERI memory safety, which enables fine-grained compartmentalization of the entire UNIX user-space, scaling to ... |
| Lossless Compression of Lookup Tables for Hardware Applications | Alireza Khataei, Kia Bazargan | 2026-09-29 | 下载 | Large lookup tables are widely used in hardware to store constant-valued arrays for applications ranging from elementary mathematical operations, such as constant-coefficient multiplication and nonlin... |
| Making Analog Training Scale: Co-Designing Mapping, Optimizer, and Converters | Zhaoxian Wu, Tayfun Gokmen, Omobayode Fagbohungbe, T. Patrick Xiao, Tianyi Chen | 2026-09-29 | 下载 | Analog in-memory computing (AIMC) offers an alternative for model training by executing matrix operations directly where weights are stored. However, scaling AIMC to train modern deep models remains a... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Enabling Efficient Client Selection in FL-as-a-Service for Multi-Application based Society 5.0 | Prachi Nandi, Sonakshi Satpathy, Timam Ghosh, Arijit Roy | 2026-09-29 | 下载 | The rapid development of the Internet of Things (IoT) has led to the generation of vast amounts of data from sensors, prompting the need for advanced learning models to analyze this data for personali... |
| RLX: A Unified Multi-Backend Tensor Compiler and Distributed Runtime in Rust | Eugene Hauptmann, Nataliya Kosmyna | 2026-09-29 | 下载 | Production machine learning (ML) stacks often split graph compilation and kernel execution across different layers and languages, making backend behavior, deployment guarantees, and performance fallba... |
| Byzantine Causal Reliable Broadcast with Constant Metadata Overhead | Purv Patel, Ajay D. Kshemkalyani | 2026-09-29 | 下载 | Asynchronous Byzantine Reliable Broadcast (BRB) is a fundamental primitive that guarantees agreement and validity in distributed systems subject to Byzantine faults, but it lacks ordering guarantees. |
| When Correct Memory Goes Wrong: Fuzzing Persistent Memory Use in LLM Agents | Yuqiao Meng, Luoxi Tang, Yingxue Zhang, Yuchen Yang, Zhaohan Xi | 2026-09-29 | 下载 | Persistent memory helps LLM agents carry information across long interactions, but correct memory can still be used incorrectly when queries change or memory states evolve. |
| Joint Effects of GPU Server Topology, Parallelism, and Congestion Control on MoE Inference: A Controlled Simulation Study | Kaikai Yuan, Rui Xi, Yu Liu | 2026-09-29 | 下载 | Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. |
| FP64 Is All You Want, INT8 Is All You Need, FP4/6/8 Is All You Have | Pratyai Mazumder, Alexandru Calotoiu, Torsten Hoefler | 2026-09-29 | 下载 | Ozaki scheme II emulates FP64 matrix products with INT8 ones through residues modulo pairwise coprime moduli, and variants for FP8 and FP4 have followed. |
| SPLASH: Switching Parallel Layouts of Attention with Seamless Handoff for LLM Serving | Chuan Liu, Shuoming Zhang, Zhicheng Li, Qianqi Sun, Ruiyuan Xu, Qiuchu Yu, Xiyu Shi, Huimin Cui, Jiacheng Zhao | 2026-09-29 | 下载 | No single way of parallelizing attention serves large language models well under all loads. Low concurrency favors tensor parallelism, many independent requests favor data-parallel attention, and long... |
| Janus: Evidence-Before-Effect Sagas and Offline-Verifiable Provenance for Agentic LLMs | Mustafa Arslan | 2026-09-29 | 下载 | Agentic large language models (LLMs) now move money through tools, yet the record of what they did is usually a trace their own process emits beside the effect. |
| DScale: Scaling Block-Diffusion Speculative Decoding with Adaptive Verification | Rongjian Chen, Minxian Xu, Zhengxin Fang, Kejiang Ye, Chengzhong Xu | 2026-09-29 | 下载 | Growing large language model applications demand efficient inference. At high concurrency, block-diffusion speculative decoding suffers from verification padding, rejected candidates, and incompatibil... |
| MEDEM: Multi-Engine DL Accelerator Design Methodology | Fareed Qararyah, Mohammad Ali Maleki, Pedro Trancoso | 2026-09-29 | 下载 | Multi-engine deep learning (DL) accelerators are becoming increasingly prevalent as they address the heterogeneity and growing complexity of modern DL workloads. |
| Encoding and Node Choices in Transversal Fault-Tolerant Distributed Quantum Computations: An Initial Study | Seng W. Loke | 2026-09-29 | 下载 | We compare and study different Bivariate-Bicycle (BB) encodings and node choices for distributed quantum operations such as transversal non-local CNOTs. |
| ARGOS: Reinforcement Learning-Driven Multidimensional Elasticity for Service Orchestration in the Computing Continuum | Javier Mateos-Bravo, Sergio Laso, Juan Luis Herrera, Ilir Murturi, Pantelis Frangoudis, Schahram Dustdar | 2026-09-29 | 下载 | Data-intensive services in the Computing Continuum must balance analytics quality, resource usage, and cost across heterogeneous nodes with limited and uneven capacity. |
| Windowed and Quantized Group-Based ADMM for Distributed Optimization in Heterogeneous Edge Networks | Gaiguo Wei, Qingying Zhang, Heqiang Wang, Yu Zhang, Xiaoxiong Zhong | 2026-09-29 | 下载 | Distributed optimization in edge networks is constrained by heterogeneous client computing capabilities and limited communication resources. We propose the Windowed and Quantized Group-Based Alternati... |
| vSkipper: Translating Dynamic Layer Skipping into LLM Serving Gains | Wei Da, Yavuz Ferhatosmanoglu, Evangelia Kalyvianaki | 2026-09-29 | 下载 | Dynamic layer skipping reduces LLM computation by allowing each token to execute only a subset of the model's layers. However, existing skippers rely on specialized generation loops and do not integra... |
| CF-LoRA: Decoupled Factor Aggregation and Adaptation-Aware Client Clustering for Federated LoRA Fine-Tuning | Mengjun Yi, Langxing Yang, Suhan Guo, Furao Shen, Jian Zhao | 2026-09-29 | 下载 | Federated LoRA fine-tuning enables parameter-efficient adaptation of pre-trained models without sharing private data, but suffers from two fundamental mismatches under heterogeneous client data: a str... |
| Cobalt: Leveraging Expert Co-activation for Efficient Distributed MoE Training | Junkang Zhou, Xinyi Liu, Fangcheng Fu | 2026-09-29 | 下载 | Mixture-of-Experts (MoE) has increasingly become a mainstream approach for scaling large language models, as it expands model capacity while keeping computation cost nearly constant. |
| Purlin: Separating Orchestration from the Datapath of Collectives | Osayamen Jonathan Aimuyo, Swapnil Gandhi, Christos Kozyrakis | 2026-09-29 | 下载 | Distributed inference depends on GPU collective communication that must keep pace with evolving hardware and specialized workloads. However, existing collective implementations often couple semantics,... |
| Efficient Agentic LLM Serving over SSD-based Sparse KV Storage | Wenhao He, Ping Zhang, Xiaohe Hu, Chutian Wang, Jinlong Hou, Yuan Cheng, Peng Sun, Fangcheng Fu | 2026-09-29 | 下载 | Agentic sessions driven by Large language models (LLMs) often alternate between model inference and tool use, accumulating long histories across successive rounds. |
| Reshaping Rollout Workloads for Asynchronous RL Post-Training on Heterogeneous Accelerators | Jiahui Li, Hao Nie, Yibo Zhu, Pengjin Xie, Yu Zhou, Xiaolong Zheng, Liang Liu, Huadong Ma | 2026-09-29 | 下载 | Reinforcement learning (RL) post-training increasingly relies on long-horizon, multi-turn rollouts. As post-training jobs outgrow a single cluster, rollout pools assembled across clusters introduce ha... |
| Federated Clustering with Unknown Local and Global Cluster Cardinalities | Mitushi Goyal, Tarun S., Riddhanya Senapathi, Arun Raman | 2026-09-29 | 下载 | Federated clustering methods that do not require the global number of clusters still assume that each client knows its local number . |
| Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference | Ruiyi Ding, Jie Li, Kang He, Ziyan Liu, Chengru Song, Yuedong Xu, Yuan Cheng | 2026-09-29 | 下载 | Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. |
| ParaAnya: Accelerating Parallel Diffusion Sampling with Plug-and-Play Output Caching | Chee-En Yu, Xiao-Xi Tan, Yi-Cheng Lin, Yun-Shao Tsai, Chee-An Yu, Hung-yi Lee | 2026-09-29 | 下载 | Diffusion models have achieved remarkable success in generative tasks, but their inherently sequential sampling process introduces a severe computational bottleneck. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Enabling Efficient Client Selection in FL-as-a-Service for Multi-Application based Society 5.0 | Prachi Nandi, Sonakshi Satpathy, Timam Ghosh, Arijit Roy | 2026-09-29 | 下载 | The rapid development of the Internet of Things (IoT) has led to the generation of vast amounts of data from sensors, prompting the need for advanced learning models to analyze this data for personali... |
| Priority-Aware Routing for Quantum Networks:Integrating Coherence-Time Constraints into Scheduling | Sadhgun Ram Dasi, Aswath Babu H | 2026-09-29 | 下载 | Quantum networks face a fundamental challenge absent in classical networks; finite memory coherence times mean that queuing delay directly degrades information quality, causing decoherence that has no... |
| Pricing IoT Data Delivered via LEO Satellites | Rahul Ramachandran, Suman Banerjee | 2026-09-29 | 下载 | IoT terminals served by LEO satellite constellations transmit data to passing satellites in discrete uplink windows. That data is delivered to buyers only when the satellite reaches a ground station. |
| Sequence Models for Layer-3 Protocol Emulation | Alix Jeannerot, Petko Petkov, Alvaro Valcarce Rial | 2026-09-29 | 下载 | This article investigates whether Layer-3 radio-protocol behavior can be represented by compact sequence models suitable for deployment inside the RAN. |
| NetLexicon: Learning Discrete Behavioral Representations for Encrypted Web Traffic Analysis | Xiangyu Gao, Tong Li, Ziqiang Wang, Yinchao Zhang, Rongbang Wu, Zhenxing Zhang, Jing Hu, Hanlin Huang, Xinle Du, Su Yao, Qi Li, Ke Xu | 2026-09-29 | 下载 | Encrypted Web traffic analysis requires effective representations of observable communication behavior. Existing pretraining methods often adapt NLP/CV objectives and sequence architectures, motivatin... |
| RingStitch: Demand-Aware Optical Stitching for Fragmented TPU Clusters | Huiru Ao, Fan Yang, Binglei Wang, Bo Liu, Jialong Li | 2026-09-29 | 下载 | Large-scale AI training clusters increasingly use optical circuit switching (OCS) to reconfigure rack-level interconnects and create elastic accelerator slices. |
| Optimal Entanglement Routing in Quantum Repeater Chains: Beyond Fixed Operation Order and Purification Schedule | Aikaterini Mandilara, Antonia Tsili, Dimitris Syvridis, Konstantinos Christodoulopoulos | 2026-09-29 | 下载 | Entanglement routing establishes entangled pairs between distant nodes of a quantum network by purifying and swapping pairs generated on elementary links. |
| Exponential Backoff: Meta-Stability and Implicit Admission Control | Richard Combes, Fabien Mathieu, Thomas Bonald | 2026-09-29 | 下载 | We analyze exponential backoff, an algorithm used to share a single communication channel in a distributed manner between several users, in a similar way as many networking standards such as 802.11. |
| DSWM: Decomposed Spatio-Temporal World Model for Demand-Driven UAV Base Station Repositioning | Shengjie Zhong, Zhongliang Zhao, Jingxuan Chen, Xianbin Cao, Xinmei Qiang, Dapeng O. Wu, Tony Q. S. Quek | 2026-09-29 | 下载 | Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot ... |
| SCORAS-MoE: Joint Compression and Resource-Adaptive Deployment of MoE-VLMs in LEO Satellite Networks | Tong Quan, Yuanlong Wan, Huasen He, Yunpeng Hou, Shuangwu Chen, Xiaofeng Jiang, Jian Yang | 2026-09-29 | 下载 | Deploying large vision-language models (VLMs) onboard satellites enables onboard data processing and reduces raw data downlink. However, onboard inference faces two resource challenges. |
| On the Lack of Periodicity of Walker Satellite Constellation Routing Tables | Chang-Sik Choi, François Baccelli | 2026-09-29 | 下载 | In a referential that rotates with Earth, the dynamics of the configurations of satellites in a Delta Walker constellation can be analyzed as a dynamical system as a function of a translation on the t... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| StateFork: Branchable Infrastructure for Agent Exploration | Jiakai Xu, Tianle Zhou, Georgios Liargkovas, Danielle Gillai, Ruizhe Fu, Patrick Shen, Eugene Wu, Kostis Kaffes | 2026-09-29 | 下载 | AI agents improve task success by exploring multiple trajectories, but for computer-use agents each trajectory modifies external environment state. |
| ContractWarden: Kernel-Enforced Damage Boundaries for AI Agents via Human-Authorized Contracts | Dongxu Cui, Zhichao Gu, Ping Zheng, Wenshuai Xi, Simeng Han, Yong Liao | 2026-09-29 | 下载 | Large language model agents can execute commands, create subprocesses, and directly access files and networks, allowing prompt injection or planning errors to become operating-system side effects. |
| Agent-Warden: eBPF-Based Kernel-Native Process-File Provenance Tracking for LLM Agents | Dongxu Cui, Zhichao Gu, Ping Zheng, Simeng Han, Yong Liao | 2026-09-29 | 下载 | LLM agents execute dynamically generated process and file operations that are often invisible to application-layer tracing. We present Agent-Warden, an extended Berkeley Packet Filter (eBPF)-based pro... |
| Efficient Linkage-Based Compartmentalization on CHERI | Dapeng Gao, John Baldwin, Jessica Clarke, Nicholas C. Connolly, Brooks Davis, Franz A. Fuchs, Alfredo Mazzinghi, Daniel Moghimi, Peter Rugg, Domagoj Stolfa, Konrad Witaszczyk, Simon W. Moore, Robert N. M. Watson | 2026-09-29 | 下载 | We present an efficient linkage-based model for in-process compartmentalization built on CHERI memory safety, which enables fine-grained compartmentalization of the entire UNIX user-space, scaling to ... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AgBench: Agentic AI Benchmarks for Personal AI Devices | Yizhou Han, Di Wu, Dhananjay Saikumar, Blesson Varghese | 2026-09-29 | 下载 | Agentic AI systems increasingly rely on cloud-hosted large language models for planning, tool use, and iterative execution, raising concerns about API cost and data exposure. |
| Formal Reasoning about Performance Models | Moussa Labbadi, Rupak Majumdar, V. R. Sathiyanarayana, Sadegh Soudjani | 2026-09-29 | 下载 | Discrete-event simulation is a standard technique for modelling and analysing the performance of computer systems, networks, and services. Although simulation tools are widely used, reasoning about th... |
| MEDEM: Multi-Engine DL Accelerator Design Methodology | Fareed Qararyah, Mohammad Ali Maleki, Pedro Trancoso | 2026-09-29 | 下载 | Multi-engine deep learning (DL) accelerators are becoming increasingly prevalent as they address the heterogeneity and growing complexity of modern DL workloads. |