2026-06-04
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Modeling, Optimizing and Exploring Multi-Die FPGA Routing Architectures | Amirhossein Poolad, Soheil Gholami Shahrouz, Andrew Boutros, Vaughn Betz | 2026-06-04 | 下载 | Die stacking has enabled 2.5D FPGAs by integrating multiple active dice on a passive silicon interposer for improved yield and capacity, and paved the way for 3D architectures that stack active dice d... |
| ITP-STDP: An Intrinsic-Timing Power-of-Two Learning Engine for On-Chip SNN Training | Haihang Xia, Xinyu Zhao, Xuecheng Wang, John Goodenough, Charith Abhayaratne, Panagiotis A. Panagiotou, Chunyi Song, Tiantai Deng | 2026-06-04 | 下载 | Spiking neural networks (SNNs) have the potential to emerge as the third generation of neural networks and have attracted increasing attention across a wide range of applications. |
| Space-CIM: Enabling Compute-In-Memory Accelerators for Thermally-Constrained Space Platforms | Sohan Salahuddin Mugdho, Md. Shahedul Hasan, Cheng Wang | 2026-06-04 | 下载 | The rapid growth in compute demand from artificial intelligence (AI) has driven a massive surge in data center construction, precipitating an energy and sustainability crisis. |
| CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMs | Mohammad Akyash, Nowfel Mashnoor, Kimia Azar, Hadi Kamali | 2026-06-04 | 下载 | Recent advances in large language models (LLMs) have enabled the automatic synthesis (generation) of register-transfer level (RTL) code from natural language instructions, offering a promising pathway... |
| FQA: A Full-Space Quantization-Driven Architecture for Hardware-Efficient Piecewise Approximation of Nonlinear Activation Functions | Chenjun Hao, Feng Yan, Hongbing Pan, Yuxuan Wang | 2026-06-04 | 下载 | In this paper, we propose a full-space quantization-driven architecture (FQA) for the hardware-efficient piecewise polynomial approximations (PPAs) of nonlinear activation functions. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| StageFrontier: Synchronization-Aware Stage Accounting for Distributed ML Training | Boram Yoon, Wei Chen, Ville Kallioniemi | 2026-06-04 | 下载 | When a distributed training job slows down, the hard part is knowing where to look. Synchronization hides the cause: a stall on one rank shows up as a wait on the others, so a data delay on a single r... |
| Towards Serverless Semi-Decentralized Federated Learning with Heterogeneous Optimizers | Su Wang, Mung Chiang, H. Vincent Poor | 2026-06-04 | 下载 | We investigate cluster formation, involving the number and composition of clusters, in decentralized federated learning (FL) with heterogeneous machine learning (ML) optimizers. |
| CarbonSim: A Lifecycle-Aware Framework for Evaluating Carbon Tradeoffs in Hardware Upgrade Decisions | Kartik Hans, Kaiwen Zhao, Stephen Lee | 2026-06-04 | 下载 | As the demand for information and communication technologies (ICT) continues to rise, the environmental impact of computing systems is becoming an increasingly critical concern. |
| On GPU Implementation for Multi-Precision Integer Division | Martin B. Marchioro, Aske N. Raahauge, Marc I. Løvenskjold, Cosmin E. Oancea, Stephen M. Watt | 2026-06-04 | 下载 | This paper presents the issues arising in implementing a fast integer division algorithm on general purpose GPUs. The algorithm uses a Newton iteration based on the shifted inverse operation, keeping ... |
| Discrete Incremental Voting: New Bounds for General Graphs and Expanders | Petra Berenbrink, Colin Cooper, Thorsten Götte, Lukas Hintze, Tomasz Radzik | 2026-06-04 | 下载 | We analyze the discrete incremental voting process (DIV) introduced by Cooper, Radzik, and Shiraga [OPODIS '23]. In this process, we consider a set of nodes connected in an undirected graph $G... |
| RadiusFPS: Efficient Farthest Point Sampling on CPUs and GPUs via Spherical Voxel Pruning | Ziyang Yu, Xiang Li, Qiong Chang, Jun Miyazaki | 2026-06-04 | 下载 | Point clouds are a primary sensory representation for robotic perception, underpinning LiDAR-based autonomous driving, simultaneous localization and mapping (SLAM), and navigation. |
| LLM-Based Porting of Optimized C++ to CUDA Through Deoptimization and Reoptimization | Daichi Mukunoki, Ryo Mikasa, Shunichiro Hayashi, Tetsuya Hoshino, Takahiro Katagiri | 2026-06-04 | 下载 | When porting high-performance computing (HPC) code from CPU to GPU, CPU-oriented optimizations may obstruct LLM-based CUDA translation. We design and evaluate a Deopt-Reopt workflow that first simplif... |
| Demystifying NVSHMEM: A System-Level Analysis on Symmetric Memory and Device-Initiated Operations in GPU Communication | Yijun Ma, Siyuan Shen, Tiancheng Chen, Akhil Langer, Jiri Kraus, Benjamin Glick, Craig Belusar, Jeff Hammond, Torsten Hoefler | 2026-06-04 | 下载 | NVSHMEM is NVIDIA's OpenSHMEM-based PGAS communication library for GPU clusters, enabling GPU-initiated, one-sided communication through symmetric memory. |
| Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM Inference | Yuansheng Chen, Yue Zhang, Xuan Mo, Weigang Wu, Jialun Li | 2026-06-04 | 下载 | With the rapid growth of interactive applications in large language model (LLM) online services, maintaining high system throughput while ensuring user-perceived latency has become a key issue in infe... |
| IN2P3 Computing Center 2024 Workload Dataset | Guillaume Cochard, Bertrand Simon | 2026-06-04 | 下载 | This paper provides and analyzes a dataset detailing the characteristics and execution data of all jobs submitted to the IN2P3 Computing Center (Villeurbanne, France), a national research and support ... |
| PoCQ: Proof of Contribution Quality as a Lightweight Blockchain Consensus for Secure Federated Learning | Sudad Abed, Nasser Sabar, Abdun Mahmood, Mohammad Jabed Morshed Chowdhury | 2026-06-04 | 下载 | Decentralized Federated Learning (FL) removes reliance on centralized coordinators but remains vulnerable to model poisoning, unreliable validation, and high validation overhead. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Natural Language Access Control (NLAC): From Help Desk Requests to Structured Policies | Jonas Wessner, Tobias Meuser, Janek Schoffit, Dennis Eisermann, Johannes Deger, Björn Scheuermann, Frank Kargl | 2026-06-04 | 下载 | Configuring network access control policies in large, complex networks is error-prone and requires significant expert effort. LLMs offer a promising interface for expressing such policies in natural l... |
| Towards Serverless Semi-Decentralized Federated Learning with Heterogeneous Optimizers | Su Wang, Mung Chiang, H. Vincent Poor | 2026-06-04 | 下载 | We investigate cluster formation, involving the number and composition of clusters, in decentralized federated learning (FL) with heterogeneous machine learning (ML) optimizers. |
| DAST: A VLM-LLM Framework for Cross-Interface Anomaly Detection in O-RAN | Francesco Spinelli, Esteban Municio, Pau Baguer, Gines Garcia-Aviles, Xavier Costa-Perez | 2026-06-04 | 下载 | O-RAN enables a disaggregated baseband stack with programmable functions that communicate over standardized open interfaces. The same openness that enables multi-vendor composition also expands the at... |
| Toward Mobile and Converged Backhaul: The Promise of Wireless Access and Backhaul | Chiara Rubaltelli, Marcello Morini, Eugenio Moro, Ilario Filippini, Antonio Capone | 2026-06-04 | 下载 | Wireless Access and Backhaul (WAB) is emerging as a key enabler for flexible and cost-efficient 5G deployments, offering a modular architecture that decouples access and backhaul while supporting mult... |
| Policy-Guided ML for Energy Savings: Cell On/Off Switching under Operator QoS Constraints in Real 5G Networks | D. Reiss, M. Catalan-Cid, D. Camps-Mur, O. Sallent | 2026-06-04 | 下载 | Energy efficiency is a critical concern in the deployment and operation of 5G networks, particularly due to the low utilization of 4G and 5G carriers during off-peak hours. |
| Quantifying the Energy-Saving and QoS Trade-Off in Traffic Offloading for Real 4G/5G Scenarios | D. Reiss, M. Catalan-Cid, D. Camps-Mur, O. Sallent | 2026-06-04 | 下载 | Despite the potential for higher energy efficiency in 5G networks, current 5G Non-Standalone (NSA) deployments often operate suboptimally due to low utilization of 4G and 5G carriers during extended p... |
| BeGREEN Intelligent Plane for AI-driven Energy Efficient O-RAN management | M. Catalan-Cid, J. Pueyo, J. Sanchez-Gonzalez, J. Gutierrez, M. Ghoraishi | 2026-06-04 | 下载 | Cellular networks are undergoing a revolutionary transform with the advent of O-RAN architectures and AI/ML solutions. O-RAN's Non-Real-Time and Near-Real Time RAN Intelligent Controllers open the doo... |
| AISC deployment in dynamic UAV-assisted MEC network: a reinforcement learning method based on heterogeneous graph attention neural network | Hanzhi Chang, Jing Bai, Xin Tang, Xiaomei Liu | 2026-06-04 | 下载 | Unmanned aerial vehicles-assisted mobile edge computing (UMEC) can execute compute-intensive and latency-critical artificial intelligence (AI) services, which can be provided by multiple UAVs collabor... |
| GS-NFS: Bandwidth-adaptive Streaming of Dynamic Gaussian Splats and Point Clouds | Rajrup Ghosh, Haodong Wang, Haoran Hong, Eduardo Pavez, Amartya Chaudhuri, Weiwu Pang, Harsha V. Madhyastha, Antonio Ortega, Ramesh Govindan | 2026-06-04 | 下载 | Dynamic 3D Gaussian Splatting (3DGS) holds great promise as a 3D video streaming technology since it can represent complex 3D scenes with high fidelity. |
| Availability-Aware and Efficiency-Driven AI Service Chain Provisioning in Multi-Domain Edge Intelligence Cloud | Hanzhi Chang, Jing Bai, Xin Tang, Xiaomei Liu, Yiming Chen | 2026-06-04 | 下载 | In a multi-domain edge intelligence cloud (MDEIC) managed by multiple network operators, AI services are delivered by chains of virtual network functions (VNFs) executed in sequence, called AI service... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AgileOS: A GPU Operating System Layer for Protected CUDA Services | Zhuoping Yang, Yiyu Shi, Alex Jones, Peipei Zhou | 2026-06-04 | 下载 | Modern GPU applications increasingly interact with storage systems, network devices, vendor libraries, and GPU-resident services rather than executing only isolated compute kernels. |
| CarbonSim: A Lifecycle-Aware Framework for Evaluating Carbon Tradeoffs in Hardware Upgrade Decisions | Kartik Hans, Kaiwen Zhao, Stephen Lee | 2026-06-04 | 下载 | As the demand for information and communication technologies (ICT) continues to rise, the environmental impact of computing systems is becoming an increasingly critical concern. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AEGIS: A Backup Reflex for Physical AI | Josef Chen | 2026-06-04 | 下载 | Long-horizon robot manipulation tends to fail gradually: one bad step degrades the state, and the policy spirals into a basin from which it cannot recover. |
| rsx: A high-performance streaming toolkit for RAD-seq sex determination | Rohit Goswami, Ruhila Goswami | 2026-06-04 | 下载 | Restriction site-associated DNA sequencing (RAD-seq) is widely used to discover sex-linked markers in non-model organisms, but large studies produce marker tables with millions of RAD tags. |
| PivCo-Huffman | Marcin Zukowski | 2026-06-04 | 下载 | Huffman encoding has been an enduring technique for 70+ years, ubiquitous in compression algorithms since its invention. In this paper we propose a new approach to Huffman coding, based on a data stru... |