Skip to content

2026-06-04 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Modeling, Optimizing and Exploring Multi-Die FPGA Routing ArchitecturesAmirhossein Poolad, Soheil Gholami Shahrouz, Andrew Boutros, Vaughn Betz2026-06-04下载Die stacking has enabled 2.5D FPGAs by integrating multiple active dice on a passive silicon interposer for improved yield and capacity, and paved the way for 3D architectures that stack active dice d...
ITP-STDP: An Intrinsic-Timing Power-of-Two Learning Engine for On-Chip SNN TrainingHaihang Xia, Xinyu Zhao, Xuecheng Wang, John Goodenough, Charith Abhayaratne, Panagiotis A. Panagiotou, Chunyi Song, Tiantai Deng2026-06-04下载Spiking neural networks (SNNs) have the potential to emerge as the third generation of neural networks and have attracted increasing attention across a wide range of applications.
Space-CIM: Enabling Compute-In-Memory Accelerators for Thermally-Constrained Space PlatformsSohan Salahuddin Mugdho, Md. Shahedul Hasan, Cheng Wang2026-06-04下载The rapid growth in compute demand from artificial intelligence (AI) has driven a massive surge in data center construction, precipitating an energy and sustainability crisis.
CASS-RTL: Correctness-Aware Subspace Steering for RTL Generation with LLMsMohammad Akyash, Nowfel Mashnoor, Kimia Azar, Hadi Kamali2026-06-04下载Recent advances in large language models (LLMs) have enabled the automatic synthesis (generation) of register-transfer level (RTL) code from natural language instructions, offering a promising pathway...
FQA: A Full-Space Quantization-Driven Architecture for Hardware-Efficient Piecewise Approximation of Nonlinear Activation FunctionsChenjun Hao, Feng Yan, Hongbing Pan, Yuxuan Wang2026-06-04下载In this paper, we propose a full-space quantization-driven architecture (FQA) for the hardware-efficient piecewise polynomial approximations (PPAs) of nonlinear activation functions.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
StageFrontier: Synchronization-Aware Stage Accounting for Distributed ML TrainingBoram Yoon, Wei Chen, Ville Kallioniemi2026-06-04下载When a distributed training job slows down, the hard part is knowing where to look. Synchronization hides the cause: a stall on one rank shows up as a wait on the others, so a data delay on a single r...
Towards Serverless Semi-Decentralized Federated Learning with Heterogeneous OptimizersSu Wang, Mung Chiang, H. Vincent Poor2026-06-04下载We investigate cluster formation, involving the number and composition of clusters, in decentralized federated learning (FL) with heterogeneous machine learning (ML) optimizers.
CarbonSim: A Lifecycle-Aware Framework for Evaluating Carbon Tradeoffs in Hardware Upgrade DecisionsKartik Hans, Kaiwen Zhao, Stephen Lee2026-06-04下载As the demand for information and communication technologies (ICT) continues to rise, the environmental impact of computing systems is becoming an increasingly critical concern.
On GPU Implementation for Multi-Precision Integer DivisionMartin B. Marchioro, Aske N. Raahauge, Marc I. Løvenskjold, Cosmin E. Oancea, Stephen M. Watt2026-06-04下载This paper presents the issues arising in implementing a fast integer division algorithm on general purpose GPUs. The algorithm uses a Newton iteration based on the shifted inverse operation, keeping ...
Discrete Incremental Voting: New Bounds for General Graphs and ExpandersPetra Berenbrink, Colin Cooper, Thorsten Götte, Lukas Hintze, Tomasz Radzik2026-06-04下载We analyze the discrete incremental voting process (DIV) introduced by Cooper, Radzik, and Shiraga [OPODIS '23]. In this process, we consider a set VV of nn nodes connected in an undirected graph $G...
RadiusFPS: Efficient Farthest Point Sampling on CPUs and GPUs via Spherical Voxel PruningZiyang Yu, Xiang Li, Qiong Chang, Jun Miyazaki2026-06-04下载Point clouds are a primary sensory representation for robotic perception, underpinning LiDAR-based autonomous driving, simultaneous localization and mapping (SLAM), and navigation.
LLM-Based Porting of Optimized C++ to CUDA Through Deoptimization and ReoptimizationDaichi Mukunoki, Ryo Mikasa, Shunichiro Hayashi, Tetsuya Hoshino, Takahiro Katagiri2026-06-04下载When porting high-performance computing (HPC) code from CPU to GPU, CPU-oriented optimizations may obstruct LLM-based CUDA translation. We design and evaluate a Deopt-Reopt workflow that first simplif...
Demystifying NVSHMEM: A System-Level Analysis on Symmetric Memory and Device-Initiated Operations in GPU CommunicationYijun Ma, Siyuan Shen, Tiancheng Chen, Akhil Langer, Jiri Kraus, Benjamin Glick, Craig Belusar, Jeff Hammond, Torsten Hoefler2026-06-04下载NVSHMEM is NVIDIA's OpenSHMEM-based PGAS communication library for GPU clusters, enabling GPU-initiated, one-sided communication through symmetric memory.
Beyond Greedy Chunking: SLO-Aware Sliding-Window Scheduling for LLM InferenceYuansheng Chen, Yue Zhang, Xuan Mo, Weigang Wu, Jialun Li2026-06-04下载With the rapid growth of interactive applications in large language model (LLM) online services, maintaining high system throughput while ensuring user-perceived latency has become a key issue in infe...
IN2P3 Computing Center 2024 Workload DatasetGuillaume Cochard, Bertrand Simon2026-06-04下载This paper provides and analyzes a dataset detailing the characteristics and execution data of all jobs submitted to the IN2P3 Computing Center (Villeurbanne, France), a national research and support ...
PoCQ: Proof of Contribution Quality as a Lightweight Blockchain Consensus for Secure Federated LearningSudad Abed, Nasser Sabar, Abdun Mahmood, Mohammad Jabed Morshed Chowdhury2026-06-04下载Decentralized Federated Learning (FL) removes reliance on centralized coordinators but remains vulnerable to model poisoning, unreliable validation, and high validation overhead.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Natural Language Access Control (NLAC): From Help Desk Requests to Structured PoliciesJonas Wessner, Tobias Meuser, Janek Schoffit, Dennis Eisermann, Johannes Deger, Björn Scheuermann, Frank Kargl2026-06-04下载Configuring network access control policies in large, complex networks is error-prone and requires significant expert effort. LLMs offer a promising interface for expressing such policies in natural l...
Towards Serverless Semi-Decentralized Federated Learning with Heterogeneous OptimizersSu Wang, Mung Chiang, H. Vincent Poor2026-06-04下载We investigate cluster formation, involving the number and composition of clusters, in decentralized federated learning (FL) with heterogeneous machine learning (ML) optimizers.
DAST: A VLM-LLM Framework for Cross-Interface Anomaly Detection in O-RANFrancesco Spinelli, Esteban Municio, Pau Baguer, Gines Garcia-Aviles, Xavier Costa-Perez2026-06-04下载O-RAN enables a disaggregated baseband stack with programmable functions that communicate over standardized open interfaces. The same openness that enables multi-vendor composition also expands the at...
Toward Mobile and Converged Backhaul: The Promise of Wireless Access and BackhaulChiara Rubaltelli, Marcello Morini, Eugenio Moro, Ilario Filippini, Antonio Capone2026-06-04下载Wireless Access and Backhaul (WAB) is emerging as a key enabler for flexible and cost-efficient 5G deployments, offering a modular architecture that decouples access and backhaul while supporting mult...
Policy-Guided ML for Energy Savings: Cell On/Off Switching under Operator QoS Constraints in Real 5G NetworksD. Reiss, M. Catalan-Cid, D. Camps-Mur, O. Sallent2026-06-04下载Energy efficiency is a critical concern in the deployment and operation of 5G networks, particularly due to the low utilization of 4G and 5G carriers during off-peak hours.
Quantifying the Energy-Saving and QoS Trade-Off in Traffic Offloading for Real 4G/5G ScenariosD. Reiss, M. Catalan-Cid, D. Camps-Mur, O. Sallent2026-06-04下载Despite the potential for higher energy efficiency in 5G networks, current 5G Non-Standalone (NSA) deployments often operate suboptimally due to low utilization of 4G and 5G carriers during extended p...
BeGREEN Intelligent Plane for AI-driven Energy Efficient O-RAN managementM. Catalan-Cid, J. Pueyo, J. Sanchez-Gonzalez, J. Gutierrez, M. Ghoraishi2026-06-04下载Cellular networks are undergoing a revolutionary transform with the advent of O-RAN architectures and AI/ML solutions. O-RAN's Non-Real-Time and Near-Real Time RAN Intelligent Controllers open the doo...
AISC deployment in dynamic UAV-assisted MEC network: a reinforcement learning method based on heterogeneous graph attention neural networkHanzhi Chang, Jing Bai, Xin Tang, Xiaomei Liu2026-06-04下载Unmanned aerial vehicles-assisted mobile edge computing (UMEC) can execute compute-intensive and latency-critical artificial intelligence (AI) services, which can be provided by multiple UAVs collabor...
GS-NFS: Bandwidth-adaptive Streaming of Dynamic Gaussian Splats and Point CloudsRajrup Ghosh, Haodong Wang, Haoran Hong, Eduardo Pavez, Amartya Chaudhuri, Weiwu Pang, Harsha V. Madhyastha, Antonio Ortega, Ramesh Govindan2026-06-04下载Dynamic 3D Gaussian Splatting (3DGS) holds great promise as a 3D video streaming technology since it can represent complex 3D scenes with high fidelity.
Availability-Aware and Efficiency-Driven AI Service Chain Provisioning in Multi-Domain Edge Intelligence CloudHanzhi Chang, Jing Bai, Xin Tang, Xiaomei Liu, Yiming Chen2026-06-04下载In a multi-domain edge intelligence cloud (MDEIC) managed by multiple network operators, AI services are delivered by chains of virtual network functions (VNFs) executed in sequence, called AI service...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
AgileOS: A GPU Operating System Layer for Protected CUDA ServicesZhuoping Yang, Yiyu Shi, Alex Jones, Peipei Zhou2026-06-04下载Modern GPU applications increasingly interact with storage systems, network devices, vendor libraries, and GPU-resident services rather than executing only isolated compute kernels.
CarbonSim: A Lifecycle-Aware Framework for Evaluating Carbon Tradeoffs in Hardware Upgrade DecisionsKartik Hans, Kaiwen Zhao, Stephen Lee2026-06-04下载As the demand for information and communication technologies (ICT) continues to rise, the environmental impact of computing systems is becoming an increasingly critical concern.

cs.PF - Performance ​

标题作者发布日期PDF摘要
AEGIS: A Backup Reflex for Physical AIJosef Chen2026-06-04下载Long-horizon robot manipulation tends to fail gradually: one bad step degrades the state, and the policy spirals into a basin from which it cannot recover.
rsx: A high-performance streaming toolkit for RAD-seq sex determinationRohit Goswami, Ruhila Goswami2026-06-04下载Restriction site-associated DNA sequencing (RAD-seq) is widely used to discover sex-linked markers in non-model organisms, but large studies produce marker tables with millions of RAD tags.
PivCo-HuffmanMarcin Zukowski2026-06-04下载Huffman encoding has been an enduring technique for 70+ years, ubiquitous in compression algorithms since its invention. In this paper we propose a new approach to Huffman coding, based on a data stru...

基于 VitePress 构建 · 使用本地搜索查找论文