Skip to content

2026-04-24 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologiesJorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz2026-04-24下载To mitigate the increasingly common underutilization of computational resources in modern GPUs, spatial sharing methods enable multiple applications to use them simultaneously.
Microarchitectural Co-Optimization for Sustained Throughput of RISC-V Multi-Lane Chaining Vector ProcessorsWeiying Wang, Zhiwei Zhang2026-04-24下载Modern RISC vector processors rely on the synergy of multi-lane parallelism and chaining to achieve high sustained throughput, yet their achieved performance often falls substantially short of the the...
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal CorrelationLong Cheng, Ritchie Zhao, Timmy Liu, Mindy Li, Xianjie Qiao, Kefeng Duan, Yu-Jung Chen, Xiaoming Chen, Bita Darvish Rouhani, June Yang2026-04-24下载Sparse-attention decoders rely on exact Top-K selection to choose the most important key-value entries for each query token. In long-context LLM serving, this Top-K stage runs once per decode query an...
Exploiting pre-optimized kernels with polyhedral transformations for CGRA compilationYuxuan Wang, María José Belda, Fernando Castro, Katzalin Olcoz, David Atienza, Giovanni Ansaloni2026-04-24下载Modern computing workloads commonly involve matrix-matrix multiplication (mmul) as a core computing pattern. Coarse-Grained Reconfigurable Arrays (CGRAs) can flexibly and efficiently support it, since...
HGQ-LUT: Fast LUT-Aware Training and Efficient Architectures for DNN InferenceChang Sun, Zhiqiang Que, Bakhtiar Zadeh, Qibin Liu, Kevin H. Alvarez, Wayne Luk, Maria Spiropulu2026-04-24下载Lookup-table (LUT) based neural networks can deliver ultra-low latency and excellent hardware efficiency on FPGAs by mapping arithmetic operations directly onto the logic primitives.
AutoINV: Automated Invariant Generation Framework for Formal Verification on High-Level Synthesis DesignsXiaofeng Zhou, Linfeng Du, Guangyu Hu, Sharad Sinha, Hongce Zhang, Wei Zhang2026-04-24下载High-level synthesis (HLS) transforms an algorithmic description of hardware from a higher abstraction (e.g., C/C++) into a register-transfer level (RTL) design, offering reduced development time and ...
GR-Evolve: Design-Adaptive Global Routing via LLM-Driven Algorithm EvolutionTaizun Jafri, Vidya A. Chhabria2026-04-24下载Modern ASIC design is becoming increasingly complex, driving up design costs while limiting productivity gains from existing EDA tools. Despite decades of progress, current tools rely on fixed heurist...
Hardware-Software Co-Design for Event-Driven SNN Deployment on Low-Cost Neuromorphic FPGAsJiwoon Lee, Souvik Chakraborty, Syed Bahauddin Alam, Cheolsoo Park2026-04-24下载Low-cost FPGA platforms can broaden access to neuromorphic systems research, but current spiking neural network (SNN) workflows remain divided between hardware-first implementations, which are difficu...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Network Edge Inference for Large Language Models: Principles, Techniques, and OpportunitiesZhixiong Chen, Bingjie Zhu, Jiangzhou Wang, Hyundong Shin, Arumugam Nallanathan, Dusit Niyato2026-04-24下载Large language models (LLMs) have advanced rapidly, emerging as versatile tools across fields thanks to their exceptional language understanding, generation, and reasoning capabilities.
Data-Free Contribution Estimation in Federated Learning using Gradient von Neumann EntropyAsim Ukaye, Mubarak Abdu-Aguye, Nurbek Tastan, Karthik Nandakumar2026-04-24下载Client contribution estimation in Federated Learning is necessary for identifying clients' importance and for providing fair rewards. Current methods often rely on server-side validation data or self-...
LaissezCloud: Continuous Resource Renegotiation for the Public CloudTejas Harith, Antoine Kaufmann2026-04-24下载Public clouds increasingly expose heterogeneous hardware, but their allocation interface remains built around rigid on-demand and spot service classes.
A comprehensive evaluation of spatial co-execution on GPUs using MPS and MIG technologiesJorge Villarrubia, Luis Costero, Francisco D. Igual, Katzalin Olcoz2026-04-24下载To mitigate the increasingly common underutilization of computational resources in modern GPUs, spatial sharing methods enable multiple applications to use them simultaneously.
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal CorrelationLong Cheng, Ritchie Zhao, Timmy Liu, Mindy Li, Xianjie Qiao, Kefeng Duan, Yu-Jung Chen, Xiaoming Chen, Bita Darvish Rouhani, June Yang2026-04-24下载Sparse-attention decoders rely on exact Top-K selection to choose the most important key-value entries for each query token. In long-context LLM serving, this Top-K stage runs once per decode query an...
Accelerating Intra-Node GPU-to-GPU Communication Through Multi-Path Transfers with CUDA GraphsAmirhossein Sojoodi, Yiltan Hassan Temucin, Amirreza Baratisedeh, Hamed Sharifian, Ahmad Afsahi2026-04-24下载Effective intra-node GPU communication is essential for optimizing performance in MPI-based HPC applications, especially when leveraging multiple communication paths.
O(K)O(K)-Approximation Coflow Scheduling in KK-Core Optical Circuit Switching NetworksXin Wang, Hong Shen, Hui Tian, Ye Tao2026-04-24下载Coflow has emerged as a fundamental application-layer abstraction in distributed systems, representing communication dependencies and enabling collaborative management of related flows to enhance job ...
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC SystemsBaodi Shan, Mauricio Araya-Polo, Barbara Chapman2026-04-24下载Distributed GPU applications increasingly rely on kernel-level, cross-node coordination to reduce launch overheads and improve compute-communication overlap, but such support is lacking.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Evaluation of the effects of 3GPP-specific beamforming and channel estimation on the 3D EIRP profile of a 5G gNBArmed Tusha, Joshua Roy Palathinkal, Monisha Ghosh2026-04-24下载Spatial domain exploitation through 3D beamforming serves as a critical technology enabler for performance enhancement in the Fifth Generation New Radio (5G NR) specification.
CosmicDancePro -- Measuring LEO satellite's orbital decay and network connectivity implications during solar stormsSuvam Basak, Amitangshu Pal, Debopam Bhattacherjee2026-04-24下载The May 2024 solar superstorm highlighted the vulnerability of rapidly expanding low Earth orbit (LEO) satellite networks to severe space weather events.
Chamelio: A Fast Shared Cloud Network Stack for Isolated Tenant-Defined ProtocolsMatheus Stolet, Simon Peter, Antoine Kaufmann2026-04-24下载Conventional cloud network virtualization sends packets through multiple guest and host layers, inflating CPU cost and tail latency. Shared host datapaths collapse this layering into one optimized pat...
Benchmarking LLM-Driven Network Configuration RepairIoannis Protogeros, Rufat Asadli, Benjamin Hoffman, Laurent Vanbever2026-04-24下载There is a rapidly growing interest in using Large Language Models (LLMs) to automate complex network operations, but their reliable adoption requires rigorous assessment of their effectiveness and sa...
OCC: Physical-Layer Assisted Congestion Control for Real-Time CommunicationsYufan Zhuang, Zili Meng, Zehong Lin, Jun Zhang2026-04-24下载Real-time communications (RTC) is a core technology for emerging applications in 6G, such as cloud gaming, teleoperation, and extended reality (XR), which require consistently low latency and high bit...
Resource-Aware Layered Intrusion Detection Allocation ModelIoan Pădurean, Béla Genge, Roland Bolboacă2026-04-24下载This paper proposes a resource-aware allocation model for layered intrusion detection in het erogeneous networks. Monitoring traffic at higher protocol layers improves the ability to detect sophistica...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Chamelio: A Fast Shared Cloud Network Stack for Isolated Tenant-Defined ProtocolsMatheus Stolet, Simon Peter, Antoine Kaufmann2026-04-24下载Conventional cloud network virtualization sends packets through multiple guest and host layers, inflating CPU cost and tail latency. Shared host datapaths collapse this layering into one optimized pat...

cs.PF - Performance ​

标题作者发布日期PDF摘要
COMPASS: A Unified Decision-Intelligence System for Navigating Performance Trade-off in HPCAnkur Lahiry, Banooqa Banday, Yugesh Bhattarai, Mohammad Zaeed, Tanzima Z. Islam2026-04-24下载HPC systems expose many configuration parameters that jointly drive competing objectives. Existing tools such as autotuners recommend good configurations but do not identify minimal changes for a near...
CosmicDancePro -- Measuring LEO satellite's orbital decay and network connectivity implications during solar stormsSuvam Basak, Amitangshu Pal, Debopam Bhattacherjee2026-04-24下载The May 2024 solar superstorm highlighted the vulnerability of rapidly expanding low Earth orbit (LEO) satellite networks to severe space weather events.
Guess-Verify-Refine: Data-Aware Top-K for Sparse-Attention Decoding on Blackwell via Temporal CorrelationLong Cheng, Ritchie Zhao, Timmy Liu, Mindy Li, Xianjie Qiao, Kefeng Duan, Yu-Jung Chen, Xiaoming Chen, Bita Darvish Rouhani, June Yang2026-04-24下载Sparse-attention decoders rely on exact Top-K selection to choose the most important key-value entries for each query token. In long-context LLM serving, this Top-K stage runs once per decode query an...

基于 VitePress 构建 · 使用本地搜索查找论文