2026-05-06
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable Arrays | Jiayi Wang, Ang Da Lu, Zhichen Zeng, Ang Li | 2026-05-06 | 下载 | While GPUs dominate massively parallel computing through the single-instruction, multiple-thread (SIMT) programming model, their underlying single-instruction, multiple-data (SIMD) execution incurs su... |
| Beyond Static Policies: Exploring Dynamic Policy Selection for Single-Thread Performance Optimization | Yanxin Zhang, Ian McDougall, Junnan Li, Shayne Wadle, Vikas Singh, Karthikeyan Sankaralingam | 2026-05-06 | 下载 | For over a decade, processor design has focused on implementing sophisticated policies for various components of the out-of-order pipeline, including cache replacement and prefetching. |
| An Open-Source Flow for Single-Phase, Edge-Triggered to Two-Phase, Non-Overlapping Clocking Conversion | Paolo Pedroso, Lee-Way Wang, Matthew Guthaus | 2026-05-06 | 下载 | Two-phase clocking offers significant advantages in timing margin and clock flexibility, yet its adoption remains limited due to the absence of automation in modern design flows. |
| Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours | The Verkor Team, Ravi Krishna, Suresh Krishna, David Chin | 2026-05-06 | 下载 | Driven by a rapid co-evolution of both harness and underlying models, LLM agents are improving at a dizzying pace. In our prior work (performed in Dec. |
| MCFlash: Bulk Bitwise Processing in 3D NAND with Dynamic Sensing and Multi-level Encoding | Habib Ur Rahman, Tharini Suresh, Sudeep Pasricha, Biswajit Ray | 2026-05-06 | 下载 | This paper presents MCFlash, a practical and immediately deployable technique for executing bulk bitwise operations directly within commercial off-the-shelf(COTS) 3D NAND flash chips. |
| Not All Faults Are Equal: Transient-Fault Sensitivity Characterization of an Open-Source RISC-V Vector Cluster | Maoyuan Cai, Amirhossein Kiamarzi, Davide Rossi, Angelo Garofalo | 2026-05-06 | 下载 | We present a transient-fault sensitivity study of the open-source RISC-V vector cluster Spatz under SET and SEU fault models. Across 100,000 fault injections on six MatMul and Widening MatMul configur... |
| AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN Architectures | Omkar B Shende, Marcello Traiola, Gayathri Ananthanarayanan | 2026-05-06 | 下载 | Deep neural network (DNN) inference at the edge demands simultaneous improvements in accuracy, computational efficiency, and energy consumption. |
| UVMarvel: an Automated LLM-aided UVM Machine for Subsystem-level RTL Verification | Junhao Ye, Dingrong Pan, Hanyuan Liu, Yuchen Hu, Jie Zhou, Ke Xu, Xinwei Fang, Xi Wang, Nan Guan, Zhe Jiang | 2026-05-06 | 下载 | Verification presents a major bottleneck in Integrated Circuit (IC) development, consuming nearly 70% of total effort. While the Universal Verification Methodology (UVM) improves reuse through structu... |
| Ultra Low-Power SDM-based Circuit-Switching for Networks-on-Chip | Meysam Zaeemi, Mehdi Modarressi | 2026-05-06 | 下载 | In many modern AI chips and multicore systems-on-chip, embedded applications exhibit predictable inter-core traffic behavior that can be characterized at design time. |
| RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNs | Hanum Ko, Sangheum Yeon, Jong Hwan Ko, Jungrae Kim | 2026-05-06 | 下载 | As DRAM scales in density and adopts 3D integration, raw fault rates increase and multi-bit errors are no longer rare. Such errors can severely impact Deep Neural Networks (DNNs): although DNNs tolera... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination | Jae-Won Chung, Zhirui Liang, Yanyong Mao, Jiasi Chen, Mosharaf Chowdhury, Vladimir Dvorkin | 2026-05-06 | 下载 | AI's growing compute demand and new datacenter buildouts present major capacity and reliability challenges for the electricity grid, leading to multi-year interconnection delays for new datacenters an... |
| Nitsum: Serving Tiered LLM Requests with Adaptive Tensor Parallelism | Vikranth Srivatsa, Zijian He, Pu Guo, Dongming Li, Yiying Zhang | 2026-05-06 | 下载 | LLM serving is increasingly multi-tenant: the same deployment must handle latency-critical interactive requests and more relaxed background workloads under a fixed GPU budget. |
| Toward a Risk Assessment Framework for Institutional DeFi: A Nine-Dimension Approach | Eva Oberholzer, Valeriy Zamaraiev | 2026-05-06 | 下载 | Decentralized finance (DeFi) protocols now intermediate over USD 100 billion in value, including regulated stablecoins and tokenized assets deployed as collateral, yet no widely adopted framework oper... |
| Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism | Sajal Dash, Feiyi Wang | 2026-05-06 | 下载 | Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hindered by large memory ... |
| Communication Offloading on SmartNIC DPUs: A Quantitative Approach | Jacob Wahlgren, Andong Hu, Roger Pearce, Maya Gokhale, Ivy Peng | 2026-05-06 | 下载 | SmartNIC Data Processing Units (DPUs) offer a promising solution for saving high-end CPU resources by offloading tasks to programmable cores near the network interface. |
| Delay-Aware Large-Small Model Collaboration over LEO Satellite Networks | Mingyu Guo, Wen Wu, Ying Wang, Songge Zhang, Liang Li | 2026-05-06 | 下载 | In this paper, we introduce a delay-aware largesmall model collaboration scheme for low Earth orbit (LEO) satellite networks, which can balance the computational load among satellites and the communic... |
| CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model Training | Yida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang, Hairui Zhao, Xingchen Liu, Yang Tian, Wenjing Huang, Zedong Liu, Yifan Chen, Jinwu Yang, Yueyuan Zhou, Qian Zhao, Haoxu Li, Tao Wang, Feng Yu, Zhan Wang, Guangming Tan, Dingwen Tao | 2026-05-06 | 下载 | As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors. |
| KEET: Explaining Performance of GPU Kernels Using LLM Agents | Joshua H. Davis, Klaudiusz Rydzy, Srinivasan Ramesh, Aadit Nilay, Daniel Nichols, Swapna Raj, Nikhil Jain, Abhinav Bhatele | 2026-05-06 | 下载 | Performance profiles of GPU kernels generated by tools such as Nsight Compute are rich in detail but are often challenging to interpret. To achieve the best performance possible on a given GPU archite... |
| One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender Serving | Wenjun Yu, Shuguang Han, Amelie Chi Zhou | 2026-05-06 | 下载 | Generative Recommender (GR) inference places embedding hot caches (EMB) and KV caches in direct competition for limited GPU HBM: allocating more memory to one improves its efficiency but degrades the ... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| When Semantic Communication Meets Queueing: Cross-Layer Latency and Task Fidelity Optimization | Yalin E. Sagduyu, Tugba Erpek | 2026-05-06 | 下载 | Semantic communication (SemCom) with learned encoder-decoder architectures enables end-to-end learning of compact task-oriented representations optimized for the wireless channel, reducing channel res... |
| Performance Characterization of dApps in Open Radio Access Networks | Conrado Boeira, Eduardo Baena, Andrea Lacava, Tommaso Melodia, Dimitrios Koutsonikolas, Israat Haque | 2026-05-06 | 下载 | Despite recommendations to deploy real-time Open Radio Access Network (O-RAN) applications (dApps) in containerized environments, existing approaches predominantly rely on bare-metal servers. |
| SILC: Lookahead Caching for Short-form Video Delivery Systems | Maleeha Masood, Shreya Kannan, Om Chabra, Deepak Vasisht, Indranil Gupta | 2026-05-06 | 下载 | Short video platforms like TikTok, Instagram Reels, and YouTube Shorts have gained immense popularity in the last few years and are responsible for a large and growing fraction of Internet traffic. |
| Age of Gossip in Ring Networks With Non-Poisson Updates | Arunabh Srivastava, Sennur Ulukus | 2026-05-06 | 下载 | We consider a network consisting of nodes connected in a ring formation and a source that generates updates according to a renewal process and disseminates them to the ring network according to a ... |
| Look Once, Beam Twice: Camera-Primed Real-Time Double-Directional mmWave Beam Management for Vehicular Connectivity | Avhishek Biswas, Apala Pramanik, Eylem Ekici, Mehmet C. Vuran | 2026-05-06 | 下载 | Millimeter-wave (mmWave) frequencies promise multi-gigabit connectivity for vehicle-to-everything (V2X) networks, but face challenges in terms of severe path loss and mobility-related beam misalignmen... |
| Traffic Chunk Sizing vs. Optical Switching Speed in Future All-Optical Satellite Networks | Sleman Mouammar, Thomas Röthig, Soheil Hosseini, Ítalo Brasileiro, Admela Jukan | 2026-05-06 | 下载 | To enable efficient resource utilization under stringent Size, Weight, and Power (SWaP) constraints through transparent and all-optical switched satellites transmission, various switching paradigms ca... |
| AFL-ICP: Enhancing Industrial Control Protocol Reliability via Specification-Guided Fuzzing | Jiaying Meng, Xuewei Feng, Qi Li, Min Liu, Ke Xu | 2026-05-06 | 下载 | Industrial Control Protocols (ICPs) are critical to the reliability and stability of industrial infrastructure, yet their security is fundamentally compromised by a specification-blindness bottleneck. |
| A Separation Between Optimal Demand-Oblivious and Demand-Aware Network Throughput | Matthias Bentert, Chen Avin, Stefan Schmid | 2026-05-06 | 下载 | The performance of distributed applications often critically depends on the interconnecting network or more specifically on its throughput: how fast data can be carried across a network. |
| Securing the Web with HSTS-Enforced | Aaron van Diepen, Adrian Zapletal, Fernando Kuipers | 2026-05-06 | 下载 | TLS stripping attacks expose sensitive web traffic by forcing secure HTTPS connections to fall back to unencrypted HTTP. At present, protection against these attacks relies on website operators explic... |
| SADE: Symptom-Aware Diagnostic Escalation for LLM-Based Network Troubleshooting | Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kosta Dekic, Arunan Sivanathan, Suranga Seneviratne | 2026-05-06 | 下载 | Large language model (LLM) agents are increasingly applied to network troubleshooting, but root-cause localization on public benchmarks remains well below practical deployment thresholds. |
| Queue-Aware and Resilient Routing in LEO Satellite Networks Using Multi-Agent Reinforcement Learning | Mudassar Liaq, Mahyar Tajeri, Peng Hu | 2026-05-06 | 下载 | With the rapid growth in data demand and stringent latency requirements of modern applications has driven significant interest in Low Earth Orbit (LEO) satellite constellations as an emerging solution... |
| Joint Optimization of Trajectory Control, Resource Allocation, and Task Offloading for Multi-UAV-Assisted IoV | Maoxin Ji, Qiong Wu, Pingyi Fan, Cui Zhang, Nan Cheng, Wen Chen, Khaled B. Letaief | 2026-05-06 | 下载 | This paper investigates a multi-Unmanned Aerial Vehicle (UAV) joint base station-assisted Internet of Vehicles (IoV) task offloading system in dense urban environments. |
| Worst-Case Discovery and Runtime Protection for RL-Based Network Controllers | Hongyu Hè, Minhao Jin, Maria Apostolaki | 2026-05-06 | 下载 | RL-based controllers achieve strong average-case performance in networking tasks such as congestion control and adaptive bitrate streaming. Yet their performance can degrade severely under network con... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Shedding Light onto Safety Integrity Level and Basic Software Constraints in a Real-World Automotive Application: Case Study with Driverator Framework | Tobias Denzinger, Matthias Becker, Peter Ulbrich | 2026-05-06 | 下载 | Automotive electronic control units (ECUs) are intricate systems with hundreds of individual functions, numerous software components, and multiple interdependent tasks. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels | Han Wang, Jintao Zhang, Kai Jiang, Haoxu Wang, Jianfei Chen, Jun Zhu | 2026-05-06 | 下载 | LLM-based Triton kernel generation has attracted significant interest, yet a fundamental empirical question remains unanswered: where does this capability break down, and why? We present KernelBench-X... |
| AGIPC: Adaptive In-Solve Algebraic Coarsening for GPU IPC | Xuan Wang, Zhaofeng Luo, Minchen Li, Taku Komura, Kemeng Huang | 2026-05-06 | 下载 | Implicit time integration is key to robustly simulating stiff materials and large deformations, but its performance is often dominated by repeatedly solving large linear systems. |
| KEET: Explaining Performance of GPU Kernels Using LLM Agents | Joshua H. Davis, Klaudiusz Rydzy, Srinivasan Ramesh, Aadit Nilay, Daniel Nichols, Swapna Raj, Nikhil Jain, Abhinav Bhatele | 2026-05-06 | 下载 | Performance profiles of GPU kernels generated by tools such as Nsight Compute are rich in detail but are often challenging to interpret. To achieve the best performance possible on a given GPU archite... |