Skip to content

2026-05-06 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
DICE: Enabling Efficient General-Purpose SIMT Execution with Statically Scheduled Coarse-Grained Reconfigurable ArraysJiayi Wang, Ang Da Lu, Zhichen Zeng, Ang Li2026-05-06下载While GPUs dominate massively parallel computing through the single-instruction, multiple-thread (SIMT) programming model, their underlying single-instruction, multiple-data (SIMD) execution incurs su...
Beyond Static Policies: Exploring Dynamic Policy Selection for Single-Thread Performance OptimizationYanxin Zhang, Ian McDougall, Junnan Li, Shayne Wadle, Vikas Singh, Karthikeyan Sankaralingam2026-05-06下载For over a decade, processor design has focused on implementing sophisticated policies for various components of the out-of-order pipeline, including cache replacement and prefetching.
An Open-Source Flow for Single-Phase, Edge-Triggered to Two-Phase, Non-Overlapping Clocking ConversionPaolo Pedroso, Lee-Way Wang, Matthew Guthaus2026-05-06下载Two-phase clocking offers significant advantages in timing margin and clock flexibility, yet its adoption remains limited due to the absence of automation in modern design flows.
Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hoursThe Verkor Team, Ravi Krishna, Suresh Krishna, David Chin2026-05-06下载Driven by a rapid co-evolution of both harness and underlying models, LLM agents are improving at a dizzying pace. In our prior work (performed in Dec.
MCFlash: Bulk Bitwise Processing in 3D NAND with Dynamic Sensing and Multi-level EncodingHabib Ur Rahman, Tharini Suresh, Sudeep Pasricha, Biswajit Ray2026-05-06下载This paper presents MCFlash, a practical and immediately deployable technique for executing bulk bitwise operations directly within commercial off-the-shelf(COTS) 3D NAND flash chips.
Not All Faults Are Equal: Transient-Fault Sensitivity Characterization of an Open-Source RISC-V Vector ClusterMaoyuan Cai, Amirhossein Kiamarzi, Davide Rossi, Angelo Garofalo2026-05-06下载We present a transient-fault sensitivity study of the open-source RISC-V vector cluster Spatz under SET and SEU fault models. Across 100,000 fault injections on six MatMul and Widening MatMul configur...
AxMoE: Characterizing the Impact of Approximate Multipliers on Mixture-of-Experts DNN ArchitecturesOmkar B Shende, Marcello Traiola, Gayathri Ananthanarayanan2026-05-06下载Deep neural network (DNN) inference at the edge demands simultaneous improvements in accuracy, computational efficiency, and energy consumption.
UVMarvel: an Automated LLM-aided UVM Machine for Subsystem-level RTL VerificationJunhao Ye, Dingrong Pan, Hanyuan Liu, Yuchen Hu, Jie Zhou, Ke Xu, Xinwei Fang, Xi Wang, Nan Guan, Zhe Jiang2026-05-06下载Verification presents a major bottleneck in Integrated Circuit (IC) development, consuming nearly 70% of total effort. While the Universal Verification Methodology (UVM) improves reuse through structu...
Ultra Low-Power SDM-based Circuit-Switching for Networks-on-ChipMeysam Zaeemi, Mehdi Modarressi2026-05-06下载In many modern AI chips and multicore systems-on-chip, embedded applications exhibit predictable inter-core traffic behavior that can be characterized at design time.
RangeGuard: Efficient, Bounded Approximate Error Correction for Reliable DNNsHanum Ko, Sangheum Yeon, Jong Hwan Ko, Jungrae Kim2026-05-06下载As DRAM scales in density and adopts 3D integration, raw fault rates increase and multi-bit errors are no longer rare. Such errors can severely impact Deep Neural Networks (DNNs): although DNNs tolera...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime CoordinationJae-Won Chung, Zhirui Liang, Yanyong Mao, Jiasi Chen, Mosharaf Chowdhury, Vladimir Dvorkin2026-05-06下载AI's growing compute demand and new datacenter buildouts present major capacity and reliability challenges for the electricity grid, leading to multi-year interconnection delays for new datacenters an...
Nitsum: Serving Tiered LLM Requests with Adaptive Tensor ParallelismVikranth Srivatsa, Zijian He, Pu Guo, Dongming Li, Yiying Zhang2026-05-06下载LLM serving is increasingly multi-tenant: the same deployment must handle latency-critical interactive requests and more relaxed background workloads under a fixed GPU budget.
Toward a Risk Assessment Framework for Institutional DeFi: A Nine-Dimension ApproachEva Oberholzer, Valeriy Zamaraiev2026-05-06下载Decentralized finance (DeFi) protocols now intermediate over USD 100 billion in value, including regulated stablecoins and tokenized assets deployed as collateral, yet no widely adopted framework oper...
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid ParallelismSajal Dash, Feiyi Wang2026-05-06下载Frontier models increasingly adopt Mixture-of-Experts (MoE) architectures to achieve large-model performance at reduced cost. However, training MoE models on HPC platforms is hindered by large memory ...
Communication Offloading on SmartNIC DPUs: A Quantitative ApproachJacob Wahlgren, Andong Hu, Roger Pearce, Maya Gokhale, Ivy Peng2026-05-06下载SmartNIC Data Processing Units (DPUs) offer a promising solution for saving high-end CPU resources by offloading tasks to programmable cores near the network interface.
Delay-Aware Large-Small Model Collaboration over LEO Satellite NetworksMingyu Guo, Wen Wu, Ying Wang, Songge Zhang, Liang Li2026-05-06下载In this paper, we introduce a delay-aware largesmall model collaboration scheme for low Earth orbit (LEO) satellite networks, which can balance the computational load among satellites and the communic...
CCL-D: A High-Precision Diagnostic System for Slow and Hang Anomalies in Large-Scale Model TrainingYida Gu, Fakang Wang, Jianhao Fu, Zhenhang Sun, Qianyu Zhang, Hairui Zhao, Xingchen Liu, Yang Tian, Wenjing Huang, Zedong Liu, Yifan Chen, Jinwu Yang, Yueyuan Zhou, Qian Zhao, Haoxu Li, Tao Wang, Feng Yu, Zhan Wang, Guangming Tan, Dingwen Tao2026-05-06下载As training scales grow, collective communication libraries (CCL) increasingly face anomalies arising from complex interactions among hardware, software, and environmental factors.
KEET: Explaining Performance of GPU Kernels Using LLM AgentsJoshua H. Davis, Klaudiusz Rydzy, Srinivasan Ramesh, Aadit Nilay, Daniel Nichols, Swapna Raj, Nikhil Jain, Abhinav Bhatele2026-05-06下载Performance profiles of GPU kernels generated by tools such as Nsight Compute are rich in detail but are often challenging to interpret. To achieve the best performance possible on a given GPU archite...
One Pool, Two Caches: Adaptive HBM Partitioning for Accelerating Generative Recommender ServingWenjun Yu, Shuguang Han, Amelie Chi Zhou2026-05-06下载Generative Recommender (GR) inference places embedding hot caches (EMB) and KV caches in direct competition for limited GPU HBM: allocating more memory to one improves its efficiency but degrades the ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
When Semantic Communication Meets Queueing: Cross-Layer Latency and Task Fidelity OptimizationYalin E. Sagduyu, Tugba Erpek2026-05-06下载Semantic communication (SemCom) with learned encoder-decoder architectures enables end-to-end learning of compact task-oriented representations optimized for the wireless channel, reducing channel res...
Performance Characterization of dApps in Open Radio Access NetworksConrado Boeira, Eduardo Baena, Andrea Lacava, Tommaso Melodia, Dimitrios Koutsonikolas, Israat Haque2026-05-06下载Despite recommendations to deploy real-time Open Radio Access Network (O-RAN) applications (dApps) in containerized environments, existing approaches predominantly rely on bare-metal servers.
SILC: Lookahead Caching for Short-form Video Delivery SystemsMaleeha Masood, Shreya Kannan, Om Chabra, Deepak Vasisht, Indranil Gupta2026-05-06下载Short video platforms like TikTok, Instagram Reels, and YouTube Shorts have gained immense popularity in the last few years and are responsible for a large and growing fraction of Internet traffic.
Age of Gossip in Ring Networks With Non-Poisson UpdatesArunabh Srivastava, Sennur Ulukus2026-05-06下载We consider a network consisting of nn nodes connected in a ring formation and a source that generates updates according to a renewal process and disseminates them to the ring network according to a ...
Look Once, Beam Twice: Camera-Primed Real-Time Double-Directional mmWave Beam Management for Vehicular ConnectivityAvhishek Biswas, Apala Pramanik, Eylem Ekici, Mehmet C. Vuran2026-05-06下载Millimeter-wave (mmWave) frequencies promise multi-gigabit connectivity for vehicle-to-everything (V2X) networks, but face challenges in terms of severe path loss and mobility-related beam misalignmen...
Traffic Chunk Sizing vs. Optical Switching Speed in Future All-Optical Satellite NetworksSleman Mouammar, Thomas Röthig, Soheil Hosseini, Ítalo Brasileiro, Admela Jukan2026-05-06下载To enable efficient resource utilization under stringent Size, Weight, and Power (SWaP) constraints through transparent and all-optical switched satellites transmission, various switching paradigms ca...
AFL-ICP: Enhancing Industrial Control Protocol Reliability via Specification-Guided FuzzingJiaying Meng, Xuewei Feng, Qi Li, Min Liu, Ke Xu2026-05-06下载Industrial Control Protocols (ICPs) are critical to the reliability and stability of industrial infrastructure, yet their security is fundamentally compromised by a specification-blindness bottleneck.
A Separation Between Optimal Demand-Oblivious and Demand-Aware Network ThroughputMatthias Bentert, Chen Avin, Stefan Schmid2026-05-06下载The performance of distributed applications often critically depends on the interconnecting network or more specifically on its throughput: how fast data can be carried across a network.
Securing the Web with HSTS-EnforcedAaron van Diepen, Adrian Zapletal, Fernando Kuipers2026-05-06下载TLS stripping attacks expose sensitive web traffic by forcing secure HTTPS connections to fall back to unencrypted HTTP. At present, protection against these attacks relies on website operators explic...
SADE: Symptom-Aware Diagnostic Escalation for LLM-Based Network TroubleshootingKuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kosta Dekic, Arunan Sivanathan, Suranga Seneviratne2026-05-06下载Large language model (LLM) agents are increasingly applied to network troubleshooting, but root-cause localization on public benchmarks remains well below practical deployment thresholds.
Queue-Aware and Resilient Routing in LEO Satellite Networks Using Multi-Agent Reinforcement LearningMudassar Liaq, Mahyar Tajeri, Peng Hu2026-05-06下载With the rapid growth in data demand and stringent latency requirements of modern applications has driven significant interest in Low Earth Orbit (LEO) satellite constellations as an emerging solution...
Joint Optimization of Trajectory Control, Resource Allocation, and Task Offloading for Multi-UAV-Assisted IoVMaoxin Ji, Qiong Wu, Pingyi Fan, Cui Zhang, Nan Cheng, Wen Chen, Khaled B. Letaief2026-05-06下载This paper investigates a multi-Unmanned Aerial Vehicle (UAV) joint base station-assisted Internet of Vehicles (IoV) task offloading system in dense urban environments.
Worst-Case Discovery and Runtime Protection for RL-Based Network ControllersHongyu Hè, Minhao Jin, Maria Apostolaki2026-05-06下载RL-based controllers achieve strong average-case performance in networking tasks such as congestion control and adaptive bitrate streaming. Yet their performance can degrade severely under network con...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Shedding Light onto Safety Integrity Level and Basic Software Constraints in a Real-World Automotive Application: Case Study with Driverator FrameworkTobias Denzinger, Matthias Becker, Peter Ulbrich2026-05-06下载Automotive electronic control units (ECUs) are intricate systems with hundreds of individual functions, numerous software components, and multiple interdependent tasks.

cs.PF - Performance ​

标题作者发布日期PDF摘要
KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU KernelsHan Wang, Jintao Zhang, Kai Jiang, Haoxu Wang, Jianfei Chen, Jun Zhu2026-05-06下载LLM-based Triton kernel generation has attracted significant interest, yet a fundamental empirical question remains unanswered: where does this capability break down, and why? We present KernelBench-X...
AGIPC: Adaptive In-Solve Algebraic Coarsening for GPU IPCXuan Wang, Zhaofeng Luo, Minchen Li, Taku Komura, Kemeng Huang2026-05-06下载Implicit time integration is key to robustly simulating stiff materials and large deformations, but its performance is often dominated by repeatedly solving large linear systems.
KEET: Explaining Performance of GPU Kernels Using LLM AgentsJoshua H. Davis, Klaudiusz Rydzy, Srinivasan Ramesh, Aadit Nilay, Daniel Nichols, Swapna Raj, Nikhil Jain, Abhinav Bhatele2026-05-06下载Performance profiles of GPU kernels generated by tools such as Nsight Compute are rich in detail but are often challenging to interpret. To achieve the best performance possible on a given GPU archite...

基于 VitePress 构建 · 使用本地搜索查找论文