2026-07-25
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| SPARC: Automated Root-Cause Analysis of Pre-Silicon Power Side-Channel Leakage in the Processor Design Flow | Andrija Nešković, Christian Ewert, Mladen Berekovic, Saleh Mulhem | 2026-07-25 | 下载 | Power-Side-Channel Leakage (PSCL) originates from architectural and micro-architectural artifacts in a processor and poses a severe threat to the confidentiality of cryptographic software. |
| Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel Dispatch | En-Ming Huang, An-Cheng Chang, Bai-Cheng Jeng, Shih-Hao Hung, H. T. Kung | 2026-07-25 | 下载 | Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kernels using static tok... |
| Magnetic Tunnel Junctions for Timekeeping in Intermittent Computing Systems | Nikola Vuk Maruszewski, Jordan Athas, Allison Fleming, Christian Duffee, Eren Yildiz, Saad Ahmed, Yaman Sangar, Pedram Khalili, Josiah Hester | 2026-07-25 | 下载 | Batteryless intermittent systems run unattended for years, but power failures erase timekeeping state, corrupting sensing, scheduling, and coordination. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference | Jianwen Xian, Zhiyuan Xu, Yuchen Li, Ziliang Lai, Kang He, Zhen Huang, Aichen Feng, Jinyan Chen, Yilin Zhang, Qinqin Chen, Chengru Song | 2026-07-25 | 下载 | Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Tensor Core computation... |
| Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool | Yan Wang, Xiulong Yuan, Kaiming Yang, Jiaxuan Peng, Pengju Lu, Mingzhen Li, Zhipeng Zhang, Chang Si, Zhixiang Ruan, Hongqing Chen, Linlang Jiang, Siyu Wang, Langshi Chen, Rui Men, Man Yuan, Guangming Tan, Yong Li, Weile Jia, Jingren Zhou | 2026-07-25 | 下载 | Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost operators, but the domi... |
| A Fixed-Point Construction of the Elementary Transcendental Functions | François Alouges, Giovanni Di Fratta, Alberto Fiorenza, Renato Fiorenza | 2026-07-25 | 下载 | We present a unified fixed-point construction of the elementary transcendental functions, encompassing the real exponential, the complex exponential (sine and cosine), and the natural logarithm. |
| A scalable online machine learning approach for Stock Recommendation | Harsh Nagarkar | 2026-07-25 | 下载 | Stock recommendation systems face the dual challenge of adapting to rapidly changing market conditions while maintaining low-latency predictions for end users. |
| Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANs | Zhihao Xu, Hao Zhong, Zeting Zhou, Yuhang Xu, Haoyu Tong, Wei Wang, Jinshan Chen, Keqiang He, Chong Zhu, Shengzhong Liu, Fan Wu, Guihai Chen | 2026-07-25 | 下载 | This paper aims to enable computation- and communication-efficient GPU sharing across devices within local area networks (LANs), facilitating ubiquitous AI inference on heterogeneous personal devices. |
| Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems | Yuchen Fan, Minghong Sun, Jikui Ma, Yunpeng Xu, Shunyu Mao, Liu He, Shunan Dong, Jiahao Yang, Yu Zhu, Xinhao Yang, Tianyan Zhong, Haoran Sun, Daoqi Liu, Zongle Huang, Xinyuan Lin, Huazhong Yang, Maokun Li, Yongpan Liu, Yu Wang, Zhenhua Zhu, Hongyang Jia, Shuwen Deng | 2026-07-25 | 下载 | AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collectives. |
| A Resource Estimation Model for the Hardware-Software Co-Design of Distributed Quantum Architectures | Raymond P. H. Wu, Chathurika Ranaweera, Sutharshan Rajasegarar, Ria Rushin Joseph, Jinho Choi, Seng W. Loke | 2026-07-25 | 下载 | In distributed quantum computing (DQC), executing monolithic quantum circuits across multiple interconnected quantum processing units (QPUs) requires dedicated communication qubits to generate and dis... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| INT8 Quantization Makes ARM Edge Inference Dispatch-Invariant | Sebastián A. Cruz Romero, Shenied E. Maldonado Guerra | 2026-07-25 | 下载 | On x86, kernel dispatch fragments the outputs of the same neural network into many equivalence classes across hardware. We ask whether the same fragmentation governs ARM edge inference, where most edg... |
| Optimized Embedded Implementation of Hyperspectral-Multispectral Image Fusion on Raspberry Pi | Salah Eddine Brezini, Okba Bekhelifi, Oussama Mezouar, Chams Eddine Choucha, Sarra Boukhacheba, Fethi Abdelatif Dali | 2026-07-25 | 下载 | Remote sensing optical images have become central to a wide range of applications. In particular, hyperspectral images, with their high spectral resolution, enable the extraction of rich information a... |