Skip to content

2026-07-25 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
SPARC: Automated Root-Cause Analysis of Pre-Silicon Power Side-Channel Leakage in the Processor Design FlowAndrija Nešković, Christian Ewert, Mladen Berekovic, Saleh Mulhem2026-07-25下载Power-Side-Channel Leakage (PSCL) originates from architectural and micro-architectural artifacts in a processor and poses a severe threat to the confidentiality of cryptographic software.
Decoding the Skew: Distribution-Aware MoE Inference with Adaptive Kernel DispatchEn-Ming Huang, An-Cheng Chang, Bai-Cheng Jeng, Shih-Hao Hung, H. T. Kung2026-07-25下载Mixture-of-Experts (MoE) inference consists of sparse expert GEMMs whose shapes vary with the runtime routing distribution. Existing serving systems typically select fused-MoE kernels using static tok...
Magnetic Tunnel Junctions for Timekeeping in Intermittent Computing SystemsNikola Vuk Maruszewski, Jordan Athas, Allison Fleming, Christian Duffee, Eren Yildiz, Saad Ahmed, Yaman Sangar, Pedram Khalili, Josiah Hester2026-07-25下载Batteryless intermittent systems run unattended for years, but power failures erase timekeeping state, corrupting sensing, scheduling, and coordination.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT InferenceJianwen Xian, Zhiyuan Xu, Yuchen Li, Ziliang Lai, Kang He, Zhen Huang, Aichen Feng, Jinyan Chen, Yilin Zhang, Qinqin Chen, Chengru Song2026-07-25下载Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Tensor Core computation...
Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence PoolYan Wang, Xiulong Yuan, Kaiming Yang, Jiaxuan Peng, Pengju Lu, Mingzhen Li, Zhipeng Zhang, Chang Si, Zhixiang Ruan, Hongqing Chen, Linlang Jiang, Siyu Wang, Langshi Chen, Rui Men, Man Yuan, Guangming Tan, Yong Li, Weile Jia, Jingren Zhou2026-07-25下载Long-context LLM training suffers from a load-balancing problem that sequence packing does not solve. Packing samples into fixed-token sequences balances memory and linear-cost operators, but the domi...
A Fixed-Point Construction of the Elementary Transcendental FunctionsFrançois Alouges, Giovanni Di Fratta, Alberto Fiorenza, Renato Fiorenza2026-07-25下载We present a unified fixed-point construction of the elementary transcendental functions, encompassing the real exponential, the complex exponential (sine and cosine), and the natural logarithm.
A scalable online machine learning approach for Stock RecommendationHarsh Nagarkar2026-07-25下载Stock recommendation systems face the dual challenge of adapting to rapidly changing market conditions while maintaining low-latency predictions for end users.
Gleam: Adaptive Network-Efficient CUDA API Remoting for Cross-Device GPU Sharing over LANsZhihao Xu, Hao Zhong, Zeting Zhou, Yuhang Xu, Haoyu Tong, Wei Wang, Jinshan Chen, Keqiang He, Chong Zhu, Shengzhong Liu, Fan Wu, Guihai Chen2026-07-25下载This paper aims to enable computation- and communication-efficient GPU sharing across devices within local area networks (LANs), facilitating ubiquitous AI inference on heterogeneous personal devices.
Application-Driven Architecture Exploration for Cross-Layer Heterogeneous SystemsYuchen Fan, Minghong Sun, Jikui Ma, Yunpeng Xu, Shunyu Mao, Liu He, Shunan Dong, Jiahao Yang, Yu Zhu, Xinhao Yang, Tianyan Zhong, Haoran Sun, Daoqi Liu, Zongle Huang, Xinyuan Lin, Huazhong Yang, Maokun Li, Yongpan Liu, Yu Wang, Zhenhua Zhu, Hongyang Jia, Shuwen Deng2026-07-25下载AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collectives.
A Resource Estimation Model for the Hardware-Software Co-Design of Distributed Quantum ArchitecturesRaymond P. H. Wu, Chathurika Ranaweera, Sutharshan Rajasegarar, Ria Rushin Joseph, Jinho Choi, Seng W. Loke2026-07-25下载In distributed quantum computing (DQC), executing monolithic quantum circuits across multiple interconnected quantum processing units (QPUs) requires dedicated communication qubits to generate and dis...

cs.PF - Performance ​

标题作者发布日期PDF摘要
INT8 Quantization Makes ARM Edge Inference Dispatch-InvariantSebastián A. Cruz Romero, Shenied E. Maldonado Guerra2026-07-25下载On x86, kernel dispatch fragments the outputs of the same neural network into many equivalence classes across hardware. We ask whether the same fragmentation governs ARM edge inference, where most edg...
Optimized Embedded Implementation of Hyperspectral-Multispectral Image Fusion on Raspberry PiSalah Eddine Brezini, Okba Bekhelifi, Oussama Mezouar, Chams Eddine Choucha, Sarra Boukhacheba, Fethi Abdelatif Dali2026-07-25下载Remote sensing optical images have become central to a wide range of applications. In particular, hyperspectral images, with their high spectral resolution, enable the extraction of rich information a...

基于 VitePress 构建 · 使用本地搜索查找论文