Skip to content

2026-07-21 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference MemoryMuhammad Husnain Mubarik, Karthik Mohan Kumar, Pedro Antonio Pena, Keshavan Varadarajan, Kunal Tyagi2026-07-21下载We characterize per-bit-position fault sensitivity in ML inference across 16 workloads -- spanning transformer-based models and attention-free CNNs -- and across three floating-point formats.
Distributed Delay-Based BIST for Mixed-Signal Circuits in Flexible ElectronicsPaula Carolina Lozano Duarte, Sule Ozev, Mehdi Tahoori2026-07-21下载Flexible electronics (FE) based on indium gallium zinc oxide thin-film transistors (IGZO-TFTs) are emerging for ultra-low-power wearable applications.
A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN InferenceAmirhossein Zarei, Shervin Vakili2026-07-21下载Efficient acceleration of convolutional neural networks (CNNs) on resource-constrained platforms remains challenging due to the irregularity of sparsity patterns and the associated hardware overhead.
High-Level Synthesis of Efficient Pipelines with Visibility ControlJungin Rhee, Minseong Jang, Jaewoo Kim, Jeehoon Kang2026-07-21下载High-level synthesis (HLS) raises the abstraction of hardware design from concurrent register-transfer level (RTL) programs to sequential programs.
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural AcceleratorsFabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann2026-07-21下载Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API.
An Efficient Fault-Tolerance Scheme for CKKS Computation on CPUsJianan Mu, Ge Yu, Tenghui Hua, Liang Kong, Jing Ye, Xing Hu, Meng Li, Xiaowei Li, Huawei Li2026-07-21下载Fully homomorphic encryption (FHE) enables computation on encrypted data, but its long ciphertext dataflow and high-dimensional modular arithmetic make it vulnerable to silent data corruption caused b...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Examining QRMI as a Unified Interface for Quantum-HPC IntegrationThomas Badts, Tim Boyle, Claudio Carvalho, Antonio Córcoles, Andrew Damin, Vadim Elisseev, Jonathan Frassineti, Daniel Gruber, Hiroshi Horii, Eun-Kyung Lee, James Machin, Sara Marzella, Mateusz Meller, Daniel Milroy, Matthieu Moreau, Munetaka Ohtani, Elisabeth Ortega-Carrasco, Doug Oucharek, Yoonho Park, Adarsh Patil, Emre M. Sahin, Gábor Samu, Seetharami Seelam, Amir Shehata, Vanessa Sochat, James Thorne, Oscar Wallis, Aleksander Wennersteen2026-07-21下载The efficient and scalable integration of quantum resources into high-performance computing (HPC) environments requires standardized mechanisms for resource management, scheduling, and workflow orches...
Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-ExpertsMinyu Cui, Anna Wingkvist, Morgan Ericsson2026-07-21下载Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trill...
SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training frameworkAkarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown2026-07-21下载Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance...
Keeping the Cache Warm Pays: Keepalive Economics for Agentic WorkloadsMaxim Khailo2026-07-21下载Frontier LLM providers cache a prompt's processed prefix so that a follow-up request sharing it pays ~10% of the input price and skips most of the prefill latency.
ARBITER: Guarded Agentic Control for SLO-Oriented Kubernetes RemediationPooyan Habibi, Alberto Leon-Garcia2026-07-21下载Maintaining service-level objectives (SLOs) on Kubernetes microservices remains difficult because autoscalers observe coarse resource metrics, recent SLO controllers often depend on custom telemetry, ...
Coherence in Control: Bridging Many-Core Mapping and Routing through Cost UnificationGuochu Xiong, Xiangzhong Luo, Weichen Liu2026-07-21下载The rapid growth of data-intensive applications increases communication demands in many-core systems, where cache coherence, while essential for correct communication and data consistency, introduces ...
A Scalable Pattern Mining Workflow for Interpretable Machine Log Analysis in High-Performance Computing EnvironmentsShilpika Shilpika, Bethany Lusch, Eric Pershey, Carlo Graziani, Venkatram Vishwanath, Michael E. Papka2026-07-21下载Modern supercomputers housed in High Performance Computing (HPC) environments generate massive volumes of log data daily, revealing intricate information and performance metrics about these complex sy...
Enabling Multi-Dimensional Distributed Trace Comparison with ContrastVaastav Anand, Rodrigo Fonseca, Jonathan Mace, Antoine Kaufmann2026-07-21下载Diagnosis using distributed traces is fundamentally a comparative task: operators seek to understand how an anomalous execution differs from expected behavior, how a deployment changes system executio...
InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite AutomataYitao Yuan, Yongchao He, Shaoke Fang, Wenfei Wu2026-07-21下载Cold starts in large language model (LLM) inference services significantly affect user experience, yet they remain inefficient due to sequential initialization and a massive number of fine-grained I/O...
A User-oriented Portable, Reproducible, and Scalable Software EcosystemAlfio Lazzaro, Utz-Uwe Haus, Sandrine Charousset, Nina Mujkanovic2026-07-21下载It is normal for scientists to perform their research on a diverse set of hardware, ranging from laptops and workstations to supercomputers and cloud resources.
Mapping Without Graphs: Learning Coherence Traffic for Task PlacementGuochu Xiong, Tianrui Ma, Weichen Liu2026-07-21下载Cache coherence is essential for communication in many-core Network-on-Chip (NoC)-based systems. As application scale and complexity increase, efficiently managing communication becomes increasingly c...
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural AcceleratorsFabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann2026-07-21下载Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API.
A Second-Moment Theory for Floating-Point Reduction TreesPiyush Sao, Narasinga Miniskar, Pedro Valero-Lara, Keita Teranishi, Sudip Seal2026-07-21下载Summation error depends on partial-sum order, which standard worst-case bounds omit. To capture this dependence, we derive an exact mean-square error (MSE) recurrence for a binary reduction tree T und...
Unstructured Hydrodynamics on Spatial Dataflow Architectures: A Joint Code and Data Decomposition ApproachPiotr Luczynski, Tal Ben-Nun, Leighton Wilson, Brian Van Essen2026-07-21下载Spatial Dataflow Architectures are an emerging hardware pattern in high-performance computing, whose mesh-connected fixed-memory processing elements are tailored for structured grid kernels with two-d...
Searching for Plans You Can Actually Build: A Realizability-Aware Full-Space Optimizer for MoE Training and ServingQuan Yuan, Jie Zhao2026-07-21下载Mixture-of-Experts (MoE) systems split a program's plan space in two: the space a cost model can rank, and the smaller space a real toolchain can actually build.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Scalable Multi-Controller Coordination in Periplus via Border-Switch Forwarding GraphsE. M. Castro Barbero, P. de las Heras Quirós, F. J. Simó Reigadas2026-07-21下载In-band SDN control planes, where control traffic shares the data-plane infrastructure, suit wide-area, resource-constrained deployments -- such as rural backbones -- that cannot afford a dedicated co...
Squeezing the Most Out of Preemption for AoI Minimization: Single-source CaseNail Akar, Mohammad Moltafet, Sennur Ulukus, Marian Codreanu, Roy D. Yates2026-07-21下载In this work, we study a single-source single-server continuous-time status update system where the updates arrive according to a Poisson process and update service times are generally distributed.
HACO: Hedged Agent Computing for Reliable LLM SystemsEnhan Li, Hongyang Du2026-07-21下载As large language model (LLM) agents move from isolated prompting to longhorizon workflows, failures increasingly arise at the role-to-instance binding boundary, where task-specific role requests must...
Structured Spectral Compression based Low-Bitrate Secure Speech Communications for Internet of Things assisted Non-Terrestrial NetworksLi Ping Qian, Zhehan Chen, Qianru Wang, Qian Wang, Yuan Wu, Xuemin Sherman Shen2026-07-21下载This paper focuses on the Low-Bitrate Secure Speech Communications based on the Structured Spectral Compression (LB-S2C2). Specifically, the Mel spectral matrix of the speech signal is first encoded a...
Online Stochastic Matchings: Stability on HypergraphsFabien Mathieu2026-07-21下载We study stochastic dynamic matching on hypergraphs: items of finitely many classes arrive over time and are removed in multisets by activating hyperedges.
NSMA: Neuro-Symbolic Manifold Alignment for Generalizable Adaptive Bitrate Streaming under Texture ShiftZhiqiang He, Zhi Liu2026-07-21下载For decades, ABR has kept two kinds of intelligence apart. Neural policies learn rich behaviors yet forget them the moment the environment changes; rules never learn, and never forget.
Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM ApproachZijiang Yan, Hao Zhou, Wael Jaafar, Jianhua Pei, Ping Wang, Halim Yanikomeroglu, Hina Tabassum2026-07-21下载The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordination of physical flight kinematics and multi-tier network handovers.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Squeezing the Most Out of Preemption for AoI Minimization: Single-source CaseNail Akar, Mohammad Moltafet, Sennur Ulukus, Marian Codreanu, Roy D. Yates2026-07-21下载In this work, we study a single-source single-server continuous-time status update system where the updates arrive according to a Poisson process and update service times are generally distributed.
Job-level Carbon and Water Footprint Estimation for HPC: Bias Assessment from Runtime to Full Life CycleXi Chen, Chris Broekema, Rob van Nieuwpoort2026-07-21下载High performance computing evaluation has traditionally focused on performance and energy, but these metrics alone cannot capture the sustainability cost of runtime configurations.
Formulation-Level Auto-Tuning for QUBO-Based Machine Learning: A Case Study Across Multiple Quantum-Inspired AnnealersNaoya Mizuki, Takahiro Katagiri, Daichi Mukunoki, Tetsuya Hoshino2026-07-21下载This paper presents an Optuna-based formulation-level auto-tuning framework for support vector machines (SVMs) implemented on multiple quantum-inspired annealers.
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural AcceleratorsFabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann2026-07-21下载Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API.

基于 VitePress 构建 · 使用本地搜索查找论文