2026-07-21
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory | Muhammad Husnain Mubarik, Karthik Mohan Kumar, Pedro Antonio Pena, Keshavan Varadarajan, Kunal Tyagi | 2026-07-21 | 下载 | We characterize per-bit-position fault sensitivity in ML inference across 16 workloads -- spanning transformer-based models and attention-free CNNs -- and across three floating-point formats. |
| Distributed Delay-Based BIST for Mixed-Signal Circuits in Flexible Electronics | Paula Carolina Lozano Duarte, Sule Ozev, Mehdi Tahoori | 2026-07-21 | 下载 | Flexible electronics (FE) based on indium gallium zinc oxide thin-film transistors (IGZO-TFTs) are emerging for ultra-low-power wearable applications. |
| A Flexible Sparsity-Aware FPGA Accelerator with Column-Wise Compression for Efficient CNN Inference | Amirhossein Zarei, Shervin Vakili | 2026-07-21 | 下载 | Efficient acceleration of convolutional neural networks (CNNs) on resource-constrained platforms remains challenging due to the irregularity of sparsity patterns and the associated hardware overhead. |
| High-Level Synthesis of Efficient Pipelines with Visibility Control | Jungin Rhee, Minseong Jang, Jaewoo Kim, Jeehoon Kang | 2026-07-21 | 下载 | High-level synthesis (HLS) raises the abstraction of hardware design from concurrent register-transfer level (RTL) programs to sequential programs. |
| BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators | Fabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann | 2026-07-21 | 下载 | Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API. |
| An Efficient Fault-Tolerance Scheme for CKKS Computation on CPUs | Jianan Mu, Ge Yu, Tenghui Hua, Liang Kong, Jing Ye, Xing Hu, Meng Li, Xiaowei Li, Huawei Li | 2026-07-21 | 下载 | Fully homomorphic encryption (FHE) enables computation on encrypted data, but its long ciphertext dataflow and high-dimensional modular arithmetic make it vulnerable to silent data corruption caused b... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Examining QRMI as a Unified Interface for Quantum-HPC Integration | Thomas Badts, Tim Boyle, Claudio Carvalho, Antonio Córcoles, Andrew Damin, Vadim Elisseev, Jonathan Frassineti, Daniel Gruber, Hiroshi Horii, Eun-Kyung Lee, James Machin, Sara Marzella, Mateusz Meller, Daniel Milroy, Matthieu Moreau, Munetaka Ohtani, Elisabeth Ortega-Carrasco, Doug Oucharek, Yoonho Park, Adarsh Patil, Emre M. Sahin, Gábor Samu, Seetharami Seelam, Amir Shehata, Vanessa Sochat, James Thorne, Oscar Wallis, Aleksander Wennersteen | 2026-07-21 | 下载 | The efficient and scalable integration of quantum resources into high-performance computing (HPC) environments requires standardized mechanisms for resource management, scheduling, and workflow orches... |
| Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts | Minyu Cui, Anna Wingkvist, Morgan Ericsson | 2026-07-21 | 下载 | Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key building block for scaling large language models (LLMs) to trill... |
| SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework | Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown | 2026-07-21 | 下载 | Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance... |
| Keeping the Cache Warm Pays: Keepalive Economics for Agentic Workloads | Maxim Khailo | 2026-07-21 | 下载 | Frontier LLM providers cache a prompt's processed prefix so that a follow-up request sharing it pays ~10% of the input price and skips most of the prefill latency. |
| ARBITER: Guarded Agentic Control for SLO-Oriented Kubernetes Remediation | Pooyan Habibi, Alberto Leon-Garcia | 2026-07-21 | 下载 | Maintaining service-level objectives (SLOs) on Kubernetes microservices remains difficult because autoscalers observe coarse resource metrics, recent SLO controllers often depend on custom telemetry, ... |
| Coherence in Control: Bridging Many-Core Mapping and Routing through Cost Unification | Guochu Xiong, Xiangzhong Luo, Weichen Liu | 2026-07-21 | 下载 | The rapid growth of data-intensive applications increases communication demands in many-core systems, where cache coherence, while essential for correct communication and data consistency, introduces ... |
| A Scalable Pattern Mining Workflow for Interpretable Machine Log Analysis in High-Performance Computing Environments | Shilpika Shilpika, Bethany Lusch, Eric Pershey, Carlo Graziani, Venkatram Vishwanath, Michael E. Papka | 2026-07-21 | 下载 | Modern supercomputers housed in High Performance Computing (HPC) environments generate massive volumes of log data daily, revealing intricate information and performance metrics about these complex sy... |
| Enabling Multi-Dimensional Distributed Trace Comparison with Contrast | Vaastav Anand, Rodrigo Fonseca, Jonathan Mace, Antoine Kaufmann | 2026-07-21 | 下载 | Diagnosis using distributed traces is fundamentally a comparative task: operators seek to understand how an anomalous execution differs from expected behavior, how a deployment changes system executio... |
| InstantInfer: Enabling Fast LLM Cold Start with Communicating Finite Automata | Yitao Yuan, Yongchao He, Shaoke Fang, Wenfei Wu | 2026-07-21 | 下载 | Cold starts in large language model (LLM) inference services significantly affect user experience, yet they remain inefficient due to sequential initialization and a massive number of fine-grained I/O... |
| A User-oriented Portable, Reproducible, and Scalable Software Ecosystem | Alfio Lazzaro, Utz-Uwe Haus, Sandrine Charousset, Nina Mujkanovic | 2026-07-21 | 下载 | It is normal for scientists to perform their research on a diverse set of hardware, ranging from laptops and workstations to supercomputers and cloud resources. |
| Mapping Without Graphs: Learning Coherence Traffic for Task Placement | Guochu Xiong, Tianrui Ma, Weichen Liu | 2026-07-21 | 下载 | Cache coherence is essential for communication in many-core Network-on-Chip (NoC)-based systems. As application scale and complexity increase, efficiently managing communication becomes increasingly c... |
| BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators | Fabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann | 2026-07-21 | 下载 | Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API. |
| A Second-Moment Theory for Floating-Point Reduction Trees | Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara, Keita Teranishi, Sudip Seal | 2026-07-21 | 下载 | Summation error depends on partial-sum order, which standard worst-case bounds omit. To capture this dependence, we derive an exact mean-square error (MSE) recurrence for a binary reduction tree T und... |
| Unstructured Hydrodynamics on Spatial Dataflow Architectures: A Joint Code and Data Decomposition Approach | Piotr Luczynski, Tal Ben-Nun, Leighton Wilson, Brian Van Essen | 2026-07-21 | 下载 | Spatial Dataflow Architectures are an emerging hardware pattern in high-performance computing, whose mesh-connected fixed-memory processing elements are tailored for structured grid kernels with two-d... |
| Searching for Plans You Can Actually Build: A Realizability-Aware Full-Space Optimizer for MoE Training and Serving | Quan Yuan, Jie Zhao | 2026-07-21 | 下载 | Mixture-of-Experts (MoE) systems split a program's plan space in two: the space a cost model can rank, and the smaller space a real toolchain can actually build. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Scalable Multi-Controller Coordination in Periplus via Border-Switch Forwarding Graphs | E. M. Castro Barbero, P. de las Heras Quirós, F. J. Simó Reigadas | 2026-07-21 | 下载 | In-band SDN control planes, where control traffic shares the data-plane infrastructure, suit wide-area, resource-constrained deployments -- such as rural backbones -- that cannot afford a dedicated co... |
| Squeezing the Most Out of Preemption for AoI Minimization: Single-source Case | Nail Akar, Mohammad Moltafet, Sennur Ulukus, Marian Codreanu, Roy D. Yates | 2026-07-21 | 下载 | In this work, we study a single-source single-server continuous-time status update system where the updates arrive according to a Poisson process and update service times are generally distributed. |
| HACO: Hedged Agent Computing for Reliable LLM Systems | Enhan Li, Hongyang Du | 2026-07-21 | 下载 | As large language model (LLM) agents move from isolated prompting to longhorizon workflows, failures increasingly arise at the role-to-instance binding boundary, where task-specific role requests must... |
| Structured Spectral Compression based Low-Bitrate Secure Speech Communications for Internet of Things assisted Non-Terrestrial Networks | Li Ping Qian, Zhehan Chen, Qianru Wang, Qian Wang, Yuan Wu, Xuemin Sherman Shen | 2026-07-21 | 下载 | This paper focuses on the Low-Bitrate Secure Speech Communications based on the Structured Spectral Compression (LB-S2C2). Specifically, the Mel spectral matrix of the speech signal is first encoded a... |
| Online Stochastic Matchings: Stability on Hypergraphs | Fabien Mathieu | 2026-07-21 | 下载 | We study stochastic dynamic matching on hypergraphs: items of finitely many classes arrive over time and are removed in multisets by activating hyperedges. |
| NSMA: Neuro-Symbolic Manifold Alignment for Generalizable Adaptive Bitrate Streaming under Texture Shift | Zhiqiang He, Zhi Liu | 2026-07-21 | 下载 | For decades, ABR has kept two kinds of intelligence apart. Neural policies learn rich behaviors yet forget them the moment the environment changes; rules never learn, and never forget. |
| Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach | Zijiang Yan, Hao Zhou, Wael Jaafar, Jianhua Pei, Ping Wang, Halim Yanikomeroglu, Hina Tabassum | 2026-07-21 | 下载 | The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordination of physical flight kinematics and multi-tier network handovers. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Squeezing the Most Out of Preemption for AoI Minimization: Single-source Case | Nail Akar, Mohammad Moltafet, Sennur Ulukus, Marian Codreanu, Roy D. Yates | 2026-07-21 | 下载 | In this work, we study a single-source single-server continuous-time status update system where the updates arrive according to a Poisson process and update service times are generally distributed. |
| Job-level Carbon and Water Footprint Estimation for HPC: Bias Assessment from Runtime to Full Life Cycle | Xi Chen, Chris Broekema, Rob van Nieuwpoort | 2026-07-21 | 下载 | High performance computing evaluation has traditionally focused on performance and energy, but these metrics alone cannot capture the sustainability cost of runtime configurations. |
| Formulation-Level Auto-Tuning for QUBO-Based Machine Learning: A Case Study Across Multiple Quantum-Inspired Annealers | Naoya Mizuki, Takahiro Katagiri, Daichi Mukunoki, Tetsuya Hoshino | 2026-07-21 | 下载 | This paper presents an Optuna-based formulation-level auto-tuning framework for support vector machines (SVMs) implemented on multiple quantum-inspired annealers. |
| BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators | Fabian Waschkowski, Prabod Rathnayaka, Lukas Wesemann | 2026-07-21 | 下载 | Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API. |