2026-07-06
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| NEMESIS: NEtlist-Driven Modeling and Equation Synthesis with Inversion-Aware SPICE Anchoring | Subhadip Ghosh, Ramesh Harjani, Sachin S. Sapatnekar | 2026-07-06 | 下载 | This work presents NEMESIS, a multimodal framework for operational transconductance amplifier (OTA) design using large language models (LLMs). |
| Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous Computing | Joel Klein, Rebecca Pelke, Roberto Laudani, Jan Moritz Joseph, Rainer Leupers | 2026-07-06 | 下载 | Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications (MVMs) in memory, making them a compelling solution for Machine Learning (ML) workloads. |
| SMART: A Machine Learning and Monte Carlo Framework for Rapid Analysis of Stochastic Transistor Aging and Process Variation in Digital Circuits | Arash Esshaghi, Siavash Es'haghi, Gholamreza Shahabadi, Alireza Moradi | 2026-07-06 | 下载 | As CMOS technology scales into the deep nanometer regime, digital circuit reliability is increasingly threatened by the combined stochastic effects of Bias Temperature Instability (BTI) and Process Va... |
| Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference | Guanyu Cai, Ruiming Tian, Lang Yang, Zhouhong Ren, Jinliang Yuan, Lingkun Li, Jiliang Wang | 2026-07-06 | 下载 | Deploying Large Language Models (LLMs) on mobile devices enhances privacy and reduces latency, but is severely bottlenecked by hardware inefficiency. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| EOM-CC Excited-State Gradients and Nonadiabatic Couplings on a Consumer GPU from a Contraction-DAG with Laplace-Transform J/K Kernels | Rubén Darío Guerrero | 2026-07-06 | 下载 | We present a unified, memory-bounded GPU realization of equation-of-motion coupled-cluster (EOM-CC) excited-state gradients and interstate nonadiabatic couplings (NACMEs) on a single 8,GB consumer GP... |
| Bounded-Memory Parallel Image Pulling for Large Container Images | Sri Saran Balaji Vellore Rajakumar, Henry Wang, Ankur Singh, James Thompson | 2026-07-06 | 下载 | AI/ML workloads increasingly run as containers, where a container image must be downloaded to the host before the workload can start. This cold image pull lands on the critical path whenever a trainin... |
| CATs: Secure Blockchain Interoperability with Cross-chain Atomic Transactions | Andreas Penzkofer, Franck Cassez | 2026-07-06 | 下载 | We propose a protocol for cross-chain atomic transactions (CATs), enabling composable atomic execution across different blockchains. The protocol addresses the key interoperability challenge of provid... |
| Adaptive Inference Batching using Policy Gradients | Ruslan Sharifullin | 2026-07-06 | 下载 | Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot ad... |
| Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous Computing | Joel Klein, Rebecca Pelke, Roberto Laudani, Jan Moritz Joseph, Rainer Leupers | 2026-07-06 | 下载 | Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications (MVMs) in memory, making them a compelling solution for Machine Learning (ML) workloads. |
| Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts Inference | Xiao Shi, Yingying Sun, Jiangsu Du, Zhiguang Chen, Yutong Lu | 2026-07-06 | 下载 | As MoE models scale to hundreds of experts, placement and pruning decisions increasingly dictate communication volume, affecting the performance of distributed inference across GPUs and nodes. |
| Watts per event: evaluating Sustainability of HEP Event Generators beyond the LHC era | Szabolcs Molnár, Gábor Bíró, Gábor Papp, Gergely Gábor Barnaföldi | 2026-07-06 | 下载 | The development, tuning and operation of Monte Carlo event generators beyond the LHC era require vast amount of resources. In this study we investigate the sustainability of these software with a cont... |
| When Words Predict Workload | Anubhab Banerjee | 2026-07-06 | 下载 | Standard distributed \ac{llm} schedulers rely on static token counts or rolling latency averages, making them susceptible to failures on statutorily constrained text. |
| TARE: Tail Aware Evaluation of HPC Job Runtime Prediction | Haili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, Rong He | 2026-07-06 | 下载 | Runtime estimates affect reservation quality, backfilling opportunities, and queue delay in HPC schedulers. Under heavy tailed workloads, however, averaging over jobs can misrepresent scheduling impac... |
| Symmetry all the way down | Ignacio Amores-Sesar, Christian Cachin, Simon Holmgaard Kamp, Juan Villacis | 2026-07-06 | 下载 | Asymmetric trust generalizes classical symmetric quorum systems by allowing each process to specify its own failure assumptions. While this flexibility enables tolerance of strictly more failure scena... |
| Terastate-per-second QUBO Brute-Force on a Single GPU: A Matrix Prefix-Suffix Decomposition | Aleksandr Maltsev, Mikhail Remnev, Alexey Kapranov, Ekaterina Krivtsova | 2026-07-06 | 下载 | This paper presents a parallel QUBO exhaustive search algorithm for dense matrices, based on a prefix-suffix decomposition and Gray code ordering. |
| No Distributed Quantum Advantage for 3-Coloring Rooted Trees and 2-Coloring Even Cycles | Pierre Fraigniaud, Frédéric Magniez, Isabella Ziccardi | 2026-07-06 | 下载 | Significant effort has been devoted over the past decade to understanding whether quantum resources can provide advantages in distributed computing, and in particular whether they can help overcome lo... |
| Performance evaluation of scheduling tasks in many-core systems utilizing processes and threads | Mejgan Dedaj, Argyro Gailla, Theofanis Ioannou, Stamatia Kastrinaki, Hermione Kimpouropoulou, Dimitrios Kontodimos, Kleopatra Kontogianni, Sotirios Kontogiannis, Michail Panagiotidis Kannas, Anastasia Papouda, Anna Maria Sidiropoulou, George Tavridis | 2026-07-06 | 下载 | This study assesses the scalability of process-based and thread-based schedulers for many-core shared-memory systems using a memory-intensive row-wise quick-sort workload on large three-dimensional te... |
| Orcaella: Hybrid Fault Tolerance with Client-Selectable Finality Latency | Lefteris Kokoris-Kogias, Alberto Sonnino | 2026-07-06 | 下载 | Classical partially synchronous state machine replication, as in PBFT, tolerates f Byzantine replicas among n at least 3f+1 using three communication steps per request. |
| Direct Model State Migration for Elastic Training of Large Language Models | Weijian Liu, Mingzhen Li, Rui Kang, Chen Sun, Guangming Tan, Weile Jia | 2026-07-06 | 下载 | Large language model (LLM) training shall adapt to dynamic resources in shared clusters to tackle the elasticity, including passive preemption and optimistic scaling. |
| Adaptive Space-efficient Collectives for Dynamic and Unstructured Sparsity on GPU Platforms | Lannie Dalton Hough, Emir Gencer, Hoffmann Muki, Abhinav Bhatele | 2026-07-06 | 下载 | High-performance collective communication primitives are necessary for a variety of high performance computing (HPC) and machine learning (ML) workloads. |
| Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action Reasoning | Jiaxing Qi, Zhongzhi Luan, Hongyu Zhang, Shaohan Huang, Carol Fung, Yongxin Tong, Hailong Yang, Depei Qian | 2026-07-06 | 下载 | Large language models (LLMs) are increasingly used to interpret operational evidence and assist incident response in cloud-native microservice systems. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Towards Quantum Network Performance Metrics: Challenges and Demonstration | Mohamed Shaban, Mariam Kiran, Muhammad Ismail | 2026-07-06 | 下载 | As quantum networks move toward practical deployment, standardized performance monitoring becomes essential. This article proposes a structured monitoring framework for quantum networks with performan... |
| A Modular O-RAN Testbed Based on SRS Open Source O-CU/O-DU and Massive Beams Modular O-RU | Fabian Göttsch, Oriol Font-Bach, Andreas Benzin, Dennis Osterland, Andre Puschmann, Felix-Christopher Lutz, Wilhelm Keusgen, Giuseppe Caire | 2026-07-06 | 下载 | In this paper, we present a modular open radio access network (O-RAN) consisting of the 5G Core, a central (O-CU) and distributed unit (O-DU) by Software Radio Systems (SRS) and an O-RAN radio unit (O... |
| RANPilot: Making AI Functionalities Robust to Dynamic O-RAN Reconfigurations | Shimin Yu, Leming Shen, Jianing Zhang, Xin Li, Xianjin Xia, Yuanqing Zheng, Yaxiong Xie | 2026-07-06 | 下载 | The Open Radio Access Network (O-RAN) promises unprecedented flexibility through its reconfigurable architecture and AI-driven control. However, this agility exposes a critical fragility: AI models tr... |
| Version-Aware Communication in Multi-Hop IoT Networks with Feedback | Erfan Delfani, Nikolaos Pappas | 2026-07-06 | 下载 | Timely communication of information in Internet of Things (IoT) networks is critical to enhancing system performance and energy efficiency by minimizing the transmission of outdated or redundant data. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Complets: Universal Compartmentalisation and Programming Model For Arm Permission Overlay Extension 2 | Vasily A. Sartakov | 2026-07-06 | 下载 | Arm Permission Overlay Extension (POE) is an intra-process isolation mechanism based on memory protection keys. This mechanism partitions virtual memory into regions whose access permissions can be re... |
| Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes | Daeyeon Son | 2026-07-06 | 下载 | On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS wants those same cores continuously. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Bounded-Memory Parallel Image Pulling for Large Container Images | Sri Saran Balaji Vellore Rajakumar, Henry Wang, Ankur Singh, James Thompson | 2026-07-06 | 下载 | AI/ML workloads increasingly run as containers, where a container image must be downloaded to the host before the workload can start. This cold image pull lands on the critical path whenever a trainin... |
| Adaptive Inference Batching using Policy Gradients | Ruslan Sharifullin | 2026-07-06 | 下载 | Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot ad... |
| TARE: Tail Aware Evaluation of HPC Job Runtime Prediction | Haili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, Rong He | 2026-07-06 | 下载 | Runtime estimates affect reservation quality, backfilling opportunities, and queue delay in HPC schedulers. Under heavy tailed workloads, however, averaging over jobs can misrepresent scheduling impac... |
| Performance evaluation of scheduling tasks in many-core systems utilizing processes and threads | Mejgan Dedaj, Argyro Gailla, Theofanis Ioannou, Stamatia Kastrinaki, Hermione Kimpouropoulou, Dimitrios Kontodimos, Kleopatra Kontogianni, Sotirios Kontogiannis, Michail Panagiotidis Kannas, Anastasia Papouda, Anna Maria Sidiropoulou, George Tavridis | 2026-07-06 | 下载 | This study assesses the scalability of process-based and thread-based schedulers for many-core shared-memory systems using a memory-intensive row-wise quick-sort workload on large three-dimensional te... |
| Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes | Daeyeon Son | 2026-07-06 | 下载 | On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS wants those same cores continuously. |