Skip to content

2026-07-06 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
NEMESIS: NEtlist-Driven Modeling and Equation Synthesis with Inversion-Aware SPICE AnchoringSubhadip Ghosh, Ramesh Harjani, Sachin S. Sapatnekar2026-07-06下载This work presents NEMESIS, a multimodal framework for operational transconductance amplifier (OTA) design using large language models (LLMs).
Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous ComputingJoel Klein, Rebecca Pelke, Roberto Laudani, Jan Moritz Joseph, Rainer Leupers2026-07-06下载Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications (MVMs) in memory, making them a compelling solution for Machine Learning (ML) workloads.
SMART: A Machine Learning and Monte Carlo Framework for Rapid Analysis of Stochastic Transistor Aging and Process Variation in Digital CircuitsArash Esshaghi, Siavash Es'haghi, Gholamreza Shahabadi, Alireza Moradi2026-07-06下载As CMOS technology scales into the deep nanometer regime, digital circuit reliability is increasingly threatened by the combined stochastic effects of Bias Temperature Instability (BTI) and Process Va...
Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM InferenceGuanyu Cai, Ruiming Tian, Lang Yang, Zhouhong Ren, Jinliang Yuan, Lingkun Li, Jiliang Wang2026-07-06下载Deploying Large Language Models (LLMs) on mobile devices enhances privacy and reduces latency, but is severely bottlenecked by hardware inefficiency.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
EOM-CC Excited-State Gradients and Nonadiabatic Couplings on a Consumer GPU from a Contraction-DAG with Laplace-Transform J/K KernelsRubén Darío Guerrero2026-07-06下载We present a unified, memory-bounded GPU realization of equation-of-motion coupled-cluster (EOM-CC) excited-state gradients and interstate nonadiabatic couplings (NACMEs) on a single 8,GB consumer GP...
Bounded-Memory Parallel Image Pulling for Large Container ImagesSri Saran Balaji Vellore Rajakumar, Henry Wang, Ankur Singh, James Thompson2026-07-06下载AI/ML workloads increasingly run as containers, where a container image must be downloaded to the host before the workload can start. This cold image pull lands on the critical path whenever a trainin...
CATs: Secure Blockchain Interoperability with Cross-chain Atomic TransactionsAndreas Penzkofer, Franck Cassez2026-07-06下载We propose a protocol for cross-chain atomic transactions (CATs), enabling composable atomic execution across different blockchains. The protocol addresses the key interoperability challenge of provid...
Adaptive Inference Batching using Policy GradientsRuslan Sharifullin2026-07-06下载Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot ad...
Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous ComputingJoel Klein, Rebecca Pelke, Roberto Laudani, Jan Moritz Joseph, Rainer Leupers2026-07-06下载Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications (MVMs) in memory, making them a compelling solution for Machine Learning (ML) workloads.
Communication-Aware Placement and Pruning for Efficient Mixture-of-Experts InferenceXiao Shi, Yingying Sun, Jiangsu Du, Zhiguang Chen, Yutong Lu2026-07-06下载As MoE models scale to hundreds of experts, placement and pruning decisions increasingly dictate communication volume, affecting the performance of distributed inference across GPUs and nodes.
Watts per event: evaluating Sustainability of HEP Event Generators beyond the LHC eraSzabolcs Molnár, Gábor Bíró, Gábor Papp, Gergely Gábor Barnaföldi2026-07-06下载The development, tuning and operation of Monte Carlo event generators beyond the LHC era require vast amount of resources. In this study we investigate the sustainability of these software with a cont...
When Words Predict WorkloadAnubhab Banerjee2026-07-06下载Standard distributed \ac{llm} schedulers rely on static token counts or rolling latency averages, making them susceptible to failures on statutorily constrained text.
TARE: Tail Aware Evaluation of HPC Job Runtime PredictionHaili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, Rong He2026-07-06下载Runtime estimates affect reservation quality, backfilling opportunities, and queue delay in HPC schedulers. Under heavy tailed workloads, however, averaging over jobs can misrepresent scheduling impac...
Symmetry all the way downIgnacio Amores-Sesar, Christian Cachin, Simon Holmgaard Kamp, Juan Villacis2026-07-06下载Asymmetric trust generalizes classical symmetric quorum systems by allowing each process to specify its own failure assumptions. While this flexibility enables tolerance of strictly more failure scena...
Terastate-per-second QUBO Brute-Force on a Single GPU: A Matrix Prefix-Suffix DecompositionAleksandr Maltsev, Mikhail Remnev, Alexey Kapranov, Ekaterina Krivtsova2026-07-06下载This paper presents a parallel QUBO exhaustive search algorithm for dense matrices, based on a prefix-suffix decomposition and Gray code ordering.
No Distributed Quantum Advantage for 3-Coloring Rooted Trees and 2-Coloring Even CyclesPierre Fraigniaud, Frédéric Magniez, Isabella Ziccardi2026-07-06下载Significant effort has been devoted over the past decade to understanding whether quantum resources can provide advantages in distributed computing, and in particular whether they can help overcome lo...
Performance evaluation of scheduling tasks in many-core systems utilizing processes and threadsMejgan Dedaj, Argyro Gailla, Theofanis Ioannou, Stamatia Kastrinaki, Hermione Kimpouropoulou, Dimitrios Kontodimos, Kleopatra Kontogianni, Sotirios Kontogiannis, Michail Panagiotidis Kannas, Anastasia Papouda, Anna Maria Sidiropoulou, George Tavridis2026-07-06下载This study assesses the scalability of process-based and thread-based schedulers for many-core shared-memory systems using a memory-intensive row-wise quick-sort workload on large three-dimensional te...
Orcaella: Hybrid Fault Tolerance with Client-Selectable Finality LatencyLefteris Kokoris-Kogias, Alberto Sonnino2026-07-06下载Classical partially synchronous state machine replication, as in PBFT, tolerates f Byzantine replicas among n at least 3f+1 using three communication steps per request.
Direct Model State Migration for Elastic Training of Large Language ModelsWeijian Liu, Mingzhen Li, Rui Kang, Chen Sun, Guangming Tan, Weile Jia2026-07-06下载Large language model (LLM) training shall adapt to dynamic resources in shared clusters to tackle the elasticity, including passive preemption and optimistic scaling.
Adaptive Space-efficient Collectives for Dynamic and Unstructured Sparsity on GPU PlatformsLannie Dalton Hough, Emir Gencer, Hoffmann Muki, Abhinav Bhatele2026-07-06下载High-performance collective communication primitives are necessary for a variety of high performance computing (HPC) and machine learning (ML) workloads.
Can LLMs Really Recover Microservice Failures? A Recovery-Aware Evaluation of Diagnosis-to-Action ReasoningJiaxing Qi, Zhongzhi Luan, Hongyu Zhang, Shaohan Huang, Carol Fung, Yongxin Tong, Hailong Yang, Depei Qian2026-07-06下载Large language models (LLMs) are increasingly used to interpret operational evidence and assist incident response in cloud-native microservice systems.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Towards Quantum Network Performance Metrics: Challenges and DemonstrationMohamed Shaban, Mariam Kiran, Muhammad Ismail2026-07-06下载As quantum networks move toward practical deployment, standardized performance monitoring becomes essential. This article proposes a structured monitoring framework for quantum networks with performan...
A Modular O-RAN Testbed Based on SRS Open Source O-CU/O-DU and Massive Beams Modular O-RUFabian Göttsch, Oriol Font-Bach, Andreas Benzin, Dennis Osterland, Andre Puschmann, Felix-Christopher Lutz, Wilhelm Keusgen, Giuseppe Caire2026-07-06下载In this paper, we present a modular open radio access network (O-RAN) consisting of the 5G Core, a central (O-CU) and distributed unit (O-DU) by Software Radio Systems (SRS) and an O-RAN radio unit (O...
RANPilot: Making AI Functionalities Robust to Dynamic O-RAN ReconfigurationsShimin Yu, Leming Shen, Jianing Zhang, Xin Li, Xianjin Xia, Yuanqing Zheng, Yaxiong Xie2026-07-06下载The Open Radio Access Network (O-RAN) promises unprecedented flexibility through its reconfigurable architecture and AI-driven control. However, this agility exposes a critical fragility: AI models tr...
Version-Aware Communication in Multi-Hop IoT Networks with FeedbackErfan Delfani, Nikolaos Pappas2026-07-06下载Timely communication of information in Internet of Things (IoT) networks is critical to enhancing system performance and energy efficiency by minimizing the transmission of outdated or redundant data.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Complets: Universal Compartmentalisation and Programming Model For Arm Permission Overlay Extension 2Vasily A. Sartakov2026-07-06下载Arm Permission Overlay Extension (POE) is an intra-process isolation mechanism based on memory protection keys. This mechanism partitions virtual memory into regions whose access permissions can be re...
Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS ProcessesDaeyeon Son2026-07-06下载On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS wants those same cores continuously.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Bounded-Memory Parallel Image Pulling for Large Container ImagesSri Saran Balaji Vellore Rajakumar, Henry Wang, Ankur Singh, James Thompson2026-07-06下载AI/ML workloads increasingly run as containers, where a container image must be downloaded to the host before the workload can start. This cold image pull lands on the critical path whenever a trainin...
Adaptive Inference Batching using Policy GradientsRuslan Sharifullin2026-07-06下载Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains static batching policies that require manual tuning and cannot ad...
TARE: Tail Aware Evaluation of HPC Job Runtime PredictionHaili Xiao, Can Wu, Shasha Lu, Xiaoning Wang, Yining Zhao, Rong He2026-07-06下载Runtime estimates affect reservation quality, backfilling opportunities, and queue delay in HPC schedulers. Under heavy tailed workloads, however, averaging over jobs can misrepresent scheduling impac...
Performance evaluation of scheduling tasks in many-core systems utilizing processes and threadsMejgan Dedaj, Argyro Gailla, Theofanis Ioannou, Stamatia Kastrinaki, Hermione Kimpouropoulou, Dimitrios Kontodimos, Kleopatra Kontogianni, Sotirios Kontogiannis, Michail Panagiotidis Kannas, Anastasia Papouda, Anna Maria Sidiropoulou, George Tavridis2026-07-06下载This study assesses the scalability of process-based and thread-based schedulers for many-core shared-memory systems using a memory-intensive row-wise quick-sort workload on large three-dimensional te...
Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS ProcessesDaeyeon Son2026-07-06下载On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS wants those same cores continuously.

基于 VitePress 构建 · 使用本地搜索查找论文