2026-09-03
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AI-Assisted Design of a Post-Quantum Cryptographic Accelerator: A Deployed-Silicon Case Study | Jungmin Park, Eunha Kim, Wooseop Kim, Seongjoon Cho, Byungho Cha | 2026-09-03 | 下载 | Post-quantum migration is mandated on published timelines, and silicon that ships with a defect cannot be patched remotely. The standard acceptance gate cannot detect an entire class of ML-DSA defects... |
| Confidence-Gated Admission for Hardware Prefetching: When the Gate Matters More Than the Predictor | Youssef Majdane, Simone Jarno Casartelli, Enrico Lopedoto | 2026-09-03 | 下载 | Learned cache prefetchers are typically evaluated against classical predictors that always issue requests, confounding the prediction model with the admission policy. |
| LevelSyn: Physical-Aware Logic Synthesis via Level-Asynchronous Graph Neural Networks | Jingyi Zhou, Zhengyuan Shi, Ziyang Zheng, Qiang Xu | 2026-09-03 | 下载 | As integrated circuit technology scales into the nanometer regime, the traditional disconnect between logic synthesis and physical design has led to significant PPA (Power, Performance, and Area) degr... |
| Huawei's τ Chip Was Supposed to Melt? | Tingbo He | 2026-09-03 | 下载 | Heat is the sharpest concern on τ scaling law --- the time-scaling principle behind Huawei's folded silicon, named for the delay τ (``Tao'') it sets out to shrink. |
| LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL | Sijie Wang, Zhiqiang Tan, Xinrui Yang, Shaohuai Shi | 2026-09-03 | 下载 | Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including DanceGRPO and FlowGRPO... |
| FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train Embedding | Jongmin Seok, Chae Eun Rhee | 2026-09-03 | 下载 | Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows acr... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Revisiting MemGuard Overhead: A Reproduction Report | Weifan Chen, Heechul Yun, Renato Mancuso | 2026-09-03 | 下载 | As an increasing number of embedded platforms incorporate multiple processing units, shared resource contention induced unpredictable execution time poses a challenge for real-time system design. |
| DPRQ: A Dynamic Programming-based Qubit Routing Algorithm for Collective Communication in Distributed Quantum Computing | Dhaval Vaidya, Ruozhou Yu | 2026-09-03 | 下载 | Distributed quantum computing (DQC) offers a promising approach to scale quantum computing by overcoming the resource limitations of a single quantum processor. |
| Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous Clusters | Milos Gravara, Andrija Stanisic, Stefan Nastic | 2026-09-03 | 下载 | Compound AI workflows are increasingly used to serve complex AI tasks by coordinating multiple AI models and software components. This approach enables deployment flexibility, as each workflow stage c... |
| Performance Study of Serverless Workloads in Confidential Virtual Machines | Rikesh Niroula, Jianchen Shan, Xiaoning Ding | 2026-09-03 | 下载 | Confidential serverless computing is rapidly emerg- ing as a critical paradigm for application domains requiring strong confidentiality guarantees, such as healthcare, finance, and machine learning. |
| Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters | Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal | 2026-09-03 | 下载 | Distributed AI training involves recurring rounds of data exchange between multiple pairs of GPU nodes. Slowdown in even one flow due to congestion can cause the entire communication round to slowdown... |
| Accelerating Atom Simulations with Variable-Block Sparse Matrix Library | Zhanghao Zhouyin, Hong Guo | 2026-09-03 | 下载 | Modern atomistic simulations increasingly employ localized orbitals to represent quantum operators, yielding sparse block matrices whose block shapes vary with chemical species and basis choice. |
| Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys | Georgios Politis, Evangelos Pappas | 2026-09-03 | 下载 | We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested. |
| Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs | Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra | 2026-09-03 | 下载 | As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. |
| Barnacle: Adaptive Multi-Leader Scheduling for DAG-Based Consensus | Zeno De Angeli, Alexandru Ianov Vitanov, Philipp Jovanovic, Lefteris Kokoris-Kogias, Alberto Sonnino, Pasindu Tennage, Igor Zablotchi | 2026-09-03 | 下载 | In DAG-based consensus, all validators propose blocks concurrently, and designated leader blocks drive transaction commit. Having multiple leader slots per round cuts queuing latency, yet production d... |
| sp-DBA: a general framework for adaptive transform-domain computation | Jingkun Jiang, Pingchuan Deng, Yang Xia | 2026-09-03 | 下载 | Transform-domain methods simplify analysis and computation, making them central to scientific computing and signal processing. However, existing adaptive strategies often introduce new data structures... |
| Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in Einsummable | Zhimin Ding, Chen-Kuan Liao, Chima Adiole, Brianna Barrow, Fangzhou Du, Yu Hsiao, Ge Huang, Yicheng Jin, Ismail Syed, Chris Jermaine | 2026-09-03 | 下载 | Distributing an AI computation across the GPUs of a multi-GPU server is one of the central problems in systems-for-AI. We present Einsummable, a prototype system that accepts a PyTorch-like descriptio... |
| JuPyLive: Seamless Migration of Jupyter Notebook Resources from Laptop to HPC | Sima Attar-Khorasani, Matthias Lieber, Siavash Ghiasvand | 2026-09-03 | 下载 | This work introduces JuPyLive, a migration mechanism that enables seamless transition of Jupyter notebooks between local resources of user's workstation and remote resources of high-performance comput... |
| FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train Embedding | Jongmin Seok, Chae Eun Rhee | 2026-09-03 | 下载 | Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows acr... |
| Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic Programming | Hao Mao, Xu Tony Liu, Shuai Lu, Peng Zhao, Wenzheng Jiang, Yuntian Chen | 2026-09-03 | 下载 | Constant optimization refines the numerical coefficients of candidate expressions in tree-based genetic programming for symbolic regression. But its per-generation cost has led modern GPU-accelerated ... |
| Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs | Jinghao Wang, Yifeng Zhang, Xiao Zhou, Yao Lu, Yihui Zhang, Xiaoyang Sun, Tianyu Wo, Xu Wang, Chunming Hu, Renyu Yang | 2026-09-03 | 下载 | Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. |
| Iapetus: Content-Aware Hierarchical Scheduling for Collaborative ViT Inference in LEO Satellite Networks | Yan Chen, Yunxiang Zhang, Guanjun Jiang, Haiquan Wang | 2026-09-03 | 下载 | Collaborative inference pools distributed resources to run compute-intensive Vision Transformers (ViTs) in satellite edge computing. Model partitioning enables such collaboration by assigning consecut... |
| Lantern: Finding Committable Transactions via Back-Propagation on DAGs | Denglong Li, Gerui Wang, Tian Guan, Mingchao Wan | 2026-09-03 | 下载 | Existing concurrency control protocols either introduce nondeterminism, resulting in a serial execution-replay dependency between primary and replica nodes, or rely on impractical prior knowledge of t... |
| A Technique for Load Shifting Low-latency Applications in Multi-Region Renewables Harvesting via SMT Core Pooling | Tharindu B. Hewage, Shashikant Ilager, Maria A. Rodriguez, Rajkumar Buyya | 2026-09-03 | 下载 | Load shifting across geographic regions to chase intermittent renewable energy availability is commonly used in reducing cloud infrastructure carbon footprint. |
| Carbon-aware Resource Management for Latency-Sensitive Cloud Computing Environments: A Taxonomy and Future Directions | Tharindu B. Hewage, Shashikant Ilager, Maria Rodriguez Read, Rajkumar Buyya | 2026-09-03 | 下载 | Proliferation of cloud-based latency-sensitive workloads requires infrastructures tuned to their workload-specific latency constraints. Today, they shape the cloud from a generalized computing platfor... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended Reality | Yoga Suhas Kuruba Manjunath, Jie Gao, Lian Zhao | 2026-09-03 | 下载 | We present ResLearn-XR, a residual learning framework for predicting eXtended Reality (XR) network traffic and estimating Quality-of-Experience (QoE) risk. |
| Uncertainty Signals for Network Intent Translation: Risk Ranking and Ambiguity Localization | Ala' A. Alsamarneh, Omar Alhussein | 2026-09-03 | 下载 | Intent-based networking realization starts by translating high-level intents into low-level network configurations. Recent approaches have shifted toward LLM-based translation. |
| Waves on the Walls: Empirical Characterization of mmWave Lateral Waves for Enhanced Indoor Coverage | Apala Pramanik, Avhishek Biswas, Sasitharan Balasubramaniam, Christos Argyropoulos, Mehmet C. Vuran | 2026-09-03 | 下载 | High-frequency millimeter-wave (mmWave) communication systems are constrained by the surrounding environment, where walls are traditionally treated as obstacles that block or reflect signals indoors. |
| Tuning Collective Patterns to Alleviate Congestion in Shared AI Clusters | Eashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal | 2026-09-03 | 下载 | Distributed AI training involves recurring rounds of data exchange between multiple pairs of GPU nodes. Slowdown in even one flow due to congestion can cause the entire communication round to slowdown... |
| Network Availability Enhancement in Low-Altitude HetNets: A Cross-Layer Design Perspective | Teng Wu, Jiandong Li, Junyu Liu, Min Sheng, Mohammadali Mohammadi, Hien Quoc Ngo, Michail Matthaiou | 2026-09-03 | 下载 | This paper proposes a computing-communication resource interchange method to enhance network availability (NA) in low-altitude heterogeneous networks (LA-HetNets). |
| Is Collision-Free Backoff Worth It in Wi-Fi? | Mohammad Yousefi, Francesc Wilhelmi, Boris Bellalta | 2026-09-03 | 下载 | The Distributed Coordination Function (DCF)---the underlying channel access protocol in Wi-Fi, based on Carrier Sense Multiple Access with Collision Avoidance (CSMA/CA) and Binary Exponential Backoff ... |
| Employing the Structural Power to Achieve Supply-Demand Balanced Payment Channel Networks | Shuyao Xiao, Shengling Wang, Hongwei Shi, Weicheng Wang, Anlin Chen | 2026-09-03 | 下载 | Blockchain technology faces scalability challenges because transactions must be validated and recorded across the network. Payment channel networks (PCNs) improve efficiency by moving transactions off... |
| From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control | Vincenzo Norman Vitale, Mohammad Solki, Antonia Maria Tulino, Andreas F. Molisch, Jaime Llorca | 2026-09-03 | 下载 | Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is essential for NextG interactive applications, yet providing strict End-to-End (E2E) peak latency guarantees remai... |
| A Semantic-Aware Multiple Access Scheme Leveraging Spatial Redundancy for Uplink-Dominant Network Services | Hamidreza Mazandarani, Masoud Shokrnezhad, Tarik Taleb | 2026-09-03 | 下载 | The transition toward semantic-aware communication offers a paradigm shift for next-generation mobile networks, promising to decouple information significance from raw data transmission. |
| An Adversarial Zero-Shot Learning Approach for Anomaly Detection in Multivariate IoT Traffic Data | Mahshid Rezakhani, Tolunay Seyfi, Fatemeh Afghah | 2026-09-03 | 下载 | Anomaly detection in Internet of Things (IoT) networks presents unique challenges due to the diversity of devices, lack of labeled data, and domain variability across environments. |
| Indirect Estimation of SINR via SSB and CSI-RS RSRP in 5G NR | Leonardo Spampinato, Mahamadou Togola, Matteo Bernabè, Azim Akhtarshenas, Lorenzo Mario Amorosa, David López-Pérez | 2026-09-03 | 下载 | Predicting user equipment (UE) performance is essential for proactive network control, resource management, and digital twin sandboxes. However, the inherent flexibility and complexity of beam-based 5... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| NACRE: Rethinking Confidential Containers through Native Architectural Support | Linke Song, Wenhao Wang, Weijie Liu, Rui Hou | 2026-09-03 | 下载 | Linux containers achieve high density and fast lifecycle operations by sharing the host kernel, but this design also lets a compromised host inspect or modify container state. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| MaxKernel: Agentic Kernel Generation for TPUs | Shangkun Wang, Nina Cai, Charles Hoong, Julian Walker, Gerson Kroiz, George Vanica, Deepak Patil, Andi Gavrilescu, Hassan Sipra, Sethu Sankaran | 2026-09-03 | 下载 | Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-... |
| PerfReasoning: How Well Do LLMs Reason on Hardware Performance? | Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang | 2026-09-03 | 下载 | Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. |
| On-board ML for Trace Gas detection in Imaging Spectroscopy data | Vít Růžička, Adam Chlus, Andrew Thorpe, David R. Thompson | 2026-09-03 | 下载 | Data collected during aerial and spaceborne imaging spectroscopy campaigns enables the detection of transient events such as trace gas emissions. |
| Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCs | Yujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra | 2026-09-03 | 下载 | As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. |
| Confidence-Gated Admission for Hardware Prefetching: When the Gate Matters More Than the Predictor | Youssef Majdane, Simone Jarno Casartelli, Enrico Lopedoto | 2026-09-03 | 下载 | Learned cache prefetchers are typically evaluated against classical predictors that always issue requests, confounding the prediction model with the admission policy. |
| RASER: Resilient Agent Scheduling and Execution Runtime for HPC Clusters | Sima Attar-Khorasani, Matthias Lieber, Siavash Ghiasvand | 2026-09-03 | 下载 | The emergence of modern agents powered by large language models has created a demand for executing long-horizon, autonomous workflows in various domains that require significant computational resource... |
| Lantern: Finding Committable Transactions via Back-Propagation on DAGs | Denglong Li, Gerui Wang, Tian Guan, Mingchao Wan | 2026-09-03 | 下载 | Existing concurrency control protocols either introduce nondeterminism, resulting in a serial execution-replay dependency between primary and replica nodes, or rely on impractical prior knowledge of t... |