Skip to content

2026-09-03 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
AI-Assisted Design of a Post-Quantum Cryptographic Accelerator: A Deployed-Silicon Case StudyJungmin Park, Eunha Kim, Wooseop Kim, Seongjoon Cho, Byungho Cha2026-09-03下载Post-quantum migration is mandated on published timelines, and silicon that ships with a defect cannot be patched remotely. The standard acceptance gate cannot detect an entire class of ML-DSA defects...
Confidence-Gated Admission for Hardware Prefetching: When the Gate Matters More Than the PredictorYoussef Majdane, Simone Jarno Casartelli, Enrico Lopedoto2026-09-03下载Learned cache prefetchers are typically evaluated against classical predictors that always issue requests, confounding the prediction model with the admission policy.
LevelSyn: Physical-Aware Logic Synthesis via Level-Asynchronous Graph Neural NetworksJingyi Zhou, Zhengyuan Shi, Ziyang Zheng, Qiang Xu2026-09-03下载As integrated circuit technology scales into the nanometer regime, the traditional disconnect between logic synthesis and physical design has led to significant PPA (Power, Performance, and Area) degr...
Huawei's τ Chip Was Supposed to Melt?Tingbo He2026-09-03下载Heat is the sharpest concern on τ scaling law --- the time-scaling principle behind Huawei's folded silicon, named for the delay τ (``Tao'') it sets out to shrink.
LeanGRPO: Eliminating Redundant Recomputation in Diffusion RLSijie Wang, Zhiqiang Tan, Xinrui Yang, Shaohuai Shi2026-09-03下载Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including DanceGRPO and FlowGRPO...
FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train EmbeddingJongmin Seok, Chae Eun Rhee2026-09-03下载Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows acr...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Revisiting MemGuard Overhead: A Reproduction ReportWeifan Chen, Heechul Yun, Renato Mancuso2026-09-03下载As an increasing number of embedded platforms incorporate multiple processing units, shared resource contention induced unpredictable execution time poses a challenge for real-time system design.
DPRQ: A Dynamic Programming-based Qubit Routing Algorithm for Collective Communication in Distributed Quantum ComputingDhaval Vaidya, Ruozhou Yu2026-09-03下载Distributed quantum computing (DQC) offers a promising approach to scale quantum computing by overcoming the resource limitations of a single quantum processor.
Atlas: Optimizing Deployment of Compound AI Workflows on Heterogeneous ClustersMilos Gravara, Andrija Stanisic, Stefan Nastic2026-09-03下载Compound AI workflows are increasingly used to serve complex AI tasks by coordinating multiple AI models and software components. This approach enables deployment flexibility, as each workflow stage c...
Performance Study of Serverless Workloads in Confidential Virtual MachinesRikesh Niroula, Jianchen Shan, Xiaoning Ding2026-09-03下载Confidential serverless computing is rapidly emerg- ing as a critical paradigm for application domains requiring strong confidentiality guarantees, such as healthcare, finance, and machine learning.
Tuning Collective Patterns to Alleviate Congestion in Shared AI ClustersEashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal2026-09-03下载Distributed AI training involves recurring rounds of data exchange between multiple pairs of GPU nodes. Slowdown in even one flow due to congestion can cause the entire communication round to slowdown...
Accelerating Atom Simulations with Variable-Block Sparse Matrix LibraryZhanghao Zhouyin, Hong Guo2026-09-03下载Modern atomistic simulations increasingly employ localized orbitals to represent quantum operators, yielding sparse block matrices whose block shapes vary with chemical species and basis choice.
Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the DecoysGeorgios Politis, Evangelos Pappas2026-09-03下载We present a systems-security case study of a two-node split-LLM training system whose privacy evaluation passed while leaving an observable channel untested.
Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCsYujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra2026-09-03下载As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges.
Barnacle: Adaptive Multi-Leader Scheduling for DAG-Based ConsensusZeno De Angeli, Alexandru Ianov Vitanov, Philipp Jovanovic, Lefteris Kokoris-Kogias, Alberto Sonnino, Pasindu Tennage, Igor Zablotchi2026-09-03下载In DAG-based consensus, all validators propose blocks concurrently, and designated leader blocks drive transaction commit. Having multiple leader slots per round cuts queuing latency, yet production d...
sp-DBA: a general framework for adaptive transform-domain computationJingkun Jiang, Pingchuan Deng, Yang Xia2026-09-03下载Transform-domain methods simplify analysis and computation, making them central to scientific computing and signal processing. However, existing adaptive strategies often introduce new data structures...
Every Kernel Is a Join: Automatic Multi-GPU Parallelism for AI Computations in EinsummableZhimin Ding, Chen-Kuan Liao, Chima Adiole, Brianna Barrow, Fangzhou Du, Yu Hsiao, Ge Huang, Yicheng Jin, Ismail Syed, Chris Jermaine2026-09-03下载Distributing an AI computation across the GPUs of a multi-GPU server is one of the central problems in systems-for-AI. We present Einsummable, a prototype system that accepts a PyTorch-like descriptio...
JuPyLive: Seamless Migration of Jupyter Notebook Resources from Laptop to HPCSima Attar-Khorasani, Matthias Lieber, Siavash Ghiasvand2026-09-03下载This work introduces JuPyLive, a migration mechanism that enables seamless transition of Jupyter notebooks between local resources of user's workstation and remote resources of high-performance comput...
FlowTT: Exploiting Computation Flow Reuse in Irregular Tensor-Train EmbeddingJongmin Seok, Chae Eun Rhee2026-09-03下载Tensor-Train (TT) decomposition effectively compresses large embedding tables in recommendation models, but TT-based embedding lookup remains inefficient because partially shared computation flows acr...
Efficient Constant Optimization for Symbolic Regression with GPU-Accelerated Tree-Based Genetic ProgrammingHao Mao, Xu Tony Liu, Shuai Lu, Peng Zhao, Wenzheng Jiang, Yuntian Chen2026-09-03下载Constant optimization refines the numerical coefficients of candidate expressions in tree-based genetic programming for symbolic regression. But its per-generation cost has led modern GPU-accelerated ...
Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUsJinghao Wang, Yifeng Zhang, Xiao Zhou, Yao Lu, Yihui Zhang, Xiaoyang Sun, Tianyu Wo, Xu Wang, Chunming Hu, Renyu Yang2026-09-03下载Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability.
Iapetus: Content-Aware Hierarchical Scheduling for Collaborative ViT Inference in LEO Satellite NetworksYan Chen, Yunxiang Zhang, Guanjun Jiang, Haiquan Wang2026-09-03下载Collaborative inference pools distributed resources to run compute-intensive Vision Transformers (ViTs) in satellite edge computing. Model partitioning enables such collaboration by assigning consecut...
Lantern: Finding Committable Transactions via Back-Propagation on DAGsDenglong Li, Gerui Wang, Tian Guan, Mingchao Wan2026-09-03下载Existing concurrency control protocols either introduce nondeterminism, resulting in a serial execution-replay dependency between primary and replica nodes, or rely on impractical prior knowledge of t...
A Technique for Load Shifting Low-latency Applications in Multi-Region Renewables Harvesting via SMT Core PoolingTharindu B. Hewage, Shashikant Ilager, Maria A. Rodriguez, Rajkumar Buyya2026-09-03下载Load shifting across geographic regions to chase intermittent renewable energy availability is commonly used in reducing cloud infrastructure carbon footprint.
Carbon-aware Resource Management for Latency-Sensitive Cloud Computing Environments: A Taxonomy and Future DirectionsTharindu B. Hewage, Shashikant Ilager, Maria Rodriguez Read, Rajkumar Buyya2026-09-03下载Proliferation of cloud-based latency-sensitive workloads requires infrastructures tuned to their workload-specific latency constraints. Today, they shape the cloud from a generalized computing platfor...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended RealityYoga Suhas Kuruba Manjunath, Jie Gao, Lian Zhao2026-09-03下载We present ResLearn-XR, a residual learning framework for predicting eXtended Reality (XR) network traffic and estimating Quality-of-Experience (QoE) risk.
Uncertainty Signals for Network Intent Translation: Risk Ranking and Ambiguity LocalizationAla' A. Alsamarneh, Omar Alhussein2026-09-03下载Intent-based networking realization starts by translating high-level intents into low-level network configurations. Recent approaches have shifted toward LLM-based translation.
Waves on the Walls: Empirical Characterization of mmWave Lateral Waves for Enhanced Indoor CoverageApala Pramanik, Avhishek Biswas, Sasitharan Balasubramaniam, Christos Argyropoulos, Mehmet C. Vuran2026-09-03下载High-frequency millimeter-wave (mmWave) communication systems are constrained by the surrounding environment, where walls are traditionally treated as obstacles that block or reflect signals indoors.
Tuning Collective Patterns to Alleviate Congestion in Shared AI ClustersEashan Gupta, Yongzhou Chen, Apoorve Mohan, Pavlos Maniotis, Abdullah Kayi, Radhika Mittal2026-09-03下载Distributed AI training involves recurring rounds of data exchange between multiple pairs of GPU nodes. Slowdown in even one flow due to congestion can cause the entire communication round to slowdown...
Network Availability Enhancement in Low-Altitude HetNets: A Cross-Layer Design PerspectiveTeng Wu, Jiandong Li, Junyu Liu, Min Sheng, Mohammadali Mohammadi, Hien Quoc Ngo, Michail Matthaiou2026-09-03下载This paper proposes a computing-communication resource interchange method to enhance network availability (NA) in low-altitude heterogeneous networks (LA-HetNets).
Is Collision-Free Backoff Worth It in Wi-Fi?Mohammad Yousefi, Francesc Wilhelmi, Boris Bellalta2026-09-03下载The Distributed Coordination Function (DCF)---the underlying channel access protocol in Wi-Fi, based on Carrier Sense Multiple Access with Collision Avoidance (CSMA/CA) and Binary Exponential Backoff ...
Employing the Structural Power to Achieve Supply-Demand Balanced Payment Channel NetworksShuyao Xiao, Shengling Wang, Hongwei Shi, Weicheng Wang, Anlin Chen2026-09-03下载Blockchain technology faces scalability challenges because transactions must be validated and recorded across the network. Payment channel networks (PCNs) improve efficiency by moving transactions off...
From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network ControlVincenzo Norman Vitale, Mohammad Solki, Antonia Maria Tulino, Andreas F. Molisch, Jaime Llorca2026-09-03下载Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is essential for NextG interactive applications, yet providing strict End-to-End (E2E) peak latency guarantees remai...
A Semantic-Aware Multiple Access Scheme Leveraging Spatial Redundancy for Uplink-Dominant Network ServicesHamidreza Mazandarani, Masoud Shokrnezhad, Tarik Taleb2026-09-03下载The transition toward semantic-aware communication offers a paradigm shift for next-generation mobile networks, promising to decouple information significance from raw data transmission.
An Adversarial Zero-Shot Learning Approach for Anomaly Detection in Multivariate IoT Traffic DataMahshid Rezakhani, Tolunay Seyfi, Fatemeh Afghah2026-09-03下载Anomaly detection in Internet of Things (IoT) networks presents unique challenges due to the diversity of devices, lack of labeled data, and domain variability across environments.
Indirect Estimation of SINR via SSB and CSI-RS RSRP in 5G NRLeonardo Spampinato, Mahamadou Togola, Matteo Bernabè, Azim Akhtarshenas, Lorenzo Mario Amorosa, David López-Pérez2026-09-03下载Predicting user equipment (UE) performance is essential for proactive network control, resource management, and digital twin sandboxes. However, the inherent flexibility and complexity of beam-based 5...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
NACRE: Rethinking Confidential Containers through Native Architectural SupportLinke Song, Wenhao Wang, Weijie Liu, Rui Hou2026-09-03下载Linux containers achieve high density and fast lifecycle operations by sharing the host kernel, but this design also lets a compromised host inspect or modify container state.

cs.PF - Performance ​

标题作者发布日期PDF摘要
MaxKernel: Agentic Kernel Generation for TPUsShangkun Wang, Nina Cai, Charles Hoong, Julian Walker, Gerson Kroiz, George Vanica, Deepak Patil, Andi Gavrilescu, Hassan Sipra, Sethu Sankaran2026-09-03下载Designing and authoring high-performance custom kernels for accelerators is a complex task that requires deep hardware-level expertise. Large Language Models (LLM) can be leveraged together with real-...
PerfReasoning: How Well Do LLMs Reason on Hardware Performance?Dan Zhao, Karthikeyan Sankaralingam, Christos Kozyrakis, Qijing Huang2026-09-03下载Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement.
On-board ML for Trace Gas detection in Imaging Spectroscopy dataVít Růžička, Adam Chlus, Andrew Thorpe, David R. Thompson2026-09-03下载Data collected during aerial and spaceborne imaging spectroscopy campaigns enables the detection of transient events such as trace gas emissions.
Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCsYujie Zhang, Huiying Lan, Ehsan Aghapour, Zhiyuan Ning, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra2026-09-03下载As edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges.
Confidence-Gated Admission for Hardware Prefetching: When the Gate Matters More Than the PredictorYoussef Majdane, Simone Jarno Casartelli, Enrico Lopedoto2026-09-03下载Learned cache prefetchers are typically evaluated against classical predictors that always issue requests, confounding the prediction model with the admission policy.
RASER: Resilient Agent Scheduling and Execution Runtime for HPC ClustersSima Attar-Khorasani, Matthias Lieber, Siavash Ghiasvand2026-09-03下载The emergence of modern agents powered by large language models has created a demand for executing long-horizon, autonomous workflows in various domains that require significant computational resource...
Lantern: Finding Committable Transactions via Back-Propagation on DAGsDenglong Li, Gerui Wang, Tian Guan, Mingchao Wan2026-09-03下载Existing concurrency control protocols either introduce nondeterminism, resulting in a serial execution-replay dependency between primary and replica nodes, or rely on impractical prior knowledge of t...

基于 VitePress 构建 · 使用本地搜索查找论文