2026-04-21
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty Estimation | Zehuan Zhang, Mark Chen, He Li, Wayne Luk | 2026-04-21 | 下载 | Complex-Valued Neural Networks (CVNNs) have significant advantages in handling tasks that involve complex numbers. However, existing CVNNs are unable to quantify predictive uncertainty. |
| Efficient Page Migration in Hybrid Memory Systems | Upasna, Venkata Kalyan Tavva | 2026-04-21 | 下载 | Heterogeneous Memory Architecture (HMA) aims to optimize memory usage by leveraging a combination of memory types, such as high-bandwidth memory (HBM), commodity DRAM, and non-volatile memory (NVM), w... |
| Co-Designing Error Mitigation and Error Detection for Logical Qubits | Rohan S. Kumar, Takahiro Tsunoda, Sophia H. Xue, Dantong Li, Robert J. Schoelkopf, Yongshan Ding | 2026-04-21 | 下载 | Near-term quantum workloads demand error management, yet the two lightest-weight techniques, Quantum Error Detection (QED) and Probabilistic Error Cancellation (PEC), have complementary cost profiles ... |
| ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration | Cagri Eryilmaz | 2026-04-21 | 下载 | Large Language Models (LLMs) show promise for generating Register-Transfer Level (RTL) code from natural language specifications, but single-shot generation achieves only 60-65% functional correctness... |
| Toward designing workload-aware Surface Code Architectures | Archisman Ghosh, Avimita Chatterjee, Swaroop Ghosh | 2026-04-21 | 下载 | Practical quantum advantage is expected to depend on fault-tolerant quantum computing, although the architectural overhead needed to support fault tolerance is still extremely high. |
| Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture Design | Chao Qian, Tianheng Ling, Gregor Schiele | 2026-04-21 | 下载 | Long Short-term Memory Networks (LSTMs) are a vital Deep Learning technique suitable for performing on-device time series analysis on local sensor data streams of embedded devices. |
| Design Rules for Extreme-Edge Scientific Computing on AI Engines | Zhenghua Ma, G Abarajithan, Dimitrios Danopoulos, Olivia Weng, Francesco Restuccia, Ryan Kastner | 2026-04-21 | 下载 | Extreme-edge scientific applications use machine learning models to analyze sensor data and make real-time decisions. Their stringent latency and throughput requirements demand small batch sizes and r... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Federated Learning over Blockchain-Enabled Cloud Infrastructure | Saloni Garg, Amit Sagtani, Kamal Kant Hiran | 2026-04-21 | 下载 | The rise of IoT devices and the uptake of cloud computing have informed a new era of data-driven intelligence. Traditional centralized machine learning models that require a large volume of data to be... |
| LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing | Yuning Xia, John Mellor-Crummey | 2026-04-21 | 下载 | More than half of the Top 500 supercomputers employ GPUs as accelerators. On GPU-accelerated platforms, developers face a key diagnostic gap: profilers show source lines where stalls occur, but not wh... |
| Equinox: Decentralized Scheduling for Hardware-Aware Orbital Intelligence | Ansel Kaplan Erol, Divya Mahajan | 2026-04-21 | 下载 | Earth-observation satellites are emerging as distributed edge platforms for time-critical tasks, yet orbital scheduling remains challenged by intermittent energy harvesting and temporal coupling where... |
| Predictive Autoscaling for Node.js on Kubernetes: Lower Latency, Right-Sized Capacity | Ivan Tymoshenko, Luca Maraschi, Matteo Collina | 2026-04-21 | 下载 | Kubernetes offers two default paths for scaling Nodejs workloads, and both have structural limitations. The Horizontal Pod Autoscaler scales on CPU utilization, which does not directly measure event l... |
| FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training | Shuyao Qi, Haoyuan Liu, Shizhen Zhao | 2026-04-21 | 下载 | Fine-grained, per-micro-batch load balancing is essential for efficient Mixture-of-Experts (MoE) training, yet every prior dynamic scheduling scheme pays for it with extra communication that is hard t... |
| ReaLB: Real-Time Load Balancing for Multimodal MoE Inference | Yingping Wang, Yi Wu, Xiangyu Wu, Junwei Cui, Weilin Cai, Zhijiang Guo, Jiayi Huang | 2026-04-21 | 下载 | Mixture-of-Experts (MoE) architectures are widely used in modern large language models and multimodal models. However, inference efficiency is often limited by highly dynamic and skewed expert workloa... |
| DPC: A Distributed Page Cache over CXL | Shai Bergman, Zhe Yang, Julien Eudine, Giorgio Negro, Onur Mutlu, Arash Tavakkol, Ji Zhang | 2026-04-21 | 下载 | Modern distributed file systems rely on uncoordinated, per node page caches that replicate hot data locally across the cluster. While ensuring fast local access, this architecture underutilizes aggreg... |
| Minimizing Intellectual Property Risks via Self-Stabilizing Algorithms | Ken Kennedy, Iman Evazzade | 2026-04-21 | 下载 | In this paper, we examine the use of self-stabilizing algorithms, operating in a hierarchical manner, to determine intellectual property risks at a macro level. |
| Optimal Routing for Federated Learning over Dynamic Satellite Networks: Tractable or Not? | Yi Zhao, Di Yuan, Tao Deng, Suzhi Cao, Ying Dong | 2026-04-21 | 下载 | Federated learning (FL) is a key paradigm for distributed model learning across decentralized data sources. Communication in each FL round typically consists of two phases: (i) distributing the global... |
| CROWDio: A Practical Mobile Crowd Computing Framework with Developer-Oriented Design, Adaptive Scheduling, and Fault Resilience | Lakshani Manamperi, Disumi Pathirana, Thiwanka Pathirana, Nipun Premarathna, Kutila Gunasekara | 2026-04-21 | 下载 | Mobile Crowd Computing (MCdC) leverages the idle computational capacity of consumer smartphones to enable distributed task processing at scale; however, widespread real-world adoption remains constrai... |
| POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and Communication | Yizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Jinhui Wei, Ziyan Zhang, Languang Gao, Zhenyu Wang, Zhiguang Chen, Yutong Lu | 2026-04-21 | 下载 | Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle-grid interaction bottlenecks and particle redistribution costs. |
| Mass Matrix Assembly on Tensor Cores for Implicit Particle-In-Cell Methods | Luca Pennati, Stefano Markidis | 2026-04-21 | 下载 | Matrix-multiply-accumulate (MMA) units, or tensor cores, are now widespread across modern computing architectures. Yet, their use for particle-grid operators remains limited. |
| A Simple Communication Scheme for Distributed Fast Multipole Methods | Srinath Kailasa | 2026-04-21 | 下载 | We present a simple hierarchical communication scheme for distributed Fast Multipole Methods (FMMs) based on MPI neighborhood collectives and uniform trees. |
| UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training | Size Zheng, Xuegui Zheng, Li-wen Chang, Jidong Zhai | 2026-04-21 | 下载 | The exponential growth in Large Language Model (LLM) parameters has transformed model training into an increasingly resource-intensive endeavor. |
| Sherpa.ai Privacy-Preserving Multi-Party Entity Alignment without Intersection Disclosure for Noisy Identifiers | Daniel M. Jimenez-Gutierrez, Enrique Zuazua, Georgios Kellaris, Joaquin Del Rio, Oleksii Sliusarenko, Xabi Uribe-Etxebarria | 2026-04-21 | 下载 | Federated Learning (FL) enables collaborative model training among multiple parties without centralizing raw data. There are two main paradigms in FL: Horizontal FL (HFL), where all participants share... |
| YAIFS: Yet (not) Another Intelligent Fog Simulator: A Framework for Agent-Driven Computing Continuum Modeling & Simulation | Isaac Lera, Carlos Guerrero | 2026-04-21 | 下载 | Simulation plays a key role in the design and evaluation of distributed systems, yet it is often treated as a static tool with limited interaction capabilities. |
| Heuristic Search Space Partitioning for Low-Latency Multi-Tenant Cloud Queries | Prashant Kumar Pathak, Chandra Biksheswaran Mouleeswaran, Rama Teja Repaka | 2026-04-21 | 下载 | Large-scale cloud security platforms must continuously query millions of structured cloud resource records distributed across thousands of tenant accounts. |
| CHRONOS: A Hardware-Assisted Phase-Decoupled Framework for Secure Federated Learning in IoT | Hung Dang | 2026-04-21 | 下载 | We propose CHRONOS, a hardware-assisted framework that decouples the cryptographic setup required for private gradient aggregation from the active training phase. |
| Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPU | Yifan Li, Giulia Guidi | 2026-04-21 | 下载 | In computational science and data analytics, many workloads involve irregular and sparse computations that are inherently difficult to optimize for modern hardware. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Differentiated Services: an Experimental vs. Simulated Case Study | Sergio Andreozzi | 2026-04-21 | 下载 | This paper aims to provide a proof of concept of the accuracy of simulations for advanced networking study. The particular target technology is the Differentiated Services (DiffServ) architecture. |
| On the Optimality of Network Topology Discovery in Single-Hop Bounded-Interference Networks | Tolunay Seyfi, Erfan Khadem, Fatemeh Afghah | 2026-04-21 | 下载 | We propose \emph{PRISM} (\textbf{Pseudorandom Residue-based Indexed Scheduling Method}), a deterministic topology-discovery framework for single-hop wireless networks with bounded interference. |
| Greedy Routing in a Sequentially Grown One-Dimensional Random Graph | Alexander Ponomarenko | 2026-04-21 | 下载 | We analyze greedy routing in a random graph G_n constructed on the vertex set V = {1, 2, ..., n} embedded in Z. Vertices are inserted according to a uniform random permutation pi, and each newly inser... |
| ZODIAC: Zero-shot Offline Diffusion for Inferring Multi-xApps Conflicts in Open Radio Access Networks | Zeyu Fang, Shu Hong, Huu Trung Thieu, Nakjung Choi, Tian Lan | 2026-04-21 | 下载 | Open Radio Access Network (O-RAN) enables network control through multi-vendor xApps operating both within and across layers, subnets, and domains, whose concurrent execution can trigger conflicts tha... |
| Active Inference-Enabled Agentic Closed-Loop ISAC with Long-Horizon Planning | Guangjin Pan, Zhuojun Tian, Mehdi Bennis, Henk Wymeersch | 2026-04-21 | 下载 | Wireless agentic systems enable agents to autonomously perceive, reason, and act. However, existing works neglect the tight coupling between sensing and control in closed-loop integrated sensing and c... |
| Revisiting and Expanding the IPv6 Network Periphery: Global-Scale Measurement and Security Analysis | Zixuan Xie, Zitao Yang, Shurui Fang, Zhaoyang Li, Wenxing Xie, Nannan Fu, Liangyu Dong, Xiang Li | 2026-04-21 | 下载 | As IPv6 deployment accelerates, understanding the evolving security posture of network peripheries becomes increasingly important. A DSN 2021 study introduced the first large-scale discovery of IPv6 n... |
| Direction-Dependent Path Loss Modeling in Olive Orchards for Precision Agriculture | Mohammad Rowhani Sistani, Katarzyna Kosek-Szott, Pierluigi Gallo | 2026-04-21 | 下载 | Wireless links deployed in orchards often exhibit significant variability in the strength of the received signal that is not adequately captured by classical distance-based propagation models. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Equinox: Decentralized Scheduling for Hardware-Aware Orbital Intelligence | Ansel Kaplan Erol, Divya Mahajan | 2026-04-21 | 下载 | Earth-observation satellites are emerging as distributed edge platforms for time-critical tasks, yet orbital scheduling remains challenged by intermittent energy harvesting and temporal coupling where... |
| An AI Agent Execution Environment to Safeguard User Data | Robert Stanley, Avi Verma, Lillian Tsai, Konstantinos Kallas, Sam Kumar | 2026-04-21 | 下载 | AI agents promise to serve as general-purpose personal assistants for their users, which requires them to have access to private user data (e.g., personal and financial information). |
| DPC: A Distributed Page Cache over CXL | Shai Bergman, Zhe Yang, Julien Eudine, Giorgio Negro, Onur Mutlu, Arash Tavakkol, Ji Zhang | 2026-04-21 | 下载 | Modern distributed file systems rely on uncoordinated, per node page caches that replicate hot data locally across the cluster. While ensuring fast local access, this architecture underutilizes aggreg... |
| Scheduling Analysis of UAV Flight Control Workloads using Raspberry Pi 5 Using PREEMPT_RT Linux | Luiz Giacomossi, Håkan Forsberg, Ivan Tomasic, Baran Çürüklü, Tommaso Cucinotta | 2026-04-21 | 下载 | Modern UAV architectures increasingly aim to unify high-level autonomy and low-level flight control on a single General-Purpose Operating System (GPOS). |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing | Yuning Xia, John Mellor-Crummey | 2026-04-21 | 下载 | More than half of the Top 500 supercomputers employ GPUs as accelerators. On GPU-accelerated platforms, developers face a key diagnostic gap: profilers show source lines where stalls occur, but not wh... |