Skip to content

2026-04-21 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Algorithm and Hardware Co-Design for Efficient Complex-Valued Uncertainty EstimationZehuan Zhang, Mark Chen, He Li, Wayne Luk2026-04-21下载Complex-Valued Neural Networks (CVNNs) have significant advantages in handling tasks that involve complex numbers. However, existing CVNNs are unable to quantify predictive uncertainty.
Efficient Page Migration in Hybrid Memory SystemsUpasna, Venkata Kalyan Tavva2026-04-21下载Heterogeneous Memory Architecture (HMA) aims to optimize memory usage by leveraging a combination of memory types, such as high-bandwidth memory (HBM), commodity DRAM, and non-volatile memory (NVM), w...
Co-Designing Error Mitigation and Error Detection for Logical QubitsRohan S. Kumar, Takahiro Tsunoda, Sophia H. Xue, Dantong Li, Robert J. Schoelkopf, Yongshan Ding2026-04-21下载Near-term quantum workloads demand error management, yet the two lightest-weight techniques, Quantum Error Detection (QED) and Probabilistic Error Cancellation (PEC), have complementary cost profiles ...
ChipCraftBrain: Validation-First RTL Generation via Multi-Agent OrchestrationCagri Eryilmaz2026-04-21下载Large Language Models (LLMs) show promise for generating Register-Transfer Level (RTL) code from natural language specifications, but single-shot generation achieves only 60-65% functional correctness...
Toward designing workload-aware Surface Code ArchitecturesArchisman Ghosh, Avimita Chatterjee, Swaroop Ghosh2026-04-21下载Practical quantum advantage is expected to depend on fault-tolerant quantum computing, although the architectural overhead needed to support fault tolerance is still extremely high.
Energy Efficient LSTM Accelerators for Embedded FPGAs through Parameterised Architecture DesignChao Qian, Tianheng Ling, Gregor Schiele2026-04-21下载Long Short-term Memory Networks (LSTMs) are a vital Deep Learning technique suitable for performing on-device time series analysis on local sensor data streams of embedded devices.
Design Rules for Extreme-Edge Scientific Computing on AI EnginesZhenghua Ma, G Abarajithan, Dimitrios Danopoulos, Olivia Weng, Francesco Restuccia, Ryan Kastner2026-04-21下载Extreme-edge scientific applications use machine learning models to analyze sensor data and make real-time decisions. Their stringent latency and throughput requirements demand small batch sizes and r...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Federated Learning over Blockchain-Enabled Cloud InfrastructureSaloni Garg, Amit Sagtani, Kamal Kant Hiran2026-04-21下载The rise of IoT devices and the uptake of cloud computing have informed a new era of data-driven intelligence. Traditional centralized machine learning models that require a large volume of data to be...
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward SlicingYuning Xia, John Mellor-Crummey2026-04-21下载More than half of the Top 500 supercomputers employ GPUs as accelerators. On GPU-accelerated platforms, developers face a key diagnostic gap: profilers show source lines where stalls occur, but not wh...
Equinox: Decentralized Scheduling for Hardware-Aware Orbital IntelligenceAnsel Kaplan Erol, Divya Mahajan2026-04-21下载Earth-observation satellites are emerging as distributed edge platforms for time-critical tasks, yet orbital scheduling remains challenged by intermittent energy harvesting and temporal coupling where...
Predictive Autoscaling for Node.js on Kubernetes: Lower Latency, Right-Sized CapacityIvan Tymoshenko, Luca Maraschi, Matteo Collina2026-04-21下载Kubernetes offers two default paths for scaling Nodejs workloads, and both have structural limitations. The Horizontal Pod Autoscaler scales on CPU utilization, which does not directly measure event l...
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed TrainingShuyao Qi, Haoyuan Liu, Shizhen Zhao2026-04-21下载Fine-grained, per-micro-batch load balancing is essential for efficient Mixture-of-Experts (MoE) training, yet every prior dynamic scheduling scheme pays for it with extra communication that is hard t...
ReaLB: Real-Time Load Balancing for Multimodal MoE InferenceYingping Wang, Yi Wu, Xiangyu Wu, Junwei Cui, Weilin Cai, Zhijiang Guo, Jiayi Huang2026-04-21下载Mixture-of-Experts (MoE) architectures are widely used in modern large language models and multimodal models. However, inference efficiency is often limited by highly dynamic and skewed expert workloa...
DPC: A Distributed Page Cache over CXLShai Bergman, Zhe Yang, Julien Eudine, Giorgio Negro, Onur Mutlu, Arash Tavakkol, Ji Zhang2026-04-21下载Modern distributed file systems rely on uncoordinated, per node page caches that replicate hot data locally across the cluster. While ensuring fast local access, this architecture underutilizes aggreg...
Minimizing Intellectual Property Risks via Self-Stabilizing AlgorithmsKen Kennedy, Iman Evazzade2026-04-21下载In this paper, we examine the use of self-stabilizing algorithms, operating in a hierarchical manner, to determine intellectual property risks at a macro level.
Optimal Routing for Federated Learning over Dynamic Satellite Networks: Tractable or Not?Yi Zhao, Di Yuan, Tao Deng, Suzhi Cao, Ying Dong2026-04-21下载Federated learning (FL) is a key paradigm for distributed model learning across decentralized data sources. Communication in each FL round typically consists of two phases: (i) distributing the global...
CROWDio: A Practical Mobile Crowd Computing Framework with Developer-Oriented Design, Adaptive Scheduling, and Fault ResilienceLakshani Manamperi, Disumi Pathirana, Thiwanka Pathirana, Nipun Premarathna, Kutila Gunasekara2026-04-21下载Mobile Crowd Computing (MCdC) leverages the idle computational capacity of consumer smartphones to enable distributed task processing at scale; however, widespread real-world adoption remains constrai...
POLAR-PIC: A Holistic Framework for Matrixized PIC with Co-Designed Compute, Layout, and CommunicationYizhuo Rao, Xingjian Cui, Shangzhi Pang, Jiabin Xie, Guangnan Feng, Jinhui Wei, Ziyan Zhang, Languang Gao, Zhenyu Wang, Zhiguang Chen, Yutong Lu2026-04-21下载Particle-in-Cell (PIC) simulations are fundamental to plasma physics but often suffer from limited scalability due to particle-grid interaction bottlenecks and particle redistribution costs.
Mass Matrix Assembly on Tensor Cores for Implicit Particle-In-Cell MethodsLuca Pennati, Stefano Markidis2026-04-21下载Matrix-multiply-accumulate (MMA) units, or tensor cores, are now widespread across modern computing architectures. Yet, their use for particle-grid operators remains limited.
A Simple Communication Scheme for Distributed Fast Multipole MethodsSrinath Kailasa2026-04-21下载We present a simple hierarchical communication scheme for distributed Fast Multipole Methods (FMMs) based on MPI neighborhood collectives and uniform trees.
UniEP: Unified Expert-Parallel MoE MegaKernel for LLM TrainingSize Zheng, Xuegui Zheng, Li-wen Chang, Jidong Zhai2026-04-21下载The exponential growth in Large Language Model (LLM) parameters has transformed model training into an increasingly resource-intensive endeavor.
Sherpa.ai Privacy-Preserving Multi-Party Entity Alignment without Intersection Disclosure for Noisy IdentifiersDaniel M. Jimenez-Gutierrez, Enrique Zuazua, Georgios Kellaris, Joaquin Del Rio, Oleksii Sliusarenko, Xabi Uribe-Etxebarria2026-04-21下载Federated Learning (FL) enables collaborative model training among multiple parties without centralizing raw data. There are two main paradigms in FL: Horizontal FL (HFL), where all participants share...
YAIFS: Yet (not) Another Intelligent Fog Simulator: A Framework for Agent-Driven Computing Continuum Modeling & SimulationIsaac Lera, Carlos Guerrero2026-04-21下载Simulation plays a key role in the design and evaluation of distributed systems, yet it is often treated as a static tool with limited interaction capabilities.
Heuristic Search Space Partitioning for Low-Latency Multi-Tenant Cloud QueriesPrashant Kumar Pathak, Chandra Biksheswaran Mouleeswaran, Rama Teja Repaka2026-04-21下载Large-scale cloud security platforms must continuously query millions of structured cloud resource records distributed across thousands of tenant accounts.
CHRONOS: A Hardware-Assisted Phase-Decoupled Framework for Secure Federated Learning in IoTHung Dang2026-04-21下载We propose CHRONOS, a hardware-assisted framework that decouples the cryptographic setup required for private gradient aggregation from the active training phase.
Ocean: Fast Estimation-Based Sparse General Matrix-Matrix Multiplication on GPUYifan Li, Giulia Guidi2026-04-21下载In computational science and data analytics, many workloads involve irregular and sparse computations that are inherently difficult to optimize for modern hardware.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Differentiated Services: an Experimental vs. Simulated Case StudySergio Andreozzi2026-04-21下载This paper aims to provide a proof of concept of the accuracy of simulations for advanced networking study. The particular target technology is the Differentiated Services (DiffServ) architecture.
On the Optimality of Network Topology Discovery in Single-Hop Bounded-Interference NetworksTolunay Seyfi, Erfan Khadem, Fatemeh Afghah2026-04-21下载We propose \emph{PRISM} (\textbf{Pseudorandom Residue-based Indexed Scheduling Method}), a deterministic topology-discovery framework for single-hop wireless networks with bounded interference.
Greedy Routing in a Sequentially Grown One-Dimensional Random GraphAlexander Ponomarenko2026-04-21下载We analyze greedy routing in a random graph G_n constructed on the vertex set V = {1, 2, ..., n} embedded in Z. Vertices are inserted according to a uniform random permutation pi, and each newly inser...
ZODIAC: Zero-shot Offline Diffusion for Inferring Multi-xApps Conflicts in Open Radio Access NetworksZeyu Fang, Shu Hong, Huu Trung Thieu, Nakjung Choi, Tian Lan2026-04-21下载Open Radio Access Network (O-RAN) enables network control through multi-vendor xApps operating both within and across layers, subnets, and domains, whose concurrent execution can trigger conflicts tha...
Active Inference-Enabled Agentic Closed-Loop ISAC with Long-Horizon PlanningGuangjin Pan, Zhuojun Tian, Mehdi Bennis, Henk Wymeersch2026-04-21下载Wireless agentic systems enable agents to autonomously perceive, reason, and act. However, existing works neglect the tight coupling between sensing and control in closed-loop integrated sensing and c...
Revisiting and Expanding the IPv6 Network Periphery: Global-Scale Measurement and Security AnalysisZixuan Xie, Zitao Yang, Shurui Fang, Zhaoyang Li, Wenxing Xie, Nannan Fu, Liangyu Dong, Xiang Li2026-04-21下载As IPv6 deployment accelerates, understanding the evolving security posture of network peripheries becomes increasingly important. A DSN 2021 study introduced the first large-scale discovery of IPv6 n...
Direction-Dependent Path Loss Modeling in Olive Orchards for Precision AgricultureMohammad Rowhani Sistani, Katarzyna Kosek-Szott, Pierluigi Gallo2026-04-21下载Wireless links deployed in orchards often exhibit significant variability in the strength of the received signal that is not adequately captured by classical distance-based propagation models.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Equinox: Decentralized Scheduling for Hardware-Aware Orbital IntelligenceAnsel Kaplan Erol, Divya Mahajan2026-04-21下载Earth-observation satellites are emerging as distributed edge platforms for time-critical tasks, yet orbital scheduling remains challenged by intermittent energy harvesting and temporal coupling where...
An AI Agent Execution Environment to Safeguard User DataRobert Stanley, Avi Verma, Lillian Tsai, Konstantinos Kallas, Sam Kumar2026-04-21下载AI agents promise to serve as general-purpose personal assistants for their users, which requires them to have access to private user data (e.g., personal and financial information).
DPC: A Distributed Page Cache over CXLShai Bergman, Zhe Yang, Julien Eudine, Giorgio Negro, Onur Mutlu, Arash Tavakkol, Ji Zhang2026-04-21下载Modern distributed file systems rely on uncoordinated, per node page caches that replicate hot data locally across the cluster. While ensuring fast local access, this architecture underutilizes aggreg...
Scheduling Analysis of UAV Flight Control Workloads using Raspberry Pi 5 Using PREEMPT_RT LinuxLuiz Giacomossi, Håkan Forsberg, Ivan Tomasic, Baran Çürüklü, Tommaso Cucinotta2026-04-21下载Modern UAV architectures increasingly aim to unify high-level autonomy and low-level flight control on a single General-Purpose Operating System (GPOS).

cs.PF - Performance ​

标题作者发布日期PDF摘要
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward SlicingYuning Xia, John Mellor-Crummey2026-04-21下载More than half of the Top 500 supercomputers employ GPUs as accelerators. On GPU-accelerated platforms, developers face a key diagnostic gap: profilers show source lines where stalls occur, but not wh...

基于 VitePress 构建 · 使用本地搜索查找论文