Skip to content

2026-05-05 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
The Anatomy of Silent Data Corruption: GPU Error Pattern Study and Modeling GuidanceChung-Hsuan Tung, Yanxiang Huang, Nirmal Saxena, Philip Shirvani, Saurabh Hukerikar, Twinkle Jain, Abhishek Tyagi, Sanjay Gongalore2026-05-05下载Silent data corruption (SDC) threatens the reliability of large-scale GPU clusters used for training large language models, yet its rarity and lack of explicit error signals make accurate high-level m...
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU ArchitecturesAaron Jarmusch, Sunita Chandrasekaran2026-05-05下载Rapidly evolving GPU architectures featuring complex memory hierarchies, matrix units, and varied precision formats continue to widen the gap between theoretical peaks and achievable performance.
täkōFormal: Enabling Robust Software for Programmable Memory Hierarchies (Extended Version)Pranav Srinivasan, Manos Kapritsos, Yatin A. Manerkar2026-05-05下载Accelerators provide large performance and energy-efficiency benefits, but can significantly change the hardware-software interface. The täkō programmable memory hierarchy accelerates data movement by...
LIPPEN: A Lightweight In-Place Pointer Encryption Architecture for Pointer IntegrityErfan Iravani, Lalit Prasad Peri, Mohannad Ismail, Charitha Tumkur Siddalingaradhya, Changwoo Min, Elif Bilge Kavun, Wenjie Xiong2026-05-05下载Memory-safety violations in C and C++ programs continue to enable sophisticated exploitation techniques such as control-flow hijacking and data-oriented attacks.
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite ComparisonRuihao Li, Andrew Jacob, Neeraja J. Yadwadkar, Lizy K. John2026-05-05下载Specialized accelerators dominate AI workloads, but CPUs remain critical for orchestrating these accelerators and running datacenter services.
Design and Implementation of BNN-Based Object Detection on FPGAXuyu Zhao, Yunpeng Wu, Mengyuan Zhu, Haoyu Huang, Xiaoyu Xu, Yanjing Li, Gaolong Zhang, Baochang Zhang2026-05-05下载This paper implements a Binary Neural Network (BNN) based YOLOv3-tiny-like object detector on a low-cost FPGA. The network takes 3203203 RGB images as input.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Coral: Cost-Efficient Multi-LLM Serving over Heterogeneous Cloud GPUsYixuan Mei, Zikun Li, Zixuan Chen, Shiqi Pan, Mengdi Wu, Xupeng Miao, Zhihao Jia, K. V. Rashmi2026-05-05下载The usage of large language models (LLMs) has grown increasingly fragmented, with no single model dominating. Meanwhile, cloud providers offer a wide range of mid-tier and older-generation GPUs that e...
GPU-Accelerated Simulations of Problems with Moving Boundaries and Fluid-Structure Interaction at Extreme ScalesSushrut Kumar, Joshua Romero, Jung-Hee Seo, Massimiliano Fatica, Rajat Mittal2026-05-05下载Computational fluid dynamics and fluid-structure interaction simulations involving moving and deforming bodies is extremely hard. In this work, we present a graphical processing unit (GPU) optimized i...
Resilient AI Supercomputer Networking using MRC and SRv6Joao Araujo, Alex Chow, Mark Handley, Ryder Lewis, Christoph Paasch, Jitendra Padhye, Michael Papamichael, Greg Steinbrecher, Amin Tootoonchian, Lihua Yuan, S. Anantharamu, Abhishek Dosi, Mohit Garg, Mahdieh Ghazi, Torsten Hoefler, Deepal Jayasinghe, Jithin Jose, Abdul Kabbani, Guohan Lu, Yang Wang, K. Doddapaneni, Murali Garimella, Vipin Jain, Yanfang Le, H. Nagulapalli, S. Narayanan, Rong Pan, Rathina Sabesan, Raghava Sivaramu, Rip Sohan, Eric Davis, Dragos Dumitrescu, Mohan Kalkunte, Bhaswar Mitra, Guglielmo Morandin, Adrian Popa, Costin Raiciu, Eric Spada, John Spillane, Niranjan Vaidya, Aviv Barnea, Idan Burstein, Elazar Cohen, Yamin Friedman, Noam Katz, Masoud Moshref, Yuval Shpigelman, Shahaf Shuler, Shy Shyman, Sayantan Sur2026-05-05下载Tail latency dominates the performance of synchronous pretraining jobs when running at very large scales. We describe a three-pronged approach: (1) a new RDMA-based transport protocol, MRC, sprays acr...
Orchestrating Serverless Applications in the Edge Cloud Space Continuum: What Breaks and What is Next?Hadi Tabatabaee Malazi, Reza Farahani, Nitinder Mohan, Schahram Dustdar2026-05-05下载Serverless computing has matured into an effective execution model for edge cloud environments, enabling function level decomposition, demand driven scaling, and workflow execution across stable, well...
ClusterLess: Deadline-Aware Serverless Workflow Orchestration on Federated Edge ClustersReza Farahani, Mario Colosi, Ilir Murturi, Stefan Nastic, Massimo Villari, Schahram Dustdar, Radu Prodan2026-05-05下载The recent convergence of edge computing, serverless execution, and Kubernetes (K8s) based container orchestration has enabled the processing of application workflows close to data sources.
Revocation-Ready CP-ABE Key Management for Blockchain-Based IoT Data SharingChun Yin Chiu2026-05-05下载Blockchain-based IoT data sharing systems increasingly adopt a hybrid architecture in which a permissioned ledger stores tamper-evident metadata while encrypted payloads are placed in content-addresse...
phys-MCP: A Control Plane for Heterogeneous Physical Neural NetworksStefan Fischer, Maliheh Hariri, Sebastian Otte2026-05-05下载Physical neural networks (PNNs) embed computation directly in material dynamics, including molecular, chemical, biological, photonic, memristive, and mechanical substrates.
Thinking fast and slow -- decision intelligence for power systemsApoorv Mathur2026-05-05下载Decision-making in power systems spans multiple timescales - from milliseconds to prevent surges, to seconds to balance frequency and protect grid assets, to minutes for real-time energy balancing, to...
ipc_shared_ptr: A Publish/Subscribe-Aware Smart Pointer for Cross-Process Object Lifetime ManagementTakahiro Ishikawa-Aso, Atsushi Yano, Koichi Imai, Takuya Azumi, Shinpei Kato2026-05-05下载True zero-copy Inter-Process Communication (IPC) in publish/subscribe (pub/sub) middleware such as Robot Operating System 2 (ROS 2) requires subscribers to reference message objects in publisher-owned...
Microbenchmark-Driven Analytical Performance Modeling Across Modern GPU ArchitecturesAaron Jarmusch, Sunita Chandrasekaran2026-05-05下载Rapidly evolving GPU architectures featuring complex memory hierarchies, matrix units, and varied precision formats continue to widen the gap between theoretical peaks and achievable performance.
Implementing True MPI Sessions and Evaluating MPI Initialization ScalabilityHui Zhou, Kenneth Raffenetti, Yanfei Guo, Michael Wilkins, Rajeev Thakur2026-05-05下载Sessions is one of the major features introduced in the MPI-4 standard. It offers an alternative to the traditional world communicator model by allowing applications to construct communicators from pr...
Surviving the Edge: Federated Learning under Networking and Resource ConstraintsMike Mwanje, Okemawo Obadofin, Theophilus Benson, Joao Barros2026-05-05下载Motivated by the growing proliferation of federated learning (FL) in edge environments, we present the first systematic characterization of transport-layer breaking points in FL systems operating unde...
A Workflow-Oriented Framework for Asynchronous Human-AI Collaboration in Hybrid and Compute-Intensive HPC EnvironmentsSergio Mendoza, Cedric Bhihe, Natalia Zamora, David Modesto, Jose Martin Bugallo Batalla, Jesus Gomez Canovas, Rafel Palomo Avellaneda, Miguel Perez Espinosa2026-05-05下载Human involvement is critical in training and deploying AI systems in high-stakes defence and security contexts. However, real-time interaction is impractical in HPC environments due to compute intens...
Lifting to tensors when compiling scientific computing workloads for AI EnginesNick Brown, Gabriel Rodriguez-Canal2026-05-05下载It has been demonstrated that specialised architectures, such as FPGAs and AMD's AI Engines (AIEs), have the potential to deliver energy and performance advantages for scientific computing.
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale DiagnosticsDragana Grbic2026-05-05下载As exascale systems reach unprecedented concurrency, traditional performance analysis tools struggle with the overhead of massive-scale telemetry.
On Solving Problems of Substantially Super-linear Complexity in No(1)N^{o(1)} Rounds in the MPC ModelAndrzej Lingas2026-05-05下载We study the possibility of designing No(1)N^{o(1)}-round protocols for problems of substantially super-linear polynomial-time (sequential) complexity in the model of Massively Parallel Computation, ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Resilient AI Supercomputer Networking using MRC and SRv6Joao Araujo, Alex Chow, Mark Handley, Ryder Lewis, Christoph Paasch, Jitendra Padhye, Michael Papamichael, Greg Steinbrecher, Amin Tootoonchian, Lihua Yuan, S. Anantharamu, Abhishek Dosi, Mohit Garg, Mahdieh Ghazi, Torsten Hoefler, Deepal Jayasinghe, Jithin Jose, Abdul Kabbani, Guohan Lu, Yang Wang, K. Doddapaneni, Murali Garimella, Vipin Jain, Yanfang Le, H. Nagulapalli, S. Narayanan, Rong Pan, Rathina Sabesan, Raghava Sivaramu, Rip Sohan, Eric Davis, Dragos Dumitrescu, Mohan Kalkunte, Bhaswar Mitra, Guglielmo Morandin, Adrian Popa, Costin Raiciu, Eric Spada, John Spillane, Niranjan Vaidya, Aviv Barnea, Idan Burstein, Elazar Cohen, Yamin Friedman, Noam Katz, Masoud Moshref, Yuval Shpigelman, Shahaf Shuler, Shy Shyman, Sayantan Sur2026-05-05下载Tail latency dominates the performance of synchronous pretraining jobs when running at very large scales. We describe a three-pronged approach: (1) a new RDMA-based transport protocol, MRC, sprays acr...
Binary Image-Based Intrusion Detection for Operational Technology Networks: Extending the SPHBI Methodology from IoT to Modbus TCPAamir Omar2026-05-05下载This paper extends the Single Packet Header Binary Image (SPHBI) intrusion detection methodology from IoT to Modbus TCP, evaluating five approaches spanning a gradient of protocol depth on the CIC Mod...
Towards a Zero-Trust Supply-Chain Assurance Rubric for ORAN RIC ApplicationsChun Yin Chiu2026-05-05下载Open RAN enables third-party xApps and rApps to be onboarded and updated at operational cadence, creating a software supply chain that spans developers, CI systems, registries, onboarding pipelines, a...
Sequential vs. Simultaneous Entanglement Swapping under Optimal Link-Layer ControlPriyam Srivastava, Akshat R. Sabavat, Siddharth Jain, Alan Scheller-Wolf, Sridhar Tayur, David Tipper, Prashant Krishnamurthy, Amy Babay, Kaushik P. Seshadreesan2026-05-05下载Connection-less, packet-switched quantum network architectures distribute entanglement across multi-hop paths through sequential entanglement swapping, in which each node acts on purely local state in...
Surviving the Edge: Federated Learning under Networking and Resource ConstraintsMike Mwanje, Okemawo Obadofin, Theophilus Benson, Joao Barros2026-05-05下载Motivated by the growing proliferation of federated learning (FL) in edge environments, we present the first systematic characterization of transport-layer breaking points in FL systems operating unde...
Nested array design of extended coprime sets for DOA estimation of non-circular signalsDongqi Chen, Kun Ye, Chuanxi Xing, Waqas Khalid, Huiping Huang2026-05-05下载In recent years, direction of arrival estimation utilizing non-circular signals has become a focal point for scholarly research. To enhance the degrees of freedom (DOF) in receiver arrays specifically...
Say the Mission, Execute the Swarm: Agent-Enhanced LLM Reasoning in the Web-of-DronesAndrea Iannoli, Lorenzo Gigli, Luca Sciullo, Angelo Trotta, Marco Di Felice2026-05-05下载Large Language Models (LLMs) are increasingly explored as high-level reasoning engines for cyber-physical systems, yet their application to real-time UAV swarm management remains challenging due to he...
SprayCheck: Finding Gray Failures in Adaptive Routing NetworksJakob Krebs, Daniel Amir, Shir Landau Feibish, Mark Silberstein2026-05-05下载Distributed machine learning (ML) training has become a dominant workload in modern data center networks, operating at massive scale with clusters comprising tens to hundreds of thousands of GPUs.
Cross-Slice Co-Location Risk-Aware SFC Provisioning in Multi-Slice LEO Satellite NetworksMohammed Mahyoub, Wael Jaafar, Sami Muhaidat, Halim Yanikomeroglu2026-05-05下载We address cross-slice co-location risk in multi-slice low Earth orbit (LEO) satellite edge networks, where virtual network functions (VNFs) from different network slices sharing the same satellite in...
Dynamic Hypergame for Task Assignment in Multi-platform Mobile Crowdsensing Under Incomplete InformationSumedh J. Dongare, Christo Kurisummoottil Thomas, Andrea Ortiz, Walid Saad, Anja Klein2026-05-05下载Mobile crowdsensing (MCS) is a promising distributed sensing paradigm for future wireless networks, where MCS platforms (MCSPs) recruit mobile units (MUs) through monetary incentives for sensing data ...
Beyond Distributive Justice: Hermeneutical Fairness in Ad DeliveryCamilla Quaresmini, Valentina Breschi, Jessica Leoni, Viola Schiaffonati, Mara Tanelli, Giulia De Pasquale2026-05-05下载Fairness in online advertising is often formalized as a distributive justice problem, aiming to ensure that impressions, opportunities, or outcomes are allocated comparably across protected groups.
DACP: A Scientific Data Access and Collaboration ProtocolZhihong Shen, Xiaojie Zhu, Zhenjing Cheng, Hao Ren, Zhaoji Liang, Changfa Lu2026-05-05下载Scientific computing is rapidly entering a data-intensive era. However, existing general-purpose network protocol stacks face limitations in eliminating data silos and improving data accessibility and...
CRT: Collision-Tolerant Residence Time for Deterministic Transmission in LEO Satellite NetworksSiqi Yang, Zonghui Li, Chaoqun You, Yue Gao2026-05-05下载Low-Earth Orbit (LEO) satellite networks are a key enabler for the 6G Non-Terrestrial Network (NTN) architecture. However, supporting time-sensitive services in LEO networks is challenging due to high...
QoS Assurance Mechanism for 5G Network Slicing Based on the Deep Reinforcement Learning PPO AlgorithmQingyang Li2026-05-05下载With the increasing diversity of 5G service types and the intensifying dynamic fluctuations of network load, achieve differentiated quality of service assurance in a network slicing environment has be...
Single-Step Six-Dimensional Movable Antenna Reconfiguration for High-Mobility IoV: Modeling, Analysis, and OptimizationMaoxin Ji, Qiong Wu, Pingyi Fan, Kezhi Wang, Wen Chen, Cui Zhang, Khaled B. Letaief2026-05-05下载The Six-Dimensional Movable Antenna (6DMA) system has emerged as a promising technology to enhance wireless capacity by fully exploiting spatial degrees of freedom.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
ipc_shared_ptr: A Publish/Subscribe-Aware Smart Pointer for Cross-Process Object Lifetime ManagementTakahiro Ishikawa-Aso, Atsushi Yano, Koichi Imai, Takuya Azumi, Shinpei Kato2026-05-05下载True zero-copy Inter-Process Communication (IPC) in publish/subscribe (pub/sub) middleware such as Robot Operating System 2 (ROS 2) requires subscribers to reference message objects in publisher-owned...
Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM ServingShi Qiu, Yifan Hu, Xintao Wang, Wenhao Zhu, Jianqin Yan, Hao Chen, Kaiqiang Xu, Kai Chen, Yiming Zhang2026-05-05下载LLM serving relies on prefix caching to improve inference performance. As growing contexts push key-value (KV) cache footprint far beyond GPU HBM and CPU DRAM capacity, KV cache is increasingly offloa...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Decentralized Edge Caching under Budget and Storage Constraints: A Game-Theoretic ApproachHamta Sedghani, Zahra Seyedi, Mauro Passacantando, Danilo Ardagna2026-05-05下载The rapid growth of mobile social networks (MSNs) has significantly increased the demand for low-latency and reliable content delivery, motivating the deployment of edge caching systems.
SPEC CPU2026: Characterization, Representativeness, and Cross-Suite ComparisonRuihao Li, Andrew Jacob, Neeraja J. Yadwadkar, Lizy K. John2026-05-05下载Specialized accelerators dominate AI workloads, but CPUs remain critical for orchestrating these accelerators and running datacenter services.
Enhancing Performance Insight at Scale: A Heterogeneous Framework for Exascale DiagnosticsDragana Grbic2026-05-05下载As exascale systems reach unprecedented concurrency, traditional performance analysis tools struggle with the overhead of massive-scale telemetry.

基于 VitePress 构建 · 使用本地搜索查找论文