2026-05-28
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode | Josef Chen | 2026-05-28 | 下载 | Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM serving: single-stream, batch-1 autoregressive de... |
| Energy-Efficient Aggregation and Minimum-Degree Spanning Trees in Radio Networks | Yi-Jun Chang, Yang Ze Guan | 2026-05-28 | 下载 | We study the aggregation problem in synchronous multi-hop radio networks with -bit messages and no collision detection. Each node initially holds a value, and the goal is to compute a globa... |
| Scheduling Mechanisms in Wireless Sensor-Actuator Networks for Multi-rate Periodic Control in Industry 4.0 | Dingwen Yuan, Luis F. Abanto-Leon, Matthias Hollick | 2026-05-28 | 下载 | This paper investigates scheduling strategies for wireless sensor-actuator networks (WSANs) in Industry 4.0 scenarios. In particular, we address the problem of real-time scheduling for multi-rate cont... |
| A Virtual Processor brings back the Free Lunch | Haymo Kutschbach | 2026-05-28 | 下载 | This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. |
| RAFI -- A Ray/Work Forwarding Infrastructure for Data Parallel Multi-Node/Multi-GPU Computing | Ingo Wald, Serkan Demirci, Alper Sahistan, Stefan Zellmann, Andrea Paris, Patrick Moran, Milan Jaros, Tatiana von Landesberger, Ugur Gudukbay, Valerio Pascucci | 2026-05-28 | 下载 | We present RaFI, a CUDA and MPI based software framework that simplifies the task of building GPU-enabled data-parallel software where rays or similar work items need to migrate between different GPUs... |
| Q-ANCHOR: Federated Quantum Learning with ZNE-guided Correction | Hoang M. Ngo, Quan Nguyen, Wanli Xing, My T. Thai | 2026-05-28 | 下载 | Quantum Federated Learning (QFL) offers a promising framework to train quantum models across distributed clients while keeping data strictly local. |
| Effective MPI: User-defined Datatypes and Cartesian Communicators for Zero-copy All-to-all Communication in Multidimensional Tori | Jesper Larsson Träff | 2026-05-28 | 下载 | We present and show how to implement a non-trivial all-to-all communication algorithm for arbitrary -dimensional tori effectively in MPI. Given a factorization of the number of processes into $... |
| Ciphera: A Decentralised Biometric Identity Framework | Ankit Kanaiyalal Prajapati, Shahzad Memon, Mohammed Mahir Rahman, Ameer Al-Nemrat | 2026-05-28 | 下载 | Centralised biometric identity systems expose users to single points of failure, opaque verification processes, and irreversible biometric compromise. |
| From Roofline to Ruggedness: Decomposing and Smoothing the GEMM Performance Landscape | Aditya Chatterjee | 2026-05-28 | 下载 | Adjacent GEMM problems that differ by a single 128-element step in N can show 30% different throughput on the same GPU. This pervasive performance ruggedness - invisible to roofline analysis and peak-... |
| CARM Tool: Cache-Aware Roofline Model Automatic Benchmarking and Application Analysis | José Morgado, Leonel Sousa, Aleksandar Ilic | 2026-05-28 | 下载 | In recent years, HPC systems and CPU architectures as their central components, have become increasingly complex, making application development and optimization quite challenging. |
| PRISM: Processing-In-Memory Sparse MTTKRP for Tensor Decomposition Acceleration | Daniel Pacheco, Leonel Sousa, Aleksandar Ilic | 2026-05-28 | 下载 | Sparse tensors are the most used representation of sparse multidimensional data. Operations that decompose them, selecting their most important features while reducing their dimension, have become pre... |
| AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training | Ling Chen, Houming Wu, Wenjie Yu | 2026-05-28 | 下载 | Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mismatch between forward and backward passes. |
| TC-MIS: Maximal Independent Set on Tensor-cores | Prajjwal Nijhara, Dip Sankar Banerjee | 2026-05-28 | 下载 | Maximal Independent Set (MIS) in a graph is a fundamental problem with applications in resource allocation, scheduling, and network optimization. |
| Design and Implementation of a Serverless MapReduce Framework for Scalable Data Pipelines | Angelos Dorotheos Chatzopoulos, Babis Andreou, Kakia Panagidi, Stathes Hadjiefthymiades | 2026-05-28 | 下载 | Modern logistics systems tend to generate continuous streams of data from sources such as GPS, IoT sensors, and logistics management systems. The aggregation, processing, and analysis of data have bec... |
| Silent Data Corruption Protection through Efficient Task Replication | Mia Reitz, Claudia Fohry | 2026-05-28 | 下载 | The trend of increasing cluster sizes of supercomputers leads to a growing susceptibility to Silent Data Corruption (SDC) that can invalidate program results. |
| Understanding and Reducing Metadata-Driven Host Overheads in Sampling-Based GNN Training | Yidong Gong, Saima Afrin, Yuchen Ma, Guannan Wang, Bin Ren, Pradeep Kumar | 2026-05-28 | 下载 | Modern deep learning workloads increasingly exhibit dynamic, metadata-driven execution, where runtime-generated information determines memory provisioning and kernel launch decisions. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Temporally Encoded Double DQN for Proactive PRB Allocation in O-RAN Enabled Industrial Networks | Elahe Delavari, Xingqi Wu, Junaid Farooq | 2026-05-28 | 下载 | Fifth-generation (5G) wireless systems are increasingly adopted in smart manufacturing to support heterogeneous industrial workloads through services such as enhanced Mobile Broadband (eMBB) and Ultra... |
| Jamming-Resilient PRB Reservation for Latency-Critical O-RAN Network Slicing | Elahe Delavari, Junaid Farooq | 2026-05-28 | 下载 | Open radio access network (O-RAN) architectures enable near real-time, software-driven control of network slicing through programmable xApps deployed on the near-real-time RAN Intelligent Controller (... |
| From Waves to Graphs: A Ray-Tracing-Inspired Neural Radio Propagation Model | Paul Almasan, Stefanos Bakirtzis, José Suárez-Varela, Andra Lutu | 2026-05-28 | 下载 | Artificial intelligence-driven radio propagation models provide agile and robust solutions for mobile network operators in their effort to ensure the optimal performance of the wireless ecosystem and ... |
| Scheduling Mechanisms in Wireless Sensor-Actuator Networks for Multi-rate Periodic Control in Industry 4.0 | Dingwen Yuan, Luis F. Abanto-Leon, Matthias Hollick | 2026-05-28 | 下载 | This paper investigates scheduling strategies for wireless sensor-actuator networks (WSANs) in Industry 4.0 scenarios. In particular, we address the problem of real-time scheduling for multi-rate cont... |
| Intent-Based Orchestration in Open RAN: An ns-3 Simulation Framework | Pouya Agheli, Grégoire Lefebvre | 2026-05-28 | 下载 | This paper presents an extensible ns-3-based simulation framework for evaluating intent-based, semantics-aware control in Open RAN architectures. |
| TraceCodec: A Compiler-Backed Neural Codec for Stateful Multi-Flow Network Traffic Traces | Junhui Ding, Xinchen Zhang, Xiaohui Xie, Shinan Liu | 2026-05-28 | 下载 | Critical networking workflows require high-fidelity packet captures (PCAPs) for testing, security analysis, and protocol validation, not just statistical flow-level summaries. |
| ARIADNE: AI-RAN Informed Link Adaptation in Digital Twin Network Environments | Maria Tsampazi, Neagin Neasamoni Santhi, Nicole Perrotta, Falko Dressler, Tommaso Melodia | 2026-05-28 | 下载 | Artificial Intelligence (AI)-powered Radio Access Network (RAN) networks have attracted significant attention from both industry and academia. |
| Network Optimization Aspects of Autonomous Vehicles: Challenges and Future Directions | Rudolf Krecht, Tamas Budai, Erno Horvath, Akos Kovacs, Nobert Marko, Miklos Unger | 2026-05-28 | 下载 | Global megatrends, such as urbanization, population growth, and emerging network solutions are accelerating the development of the Connected and Autonomous Vehicles (CAVs) industry. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| RTP-LLM: High-Performance Alibaba LLM Inference Engine | Boyu Tan, Jiarui Guo, Zongwei Lv, Hanbo Sun, Tong Yang, Kan Liu, Xinfei Shi, Zetao Hu, Yaxin Yu, Chi Zhang, Jianning Zhang, Xi Yang, Wei Zhang, Bo Cai, Silu Zhou, Xiyu Wang, Na He, Yinghao Yu, Wending Bao, Guiyang Huang, Yuxing Yuan, Juncheng Yin, Nan Wang, Lin Yang, Zechao Zhang, Lu Chen, Guoding Li, Tao Lan, Lin Qu | 2026-05-28 | 下载 | Large Language Models (LLMs) have revolutionized AI applications, but deploying them at scale presents significant challenges. We present RTP-LLM, a high-performance inference engine for industrial-sc... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Caspar: CUDA Accelerator for Symbolic Programming with Adaptive Reordering | Emil Martens, Aaron Miller, Matias Varnum, Annette Stahl | 2026-05-28 | 下载 | We present Caspar, a library that makes the power of modern GPUs more accessible in robotics and provides a state-of-the-art nonlinear GPU solver that can be applied to a wide range of different optim... |
| Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode | Josef Chen | 2026-05-28 | 下载 | Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM serving: single-stream, batch-1 autoregressive de... |
| A Virtual Processor brings back the Free Lunch | Haymo Kutschbach | 2026-05-28 | 下载 | This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. |
| MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference | Kexin Chu, Yang Zhou, Wei Zhang | 2026-05-28 | 下载 | Temperature-zero BF16 LLM inference is often treated as reproducible, yet the same request can emit different tokens when decoded alone or inside a larger batch. |
| Demystifying VEINS: A Reality Check Against Living Lab Experiments | Antonio Solida, Giovanni Gambigliani Zoccoli, Gaetano Orazio Cauchi, Filip Valgimigli, Salvatore Iandolo, Martin Klapez, Maurizio Casoni, Mirco Marchetti, Carlo Augusto Grazia | 2026-05-28 | 下载 | Safety applications in vehicle-to-everything communications and Cooperative Intelligent Transport Systems rely on reliable and timely message exchange, which in turn depends on accurate modeling of wi... |
| From Roofline to Ruggedness: Decomposing and Smoothing the GEMM Performance Landscape | Aditya Chatterjee | 2026-05-28 | 下载 | Adjacent GEMM problems that differ by a single 128-element step in N can show 30% different throughput on the same GPU. This pervasive performance ruggedness - invisible to roofline analysis and peak-... |
| Experimentation for Different Scheduling Policies on Queues: Mixed Differences-in-Q Estimators Based on Little's Law | Nanshan Jia, Ramesh Johari, Nian Si, Zeyu Zheng | 2026-05-28 | 下载 | In data centers, tasks are dispatched to various servers to evenly distribute the workload. When a data center considers implementing a new scheduling algorithm, it typically conducts an A/B test prio... |
| TC-MIS: Maximal Independent Set on Tensor-cores | Prajjwal Nijhara, Dip Sankar Banerjee | 2026-05-28 | 下载 | Maximal Independent Set (MIS) in a graph is a fundamental problem with applications in resource allocation, scheduling, and network optimization. |