2026-05-27
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory | Myeong Jun Jo | 2026-05-27 | 下载 | Large language models have achieved remarkable capabilities through scaling, and this paper does not challenge that. It instead investigates a different question: once large models already exist, can ... |
| OpenURMA: A Clean-Room Open Implementation of the Unified Bus Protocol | Bojie Li | 2026-05-27 | 下载 | Modern datacenter RDMA is bottlenecked at the network interface, not the wire. A NIC running RoCE or InfiniBand holds per-connection state for every (application, remote-endpoint) pair - hundreds of m... |
| Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon | Mohamed Amine Bergach | 2026-05-27 | 下载 | Half precision (FP16) promises to double FFT throughput on GPUs, but the prevailing view is that its 10-bit mantissa makes it unsuitable for radar-grade signal processing. |
| Nonvolatile Charge-Domain Attention with HZO Ferroelectric Capacitors: A Simulation-Based Device-to-System Evaluation | Faris Abouagour | 2026-05-27 | 下载 | Transformer decoding is constrained by both attention compute and KV-cache movement. This paper presents the Ferroelectric Charge-Domain Compute Cell (FCDC), a hafnium-zirconium-oxide (HZO) memcapacit... |
| FT-Pilot: Automated Fault-Tolerant RTL Rewriting via Vulnerability-Guided LLMs | Weixing Liu, Zizhen Liu, Jing Ye, Naixing Wang, Cheng Liu, Huawei Li, Xiaowei Li | 2026-05-27 | 下载 | As integrated circuit technologies continue to scale toward advanced process nodes, the continual reduction in node capacitance and supply voltage has made digital systems increasingly vulnerable to s... |
| HammerSim: A System-Level Tool to Model RowHammer | Kaustav Goswami, Ayaz Akram, Hari Venugopalan, Jason Lowe-Power | 2026-05-27 | 下载 | Modern architecture research relies on simulators to evaluate system security, yet analyzing emerging hardware vulnerabilities like RowHammer requires full-system visibility. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive Control | Antoonio Buo, Vittorio Cammarota, Michele Avagnale, Pierluigi Arpenti, Vincenzo Lippiello, Fabio Ruggiero | 2026-05-27 | 下载 | In the literature, actor-critic model predictive control (AC-MPC) integrates MPC with reinforcement learning to enable high-performance control of complex dynamical systems. |
| Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory | Myeong Jun Jo | 2026-05-27 | 下载 | Large language models have achieved remarkable capabilities through scaling, and this paper does not challenge that. It instead investigates a different question: once large models already exist, can ... |
| IORM: Hierarchical I/O Governance for Thousands of Consolidated Databases on Oracle Exadata | Rajarshi Chowdhury, Akshay Shah, Zakaria Alrmaih, Chenhao Guo, Anubhav Singh, Sue Lee | 2026-05-27 | 下载 | Oracle Exadata consolidates thousands of tenant databases onto shared storage infrastructure deployed at hundreds of customer sites worldwide. |
| FedQHD: Closed-Form Function-Space Federated Reinforcement Learning | Yuchen Hou, Yongshan Chen, Zhuowen Zou, Calvin Yeung, Mohsen Imani, Tian Lan, Mahdi Imani | 2026-05-27 | 下载 | Federated reinforcement learning enables decentralized agents to collaboratively improve policies or value estimates without exchanging raw trajectories. |
| SwarmHarness: Skill-Based Task Routing via Decentralized Incentive-Aligned AI Agent Networks | Edwin Jose | 2026-05-27 | 下载 | Vast quantities of compute (GPU cycles on personal workstations, idle inference servers, and edge devices between jobs) go unused because no incentive-aligned protocol exists for their owners to share... |
| Fault Tolerance of Accelerated Asynchronous Fixed-Point Iterations on Flexible Computing Infrastructure | Evan Coleman, Masha Sosonkina | 2026-05-27 | 下载 | Asynchronous iterative methods tolerate straggling processors by allowing workers to proceed with stale data, but at a cost: the iterates become inconsistent, potentially degrading convergence. |
| TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs | Miguel Graça, Aleksandar Ilic | 2026-05-27 | 下载 | State-of-the-art multiple sequence alignment (MSA) algorithms are based on progressive approaches that rely on pairwise sequence alignment (PSA) to generate guide trees to align all sequences. |
| High-Quality Multi-Constraint Hypergraph Partitioning via Greedy Rebalancing | Nikolai Maas | 2026-05-27 | 下载 | Multi-constraint hypergraph partitioning is a generalization of balanced partitioning, where the vertex set of a hypergraph is partitioned such that the inter-block connectivity of hyperedges is minim... |
| How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM Serving | Hanjiang Wu, Abhimanyu Rajeshkumar Bambhaniya, Sarbartha Banerjee, Tuhin Khare, Sudarshan Srinivasan, Suvinay Subramanian, Souvik Kundu, Madhu Kumar, Midhilesh Elavazhagan, William Won, Amir Yazdanbakhsh, Tushar Krishna | 2026-05-27 | 下载 | Modern large language model (LLM) inference has progressively disaggregated to keep pace with growing model sizes and tight TTFT and TPOT service-level objectives: from chunked-prefill aggregation, to... |
| Resource Allocation in HyperX Networks | Alejandro Cano, Cristóbal Camarero, Carmen Martínez, Ramón Beivide | 2026-05-27 | 下载 | As high-performance computing systems scale in size and complexity, efficient resource management is essential to minimize communication overhead. |
| SiDP: Memory-Efficient Data Parallelism for Offline LLM Inference | Alan Zhao, Cyril Y. He | 2026-05-27 | 下载 | The rapid adoption of large language models (LLMs) has shifted a substantial portion of inference workloads into throughput-oriented offline regimes, where fully utilizing GPU compute requires large b... |
| Throughput-Optimized Networks at Scale | Conor James Green, Mithuna Thottethodi | 2026-05-27 | 下载 | Datacenter network design plays a critical role in AI training by supporting scaling to thousands of accelerators. An open problem, designing a near-optimal throughput oriented network-topology, routi... |
| Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain | Insu Jang, Mosharaf Chowdhury | 2026-05-27 | 下载 | Multimodal LLM datasets are inherently heterogeneous, with significant data variability. Although each modality exhibits independent variability, sample-level entanglement makes it difficult to balanc... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Dynamic Entanglement Packet Scheduling for Quantum Networks | Quang-Phong Tran, Claudio Cicconetti, Marco Conti, Andrea Passarella | 2026-05-27 | 下载 | Sharing entanglement among multiple users remains a central challenge for scalable quantum networks. Recent work proposed an on-demand entanglement packet architecture in which a controller uses a Tim... |
| OpenURMA: A Clean-Room Open Implementation of the Unified Bus Protocol | Bojie Li | 2026-05-27 | 下载 | Modern datacenter RDMA is bottlenecked at the network interface, not the wire. A NIC running RoCE or InfiniBand holds per-connection state for every (application, remote-endpoint) pair - hundreds of m... |
| Efficient and Quantum-safe Internet Key Exchange Protocols for Satellite Communications | Davide De Zuane, Marco Baldi, Paolo Santini, Grégoire Anchelergues, Daniele Romano, Alessandro Cammarano, Juan José Grosso | 2026-05-27 | 下载 | This paper studies cryptographic key exchange in satellite communications, which requires specific solutions because the satellite context presents unique challenges, particularly concerning onboard r... |
| A Goal-Oriented Networking Approach for Intelligent IoT Service Deployment | Federico Tonini, Davide Borsatti, Wint Yi Poe, Riccardo Trivisonno, Walter Cerroni | 2026-05-27 | 下载 | The first 6G standardization efforts are about to start, shaping the new generation of mobile networks. The IMT-2030 extends the IMT-2020 by expanding its usage scenarios to Immersive, Massive, and Hy... |
| Automated Heuristic Design for Network Operations | Reza Namvar, José Gallego, Jose A. Ayala-Romero, Livia Elena Chatzieleftheriou, Andres Garcia-Saavedra, Albert Banchs, Marco Fiore | 2026-05-27 | 下载 | Network operation relies on heuristics to solve many tasks rapidly and efficiently across the protocol stack. These heuristics are the result of thorough human-driven design rooted in expert knowledge... |
| Kernel-Level Per-Slice UPF Latency Measurement in Containerised 5G Core Networks | Akhil Dev Mishra, Mayank Pandey | 2026-05-27 | 下载 | The 5G Core User Plane Function is responsible for packet forwarding, GTP-U decapsulation, and quality of service enforcement for every user data session. |
| Temporal Hyperbolic Graph Representation Learning for Scale-Free Internet Routing and Delay Prediction | Yi-Ling Kuo, Hao-Yu Tien, Shih-Yu Tsai | 2026-05-27 | 下载 | Predicting Internet round-trip time (RTT) is critical for routing optimization, quality-of-service (QoS) provisioning, and traffic engineering, yet remains challenging due to long-term temporal depend... |
| Throughput-Optimized Networks at Scale | Conor James Green, Mithuna Thottethodi | 2026-05-27 | 下载 | Datacenter network design plays a critical role in AI training by supporting scaling to thousands of accelerators. An open problem, designing a near-optimal throughput oriented network-topology, routi... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU Memory | Myeong Jun Jo | 2026-05-27 | 下载 | Large language models have achieved remarkable capabilities through scaling, and this paper does not challenge that. It instead investigates a different question: once large models already exist, can ... |
| Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple Silicon | Mohamed Amine Bergach | 2026-05-27 | 下载 | Half precision (FP16) promises to double FFT throughput on GPUs, but the prevailing view is that its 10-bit mantissa makes it unsuitable for radar-grade signal processing. |