Skip to content

2026-05-27 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU MemoryMyeong Jun Jo2026-05-27下载Large language models have achieved remarkable capabilities through scaling, and this paper does not challenge that. It instead investigates a different question: once large models already exist, can ...
OpenURMA: A Clean-Room Open Implementation of the Unified Bus ProtocolBojie Li2026-05-27下载Modern datacenter RDMA is bottlenecked at the network interface, not the wire. A NIC running RoCE or InfiniBand holds per-connection state for every (application, remote-endpoint) pair - hundreds of m...
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple SiliconMohamed Amine Bergach2026-05-27下载Half precision (FP16) promises to double FFT throughput on GPUs, but the prevailing view is that its 10-bit mantissa makes it unsuitable for radar-grade signal processing.
Nonvolatile Charge-Domain Attention with HZO Ferroelectric Capacitors: A Simulation-Based Device-to-System EvaluationFaris Abouagour2026-05-27下载Transformer decoding is constrained by both attention compute and KV-cache movement. This paper presents the Ferroelectric Charge-Domain Compute Cell (FCDC), a hafnium-zirconium-oxide (HZO) memcapacit...
FT-Pilot: Automated Fault-Tolerant RTL Rewriting via Vulnerability-Guided LLMsWeixing Liu, Zizhen Liu, Jing Ye, Naixing Wang, Cheng Liu, Huawei Li, Xiaowei Li2026-05-27下载As integrated circuit technologies continue to scale toward advanced process nodes, the continual reduction in node capacitance and supply voltage has made digital systems increasingly vulnerable to s...
HammerSim: A System-Level Tool to Model RowHammerKaustav Goswami, Ayaz Akram, Hari Venugopalan, Jason Lowe-Power2026-05-27下载Modern architecture research relies on simulators to evaluate system security, yet analyzing emerging hardware vulnerabilities like RowHammer requires full-system visibility.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
CA-AC-MPC: CUDA-Accelerated Actor-Critic Model Predictive ControlAntoonio Buo, Vittorio Cammarota, Michele Avagnale, Pierluigi Arpenti, Vincenzo Lippiello, Fabio Ruggiero2026-05-27下载In the literature, actor-critic model predictive control (AC-MPC) integrates MPC with reinforcement learning to enable high-performance control of complex dynamical systems.
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU MemoryMyeong Jun Jo2026-05-27下载Large language models have achieved remarkable capabilities through scaling, and this paper does not challenge that. It instead investigates a different question: once large models already exist, can ...
IORM: Hierarchical I/O Governance for Thousands of Consolidated Databases on Oracle ExadataRajarshi Chowdhury, Akshay Shah, Zakaria Alrmaih, Chenhao Guo, Anubhav Singh, Sue Lee2026-05-27下载Oracle Exadata consolidates thousands of tenant databases onto shared storage infrastructure deployed at hundreds of customer sites worldwide.
FedQHD: Closed-Form Function-Space Federated Reinforcement LearningYuchen Hou, Yongshan Chen, Zhuowen Zou, Calvin Yeung, Mohsen Imani, Tian Lan, Mahdi Imani2026-05-27下载Federated reinforcement learning enables decentralized agents to collaboratively improve policies or value estimates without exchanging raw trajectories.
SwarmHarness: Skill-Based Task Routing via Decentralized Incentive-Aligned AI Agent NetworksEdwin Jose2026-05-27下载Vast quantities of compute (GPU cycles on personal workstations, idle inference servers, and edge devices between jobs) go unused because no incentive-aligned protocol exists for their owners to share...
Fault Tolerance of Accelerated Asynchronous Fixed-Point Iterations on Flexible Computing InfrastructureEvan Coleman, Masha Sosonkina2026-05-27下载Asynchronous iterative methods tolerate straggling processors by allowing workers to proceed with stale data, but at a cost: the iterates become inconsistent, potentially degrading convergence.
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUsMiguel Graça, Aleksandar Ilic2026-05-27下载State-of-the-art multiple sequence alignment (MSA) algorithms are based on progressive approaches that rely on pairwise sequence alignment (PSA) to generate guide trees to align all sequences.
High-Quality Multi-Constraint Hypergraph Partitioning via Greedy RebalancingNikolai Maas2026-05-27下载Multi-constraint hypergraph partitioning is a generalization of balanced partitioning, where the vertex set of a hypergraph is partitioned such that the inter-block connectivity of hyperedges is minim...
How Far Can Disaggregation Go? A Design-Space Exploration of Attention-FFN Disaggregation for Efficient MoE LLM ServingHanjiang Wu, Abhimanyu Rajeshkumar Bambhaniya, Sarbartha Banerjee, Tuhin Khare, Sudarshan Srinivasan, Suvinay Subramanian, Souvik Kundu, Madhu Kumar, Midhilesh Elavazhagan, William Won, Amir Yazdanbakhsh, Tushar Krishna2026-05-27下载Modern large language model (LLM) inference has progressively disaggregated to keep pace with growing model sizes and tight TTFT and TPOT service-level objectives: from chunked-prefill aggregation, to...
Resource Allocation in HyperX NetworksAlejandro Cano, Cristóbal Camarero, Carmen Martínez, Ramón Beivide2026-05-27下载As high-performance computing systems scale in size and complexity, efficient resource management is essential to minimize communication overhead.
SiDP: Memory-Efficient Data Parallelism for Offline LLM InferenceAlan Zhao, Cyril Y. He2026-05-27下载The rapid adoption of large language models (LLMs) has shifted a substantial portion of inference workloads into throughput-oriented offline regimes, where fully utilizing GPU compute requires large b...
Throughput-Optimized Networks at ScaleConor James Green, Mithuna Thottethodi2026-05-27下载Datacenter network design plays a critical role in AI training by supporting scaling to thousands of accelerators. An open problem, designing a near-optimal throughput oriented network-topology, routi...
Addressing Variable Heterogeneity in Distributed Multimodal Training with EntrainInsu Jang, Mosharaf Chowdhury2026-05-27下载Multimodal LLM datasets are inherently heterogeneous, with significant data variability. Although each modality exhibits independent variability, sample-level entanglement makes it difficult to balanc...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Dynamic Entanglement Packet Scheduling for Quantum NetworksQuang-Phong Tran, Claudio Cicconetti, Marco Conti, Andrea Passarella2026-05-27下载Sharing entanglement among multiple users remains a central challenge for scalable quantum networks. Recent work proposed an on-demand entanglement packet architecture in which a controller uses a Tim...
OpenURMA: A Clean-Room Open Implementation of the Unified Bus ProtocolBojie Li2026-05-27下载Modern datacenter RDMA is bottlenecked at the network interface, not the wire. A NIC running RoCE or InfiniBand holds per-connection state for every (application, remote-endpoint) pair - hundreds of m...
Efficient and Quantum-safe Internet Key Exchange Protocols for Satellite CommunicationsDavide De Zuane, Marco Baldi, Paolo Santini, Grégoire Anchelergues, Daniele Romano, Alessandro Cammarano, Juan José Grosso2026-05-27下载This paper studies cryptographic key exchange in satellite communications, which requires specific solutions because the satellite context presents unique challenges, particularly concerning onboard r...
A Goal-Oriented Networking Approach for Intelligent IoT Service DeploymentFederico Tonini, Davide Borsatti, Wint Yi Poe, Riccardo Trivisonno, Walter Cerroni2026-05-27下载The first 6G standardization efforts are about to start, shaping the new generation of mobile networks. The IMT-2030 extends the IMT-2020 by expanding its usage scenarios to Immersive, Massive, and Hy...
Automated Heuristic Design for Network OperationsReza Namvar, José Gallego, Jose A. Ayala-Romero, Livia Elena Chatzieleftheriou, Andres Garcia-Saavedra, Albert Banchs, Marco Fiore2026-05-27下载Network operation relies on heuristics to solve many tasks rapidly and efficiently across the protocol stack. These heuristics are the result of thorough human-driven design rooted in expert knowledge...
Kernel-Level Per-Slice UPF Latency Measurement in Containerised 5G Core NetworksAkhil Dev Mishra, Mayank Pandey2026-05-27下载The 5G Core User Plane Function is responsible for packet forwarding, GTP-U decapsulation, and quality of service enforcement for every user data session.
Temporal Hyperbolic Graph Representation Learning for Scale-Free Internet Routing and Delay PredictionYi-Ling Kuo, Hao-Yu Tien, Shih-Yu Tsai2026-05-27下载Predicting Internet round-trip time (RTT) is critical for routing optimization, quality-of-service (QoS) provisioning, and traffic engineering, yet remains challenging due to long-term temporal depend...
Throughput-Optimized Networks at ScaleConor James Green, Mithuna Thottethodi2026-05-27下载Datacenter network design plays a critical role in AI training by supporting scaling to thousands of accelerators. An open problem, designing a near-optimal throughput oriented network-topology, routi...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Rotary GPU: Exploring Local Execution Paths for Large Mixture-of-Experts Models Under Limited GPU MemoryMyeong Jun Jo2026-05-27下载Large language models have achieved remarkable capabilities through scaling, and this paper does not challenge that. It instead investigates a different question: once large models already exist, can ...
Range, Not Precision: Block-Floating-Point Half-Precision FFT and SAR Imaging on Apple SiliconMohamed Amine Bergach2026-05-27下载Half precision (FP16) promises to double FFT throughput on GPUs, but the prevailing view is that its 10-bit mantissa makes it unsuitable for radar-grade signal processing.

基于 VitePress 构建 · 使用本地搜索查找论文