2026-04-16
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| EasyRider: Mitigating Power Transients in Datacenter-Scale Training Workloads | Dillon Jensen, Obi Nnorom, Grant Wilkins, Hugo Budd, Ram Rajagopal, Juan Rivas-Davila, Phil Levis | 2026-04-16 | 下载 | Large-scale AI model training workloads use thousands of GPUs operating in tightly synchronized loops. During synchronous communication, start-up, shut-down, and checkpointing, GPU power consumption c... |
| Democratization of Real-time Multi-Spectral Photoacoustic Imaging: Open-Sourced System Architecture for OPOTEK Phocus & Verasonics Vantage Combination | Ryo Murakami, Yichuan Tang, Haichong K. Zhang | 2026-04-16 | 下载 | Real-time multi-spectral photoacoustic imaging (RT-mPAI) often suffers from synchronization instabilities when interfacing fast-tuning lasers with data acquisition platforms executing on non-real-time... |
| SCENIC: Stream Computation-Enhanced SmartNIC | Benjamin Ramhorst, Maximilian Jakob Heer, Luhao Liu, Heejae Kim, Jonas Dann, Jin-Soo Kim, Gustavo Alonso | 2026-04-16 | 下载 | Although modern, AI-centric datacenters heavily rely on SmartNICs, existing devices impose a hard trade-off. Commercial SmartNICs provide high bandwidth and easy software integration, but offer limite... |
| Autonomous Evolution of EDA Tools: Multi-Agent Self-Evolved ABC | Cunxi Yu, Haoxing Ren | 2026-04-16 | 下载 | This paper introduces the first \emph{self-evolving} logic synthesis framework, which leverages Large Language Model (LLM) agents to autonomously improve the source code of \textsc{ABC}, the widely ad... |
| Dr.~RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-Improvement | Wenji Fang, Yao Lu, Shang Liu, Jing Wang, Ziyan Guo, Junxian He, Fengbin Tu, Zhiyao Xie | 2026-04-16 | 下载 | Recent advances in large language models (LLMs) have sparked growing interest in automatic RTL optimization for better performance, power, and area (PPA). |
| Accelerating CRONet on AMD Versal AIE-ML Engines | Kaustubh Mhatre, Vedant Tewari, Aditya Ray, Farhan Khan, Ridwan Olabiyi, Ashif Iquebal, Aman Arora | 2026-04-16 | 下载 | Topology optimization is a computational method used to determine the optimal material distribution within a prescribed design domain, aiming to minimize structural weight while satisfying load and bo... |
| Scaling Photonic Tensor Cores with Unary and Homodyne Designs | Oluwaseun Alo, Ishan Thakkar | 2026-04-16 | 下载 | We analyze five photonic microring tensor core designs with a common optical power model. The results show that circuit ordering, unary encoding, and homodyne accumulation shape scalability, with the ... |
| Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench Automation | Mu-Chi Chen, Po-Hsuan Huang, Yu-Hung Kao, Yen-Fu Liu, Yu-Kai Hung, Cheng Liang, Shao-Chun Ho, Chia-Heng Tu, Shih-Hao Hung | 2026-04-16 | 下载 | Recent advances in large language models have improved code generation, but their use in hardware description languages is still limited. Moreover, training data and testbenches for these models are o... |
| ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving | Yuseon Choi, Jingu Lee, Jungjun Oh, Sunjoo Whang, Byeongcheol Kim, Minsung Kim, Hoi-Jun Yoo, Sangjin Kim | 2026-04-16 | 下载 | Mixture-of-Experts (MoE) models have become the dominant architecture for large-scale language models, yet on-premises serving remains fundamentally memory-bound as batching turns sparse per-token com... |
| DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance | Kathiravan Palaniappan | 2026-04-16 | 下载 | Modern datacenters increasingly rely on low-power, single-slot inference accelerators to balance performance, energy efficiency, and rack density constraints. |
| VeriGraphi: A Multi-Agent Framework of Hierarchical RTL Generation for Large Hardware Designs | Sazzadul Islam, Tasnim Tabassum, Hao Zheng | 2026-04-16 | 下载 | Generating synthesizable Verilog for large, hierarchical hardware designs remains a significant challenge for large language models (LLMs), which struggle to replicate the structured reasoning that hu... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Optimizing Stochastic Gradient Push under Broadcast Communications | Tuan Nguyen, Ting He | 2026-04-16 | 下载 | We consider the problem of minimizing the convergence time for decentralized federated learning (DFL) in wireless networks under broadcast communications, with focus on mixing matrix design. |
| Wave-Based Dispatch for Circuit Cutting in Hybrid HPC--Quantum Systems | Ricard S. García-Raigada, Josep Jorba, Sergio Iserte | 2026-04-16 | 下载 | Hybrid High-performance Computing (HPC)-quantum workloads based on circuit cutting decompose large quantum circuits into independent fragments, but existing frameworks tightly couple cutting logic to ... |
| Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines | Marcel Wagenländer, Otto White, Britannio Jarrett, Pedro Silvestre, Yanda Tao, Guo Li, Huanzhou Zhu, Llúis Vilanova, Peter Pietzuch | 2026-04-16 | 下载 | Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving such workflows at a target throughput with low latency is challenging because they c... |
| SCENIC: Stream Computation-Enhanced SmartNIC | Benjamin Ramhorst, Maximilian Jakob Heer, Luhao Liu, Heejae Kim, Jonas Dann, Jin-Soo Kim, Gustavo Alonso | 2026-04-16 | 下载 | Although modern, AI-centric datacenters heavily rely on SmartNICs, existing devices impose a hard trade-off. Commercial SmartNICs provide high bandwidth and easy software integration, but offer limite... |
| Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter | Ruoyu Qin, Weiran He, Yaoyu Wang, Zheming Li, Xinran Xu, Yongwei Wu, Weimin Zheng, Mingxing Zhang | 2026-04-16 | 下载 | Prefill-decode (PD) disaggregation has become the standard architecture for large-scale LLM serving, but in practice its deployment boundary is still determined by KVCache transfer. |
| Efficient calculation of available space for multi-NUMA virtual machines | Andrei Gudkov, Elizaveta Ponomareva, Alexis Pospelov | 2026-04-16 | 下载 | Increasing demand for computational power has led cloud providers to employ multi-NUMA servers and offer multi-NUMA virtual machines to their customers. |
| Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving | Tingyang Sun, Ting He, I-Hong Hou | 2026-04-16 | 下载 | As a current trend in Artificial Intelligence (AI), large foundation models are increasingly employed as the core of AI services. However, even after training, serving such models at scale remains a c... |
| Cooperate to Compete: Strategic Data Generation and Incentivization Framework for Coopetitive Cross-Silo Federated Learning | Thanh Linh Nguyen, Nguyen Van Huynh, Quoc-Viet Pham | 2026-04-16 | 下载 | In data-sensitive domains such as healthcare, cross-silo federated learning (CFL) allows organizations to collaboratively train AI models without sharing raw data. |
| Exploiting Correlations in Federated Learning: Opportunities and Practical Limitations | Adrian Edin, Michel Kieffer, Mikael Johansson, Zheng Chen | 2026-04-16 | 下载 | The communication bottleneck in federated learning (FL) has spurred extensive research into techniques to reduce the volume of data exchanged between client devices and the central parameter server. |
| ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises Serving | Yuseon Choi, Jingu Lee, Jungjun Oh, Sunjoo Whang, Byeongcheol Kim, Minsung Kim, Hoi-Jun Yoo, Sangjin Kim | 2026-04-16 | 下载 | Mixture-of-Experts (MoE) models have become the dominant architecture for large-scale language models, yet on-premises serving remains fundamentally memory-bound as batching turns sparse per-token com... |
| AgileLog: A Forkable Shared Log for Agents on Data Streams | Shreesha G. Bhat, Tony Hong, Michael Noguera, Ramnatthan Alagappan, Aishwarya Ganesan | 2026-04-16 | 下载 | In modern data-streaming systems, alongside traditional programs, a new type of entity has emerged that can interact with streaming data: AI agents. |
| CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism | Bin Ma, Xingjian Ding, Tekin Bicer, Pengfei Su, Dong Li | 2026-04-16 | 下载 | Diffusion Transformers (DiTs) are increasingly adopted in scientific computing, yet growing model sizes and resolutions make distributed multi-GPU inference essential. |
| Fast Concurrent Primitives Despite Contention | Michael A. Bender, Guy E. Blelloch, Martin Farach-Colton, Yang Hu, Rob Johnson, Rotem Oshman, Renfei Zhou | 2026-04-16 | 下载 | We study the problem of constructing concurrent objects in a setting where processes run in parallel and interact through a shared memory that is subject to write contention. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Inter-Satellite Link Optimization for Low-Latency Global Networking | Arman Mollakhani, Jerayu Tiamraj, Shu-Jie Cao, Dongning Guo | 2026-04-16 | 下载 | Large-scale low-Earth-orbit satellite constellations offer a promising platform for global low-latency networking, aided by faster propagation in free space than in fiber and copper. |
| A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area network | Mehdi Hosseinzadeh, Roohallah Alizadehsani, Amin Beheshti, Hamid Alinejad-Roknyd, Lu Chen, Mohammad Sadegh Yousefpoor, Efat Yousefpoor, Muneera Altayeb, Thantrira Porntaveetus, Sadia Din | 2026-04-16 | 下载 | The Internet of Medical Things (IoMT) enables intelligent healthcare services but faces challenges such as dynamic topology, energy constraints, and diverse QoS requirements. |
| Expanding into Reality: Random Graphs for Datacenter Networks | Giacomo Bernardi, Ratul Mahajan, C. Seshadhri, Enrico Carlesso, Chinchu Merine Joseph, Saurabh Kumar, Pavan Manikonda, Luiza Popa, Randy Ram, Steven Robinson, Elizabeth Tennent | 2026-04-16 | 下载 | We design and deploy at Amazon the first production datacenter fabrics based on random graphs. While the cost and fault-tolerance benefits of such topologies have been long known, their practical real... |
| SCENIC: Stream Computation-Enhanced SmartNIC | Benjamin Ramhorst, Maximilian Jakob Heer, Luhao Liu, Heejae Kim, Jonas Dann, Jin-Soo Kim, Gustavo Alonso | 2026-04-16 | 下载 | Although modern, AI-centric datacenters heavily rely on SmartNICs, existing devices impose a hard trade-off. Commercial SmartNICs provide high bandwidth and easy software integration, but offer limite... |
| MLDAS: Machine Learning Dynamic Algorithm Selection for Software-Defined Networking Security | Pablo Benlloch, Oscar Romero, Antonio Leon, Jaime Lloret | 2026-04-16 | 下载 | Network security is a critical concern in the digital landscape of today, with users demanding secure browsing experiences and protection of their personal data. |
| Learning Ad Hoc Network Dynamics via Graph-Structured World Models | Can Karacelebi, Yusuf Talha Sahin, Elif Surer, Ertan Onur | 2026-04-16 | 下载 | Ad hoc wireless networks exhibit complex, innate and coupled dynamics: node mobility, energy depletion and topology change that are difficult to model analytically. |
| Towards Trustworthy 6G Network Digital Twins: A Framework for Validating Counterfactual What-If Analysis in Edge Computing Resources | Julian Jimenez Agudelo, Paola Soto, Ayat Zaki-Hindi, Jean-Sébastien Sottet, Sébastien Faye, Nina Slamnik-Kriještorac, Johann Marquez-Barja, Miguel Camelo Botero | 2026-04-16 | 下载 | Network Digital Twins (NDTs) enable safe what-if analysis for 6G cloud-edge infrastructures, but adoption is often limited by fragmented workflows from telemetry to validation. |
| Switching Efficiency: A Novel Framework for Dissecting AI Data Center Network Efficiency | Niangen Ye, Jiawen Zhu, Baojun Chen, Dong Wang, Jiang Sun, Weiqiang Sun, Weisheng Hu | 2026-04-16 | 下载 | Communication is pivotal in LLM training, and a thorough analysis of the communication efficiency of AI data center (AIDC) network is essential for guiding the design of these capital-intensive cluste... |
| PlanB: Efficient Software IPv6 Lookup with Linearized -Tree | Zhihao Zhang, Lanzheng Liu, Chen Chen, Huiba Li, Jiwu Shu, Windsor Hsu, Yiming Zhang | 2026-04-16 | 下载 | IP lookup via Longest Prefix Match (LPM) is critical for packet forwarding. Unfortunately, conventional lookup algorithms are inefficient for IPv6 Forwarding Information Bases (FIBs), which are charac... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPU | Jevin Jiang, Ying Chen, Blake A. Hechtman, Fenghui Zhang, Yarong Mu | 2026-04-16 | 下载 | Large Language Model (LLM) deployment is increasingly shifting to cost-efficient accelerators like Google's Tensor Processing Units (TPUs), prioritizing both performance and total cost of ownership (T... |
| Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model Serving | Tingyang Sun, Ting He, I-Hong Hou | 2026-04-16 | 下载 | As a current trend in Artificial Intelligence (AI), large foundation models are increasingly employed as the core of AI services. However, even after training, serving such models at scale remains a c... |
| DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural Performance | Kathiravan Palaniappan | 2026-04-16 | 下载 | Modern datacenters increasingly rely on low-power, single-slot inference accelerators to balance performance, energy efficiency, and rack density constraints. |