Skip to content

2026-04-16 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
EasyRider: Mitigating Power Transients in Datacenter-Scale Training WorkloadsDillon Jensen, Obi Nnorom, Grant Wilkins, Hugo Budd, Ram Rajagopal, Juan Rivas-Davila, Phil Levis2026-04-16下载Large-scale AI model training workloads use thousands of GPUs operating in tightly synchronized loops. During synchronous communication, start-up, shut-down, and checkpointing, GPU power consumption c...
Democratization of Real-time Multi-Spectral Photoacoustic Imaging: Open-Sourced System Architecture for OPOTEK Phocus & Verasonics Vantage CombinationRyo Murakami, Yichuan Tang, Haichong K. Zhang2026-04-16下载Real-time multi-spectral photoacoustic imaging (RT-mPAI) often suffers from synchronization instabilities when interfacing fast-tuning lasers with data acquisition platforms executing on non-real-time...
SCENIC: Stream Computation-Enhanced SmartNICBenjamin Ramhorst, Maximilian Jakob Heer, Luhao Liu, Heejae Kim, Jonas Dann, Jin-Soo Kim, Gustavo Alonso2026-04-16下载Although modern, AI-centric datacenters heavily rely on SmartNICs, existing devices impose a hard trade-off. Commercial SmartNICs provide high bandwidth and easy software integration, but offer limite...
Autonomous Evolution of EDA Tools: Multi-Agent Self-Evolved ABCCunxi Yu, Haoxing Ren2026-04-16下载This paper introduces the first \emph{self-evolving} logic synthesis framework, which leverages Large Language Model (LLM) agents to autonomously improve the source code of \textsc{ABC}, the widely ad...
Dr.~RTL: Autonomous Agentic RTL Optimization through Tool-Grounded Self-ImprovementWenji Fang, Yao Lu, Shang Liu, Jing Wang, Ziyan Guo, Junxian He, Fengbin Tu, Zhiyao Xie2026-04-16下载Recent advances in large language models (LLMs) have sparked growing interest in automatic RTL optimization for better performance, power, and area (PPA).
Accelerating CRONet on AMD Versal AIE-ML EnginesKaustubh Mhatre, Vedant Tewari, Aditya Ray, Farhan Khan, Ridwan Olabiyi, Ashif Iquebal, Aman Arora2026-04-16下载Topology optimization is a computational method used to determine the optimal material distribution within a prescribed design domain, aiming to minimize structural weight while satisfying load and bo...
Scaling Photonic Tensor Cores with Unary and Homodyne DesignsOluwaseun Alo, Ishan Thakkar2026-04-16下载We analyze five photonic microring tensor core designs with a common optical power model. The results show that circuit ordering, unary encoding, and homodyne accumulation shape scalability, with the ...
Exploring LLM-based Verilog Code Generation with Data-Efficient Fine-Tuning and Testbench AutomationMu-Chi Chen, Po-Hsuan Huang, Yu-Hung Kao, Yen-Fu Liu, Yu-Kai Hung, Cheng Liang, Shao-Chun Ho, Chia-Heng Tu, Shih-Hao Hung2026-04-16下载Recent advances in large language models have improved code generation, but their use in hardware description languages is still limited. Moreover, training data and testbenches for these models are o...
ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises ServingYuseon Choi, Jingu Lee, Jungjun Oh, Sunjoo Whang, Byeongcheol Kim, Minsung Kim, Hoi-Jun Yoo, Sangjin Kim2026-04-16下载Mixture-of-Experts (MoE) models have become the dominant architecture for large-scale language models, yet on-premises serving remains fundamentally memory-bound as batching turns sparse per-token com...
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural PerformanceKathiravan Palaniappan2026-04-16下载Modern datacenters increasingly rely on low-power, single-slot inference accelerators to balance performance, energy efficiency, and rack density constraints.
VeriGraphi: A Multi-Agent Framework of Hierarchical RTL Generation for Large Hardware DesignsSazzadul Islam, Tasnim Tabassum, Hao Zheng2026-04-16下载Generating synthesizable Verilog for large, hierarchical hardware designs remains a significant challenge for large language models (LLMs), which struggle to replicate the structured reasoning that hu...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Optimizing Stochastic Gradient Push under Broadcast CommunicationsTuan Nguyen, Ting He2026-04-16下载We consider the problem of minimizing the convergence time for decentralized federated learning (DFL) in wireless networks under broadcast communications, with focus on mixing matrix design.
Wave-Based Dispatch for Circuit Cutting in Hybrid HPC--Quantum SystemsRicard S. García-Raigada, Josep Jorba, Sergio Iserte2026-04-16下载Hybrid High-performance Computing (HPC)-quantum workloads based on circuit cutting decompose large quantum circuits into independent fragments, but existing frameworks tightly couple cutting logic to ...
Scepsy: Serving Agentic Workflows Using Aggregate LLM PipelinesMarcel Wagenländer, Otto White, Britannio Jarrett, Pedro Silvestre, Yanda Tao, Guo Li, Huanzhou Zhu, Llúis Vilanova, Peter Pietzuch2026-04-16下载Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving such workflows at a target throughput with low latency is challenging because they c...
SCENIC: Stream Computation-Enhanced SmartNICBenjamin Ramhorst, Maximilian Jakob Heer, Luhao Liu, Heejae Kim, Jonas Dann, Jin-Soo Kim, Gustavo Alonso2026-04-16下载Although modern, AI-centric datacenters heavily rely on SmartNICs, existing devices impose a hard trade-off. Commercial SmartNICs provide high bandwidth and easy software integration, but offer limite...
Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-DatacenterRuoyu Qin, Weiran He, Yaoyu Wang, Zheming Li, Xinran Xu, Yongwei Wu, Weimin Zheng, Mingxing Zhang2026-04-16下载Prefill-decode (PD) disaggregation has become the standard architecture for large-scale LLM serving, but in practice its deployment boundary is still determined by KVCache transfer.
Efficient calculation of available space for multi-NUMA virtual machinesAndrei Gudkov, Elizaveta Ponomareva, Alexis Pospelov2026-04-16下载Increasing demand for computational power has led cloud providers to employ multi-NUMA servers and offer multi-NUMA virtual machines to their customers.
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model ServingTingyang Sun, Ting He, I-Hong Hou2026-04-16下载As a current trend in Artificial Intelligence (AI), large foundation models are increasingly employed as the core of AI services. However, even after training, serving such models at scale remains a c...
Cooperate to Compete: Strategic Data Generation and Incentivization Framework for Coopetitive Cross-Silo Federated LearningThanh Linh Nguyen, Nguyen Van Huynh, Quoc-Viet Pham2026-04-16下载In data-sensitive domains such as healthcare, cross-silo federated learning (CFL) allows organizations to collaboratively train AI models without sharing raw data.
Exploiting Correlations in Federated Learning: Opportunities and Practical LimitationsAdrian Edin, Michel Kieffer, Mikael Johansson, Zheng Chen2026-04-16下载The communication bottleneck in federated learning (FL) has spurred extensive research into techniques to reduce the volume of data exchanged between client devices and the central parameter server.
ELMoE-3D: Leveraging Intrinsic Elasticity of MoE for Hybrid-Bonding-Enabled Self-Speculative Decoding in On-Premises ServingYuseon Choi, Jingu Lee, Jungjun Oh, Sunjoo Whang, Byeongcheol Kim, Minsung Kim, Hoi-Jun Yoo, Sangjin Kim2026-04-16下载Mixture-of-Experts (MoE) models have become the dominant architecture for large-scale language models, yet on-premises serving remains fundamentally memory-bound as batching turns sparse per-token com...
AgileLog: A Forkable Shared Log for Agents on Data StreamsShreesha G. Bhat, Tony Hong, Michael Noguera, Ramnatthan Alagappan, Aishwarya Ganesan2026-04-16下载In modern data-streaming systems, alongside traditional programs, a new type of entity has emerged that can interact with streaming data: AI agents.
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence ParallelismBin Ma, Xingjian Ding, Tekin Bicer, Pengfei Su, Dong Li2026-04-16下载Diffusion Transformers (DiTs) are increasingly adopted in scientific computing, yet growing model sizes and resolutions make distributed multi-GPU inference essential.
Fast Concurrent Primitives Despite ContentionMichael A. Bender, Guy E. Blelloch, Martin Farach-Colton, Yang Hu, Rob Johnson, Rotem Oshman, Renfei Zhou2026-04-16下载We study the problem of constructing concurrent objects in a setting where PP processes run in parallel and interact through a shared memory that is subject to write contention.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Inter-Satellite Link Optimization for Low-Latency Global NetworkingArman Mollakhani, Jerayu Tiamraj, Shu-Jie Cao, Dongning Guo2026-04-16下载Large-scale low-Earth-orbit satellite constellations offer a promising platform for global low-latency networking, aided by faster propagation in free space than in fiber and copper.
A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area networkMehdi Hosseinzadeh, Roohallah Alizadehsani, Amin Beheshti, Hamid Alinejad-Roknyd, Lu Chen, Mohammad Sadegh Yousefpoor, Efat Yousefpoor, Muneera Altayeb, Thantrira Porntaveetus, Sadia Din2026-04-16下载The Internet of Medical Things (IoMT) enables intelligent healthcare services but faces challenges such as dynamic topology, energy constraints, and diverse QoS requirements.
Expanding into Reality: Random Graphs for Datacenter NetworksGiacomo Bernardi, Ratul Mahajan, C. Seshadhri, Enrico Carlesso, Chinchu Merine Joseph, Saurabh Kumar, Pavan Manikonda, Luiza Popa, Randy Ram, Steven Robinson, Elizabeth Tennent2026-04-16下载We design and deploy at Amazon the first production datacenter fabrics based on random graphs. While the cost and fault-tolerance benefits of such topologies have been long known, their practical real...
SCENIC: Stream Computation-Enhanced SmartNICBenjamin Ramhorst, Maximilian Jakob Heer, Luhao Liu, Heejae Kim, Jonas Dann, Jin-Soo Kim, Gustavo Alonso2026-04-16下载Although modern, AI-centric datacenters heavily rely on SmartNICs, existing devices impose a hard trade-off. Commercial SmartNICs provide high bandwidth and easy software integration, but offer limite...
MLDAS: Machine Learning Dynamic Algorithm Selection for Software-Defined Networking SecurityPablo Benlloch, Oscar Romero, Antonio Leon, Jaime Lloret2026-04-16下载Network security is a critical concern in the digital landscape of today, with users demanding secure browsing experiences and protection of their personal data.
Learning Ad Hoc Network Dynamics via Graph-Structured World ModelsCan Karacelebi, Yusuf Talha Sahin, Elif Surer, Ertan Onur2026-04-16下载Ad hoc wireless networks exhibit complex, innate and coupled dynamics: node mobility, energy depletion and topology change that are difficult to model analytically.
Towards Trustworthy 6G Network Digital Twins: A Framework for Validating Counterfactual What-If Analysis in Edge Computing ResourcesJulian Jimenez Agudelo, Paola Soto, Ayat Zaki-Hindi, Jean-Sébastien Sottet, Sébastien Faye, Nina Slamnik-Kriještorac, Johann Marquez-Barja, Miguel Camelo Botero2026-04-16下载Network Digital Twins (NDTs) enable safe what-if analysis for 6G cloud-edge infrastructures, but adoption is often limited by fragmented workflows from telemetry to validation.
Switching Efficiency: A Novel Framework for Dissecting AI Data Center Network EfficiencyNiangen Ye, Jiawen Zhu, Baojun Chen, Dong Wang, Jiang Sun, Weiqiang Sun, Weisheng Hu2026-04-16下载Communication is pivotal in LLM training, and a thorough analysis of the communication efficiency of AI data center (AIDC) network is essential for guiding the design of these capital-intensive cluste...
PlanB: Efficient Software IPv6 Lookup with Linearized B+B^+-TreeZhihao Zhang, Lanzheng Liu, Chen Chen, Huiba Li, Jiwu Shu, Windsor Hsu, Yiming Zhang2026-04-16下载IP lookup via Longest Prefix Match (LPM) is critical for packet forwarding. Unfortunately, conventional lookup algorithms are inefficient for IPv6 Forwarding Information Bases (FIBs), which are charac...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Ragged Paged Attention: A High-Performance and Flexible LLM Inference Kernel for TPUJevin Jiang, Ying Chen, Blake A. Hechtman, Fenghui Zhang, Yarong Mu2026-04-16下载Large Language Model (LLM) deployment is increasingly shifting to cost-efficient accelerators like Google's Tensor Processing Units (TPUs), prioritizing both performance and total cost of ownership (T...
Serving Chain-structured Jobs with Large Memory Footprints with Application to Large Foundation Model ServingTingyang Sun, Ting He, I-Hong Hou2026-04-16下载As a current trend in Artificial Intelligence (AI), large foundation models are increasingly employed as the core of AI services. However, even after training, serving such models at scale remains a c...
DEEP-GAP: Deep-learning Evaluation of Execution Parallelism in GPU Architectural PerformanceKathiravan Palaniappan2026-04-16下载Modern datacenters increasingly rely on low-power, single-slot inference accelerators to balance performance, energy efficiency, and rack density constraints.

基于 VitePress 构建 · 使用本地搜索查找论文