Skip to content

2026-08-23 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM InferenceIn-Jun Jung, Jaeha Min, Joo-Young Kim2026-08-23下载The rapid evolution of hybrid large language models (LLMs), which interleave grouped-query-attention (GQA), state-space model (SSM), and Mixture-of-Experts (MoE) layers, introduces two fundamental cha...
Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI EraJunrui Pan, Weili An, Cesar Avalos Baddouh, Christin David Bose, Ni Kang, Aaron Barnes, Ahmad Alawneh, Fangjia Shen, Yechen Liu, Anusuya Nallathambi, Atthin Chandrashekar, Timothy G. Rogers2026-08-23下载The rapid evolution of machine learning workloads has fundamentally transformed GPU hardware, driving architectures toward Multi-Chip Module (MCM) topologies, asynchronous execution primitives, and pe...
Precision-Aware Variable Bit Processing Elements for Hardware-Efficient Systolic Array DesignsDantu Nandini Devi, Madhav Rao2026-08-23下载Systolic arrays (SAs) have emerged as prominent hardware accelerators for matrix operations in deep learning, while floating point number formats enable precision control across computational domains.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta PrefetchingNobel Dhar, Md Romyull Islam, Xuechen Zhang, Gongjin Sun, Sahidul Islam, Bobin Deng, Kun Suo2026-08-23下载Deploying large language models on edge devices is increasingly limited by a widening gap between model size and available memory. Existing approaches such as quantization, smaller models, and offload...
GCA: Global Centroid Alignment in Federated LearningJong-Ik Park, Harry Jiang, Logan Blakely, Georgios Fragkos, Shamina Hossain-McKenzie, Carlee Joe-Wong2026-08-23下载Autoencoder (AE)-based federated learning (FL) is attractive for anomaly detection when clients have limited local data. However, conventional FL exchanges AE parameters or gradients, incurring substa...
Model-Consistent Byzantine-Resilient Decentralized Federated Learning for Collaborative MissionsYue Li, Sudip Bhujel, Cameron Lira, Ning Wang, Yang Xiao2026-08-23下载Decentralized federated learning (DFL) is a promising paradigm for autonomous nodes to collaboratively train AI models without relying on a central server.
Understanding the Synchronization Tax in GPU Scale-Up DomainsArjun Devraj, Lindsey Bowen, Rachee Singh2026-08-23下载GPU scale-up domains have become the building block of modern machine learning infrastructure, and their design follows a clear trajectory of exponential growth in both interconnect bandwidth and doma...
Non-Leaking Concurrent ObjectsHagit Attiya, Rotem Oshman, Noa Schiller, Corentin Travers2026-08-23下载Abstract specifications of concurrent objects determine which values operations may return, but they also implicitly constrain which information operations may know, for example the arguments of other...
Scalable Exact Path Selection via Structure-Aware Search for Virtual Payment ChannelsJiangnan Luo, Zhebei Shen, Yuan Zhang, Sheng Zhong2026-08-23下载Virtual Payment Channels (VPCs) enable efficient off-chain transactions in Payment Channel Networks (PCNs), but their performance depends on selecting high-quality underlying paths.
Fleet-Scale Pod Deployment with VPC-Native Networking in Managed KubernetesSri Saran Balaji Vellore Rajakumar, Jayanth Varavani, Murat Parlakisik, Pavani Panakanti2026-08-23下载Managed Kubernetes pod deployment is often presented as a choice between VPC-native and overlay networking. Public EKS documentation describes prefix-mode capacity and configuration, while prior CNI s...
Unveiling the Depth-Performance Dilemma in Split-Federated Fine-tuning of LLMsHariharan Ramesh, Someshwaran Murugaiyan, Jyotikrishna Dass2026-08-23下载Split Federated Fine-tuning (SFF) is a promising paradigm for scaling Large Language Models (LLMs) by partitioning model depth between resource-constrained clients and a centralized server.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Advanced LLM-Enhanced Intent-Based 5G Network Management using Dynamic Semantic RoutesThomas Benton Townsend, Dimitrios Michael Manias2026-08-23下载As the use of Artificial Intelligence (AI) and Large Language Models (LLMs) is becoming common in everyday applications, their ability to interpret natural language has increased significantly.
STAR-GS: Truthful and Visibility-Aware Resource Scheduling for Ground Station as a ServiceZhiying Wang, Xiaojian Wang, Huayue Gu, Zhishan Guo, Ruozhou Yu2026-08-23下载The rapid growth of Low Earth Orbit satellite constellations has created increasing demand for efficient and scalable downlink services. Ground Station as a Service (GSaaS) provides an on-demand acces...
Fleet-Scale Pod Deployment with VPC-Native Networking in Managed KubernetesSri Saran Balaji Vellore Rajakumar, Jayanth Varavani, Murat Parlakisik, Pavani Panakanti2026-08-23下载Managed Kubernetes pod deployment is often presented as a choice between VPC-native and overlay networking. Public EKS documentation describes prefix-mode capacity and configuration, while prior CNI s...

基于 VitePress 构建 · 使用本地搜索查找论文