2026-08-23
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| NOVA: Technology-Architecture Co-Design of Near-Memory Processing for Attention-SSM-MoE Hybrid LLM Inference | In-Jun Jung, Jaeha Min, Joo-Young Kim | 2026-08-23 | 下载 | The rapid evolution of hybrid large language models (LLMs), which interleave grouped-query-attention (GQA), state-space model (SSM), and Mixture-of-Experts (MoE) layers, introduces two fundamental cha... |
| Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era | Junrui Pan, Weili An, Cesar Avalos Baddouh, Christin David Bose, Ni Kang, Aaron Barnes, Ahmad Alawneh, Fangjia Shen, Yechen Liu, Anusuya Nallathambi, Atthin Chandrashekar, Timothy G. Rogers | 2026-08-23 | 下载 | The rapid evolution of machine learning workloads has fundamentally transformed GPU hardware, driving architectures toward Multi-Chip Module (MCM) topologies, asynchronous execution primitives, and pe... |
| Precision-Aware Variable Bit Processing Elements for Hardware-Efficient Systolic Array Designs | Dantu Nandini Devi, Madhav Rao | 2026-08-23 | 下载 | Systolic arrays (SAs) have emerged as prominent hardware accelerators for matrix operations in deep learning, while floating point number formats enable precision control across computational domains. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| NeuroPrefetcher: Storage-Aware Sparse LLM Inference via Delta Prefetching | Nobel Dhar, Md Romyull Islam, Xuechen Zhang, Gongjin Sun, Sahidul Islam, Bobin Deng, Kun Suo | 2026-08-23 | 下载 | Deploying large language models on edge devices is increasingly limited by a widening gap between model size and available memory. Existing approaches such as quantization, smaller models, and offload... |
| GCA: Global Centroid Alignment in Federated Learning | Jong-Ik Park, Harry Jiang, Logan Blakely, Georgios Fragkos, Shamina Hossain-McKenzie, Carlee Joe-Wong | 2026-08-23 | 下载 | Autoencoder (AE)-based federated learning (FL) is attractive for anomaly detection when clients have limited local data. However, conventional FL exchanges AE parameters or gradients, incurring substa... |
| Model-Consistent Byzantine-Resilient Decentralized Federated Learning for Collaborative Missions | Yue Li, Sudip Bhujel, Cameron Lira, Ning Wang, Yang Xiao | 2026-08-23 | 下载 | Decentralized federated learning (DFL) is a promising paradigm for autonomous nodes to collaboratively train AI models without relying on a central server. |
| Understanding the Synchronization Tax in GPU Scale-Up Domains | Arjun Devraj, Lindsey Bowen, Rachee Singh | 2026-08-23 | 下载 | GPU scale-up domains have become the building block of modern machine learning infrastructure, and their design follows a clear trajectory of exponential growth in both interconnect bandwidth and doma... |
| Non-Leaking Concurrent Objects | Hagit Attiya, Rotem Oshman, Noa Schiller, Corentin Travers | 2026-08-23 | 下载 | Abstract specifications of concurrent objects determine which values operations may return, but they also implicitly constrain which information operations may know, for example the arguments of other... |
| Scalable Exact Path Selection via Structure-Aware Search for Virtual Payment Channels | Jiangnan Luo, Zhebei Shen, Yuan Zhang, Sheng Zhong | 2026-08-23 | 下载 | Virtual Payment Channels (VPCs) enable efficient off-chain transactions in Payment Channel Networks (PCNs), but their performance depends on selecting high-quality underlying paths. |
| Fleet-Scale Pod Deployment with VPC-Native Networking in Managed Kubernetes | Sri Saran Balaji Vellore Rajakumar, Jayanth Varavani, Murat Parlakisik, Pavani Panakanti | 2026-08-23 | 下载 | Managed Kubernetes pod deployment is often presented as a choice between VPC-native and overlay networking. Public EKS documentation describes prefix-mode capacity and configuration, while prior CNI s... |
| Unveiling the Depth-Performance Dilemma in Split-Federated Fine-tuning of LLMs | Hariharan Ramesh, Someshwaran Murugaiyan, Jyotikrishna Dass | 2026-08-23 | 下载 | Split Federated Fine-tuning (SFF) is a promising paradigm for scaling Large Language Models (LLMs) by partitioning model depth between resource-constrained clients and a centralized server. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Advanced LLM-Enhanced Intent-Based 5G Network Management using Dynamic Semantic Routes | Thomas Benton Townsend, Dimitrios Michael Manias | 2026-08-23 | 下载 | As the use of Artificial Intelligence (AI) and Large Language Models (LLMs) is becoming common in everyday applications, their ability to interpret natural language has increased significantly. |
| STAR-GS: Truthful and Visibility-Aware Resource Scheduling for Ground Station as a Service | Zhiying Wang, Xiaojian Wang, Huayue Gu, Zhishan Guo, Ruozhou Yu | 2026-08-23 | 下载 | The rapid growth of Low Earth Orbit satellite constellations has created increasing demand for efficient and scalable downlink services. Ground Station as a Service (GSaaS) provides an on-demand acces... |
| Fleet-Scale Pod Deployment with VPC-Native Networking in Managed Kubernetes | Sri Saran Balaji Vellore Rajakumar, Jayanth Varavani, Murat Parlakisik, Pavani Panakanti | 2026-08-23 | 下载 | Managed Kubernetes pod deployment is often presented as a choice between VPC-native and overlay networking. Public EKS documentation describes prefix-mode capacity and configuration, while prior CNI s... |