2026-05-24
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry | Bo Lv, Zhiheng Xu, KeDong Xiu, Ruyi Ding, Tianhang Zheng, Zhibo Wang, Kui Ren | 2026-05-24 | 下载 | Mixture-of-Experts (MoE) architectures have become an increasingly important paradigm for scaling Large Language Models (LLMs). As MoE models are increasingly deployed in real-world services, safety a... |
| XL-HD: Extended Learning in Hyperdimensional Computing via Deterministic Projections for In-Memory Accelerators | Sabrina Hassan Moon, Abu Kaisar Mohammad Masum, Sercan Aygun, Dayane Reis | 2026-05-24 | 下载 | Hyperdimensional computing (HDC) is a promising approach for energy-efficient edge machine learning (ML), where low latency, low power, and tight memory budgets are essential. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics | Diogo Landau, Jorge G. Barbosa, Nishant Saurabh | 2026-05-24 | 下载 | Online Data-Intensive applications face performance degradation from load variability and resource interference. While Thread State Analysis (TSA) based approaches enable identifying constrained subsy... |
| DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud-HPC-Edge Compute Continuum | Aasish Kumar Sharma, Felix Stein, Mirac Aydin, Michael Bidollahkhani, Sachin P. Nanavati, Mohsen Seyedkazemi Ardebili, Giorgi Mamulashvili, Mojtaba Akbari, Jonathan Decker, Zoya Masih, Julian M. Kunkel | 2026-05-24 | 下载 | This paper presents the DECICE project (Device Edge Cloud Intelligent Collaboration framEwork), a Horizon Europe Research and Innovation Action (Grant No. |
| Kavier: Exploring Performance, Sustainability, and Efficiency of LLM Ecosystems under Inference through Cache-Aware Discrete-Event Simulation | Radu Nicolae, Alexandru Iosup, Animesh Trivedi, Jesse Donkervliet | 2026-05-24 | 下载 | Large Language Models (LLMs) are widely used by our increasingly digitalized society, but raise sustainability, performance, and financial concerns, especially as inference workloads grow. |
| Optimus: Elastic Decoding for Efficient Diffusion LLM Serving | Chiyue Wei, Cong Guo, Bowen Duan, Junyao Zhang, Haoxuan Shan, Yifei Wang, Yangjie Zhou, Hai "Helen" Li, Danyang Zhuo, Yiran Chen | 2026-05-24 | 下载 | Large language model (LLM) serving is fundamentally limited by inefficient hardware utilization. Autoregressive (AR) decoding underutilizes GPUs due to its strictly sequential execution, while diffusi... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| K8S Power Irrigation: Deep Reinforcement Learning for Performance-Aware Power Efficiency of Kubernetes Cloud-Native Microservices | Zouhir Bellal, Laaziz Lahlou, Nadjia Kara, Timothy Murphy, Tan Phat Nguyen | 2026-05-24 | 下载 | Modern cloud platforms are facing a sharp increase in power demand driven by the rapid adoption of AI-powered applications, making power optimization urgent under net-zero commitments and sustainabili... |
| Securing High-Performance Data Transfers: Implementing AES Encryption in RDMA Systems | Erik Bångsbo, Zakaria Hersi, Anna Benktson, Stefan Holmgren, Romaric Duvignau | 2026-05-24 | 下载 | Remote Direct Memory Access (RDMA) is a key enabler of high-performance systems, offering low latency, high throughput, and reduced CPU overhead by allowing direct memory-to-memory transfers between m... |
| Scaling up Energy-Aware Multi-Agent Reinforcement Learning for Mission-Oriented Drone Networks with Individual Reward | Changling Li, Ying Li | 2026-05-24 | 下载 | Multi-agent reinforcement learning (MARL) has shown wide applicability in collaborative systems such as autonomous driving and smart cities for its ability of learning through interaction. |
| Clustering as Reasoning: A -Means Interpretation of Chain-of-Thought Graph Learning | Xuanting Xie, Zhaochen Guo, Bingheng Li, Xingtong Yu, Zhifei Liao, Zhao Kang, Yuan Fang | 2026-05-24 | 下载 | Chain-of-Thought (CoT) prompting has shown promise in enhancing the reasoning capabilities of large language models (LLMs) on text-attributed graphs (TAGs). |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Beyond Thread States: Diagnosing Performance Degradation with eBPF and Thread Dynamics | Diogo Landau, Jorge G. Barbosa, Nishant Saurabh | 2026-05-24 | 下载 | Online Data-Intensive applications face performance degradation from load variability and resource interference. While Thread State Analysis (TSA) based approaches enable identifying constrained subsy... |