2026-06-10
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable Devices | Farough Shayeste Roodi, Parham Zilouchian Moghaddam, Mahdi Mohammadi-nasab, Mehdi Modarressi, Mostafa Ersali Salehi Nasab, Masoud Daneshtalab | 2026-06-10 | 下载 | Wearable healthcare devices are the fastest-growing Internet of Things (IoT) sector. Many automated healthcare services rely on two crucial biological signals, namely ECG and EEG, which reflect the ac... |
| Eidola: Modeling Multi-GPU Network Communication Traffic in Distributed AI Workloads | Ranganath R. Selagamsetty, Matthew Poremba, Bradford M. Beckmann, Joshua San Miguel, Mikko H. Lipasti | 2026-06-10 | 下载 | As distributed AI workloads grow in scale, multi-GPU systems have become essential for training large models. Although techniques like kernel fusion and overlapping communication with computation help... |
| Partitioned Tags, Shared Data: Reconciling Strict Cache Isolation with Write-Shared Coherence | Kartik Ramkrishnan, Stephen McCamant, Antonia Zhai, Pen Chung Yew | 2026-06-10 | 下载 | Cache partitioning is among the strongest structural defenses against eviction-based cache side channels, yet a decade-old design issue has blocked its widespread deployment in secure shared-OS settin... |
| BenDi: An Energy-Efficient Quasi-Stochastic Systolic Architecture for Edge Bioelectronics | Bochen Ye, Yihan Pan, Shady Agwa, Themis Prodromakis | 2026-06-10 | 下载 | Continuous long-term monitoring and diagnosis of biomedical signals, such as electrocardiograms (ECGs), can help mitigate an increasing threat to public health. |
| Making Locality-aware GEMM Compatible with Page-Granularity Placement on Chiplet GPUs | Euijun Chung, Jae Hyung Ju, Hyesoon Kim | 2026-06-10 | 下载 | Multi-chiplet GPUs scale compute throughput and high-bandwidth memory (HBM) capacity, but their non-uniform memory system makes locality between chiplets and their data critical to the GPU's performan... |
| A Fast Locality Simulator for GEMM Design-Space Exploration on Multi-Chiplet GPUs | Euijun Chung, Hyesoon Kim | 2026-06-10 | 下载 | Multi-chiplet GPUs split memory into local and remote HBM regions across a silicon interposer, and reducing the remote HBM traffic is crucial for the performance and energy efficiency of multi-chiplet... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| On the Limits of Performance Portability in Directive-Based GPU Programming | Alessandro Romeo, Nitin Shukla, Stefano Truzzi, Alessio Suriano, Andrea Mignone | 2026-06-10 | 下载 | The transition of scientific applications to GPU-accelerated exascale systems is constrained by trade-offs between performance, portability, and productivity. |
| M*: A Modular, Extensible, Serving System for Multimodal Models | Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang | 2026-06-10 | 下载 | We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and... |
| A Communication Complexity Lower Bound for Nonuniformly Convex Consensus Optimization | Demyan Yarmoshik, Maxim Klimenko | 2026-06-10 | 下载 | We study the communication complexity of convex decentralized optimization over time-varying networks, where nodes hold private functions and must agree on the global minimizer using only synchron... |
| Eidola: Modeling Multi-GPU Network Communication Traffic in Distributed AI Workloads | Ranganath R. Selagamsetty, Matthew Poremba, Bradford M. Beckmann, Joshua San Miguel, Mikko H. Lipasti | 2026-06-10 | 下载 | As distributed AI workloads grow in scale, multi-GPU systems have become essential for training large models. Although techniques like kernel fusion and overlapping communication with computation help... |
| ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid Memories | Hakbeom Jang, Younghoon Min, Sunwoong Kim, Taeyoung Ahn, Hanyee Kim, Youngpyo Joo, Hoshik Kim, Jongryool Kim | 2026-06-10 | 下载 | The rapid shift toward agentic and long-context workloads in Large Language Models (LLMs) is pushing the industry beyond the capacity of individual servers toward disaggregated shared storage to handl... |
| Fair Comparison of Scheduling Algorithms on Heterogeneous Edge Clusters: A Continuous Adaptive Benchmark | Zihang Wang, Boris Sedlak, Juan Luis Herrera, Schahram Dustdar | 2026-06-10 | 下载 | Modern Artificial Intelligence (AI) workloads deployed across the heterogeneous tiers of an edge--cloud continuum must satisfy multi-dimensional Service Level Objectives (SLOs) over latency, throughpu... |
| Efficient and Robust Online Learning to Rank in Decentralized Systems | Marcel Gregoriadis, Martijn de Vos, Sayan Biswas, Anne-Marie Kermarrec, Johan Pouwelse | 2026-06-10 | 下载 | In Online Learning to Rank (OLTR), ranking models are trained directly from live user interactions, but existing systems rely on a trusted central server to collect and process these interactions. |
| The PM-EdgeMap: Towards Real-Time Process Mining on the Edge-Cloud Continuum | Hendrik Reiter, Christian Imenkamp, Olaf Landsiedel, Andrea Maldonado, Patrick Rathje, Wilhelm Hasselbring | 2026-06-10 | 下载 | Smart factories are evolving into Cyber-Physical Systems (CPS), demanding increased autonomy. This necessitates real-time decision making, facilitated by insights derived from sensor data. |
| Near-Optimal Distributed 2-Ruling Sets on Graphs with Low Arboricity | Malte Baumecker, Rustam Latypov, Yannic Maus, Jara Uitto | 2026-06-10 | 下载 | Given a graph , a β-ruling set is a subset of nodes that is independent, and each node in is at distance at most β from some node in . |
| From Fork-Join to Asynchronous Tasks: Parallelizing Tiled Cholesky Decomposition with OpenMP and HPX | Alexander Strack, Alexander Van Craen, Dirk Pflüger | 2026-06-10 | 下载 | Fork-join parallelism, popularized by OpenMP, remains the dominant model for shared-memory parallel programming, but its implicit synchronization barriers can penalize algorithms with inhomogeneous wo... |
| Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training | Yuming Zhou, Haoyang Li, Sheng Lin, Yanfeng Zhao, Tong Zhao, Xupeng Miao, Jie Jiang, Fangcheng Fu, Bin Cui | 2026-06-10 | 下载 | Mixture-of-Experts (MoE) and reinforcement learning (RL) post-training now dominate large language model (LLM) development, yet expert load imbalance remains a critical challenge. |
| Optimizing Cloud Deployment: Blending of IaaS and FaaS for Microservice Architecture | Nikhil Kapoor, Sougata Mukherjea | 2026-06-10 | 下载 | The rapid evolution of cloud computing has resulted in the adoption of hybrid deployments that blend Infrastructure-as-a-Service (IaaS) and Function-as-a-Service (FaaS) service models to optimize reso... |
| Consensus Time in 3-Majority and 2-Choices Is Determined by the Maximum Initial Opinion Density | Niccolò D Archivio | 2026-06-10 | 下载 | We establish the correct parameter governing the convergence time of the 3-Majority and 2-Choices dynamics on the complete graph in the synchronous model. |
| MHOT: Height-Optimized Authenticated Data Structure for Blockchain State Commitment | Sipeng Xie, Qianhong Wu, Minghang Li, Qiyuan Gao, Bo Qin, Qin Wang | 2026-06-10 | 下载 | State root computation dominates (78%) blockchain block processing time. Ethereum's canonical authenticated data structure, i.e., Merkle Patricia Trie (MPT), suffers from severe tree-height growth and... |
| Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation | Chitral Patil | 2026-06-10 | 下载 | Every public LLM cost calculator we surveyed treats GPU utilization as a fixed input -- entered by the user, baked in as a preset, or silently assumed at 100% -- never measured against the operator's ... |
| Sovereign Assurance Boundary: Certificate-Bound Admission for Agentic Infrastructure | Jun He, Deying Yu | 2026-06-10 | 下载 | Agentic infrastructure introduces a critical control-plane authorization problem: non-deterministic reasoning systems can propose high-stakes mutations to production resources, yet existing security m... |
| Tensor-Network-Based Distributed Quantum Dynamics on Independent Quantum Computers | Anurag Dwivedi, Melissa C. Revelle, Daniel S. Lobser, Brian K. McFarland, Edward C. Tortorici, Christopher G. Yale, Susan M. Clark, Philip Richerme, Srinivasan S. Iyengar | 2026-06-10 | 下载 | We present an approach based on tensor networks for distributed quantum computing simulation of chemical wavepacket dynamics in a continuous variable representation. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Free-Placement Optimization of Ground Station Locations for Low-Earth Orbit Satellites | Grace Ra Kim, Duncan Eddy, Vedant Srinivas, Mykel J. Kochenderfer | 2026-06-10 | 下载 | Rapidly expanding low Earth orbit satellite constellations are placing increasing demands on terrestrial ground networks, motivating the development of more efficient ground station network designs. |
| Greenness-Driven Scheduling in Far Edge Kubernetes: A CODECO Evaluation | Kaikang Huang, Dalal Ali, Rute C. Sofia | 2026-06-10 | 下载 | Energy consumption is an increasing concern in IoT-Edge-Cloud infrastructures, where containerized application orchestration must balance performance with sustainability. |
| Exploratory Analysis of Wi-Fi 6 Dynamic Resource Unit Sharing in Small-Scale Network Scenarios | Sai Mada, Anna Baron, Luigi Martino, Rute C. Sofia | 2026-06-10 | 下载 | This paper investigates dynamic Resource Unit (RU) allocation strategies for Wi-Fi~6 (IEEE 802.11ax) networks integrated with Time-Sensitive Networking (TSN), targeting the limitations of static RU sc... |
| LLM-Enabled NWDAF: A Step Toward AI-Native 6G Network Intelligence | Henok Daniel, Omar Alhussein, Cheng Li, Jie Liang, Ernesto Damiani | 2026-06-10 | 下载 | The Network Data Analytics Function (NWDAF) is central to enabling zero-touch network management in fifth-generation (5G) networks by supporting real-time analytics and closed-loop automation. |
| SwarmSense-DNN: A Trustworthy and Decentralized Neural Framework for Proactive Anomaly Defense in Consumer IoT | Jing Yang, Vijay Govindarajan, Saad Arif, Xu Xu, Mohamed Kallel, Zaffar Ahmed Shaikh, Zhe Liu, Chunhong Yuan, Lip Yee Por | 2026-06-10 | 下载 | The rapid growth of consumer IoT devices has introduced unprecedented challenges in trustworthy anomaly detection against AI-enabled cyber threats, requiring real-time, privacy-preserving, and scalabl... |
| A VPN-as-a-Service Tailored Enabler for Computing-constrained Environments | Carolina Fernández-Martínez, César Cajas Parra, Shuaib Siddiqui | 2026-06-10 | 下载 | Industry has embraced Zero Trust (ZT) architectural tenets and implementations for cloud-native environments, following stricter security requirements to both internal and external tenants. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Real-Time Language Model Jamming: A Case Study for Live Music Accompaniment Generation | Bowen Zheng, Andrew H. Yang, Jiaqi Ruan, Jia He, Xinyue Li, Yuan-Hsin Chen, Ziyu Wang, Xiaosong Ma | 2026-06-10 | 下载 | Language models (LMs) have become one of the most prominent paradigms in modern generative modeling. While making them faster has been the main focus of real-time deployment, speed alone is not enough... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| nomp: A Framework for Building Domain Specific Compilers | Thilina Ratnayaka, Kaushik Kulkarni, Nipuna Fernando, Pubudu Hewavitharana, Hirumal Priyashan, Poorna Gunathilaka, Nagitha Abeywickrema, Ravindu Hirimuthugoda, Tarun Prabhu, Kirshanthan Sundararajah, Sanath Jayasena | 2026-06-10 | 下载 | The low-level GPU programming models (CUDA, HIP, OpenCL, etc.) provide detailed control of the data flow and execution plan of a program in order to extract close-to-metal performance. |
| The Brain That Goes Quiet: Serving a Large Model's Knowledge at 131 Tokens per Second on an 8 GB Laptop by Removing the Large Model from the Runtime Path | Myeong Jun Jo | 2026-06-10 | 下载 | In earlier work I showed that a 35B-class Mixture-of-Experts model can be loaded and executed on a consumer laptop with 8 GB of GPU memory. That result solved a placement problem and immediately expos... |
| From Fork-Join to Asynchronous Tasks: Parallelizing Tiled Cholesky Decomposition with OpenMP and HPX | Alexander Strack, Alexander Van Craen, Dirk Pflüger | 2026-06-10 | 下载 | Fork-join parallelism, popularized by OpenMP, remains the dominant model for shared-memory parallel programming, but its implicit synchronization barriers can penalize algorithms with inhomogeneous wo... |
| Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation | Chitral Patil | 2026-06-10 | 下载 | Every public LLM cost calculator we surveyed treats GPU utilization as a fixed input -- entered by the user, baked in as a preset, or silently assumed at 100% -- never measured against the operator's ... |
| XPR: An Extensible Cross-Platform Point-Based Differentiable Renderer | Steve Rhyner, Sankeerth Durvasula, Aleksandr Kovalev, Hansel Jia, Adrian Zhao, Mrutunjayya Mrutunjayya, Nilesh Ahuja, Selvakumar Panneer, Christina Giannoula, Nandita Vijaykumar | 2026-06-10 | 下载 | Point-based differentiable rendering underpins modern 3D reconstruction, novel-view synthesis, and learning-based graphics pipelines, but developing new rendering methods often requires extensive low-... |