2026-05-30
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| LP5X-PIM Sim: A High-Fidelity HW/SW Integrated Simulator for LPDDR5X-PIM | SangHoon Cha, Jaewan Choi, Byeongho Kim, Yoonah Paik, Sukhan Lee, Kyomin Sohn | 2026-05-30 | 下载 | This tech note describes the architecture and execution results of the LPDDR5X-PIM simulator, developed by Samsung Electronics. Based on the latest research and internal specifications, the simulator ... |
| Regular-Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures | Dazhi Yang, Shafayat Mowla Anik, Byeong Kil Lee, Jeeho Ryoo | 2026-05-30 | 下载 | Recent diffusion accelerators exploit activation sparsity by skipping near-zero GELU outputs, reporting 52--85% element-level sparsity. However, systolic-array hardware processes activations at column... |
| Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures | Shafayat Mowla Anik, Yongchan Jung, Jeeho Ryoo, Byeong Kil Lee | 2026-05-30 | 下载 | GPU workloads with large memory footprints frequently suffer from redundant L2 TLB misses in which a recently evicted translation is immediately re-walked at full page-walk cost. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving | Seokjin Go, Marko Scrbak, Ephrem Wu, Srilatha Manne, Divya Mahajan | 2026-05-30 | 下载 | In distributed Mixture-of-Experts (MoE) inference, input-dependent token routing interacts with GPU performance variability to create persistent stragglers under synchronized execution, where the slow... |
| The Cartan-Topos Protocol: A Unified Geometric and Categorical Framework for Resilient Multi-Agent Coordination | Manuel Hernández, Eduardo Sánchez-Soto | 2026-05-30 | 下载 | Multi-agent coordination faces a fundamental divide between continuous Euclidean consensus, which fails under non-integrable constraints, and discrete symbolic logic, which collapses under open-world ... |
| ScanWeaver: Compiler-Driven Parallelization of Affine Recurrences via Associative Scan Lowering | Qiying Wu, Pavel Zolnikov | 2026-05-30 | 下载 | Selective state-space models such as Mamba highlight the practical importance of input-dependent scan recurrences, which preserve linear-time sequence modeling while improving language modeling capabi... |
| Edge-Based QoS-Aware Adaptive Task Placement: A Closed-Loop Control in Multi-Robot Systems | Thien Tran, Jonathan Kua, Thuong Hoang, Minh Tran, Honghao Lyu, Jiong Jin | 2026-05-30 | 下载 | Multi-robot systems (MRS) increasingly offload compute-intensive perception tasks to edge nodes to meet strict time-sensitive Quality-of-Service (QoS) constraints. |
| Joint Optimization of Qubit Leasing and Quantum Circuit Distribution | Anoushka Dey, Gaurav S. Kasbekar | 2026-05-30 | 下载 | We consider an agent, who would like to execute a given quantum circuit using resources leased from a set of quantum computers (QCs) connected by a quantum network. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Edge-Based QoS-Aware Adaptive Task Placement: A Closed-Loop Control in Multi-Robot Systems | Thien Tran, Jonathan Kua, Thuong Hoang, Minh Tran, Honghao Lyu, Jiong Jin | 2026-05-30 | 下载 | Multi-robot systems (MRS) increasingly offload compute-intensive perception tasks to edge nodes to meet strict time-sensitive Quality-of-Service (QoS) constraints. |
| Joint Optimization of Qubit Leasing and Quantum Circuit Distribution | Anoushka Dey, Gaurav S. Kasbekar | 2026-05-30 | 下载 | We consider an agent, who would like to execute a given quantum circuit using resources leased from a set of quantum computers (QCs) connected by a quantum network. |
| XOR Bidding and Knapsack Formulations for HPC Network Resource Allocation | Abrar Hossain, Kishwar Ahmed | 2026-05-30 | 下载 | Modern High Performance Computing (HPC) centers face growing challenges in ingesting large and diverse data streams. These issues often create bottlenecks that limit bandwidth utilization and delay sc... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI | Tian Xia, Hanchen Li, Zhifei Li, Xiaokun Chen, Hao Kang, Yifan Qiao, Yi Xu, Ion Stoica | 2026-05-30 | 下载 | Modern LLM serving systems increasingly host agentic workloads, whose sessions issue tens of model invocations interleaved with tool calls, accumulating KV cache that can be reused across steps. |
| Edge-Based QoS-Aware Adaptive Task Placement: A Closed-Loop Control in Multi-Robot Systems | Thien Tran, Jonathan Kua, Thuong Hoang, Minh Tran, Honghao Lyu, Jiong Jin | 2026-05-30 | 下载 | Multi-robot systems (MRS) increasingly offload compute-intensive perception tasks to edge nodes to meet strict time-sensitive Quality-of-Service (QoS) constraints. |
| Beyond Edge Coverage: Per-Task Data-Flow Extraction at Kernel Function Boundaries via LLVM | Yunseong Kim | 2026-05-30 | 下载 | Coverage-guided kernel fuzzers such as syzkaller rely on edge coverage (trace-pc) as their sole feedback signal. This context-blind approach cannot distinguish execution paths that differ only in argu... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Regular-Activation Concentration: Characterizing Column-Level Output Sparsity Across Diffusion Model Architectures | Dazhi Yang, Shafayat Mowla Anik, Byeong Kil Lee, Jeeho Ryoo | 2026-05-30 | 下载 | Recent diffusion accelerators exploit activation sparsity by skipping near-zero GELU outputs, reporting 52--85% element-level sparsity. However, systolic-array hardware processes activations at column... |
| Regular-Dead on Arrival: Characterizing and Protecting Against Dead-Entry TLB Misses in GPU Microarchitectures | Shafayat Mowla Anik, Yongchan Jung, Jeeho Ryoo, Byeong Kil Lee | 2026-05-30 | 下载 | GPU workloads with large memory footprints frequently suffer from redundant L2 TLB misses in which a recently evicted translation is immediately re-walked at full page-walk cost. |
| Maximizing Compute Capacity in AI Data Centers through Cooling, Energy Storage, and Computing Adaptation | Shaolei Ren, Mohammad A. Islam, Adam Wierman | 2026-05-30 | 下载 | The deployment of artificial intelligence is increasingly constrained by limited site-level power capacity, which must support both compute systems and non-compute systems (primarily cooling) at all t... |