2026-05-09
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes | Willy Fitra Hendria | 2026-05-09 | 下载 | Autoregressive inference is typically assumed to scale predictably with decoding length, and key-value (KV) caching is widely regarded as a universally beneficial optimization for accelerating decodin... |
| HyDRA: Deadline and Reuse-Aware Cacheability for Hardware Accelerators | Ayushi Agarwal, Anannya Mathur, Preeti Ranjan Panda | 2026-05-09 | 下载 | The system-level cache is a critical resource shared by processor cores and domain-specific accelerators in heterogeneous systems on chips (SoCs). |
| Low-Complexity Beamspace Channel Denoiser for mmWave Massive MIMO with Low-Resolution ADCs | Hanyoung Park, Eunho Kim, Ji-Woong Choi | 2026-05-09 | 下载 | In this paper, we propose a low-complexity beamspace channel denoising algorithm for millimeter-wave (mmWave) massive multi-input multi-output (MIMO) systems with low-resolution analog-to-digital conv... |
| A Reconfigurable Multiplier Architecture for Error-Resilient Applications in RISC-V Core | Pragun Jaswal, L. Hemanth Krishna, B. Srinivasu | 2026-05-09 | 下载 | Neural Networks (NNs) have been widely adopted due to their outstanding efficacy and adaptability across computer vision and deep learning applications. |
| Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation | Chih-Hua Ke | 2026-05-09 | 下载 | DDR5 SDRAM partitions each 64-bit memory channel into two independent 32-bit sub-channels. A DIMM populating only one sub-channel halves the die count required for a given module, enabling 8 GB module... |
| DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing | Yuhan Zhang, Zhou Wang, Zhou Shu, Jiuren Zhou, Yanqing Xu, Xiaonan Tang, Shushan Qiao, Tianchun Ye, Yang Liu, Anil A. Bharath, Emm Mic Drakakis | 2026-05-09 | 下载 | In recent years, DeepSeek has achieved strong inference performance but remains hard to deploy on energy-constrained edge devices. This paper presents the DeepSeek Processing Element (DSPE), an edge-o... |
| FLARE: One-Shot PE-Level Fault Localization in Systolic Arrays via Algebraic Test Vectors | Logashree Venkatasubramanian, Zishen Wan, Viveck Cadambe | 2026-05-09 | 下载 | Systolic arrays are the dominant compute fabric for neural network inference. Prior work has addressed column-level fault detection efficiently with uniform test patterns, but row-level (PE-level) fau... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Light Cone Consistency: Toward a Unified Theory of Consistency in Message-Passing Systems | Rob Landers, Kaben Kramer | 2026-05-09 | 下载 | Every distributed system -- databases, networks, postal services, CPU caches -- is a message-passing system. Every message-passing system is a growing causal log observed by a set of observers. |
| MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production | Chunyu Xue, Yangrui Chen, Jianyu Jiang, Ningxin Zheng, Junda Feng, Jingji Chen, Shixiong Zhao, Shen Yan, Yi Lin, Lei Shi, Zanbo Wang, Lishu Luo, Faming Wu, Haibin Lin, Xin Liu, Yanghua Peng, Quan Chen | 2026-05-09 | 下载 | As the foundational component of versatile AI applications, training an multimodal large language model (MLLM) relies on multimodal datasets with dynamic modality mixture proportions and sample length... |
| Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction | Zhirayr Tovmasyan, Artavazd Maranjyan, Peter Richtárik | 2026-05-09 | 下载 | Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network delays, and system-level instabilities. |
| FedGMI: Generative Model-Driven Federated Learning for Probabilistic Mixture Inference | Qijun Hou, Yuchen Shi, Pingyi Fan, Khaled B. Letaief | 2026-05-09 | 下载 | Federated Learning (FL) facilitates collaborative model training across decentralized clients while preserving data privacy by avoiding raw data exchange. |
| TS-Verkle: A TypeScript Native Verkle Library With On-chain Verifier | Zhikai Li, Xuekai Liu, Boyuan Xu, Eric Chen, Bhaskar Krishnamachari | 2026-05-09 | 下载 | Blockchain systems face significant scalability challenges due to growing data volumes and increasing transaction demands, necessitating more efficient data structures and verification mechanisms. |
| PAAC: Privacy-Aware Agentic Device-Cloud Collaboration | Liangqi Yuan, Wenzhi Fang, Shiqiang Wang, Christopher G. Brinton | 2026-05-09 | 下载 | Large language model (LLM) agents face a structural tension: cloud agents provide strong reasoning but expose user data, while on-device agents preserve privacy at the cost of overall capability. |
| Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data Reduction | Jinxiao Zhang, Runmin Dong, Xiyong Wu, Xihan Huang, Shenggan Cheng, Yunkai Yang, Zheng Zhou, Yunpu Xu, Zhaoyang Luo, Miao Yang, Fan Wei, Mengxuan Chen, Yang You, Juepeng Zheng, Weijia Li, Yutong Lu, Haohuan Fu | 2026-05-09 | 下载 | Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and transmission tool rather than a new way to use dat... |
| Large Language Models over Networks: Collaborative Intelligence under Resource Constraints | Liangqi Yuan, Wenzhi Fang, Shiqiang Wang, H. Vincent Poor, Christopher G. Brinton | 2026-05-09 | 下载 | Large language models (LLMs) are transforming society, powering applications from smartphone assistants to autonomous driving. Yet cloud-based LLM services alone cannot serve a growing class of applic... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Semantics-Aware Communication:A Differentiated Allocation Perspective | Fangming Zhao, Nikolaos Pappas, Howard H. Yang | 2026-05-09 | 下载 | We study the joint optimization of timeliness and reliability in semantics-aware Wireless Networked Control Systems (WNCS) under computation resource constraints. |
| Locational Pricing for Generative-AI Services via Token-Flow Market Clearing | Shaohui Liu | 2026-05-09 | 下载 | GenAI services are in an early yet fast expanding phase. Providers compete on model capability and service quality, while the underlying infrastructure remains expensive and heterogeneous across regio... |
| LUDB++: Enabling LUDB for the Analysis of Shaped Feedforward FIFO Networks using Network Calculus | Alexander Scheffler | 2026-05-09 | 下载 | This paper discusses how latency guarantees for non-cyclic (feedforward) First-In-First-Out (FIFO) networks with shapers can be computed within the Network Calculus (NC) framework. |
| Technical Report: A Hierarchical Dynamically Weighting Deep Reinforcement Learning Method for Multi-UAV Multi-Task Coordination | Xindi Wang, Haining Li, Tao Ding, Bolin Cai | 2026-05-09 | 下载 | This paper investigates the multi-UAV multi-task coordination problem in infrastructure-less emergency scenarios, where UAVs collaboratively are required to jointly perform aerial image acquisition an... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution Regimes | Willy Fitra Hendria | 2026-05-09 | 下载 | Autoregressive inference is typically assumed to scale predictably with decoding length, and key-value (KV) caching is widely regarded as a universally beneficial optimization for accelerating decodin... |
| A Controlled Study of Memory Hierarchy Transitions in Quantum Circuit Simulation on Apple M4 Pro Unified Memory Architecture | Gyan Pratipat | 2026-05-09 | 下载 | State-vector quantum circuit simulation is memory-bandwidth bound, yet the interaction between memory hierarchy, access pattern, and hardware parallelism remains incompletely characterized. |
| Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data Loaders | Vladimir Iglovikov | 2026-05-09 | 下载 | JPEG decode is routine ML infrastructure, but Python decoder choices are often justified by single-process, single-thread microbenchmarks. We audit this evaluation assumption with twelve Python-access... |
| Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and Standardisation | Chih-Hua Ke | 2026-05-09 | 下载 | DDR5 SDRAM partitions each 64-bit memory channel into two independent 32-bit sub-channels. A DIMM populating only one sub-channel halves the die count required for a given module, enabling 8 GB module... |