Skip to content

2026-05-09 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution RegimesWilly Fitra Hendria2026-05-09下载Autoregressive inference is typically assumed to scale predictably with decoding length, and key-value (KV) caching is widely regarded as a universally beneficial optimization for accelerating decodin...
HyDRA: Deadline and Reuse-Aware Cacheability for Hardware AcceleratorsAyushi Agarwal, Anannya Mathur, Preeti Ranjan Panda2026-05-09下载The system-level cache is a critical resource shared by processor cores and domain-specific accelerators in heterogeneous systems on chips (SoCs).
Low-Complexity Beamspace Channel Denoiser for mmWave Massive MIMO with Low-Resolution ADCsHanyoung Park, Eunho Kim, Ji-Woong Choi2026-05-09下载In this paper, we propose a low-complexity beamspace channel denoising algorithm for millimeter-wave (mmWave) massive multi-input multi-output (MIMO) systems with low-resolution analog-to-digital conv...
A Reconfigurable Multiplier Architecture for Error-Resilient Applications in RISC-V CorePragun Jaswal, L. Hemanth Krishna, B. Srinivasu2026-05-09下载Neural Networks (NNs) have been widely adopted due to their outstanding efficacy and adaptability across computer vision and deep learning applications.
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and StandardisationChih-Hua Ke2026-05-09下载DDR5 SDRAM partitions each 64-bit memory channel into two independent 32-bit sub-channels. A DIMM populating only one sub-channel halves the die count required for a given module, enabling 8 GB module...
DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit ProcessingYuhan Zhang, Zhou Wang, Zhou Shu, Jiuren Zhou, Yanqing Xu, Xiaonan Tang, Shushan Qiao, Tianchun Ye, Yang Liu, Anil A. Bharath, Emm Mic Drakakis2026-05-09下载In recent years, DeepSeek has achieved strong inference performance but remains hard to deploy on energy-constrained edge devices. This paper presents the DeepSeek Processing Element (DSPE), an edge-o...
FLARE: One-Shot PE-Level Fault Localization in Systolic Arrays via Algebraic Test VectorsLogashree Venkatasubramanian, Zishen Wan, Viveck Cadambe2026-05-09下载Systolic arrays are the dominant compute fabric for neural network inference. Prior work has addressed column-level fault detection efficiently with uniform test patterns, but row-level (PE-level) fau...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Light Cone Consistency: Toward a Unified Theory of Consistency in Message-Passing SystemsRob Landers, Kaben Kramer2026-05-09下载Every distributed system -- databases, networks, postal services, CPU caches -- is a message-passing system. Every message-passing system is a growing causal log observed by a set of observers.
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in ProductionChunyu Xue, Yangrui Chen, Jianyu Jiang, Ningxin Zheng, Junda Feng, Jingji Chen, Shixiong Zhao, Shen Yan, Yi Lin, Lei Shi, Zanbo Wang, Lishu Luo, Faming Wu, Haibin Lin, Xin Liu, Yanghua Peng, Quan Chen2026-05-09下载As the foundational component of versatile AI applications, training an multimodal large language model (MLLM) relies on multimodal datasets with dynamic modality mixture proportions and sample length...
Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance ReductionZhirayr Tovmasyan, Artavazd Maranjyan, Peter Richtárik2026-05-09下载Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network delays, and system-level instabilities.
FedGMI: Generative Model-Driven Federated Learning for Probabilistic Mixture InferenceQijun Hou, Yuchen Shi, Pingyi Fan, Khaled B. Letaief2026-05-09下载Federated Learning (FL) facilitates collaborative model training across decentralized clients while preserving data privacy by avoiding raw data exchange.
TS-Verkle: A TypeScript Native Verkle Library With On-chain VerifierZhikai Li, Xuekai Liu, Boyuan Xu, Eric Chen, Bhaskar Krishnamachari2026-05-09下载Blockchain systems face significant scalability challenges due to growing data volumes and increasing transaction demands, necessitating more efficient data structures and verification mechanisms.
PAAC: Privacy-Aware Agentic Device-Cloud CollaborationLiangqi Yuan, Wenzhi Fang, Shiqiang Wang, Christopher G. Brinton2026-05-09下载Large language model (LLM) agents face a structural tension: cloud agents provide strong reasoning but expose user data, while on-device agents preserve privacy at the cost of overall capability.
Transforming the Use of Earth Observation Data: Exascale Training of a Generative Compression Model with Historical Priors for up to 10,000x Data ReductionJinxiao Zhang, Runmin Dong, Xiyong Wu, Xihan Huang, Shenggan Cheng, Yunkai Yang, Zheng Zhou, Yunpu Xu, Zhaoyang Luo, Miao Yang, Fan Wei, Mengxuan Chen, Yang You, Juepeng Zheng, Weijia Li, Yutong Lu, Haohuan Fu2026-05-09下载Earth observation is becoming one of the largest data-producing activities in science, yet current pipelines still treat compression as a storage and transmission tool rather than a new way to use dat...
Large Language Models over Networks: Collaborative Intelligence under Resource ConstraintsLiangqi Yuan, Wenzhi Fang, Shiqiang Wang, H. Vincent Poor, Christopher G. Brinton2026-05-09下载Large language models (LLMs) are transforming society, powering applications from smartphone assistants to autonomous driving. Yet cloud-based LLM services alone cannot serve a growing class of applic...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Semantics-Aware Communication:A Differentiated Allocation PerspectiveFangming Zhao, Nikolaos Pappas, Howard H. Yang2026-05-09下载We study the joint optimization of timeliness and reliability in semantics-aware Wireless Networked Control Systems (WNCS) under computation resource constraints.
Locational Pricing for Generative-AI Services via Token-Flow Market ClearingShaohui Liu2026-05-09下载GenAI services are in an early yet fast expanding phase. Providers compete on model capability and service quality, while the underlying infrastructure remains expensive and heterogeneous across regio...
LUDB++: Enabling LUDB for the Analysis of Shaped Feedforward FIFO Networks using Network CalculusAlexander Scheffler2026-05-09下载This paper discusses how latency guarantees for non-cyclic (feedforward) First-In-First-Out (FIFO) networks with shapers can be computed within the Network Calculus (NC) framework.
Technical Report: A Hierarchical Dynamically Weighting Deep Reinforcement Learning Method for Multi-UAV Multi-Task CoordinationXindi Wang, Haining Li, Tao Ding, Bolin Cai2026-05-09下载This paper investigates the multi-UAV multi-task coordination problem in infrastructure-less emergency scenarios, where UAVs collaboratively are required to jointly perform aerial image acquisition an...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Non-Monotonic Latency in Apple MPS Decoding: KV Cache Interactions and Execution RegimesWilly Fitra Hendria2026-05-09下载Autoregressive inference is typically assumed to scale predictably with decoding length, and key-value (KV) caching is widely regarded as a universally beneficial optimization for accelerating decodin...
A Controlled Study of Memory Hierarchy Transitions in Quantum Circuit Simulation on Apple M4 Pro Unified Memory ArchitectureGyan Pratipat2026-05-09下载State-vector quantum circuit simulation is memory-bandwidth bound, yet the interaction between memory hierarchy, access pattern, and hardware parallelism remains incompletely characterized.
Single-Thread JPEG Decoder Benchmarks Mis-Evaluate ML Data LoadersVladimir Iglovikov2026-05-09下载JPEG decode is routine ML infrastructure, but Python decoder choices are often justified by single-process, single-thread microbenchmarks. We audit this evaluation assumption with twelve Python-access...
Single 32-bit Sub-Channel DDR5 DIMMs: Architecture, Performance Bounds, and StandardisationChih-Hua Ke2026-05-09下载DDR5 SDRAM partitions each 64-bit memory channel into two independent 32-bit sub-channels. A DIMM populating only one sub-channel halves the die count required for a given module, enabling 8 GB module...

基于 VitePress 构建 · 使用本地搜索查找论文