Skip to content

2026-05-22 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
EVA: Accelerating LLM Decoding via an Efficient Vector Quantization ArchitectureBowen Duan, Cong Guo, Chiyue Wei, Haoxuan Shan, Yuzhe Fu, Xinhua Chen, Yifan Xu, Ziyue Zhang, Changchun Zhou, Hai Li, Yiran Chen2026-05-22下载Large Language Models (LLMs) have achieved impressive performance across diverse domains but remain inefficient during the autoregressive decoding phase.
DORA: Dataflow-Instruction Orchestration Architecture for DNN AccelerationXingzhen Chen, Zhuoping Yang, Jinming Zhuang, Shixin Ji, Sarah Schultz, Zheng Dong, Weisong Shi, Peipei Zhou2026-05-22下载As deep neural networks develop significantly more diverse and complex, achieving high performance and efficiency on complicated DNN models faces pressing challenges.
UniSpike: Accelerating Spiking Neural Networks on Neuromorphic Systems via Eliminating Address RedundancyQinghui Xing, Zhuo Chen, Xin Du, Ouwen Jin, Ming Zhang, Pan Lv, Ying Li, Shuiguang Deng, Gang Pan2026-05-22下载Many-core neuromorphic systems accelerate Spiking Neural Networks (SNNs), yet their packet-based spike communication can spend substantial traffic and energy repeatedly transmitting destination addres...
To Overlay or to Customize? Revisiting Architectural Choices in Heterogeneous SystemsXingzhen Chen, Shixin Ji, Zheng Dong, Peipei Zhou2026-05-22下载In this work, we present a systematic study of this trade-off from a deployment-centric perspective, focusing on an autonomous driving scenario.
DAE4HLS: Exposing Memory-Level Parallelism for High-Level Synthesis using Explicit DecouplingDavid Metz, Magnus Själander2026-05-22下载High-level synthesis (HLS) performs well for simple memory access patterns, such as for sequential accesses that can be turned into bursts, or for memory accesses into small datasets that can be store...
NASiC: 3D NAND-based CAM-Selected Multibit CIM Architecture for Efficient On-Device Mixture-of-Experts LLM InferenceWeikai Xu, Meng Li, Shuzhang Zhong, Tianyang Luo, Dongxue Zhao, Ling Liang, Zongwei Wang, Qianqian Huang, Yimao Cai, Ru Huang2026-05-22下载The Mixture-of-Experts (MoE) models have emerged as the state-of-the-art paradigm for scaling up large language models (LLMs) without proportionally increased computational cost.
MASQ: Accelerating Masked Diffusion via Stage-Wise Multi-Precision QuantizationSeeyeon Kim, Jaehun Lee, Sungyeob Yoo, Joo-Young Kim2026-05-22下载Masked diffusion enables region-specific image synthesis but suffers from computational redundancy, since the entire image is processed each timestep even though only the masked region requires genera...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Resident KV Claims: A Conformance Contract for Future Reuse under Active KV PressureLukas Stepanek2026-05-22下载KV-cache reuse mechanisms increasingly expose priority, duration, offload, routing hints, scheduler modes, and event streams. These mechanisms help preserve reusable prefixes, but they do not by thems...
Polar: Agentic RL on Any Harness at ScaleBinfeng Xu, Hao Zhang, Shaokun Zhang, Songyang Han, Mingjie Liu, Jian Hu, Shizhe Diao, Zhenghui Jin, Yunheng Zou, Michael Demoret, Jan Kautz, Yi Dong2026-05-22下载Reinforcement learning for language agents increasingly depends on custom harnesses that manage long-running context, multi-turn tool use and multi-agent orchestration.
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference BenchmarksAshok Chandrasekar, Jason Kramberger2026-05-22下载As Large Language Models (LLMs) transition from research environments to production deployments, evaluating their performance against strict Service Level Objectives (SLOs) has become critical.
The Time is Here for Just-in-Time Systems: Challenges and OpportunitiesShu Liu, Alexander Krentsel, Shubham Agarwal, Mert Cemri, Ziming Mao, Soujanya Ponnapalli, Alexandros G. Dimakis, Sylvia Ratnasamy, Matei Zaharia, Aditya Parameswaran, Ion Stoica2026-05-22下载Core systems like key-value stores have historically taken years to build, and are designed to be general so as to amortize cost across deployments, paying a significant performance cost.
Enhancing Energy Efficiency in Scientific Workflows through CFD based PIVAEsAli Zahir, Ashiq Anjum, Mark Wilkinson, Jeyan Thiyagalingam2026-05-22下载The growing complexity and scale of scientific workflows in high performance computing (HPC) environments have led to significant challenges in managing energy consumption without compromising computa...
SDNator is Not Another SDN Controller: Enabling Extensible Data-Driven Control in Cyber-Physical SystemsY. Lin, R. Zhang, E. Balta, X. Zhu, J. Zhang, K. Barton, D. Tilbury, Z. Mao2026-05-22下载An SDN-like centralized control architecture is increasingly popular and has been widely explored in cyber-physical systems (CPS) such as manufacturing, internet-of-things, and autonomous vehicle syst...
A Pragmatic Approach to Learned Indexing in RocksDB: Targeted Optimizations with Minimal System ModificationShubham Vashisth, Olivier Michaud, Bettina Kemme, Oana Balmau2026-05-22下载Learned indexes have emerged as a promising alternative to traditional index structures, offering higher throughput and lower memory usage by approximating the cumulative key distribution function wit...
HyperParallel-MoE: Multi-Core Interleaved Scheduling for Fast MoE Training on Ascend NPUsZewen Jin, Congkun Ai, Guangpeng Zhang, Hanbo Zhang, Haoran Wang, Shihan Xiao, Da Lei, Xuefeng Jin, Teng Su, Cheng Li2026-05-22下载Modern Mixture-of-Experts (MoE) models increasingly rely on large-scale AI accelerator clusters for efficient training. Ascend NPUs expose heterogeneous on-chip compute resources, including matrix-ori...
Flare: Leveraging Serverless Elasticity to Absorb Microservice Load SpikesDilina Dehigama, Shyam Jesalpura, David Schall, Antonios Katsarakis, Marios Kogias, Rakesh Kumar, Boris Grot2026-05-22下载Online services strive to maintain application responsiveness even when the traffic is unpredictable and fluctuating. Today's online services are commonly deployed as chains of microservices, each mic...
AMP: Arc Multi-Proposer Protocol with Bounded Inclusion GuaranteesDaniel Cason, Gordon Liao, Sergio Mena, Nenad Milošević, Adi Seredinschi, Alessandro Sforzin, João Sousa, Preston Vander Vos2026-05-22下载Blockchain systems that settle financial transactions face a structural tension: the single validator that assembles each block holds unilateral power over transaction inclusion and ordering.
Herring: Parallel Batch-Order-Fairness on DAG-based Blockchain ConsensusMarko Putnik, Jérémie Decouchant2026-05-22下载Transaction ordering attacks extract billions of dollars annually from decentralized finance users in the form of Maximal Extractable Value (MEV).
Multi-Factor Trust-Driven Secure Communication Model for Cloud-Based Digital TwinsDeepika Saxena, Ashutosh Kumar Singh2026-05-22下载Cloud-based Digital Twin (DT) platforms enable real-time monitoring, simulation, and collaborative decision-making across distributed clients.
Multi-Round Visibility: A Post-Consensus Ordering Layer for DAG-Based BFTPengkun Ren, Dong Hai, Nasrin Sohrabi, Zahir Tari2026-05-22下载Directed acyclic graph (DAG)-based Byzantine Fault-Tolerant (BFT) protocols achieve high throughput by decoupling dissemination from agreement and allowing many vertices to be committed concurrently.
AlignedServe: Orchestrating Prefix-aware Batching to Build a High-throughput and Computing-efficient LLM Serving SystemFengyao Bai, Hongbin Zhang, Zhitao Chen, Jiangsu Du, Zhiguang Chen, Yutong Lu2026-05-22下载High-throughput inference serving is essential for applications built on large language models (LLMs). Existing serving frameworks reduce request-level and batch-level bubbles through batching and sch...
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy FarmsTella Rajashekhar Reddy, Atharva Deshmukh, Liangcheng Yu, Chaojie Zhang, Mike Shepperd, Rohan Gandhi, Anjaly Parayil, Srinivasan Iyengar, Ajay Manchepalli, Debopam Bhattacherjee2026-05-22下载AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up. Grid expansion comes with high capital expenditure and long-distance transmission losses...
Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI SystemsAasish Kumar Sharma, Julian M. Kunkel2026-05-22下载AI-enabled services deployed in critical digital infrastructure are subject to governance obligations spanning transparency, accountability, fairness, and traceability.
Inductive Deductive Synthesis: Enabling AI to Generate Formally Verified SystemsShubham Agarwal, Alexander Krentsel, Shu Liu, Mert Cemri, Audrey Cheng, Rui Meng, Tomas Pfister, Chun-Liang Li, Sylvia Ratnasamy, Aditya Parameswaran, Matei Zaharia, Ion Stoica, Mohsen Lesani2026-05-22下载AI agents increasingly excel at generating, testing, and refining code. However, they fall short on tasks requiring formal guarantees of full coverage that testing alone cannot provide.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
RxGS: Receiver-Generalizable 3D Gaussian Splatting for Radio-Frequency Data SynthesisKang Yang, Mani Srivastava2026-05-22下载Radio-frequency (RF) data synthesis predicts the received signal given transmitter and receiver positions, and is essential for wireless applications.
Ant Backpressure Routing for Dynamic Wireless Multi-hop Networks with Mixed Traffic PatternsNegar Erfaniantaghvayi, Zhongyuan Zhao, Kevin Chan, Ananthram Swami, Santiago Segarra2026-05-22下载Backpressure (BP) routing and its shortest-path biased variant (SP-BP) provide powerful congestion-aware multipath resource allocation for wireless multi-hop networks, but they rely on per-commodity q...
BShare: Packet Queueing Delay-Driven Buffer Sharing for Datacenter SwitchesKrishna Agarwal, Muhamad Rizka Maulana, Vamsi Addanki, Habib Mostafaei2026-05-22下载Modern datacenter switches share packet buffers across ports to boost overall throughput and reduce packet loss. However, as buffer availability per-port-per-bandwidth unit continues to decrease, exis...
SDNator is Not Another SDN Controller: Enabling Extensible Data-Driven Control in Cyber-Physical SystemsY. Lin, R. Zhang, E. Balta, X. Zhu, J. Zhang, K. Barton, D. Tilbury, Z. Mao2026-05-22下载An SDN-like centralized control architecture is increasingly popular and has been widely explored in cyber-physical systems (CPS) such as manufacturing, internet-of-things, and autonomous vehicle syst...
SafeSABR: Risk-Calibrated Adaptive Bitrate Streaming over Starlink NetworksHongjun Xie, Jiahang Zhu, Zhiming Shao, Chao Fan, Zenghui Zhang, Genke Yang, Pengcheng Luo2026-05-22下载Starlink, as a representative low Earth orbit (LEO) satellite broadband system, makes high-bitrate video streaming possible in regions where terrestrial broadband is unavailable.
Sea Trial Validation of the ROS-DESERT Middleware with Autonomous Underwater VehiclesDavide Cosimo, Davide Costa, Riccardo Costanzi, Filippo Campagnaro, Andrea Caiti, Michele Zorzi2026-05-22下载This paper presents a modular software architecture that enables environmental-aware coordination of heterogeneous Autonomous Underwater Vehicles (AUVs) to improve underwater acoustic connectivity.
Experimental Evaluation of LPWAN Technologies: mioty, LoRaWAN, Sigfox, NB-IoT, and LTE-M in Deep Indoor EnvironmentsChristof Röhrig, Benz Cramer2026-05-22下载Low Power Wide Area Networks (LPWAN) are often used in applications such as Smart City, Smart Buildings and Smart Metering. Energy meters are often located in underground spaces that are difficult to ...
XWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy FarmsTella Rajashekhar Reddy, Atharva Deshmukh, Liangcheng Yu, Chaojie Zhang, Mike Shepperd, Rohan Gandhi, Anjaly Parayil, Srinivasan Iyengar, Ajay Manchepalli, Debopam Bhattacherjee2026-05-22下载AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up. Grid expansion comes with high capital expenditure and long-distance transmission losses...
Purification Strategy Optimization for Entanglement Routing in Quantum NetworksJavier Vecino Peñas, Ana Fernández-Vilas, Rebeca P. Díaz-Redondo, Sergio Gándara Gándara, Manuel Fernández-Veiga2026-05-22下载Quantum networks rely on the efficient distribution of entanglement to enable long-distance quantum communication and information processing. A key challenge in these networks is the design of routing...
On the Performance of DCF in Full Duplex WLANs with Hidden TerminalsAnastasios C. Politis, Constantinos S. Hilas, Hristos T. Anastassiu2026-05-22下载Full Duplex (FD) technology is considered as one of the next big leap in the evolution of modern WLANs. Allowing a node to simultaneously transmit a data frame while in receive mode, can theoretically...
Orchestrating Data Collection and Computation in Green IoT NetworksJunfei Zhan, Tengjiao He, Kwan-Wu Chin, Benyu Chen, Fei Song2026-05-22下载Future Internet of things (IoT) networks will host applications that involve data collection and computation tasks on one or more servers. To this end, this paper proposes the first mixed integer line...
Combined Radar and Magnetometer Sensor Network with LoRa-Mediated Awareness for Wildlife-Vehicle Collision Prevention: A Monte Carlo AnalysisSergii Makovetskyi, Lars Thomsen2026-05-22下载Wildlife-vehicle collisions (WVCs) cause approximately 570 human fatalities in Canada per 20-year cohort, with Alberta accounting for 22% of these and incurring an estimated CAD $300,000 per day in di...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Cost-Effective Model Evaluation with Meta-LearningTrinh Pham, Viet Huynh, Hongzhi Yin, Quoc Viet Hung Nguyen, Thanh Tam Nguyen2026-05-22下载The rapid growth of machine learning has produced an ever-expanding ecosystem of models, making it increasingly challenging to verify the reliability of newly released models on unseen, unlabeled data...

基于 VitePress 构建 · 使用本地搜索查找论文