Skip to content

2026-05-08 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
AccelSync: Verifying Synchronization Coverage in Accelerator Pipeline ProgramsHangcheng An, Rui Wang, Depei Qian2026-05-08下载AI accelerator operators are compiled into multi-stage pipeline programs where DMA, vector, matrix, and scalar units execute concurrently on shared on-chip buffers.
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System ModelingYinrong Li, Zexin Fu, Yichao Zhang, Germain Haugou, Chi Zhang, Marco Bertuletti, Bowen Wang, Luca Benini2026-05-08下载Modern large language model workloads put increasing demands on parallel compute capability and on-chip memory capacity, while also stressing fine-grained data movement and synchronization.
Post-Moore Technologies for Plasma Simulation: A Community RoadmapLuca Pennati, Erik M. Åsgrim, Jeremy J. Williams, Stefan Costea, David Tskhakaya, Leon Kos, Ales Podolnik, Yi Ju, Tapish Narwal, Julian Lenz, Michael Bussmann, Urs Ganse, Minna Palmroth, Kallia Chronaki, Vassilis Papaefstathiou, Etienne Renault, Felix Jung, Martin Schulz, Valentin Seitz, Marta Garcia-Gasulla, Filippo Mantovani, Frank Jenko, Erwin Laure, Stefano Markidis2026-05-08下载Plasma simulations are among the most computationally demanding scientific workloads, combining high-dimensional kinetic evolution, particle-mesh coupling, field solves, and data-intensive communicati...
Graph Computation Meets Circuit Algebra: A Task-Aligned Analysis of Graph Neural Networks for Electronic Design AutomationHyunmog Kim2026-05-08下载EDA problems are graph-structured, but not all graph-structured problems call for the same GNN computation. We argue that successful GNN-for-EDA methods are those whose propagation, aggregation, and s...
Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNsMohammad Hasan Ahmadilivani, Marten Roots, Marco Restifo, Sven-Markus Loorits, Luca Di Mauro, Jaan Raik2026-05-08下载Modern Deep Learning (DL) workloads are increasingly deployed in safety-critical domains, such as automotive systems and hyperscale data centers, where transient hardware faults pose a serious threat ...
TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and ClassificationVijay Pratap Sharma, Mukul Lokhande, Ratko Pilipovic, Omkar Kokane, Santosh Kumar Vishvakarma2026-05-08下载This work presents TREA, a low-precision time-multiplexed and resource-efficient edge-AI accelerator for object detection and classification, targeting stringent area-power-latency constraints of edge...
TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI EnginesJiayi Wang, Maohua Nie, Sin-Chen Lin, C. -J. Richard Shi, Ang Li2026-05-08下载Commercial FPGAs, such as AMD Versal devices, increasingly incorporate AI engines that exploit low-precision packed-SIMD fused multiply-accumulate (FMA) to achieve proportional throughput gains.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
QUANTAS 2 An Abstract, Concrete and Byzantine SimulatorMikhail Nesterenko, Joseph Oglio2026-05-08下载We present QUANTAS 2: a new distributed algorithm simulator and quantitative performance analysis tool. We use the original QUANTAS as a foundation.
MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a ServiceTimothy Tin Long Yu, Gursimran Singh, Ge Shi, Hanieh Sadri, Yong Zhang, Zhenan Fan2026-05-08下载Reinforcement Learning from Verifiable Rewards (RLVR) has significantly improved the reasoning capabilities of large language models (LLMs), particularly in multi-turn agentic settings involving envir...
Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCPYilong Zhao, Xiaonan Nie, Kan Zhu, Shuang Ma, Zhichao Lai, Hongxiang Hao, Yang Zhou, Baris Kasikci, Ion Stoica2026-05-08下载Context parallelism (CP) has been widely adopted to support the growing context length in foundation model pretraining. However, existing designs fail to handle the large variation in sequence length ...
FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage OrchestrationZhengding Hu, Mingge Lu, Zhen Wang, Jixuan Ruan, Chang Chen, Zaifeng Pan, Yue Guan, Ruiyi Wang, Zhongkai Yu, Chao Zhang, Yufei Ding2026-05-08下载LLM-based evolution has emerged as a promising way to improve agents by refining non-parametric artifacts, but its wall-clock cost remains a major bottleneck.
Private Vertical Federated Inference for Time-SeriesLucas Fenaux, Larris Xie, Aditya Bang, Alex Zhang, Kevin Wilson, Florian Kerschbaum2026-05-08下载Institutions may benefit from collaborative inference on time-series data. In settings where privacy is necessary, multi-party computation (MPC) is a straightforward approach to providing strong guara...
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference SimulationJoon Ha Kim, Geon-Woo Kim, Anoop Rachakonda, Daehyeok Kim2026-05-08下载Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, since no single choice performs best across all wor...
FLAM: Evaluating Model Performance with Aggregatable Measures in Federated LearningFabian Stricker, Jose A. Peregrina, David Bermbach, Christian Zirpins2026-05-08下载Performance evaluation is essential for assessing the quality of machine learning (ML) models and guiding deployment decisions. In federated learning (FL), assessing the performance is challenging bec...
Stencil Computations on Cerebras Wafer-Scale EngineElia Belli, Daniele De Sensi2026-05-08下载Stencil computations are a fundamental kernel in scientific computing, critical for simulations in domains such as fluid dynamics and climate modeling.
\mathsf{VISTA}: Decentralized Machine Learning in Adversary Dominated EnvironmentsHanzaleh Akbari Nodehi, Parsa Moradi, Soheil Mohajer, Mohammad Ali Maddah-Ali2026-05-08下载Decentralized machine learning often relies on outsourcing computations, such as gradient evaluations, to untrusted worker nodes. Existing robust aggregation methods can mitigate malicious behavior un...
Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System ModelingYinrong Li, Zexin Fu, Yichao Zhang, Germain Haugou, Chi Zhang, Marco Bertuletti, Bowen Wang, Luca Benini2026-05-08下载Modern large language model workloads put increasing demands on parallel compute capability and on-chip memory capacity, while also stressing fine-grained data movement and synchronization.
A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language ModelsAjay Navilarekal Rajgopal, Nikolai Solmsdorf2026-05-08下载Large Language Models (LLMs) continue to demonstrate superior performance with increasing scale, yet training models with billions to trillions of parameters requires staggering computational resource...
Stencil Computations on Tenstorrent WormholeLorenzo Piarulli, Daniele De Sensi2026-05-08下载As investment in AI-focused accelerators grows and their deployment in supercomputing facilities expands, understanding whether these architectures can efficiently support traditional scientific kerne...
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous HardwareYan Liang, Youhe Jiang, Ran Yan, Binhang Yuan, Wei Wang, Chuan Wu2026-05-08下载Long-context training of large language models (LLMs) is commonly distributed with Context Parallelism (CP) and Head Parallelism (HP), but existing training systems largely assume homogeneous GPU mesh...
Deadline-Driven Hierarchical Agentic Resource Sharing for AI Services and RAN Functions in AI-RANHaiyuan Li, Yulei Wu, Dimitra Simeonidou2026-05-08下载AI-RAN consolidates AI services and Radio Access Network (RAN) functions onto a unified, GPU-accelerated infrastructure at the network edge. However, compute sharing between real-time RAN functions an...
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV CachingZhan Zhao, Yuxin Wang, Amelie Chi Zhou2026-05-08下载Large Language Models (LLMs) are transforming recommendation from ranking into a generative task, but industrial deployment remains limited by the high latency of processing long, personalized prompts...
UMEDA: Unified Multi-modal Efficient Data Fusion for Privacy-Preserving Graph Federated Learning via Spectral-Gated Attention and Diffusion-Based Operator AlignmentShih-Yu Lai, Hirozumi Yamaguchi, Shang-Tse Chen, Yu-Lun Liu, Bing-Yu Chen2026-05-08下载Device-free localization trains models from heterogeneous wireless and visual sensors (e.g., Wi-Fi, LiDAR) distributed across edge devices. Federated learning offers a privacy-respecting framework, bu...
MERBIT: A GPU-Based SpMV Method for Iterative WorkloadsQi Zhang, Zhengan Yao, Zhenglu Jiang, Zan-Bo Zhang2026-05-08下载Sparse Matrix-Vector Multiplication (SpMV) is the cornerstone in many iterative workloads, including large-scale graph analytics and sparse iterative solvers.
SparseRL-Sync: Lossless Weight Synchronization with ~100x Less CommunicationLucas Hu, Ranchi Zhao, Isaac Zhu, Zach Zhang, Hscos Zhang, Hugh Yin, Jason Zhao2026-05-08下载In large-scale reinforcement learning (RL) systems with decoupled Trainer-Rollout execution, the Trainer must regularly synchronize policy weights to the Rollout side to limit policy staleness.
TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and ClassificationVijay Pratap Sharma, Mukul Lokhande, Ratko Pilipovic, Omkar Kokane, Santosh Kumar Vishvakarma2026-05-08下载This work presents TREA, a low-precision time-multiplexed and resource-efficient edge-AI accelerator for object detection and classification, targeting stringent area-power-latency constraints of edge...
Resource-Element Energy Difference for Noncoherent Over-the-Air Federated LearningHao Chen, Zavareh Bozorgasl2026-05-08下载Over-the-air federated learning (OTA-FL) reduces uplink latency by exploiting waveform superposition, but conventional analog aggregation schemes typically require instantaneous channel state informat...
FATE: Future-State-Aware Scheduling for Heterogeneous LLM WorkflowsZirui Huang, Yi-Xiang Hu, Feng Wu, Xiangyang Li2026-05-08下载Large language model (LLM) applications are increasingly executed as heterogeneous multi-stage workflows rather than isolated inference calls.
Execution Envelopes: A Shared Admission Contract for Backend AI Execution RequestsKrti Tallam2026-05-08下载Enterprise AI backends increasingly admit heterogeneous execution requests across model deployment, inference, evaluation, data movement, and agentic workflows.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Graph Representation Learning Augmented Model Manipulation on Federated Fine-Tuning of LLMsHanlin Cai, Kai Li, Houtianfu Wang, Haofan Dong, Yichen Li, Falko Dressler, Ozgur B. Akan2026-05-08下载Federated fine-tuning (FFT) has emerged as a privacy-preserving paradigm for collaboratively adapting large language models (LLMs). Built upon federated learning, FFT enables distributed agents to joi...
Suitability of the Data Distribution Service for Next-Generation Ethernet-Based Agricultural Machinery NetworkingSamuel Brodie, Henri Hornburg, Daniel Ostermeier, Maksim Pavlov, Timo Oksanen2026-05-08下载The current state of the art in the agricultural industry for inter-manufacturer, plug-and-play communications is the ISO 11783 standard series, which mandates the use of 250 Kb/s CAN bus.
Deadline-Driven Hierarchical Agentic Resource Sharing for AI Services and RAN Functions in AI-RANHaiyuan Li, Yulei Wu, Dimitra Simeonidou2026-05-08下载AI-RAN consolidates AI services and Radio Access Network (RAN) functions onto a unified, GPU-accelerated infrastructure at the network edge. However, compute sharing between real-time RAN functions an...
Unconsented Sensing: A Sociotechnical Governance Framework for 6G ISACAnass Sedrati2026-05-08下载The forthcoming deployment of 6G Integrated Sensing and Communication (ISAC) will transform cellular infrastructure into pervasive, continuous environmental and biometric sensing grids.
From Map-and-Encap to BIER: Observations on Network Routing ScalabilityTianyuan Yu, Lan Wang, Beichuan Zhang, Lixia Zhang2026-05-08下载The TCP/IP protocol stack uses IP addresses for two distinct roles: identifying hosts and locating their attachment points in the network topology.

cs.PF - Performance ​

标题作者发布日期PDF摘要
CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAGPengzhou Chen, Tao Chen2026-05-08下载Retrieval-Augmented Generation (RAG) is sensitive to the vast hyperparameters of the retriever and generator, yet optimizing them using given queries is a challenging task due to the complex interacti...
FlashSVD v1.5: Making Low-Rank Transformers Inference Actually FastWenhao Wu, Zishan Shao, Kangning Cui, Jinhee Kim, Yixiao Wang, Hancheng Ye, Danyang Zhuo, Yiran Chen2026-05-08下载SVD-based Low-rank compression reduces transformer parameters and nominal FLOPs, but these savings often translate poorly into real LLM serving speedups.
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context InferenceFeiyu Yao, Zhixiong Niu, Xiaqing Li, Yongqiang Xiong, Juan Fang, Qian Wang2026-05-08下载Long-context inference increasingly operates over CPU-resident KV caches, either because decoding-time KV states exceed GPU memory capacity or because disaggregated prefill-decode systems place KV dat...
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM SystemsSiyu Wu, Yulong Ye, Zezhen Xiang, Pengzhou Chen, Gangda Xiong, Tao Chen2026-05-08下载Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperparameter optimization (HPO) for the AutoML community.

基于 VitePress 构建 · 使用本地搜索查找论文