2026-05-08
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AccelSync: Verifying Synchronization Coverage in Accelerator Pipeline Programs | Hangcheng An, Rui Wang, Depei Qian | 2026-05-08 | 下载 | AI accelerator operators are compiled into multi-stage pipeline programs where DMA, vector, matrix, and scalar units execute concurrently on shared on-chip buffers. |
| Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling | Yinrong Li, Zexin Fu, Yichao Zhang, Germain Haugou, Chi Zhang, Marco Bertuletti, Bowen Wang, Luca Benini | 2026-05-08 | 下载 | Modern large language model workloads put increasing demands on parallel compute capability and on-chip memory capacity, while also stressing fine-grained data movement and synchronization. |
| Post-Moore Technologies for Plasma Simulation: A Community Roadmap | Luca Pennati, Erik M. Åsgrim, Jeremy J. Williams, Stefan Costea, David Tskhakaya, Leon Kos, Ales Podolnik, Yi Ju, Tapish Narwal, Julian Lenz, Michael Bussmann, Urs Ganse, Minna Palmroth, Kallia Chronaki, Vassilis Papaefstathiou, Etienne Renault, Felix Jung, Martin Schulz, Valentin Seitz, Marta Garcia-Gasulla, Filippo Mantovani, Frank Jenko, Erwin Laure, Stefano Markidis | 2026-05-08 | 下载 | Plasma simulations are among the most computationally demanding scientific workloads, combining high-dimensional kinetic evolution, particle-mesh coupling, field solves, and data-intensive communicati... |
| Graph Computation Meets Circuit Algebra: A Task-Aligned Analysis of Graph Neural Networks for Electronic Design Automation | Hyunmog Kim | 2026-05-08 | 下载 | EDA problems are graph-structured, but not all graph-structured problems call for the same GNN computation. We argue that successful GNN-for-EDA methods are those whose propagation, aggregation, and s... |
| Effective and Memory-Efficient Alternatives to ECC for Reliable Large-Scale DNNs | Mohammad Hasan Ahmadilivani, Marten Roots, Marco Restifo, Sven-Markus Loorits, Luca Di Mauro, Jaan Raik | 2026-05-08 | 下载 | Modern Deep Learning (DL) workloads are increasingly deployed in safety-critical domains, such as automotive systems and hyperscale data centers, where transient hardware faults pose a serious threat ... |
| TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and Classification | Vijay Pratap Sharma, Mukul Lokhande, Ratko Pilipovic, Omkar Kokane, Santosh Kumar Vishvakarma | 2026-05-08 | 下载 | This work presents TREA, a low-precision time-multiplexed and resource-efficient edge-AI accelerator for object detection and classification, targeting stringent area-power-latency constraints of edge... |
| TransDot: An Area-efficient Reconfigurable Floating-Point Unit for Trans-Precision Dot-Product Accumulation for FPGA AI Engines | Jiayi Wang, Maohua Nie, Sin-Chen Lin, C. -J. Richard Shi, Ang Li | 2026-05-08 | 下载 | Commercial FPGAs, such as AMD Versal devices, increasingly incorporate AI engines that exploit low-precision packed-SIMD fused multiply-accumulate (FMA) to achieve proportional throughput gains. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| QUANTAS 2 An Abstract, Concrete and Byzantine Simulator | Mikhail Nesterenko, Joseph Oglio | 2026-05-08 | 下载 | We present QUANTAS 2: a new distributed algorithm simulator and quantitative performance analysis tool. We use the original QUANTAS as a foundation. |
| MARLaaS: Multi-Tenant Asynchronous Reinforcement Learning as a Service | Timothy Tin Long Yu, Gursimran Singh, Ge Shi, Hanieh Sadri, Yong Zhang, Zhenan Fan | 2026-05-08 | 下载 | Reinforcement Learning from Verifiable Rewards (RLVR) has significantly improved the reasoning capabilities of large language models (LLMs), particularly in multi-turn agentic settings involving envir... |
| Unleashing Scalable Context Parallelism for Foundation Models Pre-Training via FCP | Yilong Zhao, Xiaonan Nie, Kan Zhu, Shuang Ma, Zhichao Lai, Hongxiang Hao, Yang Zhou, Baris Kasikci, Ion Stoica | 2026-05-08 | 下载 | Context parallelism (CP) has been widely adopted to support the growing context length in foundation model pretraining. However, existing designs fail to handle the large variation in sequence length ... |
| FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestration | Zhengding Hu, Mingge Lu, Zhen Wang, Jixuan Ruan, Chang Chen, Zaifeng Pan, Yue Guan, Ruiyi Wang, Zhongkai Yu, Chao Zhang, Yufei Ding | 2026-05-08 | 下载 | LLM-based evolution has emerged as a promising way to improve agents by refining non-parametric artifacts, but its wall-clock cost remains a major bottleneck. |
| Private Vertical Federated Inference for Time-Series | Lucas Fenaux, Larris Xie, Aditya Bang, Alex Zhang, Kevin Wilson, Florian Kerschbaum | 2026-05-08 | 下载 | Institutions may benefit from collaborative inference on time-series data. In settings where privacy is necessary, multi-party computation (MPC) is a straightforward approach to providing strong guara... |
| Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation | Joon Ha Kim, Geon-Woo Kim, Anoop Rachakonda, Daehyeok Kim | 2026-05-08 | 下载 | Selecting the optimal LLM inference configuration requires evaluation across hardware, serving engines, attention backends, and model architectures, since no single choice performs best across all wor... |
| FLAM: Evaluating Model Performance with Aggregatable Measures in Federated Learning | Fabian Stricker, Jose A. Peregrina, David Bermbach, Christian Zirpins | 2026-05-08 | 下载 | Performance evaluation is essential for assessing the quality of machine learning (ML) models and guiding deployment decisions. In federated learning (FL), assessing the performance is challenging bec... |
| Stencil Computations on Cerebras Wafer-Scale Engine | Elia Belli, Daniele De Sensi | 2026-05-08 | 下载 | Stencil computations are a fundamental kernel in scientific computing, critical for simulations in domains such as fluid dynamics and climate modeling. |
| \mathsf{VISTA}: Decentralized Machine Learning in Adversary Dominated Environments | Hanzaleh Akbari Nodehi, Parsa Moradi, Soheil Mohajer, Mohammad Ali Maddah-Ali | 2026-05-08 | 下载 | Decentralized machine learning often relies on outsourcing computations, such as gradient evaluations, to untrusted worker nodes. Existing robust aggregation methods can mitigate malicious behavior un... |
| Accelerating Precise End-to-End Simulation: Latency-Sensitive Many-core System Modeling | Yinrong Li, Zexin Fu, Yichao Zhang, Germain Haugou, Chi Zhang, Marco Bertuletti, Bowen Wang, Luca Benini | 2026-05-08 | 下载 | Modern large language model workloads put increasing demands on parallel compute capability and on-chip memory capacity, while also stressing fine-grained data movement and synchronization. |
| A Scalable Recipe on SuperMUC-NG Phase 2: Efficient Large-Scale Training of Language Models | Ajay Navilarekal Rajgopal, Nikolai Solmsdorf | 2026-05-08 | 下载 | Large Language Models (LLMs) continue to demonstrate superior performance with increasing scale, yet training models with billions to trillions of parameters requires staggering computational resource... |
| Stencil Computations on Tenstorrent Wormhole | Lorenzo Piarulli, Daniele De Sensi | 2026-05-08 | 下载 | As investment in AI-focused accelerators grows and their deployment in supercomputing facilities expands, understanding whether these architectures can efficiently support traditional scientific kerne... |
| HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware | Yan Liang, Youhe Jiang, Ran Yan, Binhang Yuan, Wei Wang, Chuan Wu | 2026-05-08 | 下载 | Long-context training of large language models (LLMs) is commonly distributed with Context Parallelism (CP) and Head Parallelism (HP), but existing training systems largely assume homogeneous GPU mesh... |
| Deadline-Driven Hierarchical Agentic Resource Sharing for AI Services and RAN Functions in AI-RAN | Haiyuan Li, Yulei Wu, Dimitra Simeonidou | 2026-05-08 | 下载 | AI-RAN consolidates AI services and Radio Access Network (RAN) functions onto a unified, GPU-accelerated infrastructure at the network edge. However, compute sharing between real-time RAN functions an... |
| RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching | Zhan Zhao, Yuxin Wang, Amelie Chi Zhou | 2026-05-08 | 下载 | Large Language Models (LLMs) are transforming recommendation from ranking into a generative task, but industrial deployment remains limited by the high latency of processing long, personalized prompts... |
| UMEDA: Unified Multi-modal Efficient Data Fusion for Privacy-Preserving Graph Federated Learning via Spectral-Gated Attention and Diffusion-Based Operator Alignment | Shih-Yu Lai, Hirozumi Yamaguchi, Shang-Tse Chen, Yu-Lun Liu, Bing-Yu Chen | 2026-05-08 | 下载 | Device-free localization trains models from heterogeneous wireless and visual sensors (e.g., Wi-Fi, LiDAR) distributed across edge devices. Federated learning offers a privacy-respecting framework, bu... |
| MERBIT: A GPU-Based SpMV Method for Iterative Workloads | Qi Zhang, Zhengan Yao, Zhenglu Jiang, Zan-Bo Zhang | 2026-05-08 | 下载 | Sparse Matrix-Vector Multiplication (SpMV) is the cornerstone in many iterative workloads, including large-scale graph analytics and sparse iterative solvers. |
| SparseRL-Sync: Lossless Weight Synchronization with ~100x Less Communication | Lucas Hu, Ranchi Zhao, Isaac Zhu, Zach Zhang, Hscos Zhang, Hugh Yin, Jason Zhao | 2026-05-08 | 下载 | In large-scale reinforcement learning (RL) systems with decoupled Trainer-Rollout execution, the Trainer must regularly synchronize policy weights to the Rollout side to limit policy staleness. |
| TREA: Low-precision Time-Multiplexed, Resource-Efficient Edge Accelerator for Object Detection and Classification | Vijay Pratap Sharma, Mukul Lokhande, Ratko Pilipovic, Omkar Kokane, Santosh Kumar Vishvakarma | 2026-05-08 | 下载 | This work presents TREA, a low-precision time-multiplexed and resource-efficient edge-AI accelerator for object detection and classification, targeting stringent area-power-latency constraints of edge... |
| Resource-Element Energy Difference for Noncoherent Over-the-Air Federated Learning | Hao Chen, Zavareh Bozorgasl | 2026-05-08 | 下载 | Over-the-air federated learning (OTA-FL) reduces uplink latency by exploiting waveform superposition, but conventional analog aggregation schemes typically require instantaneous channel state informat... |
| FATE: Future-State-Aware Scheduling for Heterogeneous LLM Workflows | Zirui Huang, Yi-Xiang Hu, Feng Wu, Xiangyang Li | 2026-05-08 | 下载 | Large language model (LLM) applications are increasingly executed as heterogeneous multi-stage workflows rather than isolated inference calls. |
| Execution Envelopes: A Shared Admission Contract for Backend AI Execution Requests | Krti Tallam | 2026-05-08 | 下载 | Enterprise AI backends increasingly admit heterogeneous execution requests across model deployment, inference, evaluation, data movement, and agentic workflows. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Graph Representation Learning Augmented Model Manipulation on Federated Fine-Tuning of LLMs | Hanlin Cai, Kai Li, Houtianfu Wang, Haofan Dong, Yichen Li, Falko Dressler, Ozgur B. Akan | 2026-05-08 | 下载 | Federated fine-tuning (FFT) has emerged as a privacy-preserving paradigm for collaboratively adapting large language models (LLMs). Built upon federated learning, FFT enables distributed agents to joi... |
| Suitability of the Data Distribution Service for Next-Generation Ethernet-Based Agricultural Machinery Networking | Samuel Brodie, Henri Hornburg, Daniel Ostermeier, Maksim Pavlov, Timo Oksanen | 2026-05-08 | 下载 | The current state of the art in the agricultural industry for inter-manufacturer, plug-and-play communications is the ISO 11783 standard series, which mandates the use of 250 Kb/s CAN bus. |
| Deadline-Driven Hierarchical Agentic Resource Sharing for AI Services and RAN Functions in AI-RAN | Haiyuan Li, Yulei Wu, Dimitra Simeonidou | 2026-05-08 | 下载 | AI-RAN consolidates AI services and Radio Access Network (RAN) functions onto a unified, GPU-accelerated infrastructure at the network edge. However, compute sharing between real-time RAN functions an... |
| Unconsented Sensing: A Sociotechnical Governance Framework for 6G ISAC | Anass Sedrati | 2026-05-08 | 下载 | The forthcoming deployment of 6G Integrated Sensing and Communication (ISAC) will transform cellular infrastructure into pervasive, continuous environmental and biometric sensing grids. |
| From Map-and-Encap to BIER: Observations on Network Routing Scalability | Tianyuan Yu, Lan Wang, Beichuan Zhang, Lixia Zhang | 2026-05-08 | 下载 | The TCP/IP protocol stack uses IP addresses for two distinct roles: identifying hosts and locating their attachment points in the network topology. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG | Pengzhou Chen, Tao Chen | 2026-05-08 | 下载 | Retrieval-Augmented Generation (RAG) is sensitive to the vast hyperparameters of the retriever and generator, yet optimizing them using given queries is a challenging task due to the complex interacti... |
| FlashSVD v1.5: Making Low-Rank Transformers Inference Actually Fast | Wenhao Wu, Zishan Shao, Kangning Cui, Jinhee Kim, Yixiao Wang, Hancheng Ye, Danyang Zhuo, Yiran Chen | 2026-05-08 | 下载 | SVD-based Low-rank compression reduces transformer parameters and nominal FLOPs, but these savings often translate poorly into real LLM serving speedups. |
| An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference | Feiyu Yao, Zhixiong Niu, Xiaqing Li, Yongqiang Xiong, Juan Fang, Qian Wang | 2026-05-08 | 下载 | Long-context inference increasingly operates over CPU-resident KV caches, either because decoding-time KV states exceed GPU memory capacity or because disaggregated prefill-decode systems place KV dat... |
| LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems | Siyu Wu, Yulong Ye, Zezhen Xiang, Pengzhou Chen, Gangda Xiong, Tao Chen | 2026-05-08 | 下载 | Large Language Model (LLM) systems have been the frontier of AI in many application domains, leading to new challenges and opportunities for hyperparameter optimization (HPO) for the AutoML community. |