Skip to content

2026-09-21 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
GRADE-RTL: Evaluating LLM-Generated RTL Beyond CompilationHepziba Susan, Shivaranjani G. R., Malik Imran, Muhammad Rashid, Sumathi Gokulanathan, Zain Ul Abideen2026-09-21下载Large language models (LLMs) can generate register-transfer-level (RTL) code from natural-language specifications, but compilation alone does not establish structural completeness, functional correctn...
Toward Multi-kW Power Delivery Methodologies for Advanced 3D Heterogeneous IntegrationPeiyi Yue, Hangyu Zhang, Ratul Das, Ramesh Harjani, Sachin S. Sapatnekar2026-09-21下载The demands of modern applications require the construction of ever more complex integrated systems, with AI applications in particular serving as a significant driver for increased system size, poten...
SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled ArchitectureGabriele Tombesi, William Baisi, Je Yang, Elisavet Lydia Alvanaki, Kevin Lee, Michael Lippe, Biruk Seyoum, Luca P. Carloni2026-09-21下载LLM inference on edge devices is constrained by computational and memory resources, making efficient autoregressive decoding challenging. Speculative decoding alleviates this bottleneck by generating ...
NPU Accelerator: Quantized Real-Time Vehicle Detection on PYNQ-Z1 Using FINNDaniel Gutierrez, Antonio Cuesta, Jorge Fe, Bruno Gutierrez, Rashed Al Koutayni2026-09-21下载This paper presents the design, optimization, implementation, and on-board validation of a neural processing unit (NPU) accelerator for real-time vehicle detection on the resource-constrained Xilinx Z...
AWE: Adaptive Weight Encoding for Exact Integer Matrix Products with Fewer GEMMs on FP4 Tensor CoresShun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri2026-09-21下载Emulation of high-accuracy floating-point matrix multiplication, as in the Ozaki scheme, splits the inputs into low-precision components and multiplies them pairwise.
ScaleMPA: Rethinking Scalable RRT* Acceleration With a Grid-Native RepresentationZilong Wang, Yuzhou Chen, Xinyue He, Chen Zhang, Guanghui He2026-09-21下载Real-time motion planning remains challenging in large and high-dimensional environments. Prior acceleration of RRT* follows tree-centric state organization, which reduces per-query cost but preserves...
Circuit-Architecture-Training Co-Design with Regenerative-SA Similarity Sensing for Aggressive SAR Skipping in Analog Compute-in-MemoryYufei Liu, Shuang Liu, Junjie Wang2026-09-21下载This work presents a circuit-architecture-training co-design framework that exploits sense-amplifier (SA) regeneration to detect analog-output similarity and reduce SAR comparisons in compute-in-memor...
Dissecting How Die Scaling Breaks GPU Fine-grained SchedulingXiaoze Fan, Jianhao Wang, Weihao Cui, Han Zhao, Zhuobin Huang, Yangjie Zhou, Yuxian Qiu, Shixuan Sun, Bingsheng He, Quan Chen, Minyi Guo2026-09-21下载Modern GPUs are no longer physically symmetric. Die scaling leads to both manufacturing-driven floorsweeping and cache and memory partitioning.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Adaptive and Cost-Efficient Joint Scheduling of UAV Routes and Analytics with Transit-Borne FogSuman Raj, Arindam Khanda, Gagana M D, Yogesh Simmhan, Sajal K. Das2026-09-21下载Unmanned Aerial Vehicles (UAVs) performing deadline-bound analytics over large rural areas cannot reliably offload workloads to sparse cellular base stations.
Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future DirectionsNiloofar Gholipour, Marcos Assuncao, Gursimran Singh, Timothy Yu, Rajkumar Buyya, Julien Gascon-Samson, Zhenan Fan, Yong Zhang, Xiaojie Xu, Yaqiang Yao, Xiaolong Bai2026-09-21下载Reasoning-oriented reinforcement learning enables large language models to solve mathematical, coding, and other multi-step tasks, but shifts a substantial portion of the training cost to rollout, whe...
Fast Recovery for LLM Serving via Decoupled Device Memory Lifetime in DynamoSchwinn Saereesitthipitak, Mohammed Abdulwahhab, Hannah Zhang, Dan Feigin, Neelay Shah, Maksim Khadkevich, Itay Neeman, Vikram Sharma Mailthody, Wen-mei W. Hwu2026-09-21下载Large language model (LLM) inference replicas run across tightly coupled GPUs and serve traffic continuously for weeks. Hardware and software failures are therefore inevitable, and one worker failure ...
ZeroGate: Trust-Preserving Fast Paths for Governed AI Agent RuntimesZexun Wang2026-09-21下载Moving authorization earlier can shorten an agent's dispatch boundary without removing authorization work. It can also admit an action whose payload, authority, or relevant state has changed.
WeightBridge: An Efficient Weight Transfer Library for Reinforcement LearningXuanlin Jiang, Samuel Hsia, Michael Kuchnik, Zachary DeVito, Minlan Yu, Carole-Jean Wu2026-09-21下载Weight transfer - the propagation of updated parameters from trainers to rollout generators - is becoming an important performance bottleneck in reinforcement learning (RL) systems for LLMs.
Data center cooling choices shift water impacts across the grid: An integrated water-energy model for sustainable data center developmentGarrett Alston, Nancy Love, Rabab Haider2026-09-21下载Data centers are being developed at an unprecedented pace, yet their energy and water impacts, and the spatial and temporal distribution of these impacts, remain poorly characterized.
Cloud, Edge, or Split? Profiling Onboard and Split Vision-Language Model Deployment for Drone AIZoha Azimi, Reza Farahani, Schahram Dustdar, Christian Timmerer2026-09-21下载Vision-Language Models (VLMs) enable edge devices like unmanned aerial vehicles (UAVs) to interpret visual observations and reason about complex environments using natural-language instructions.
Who Pays for the KV Cache? Attributing Shared AI Inference Spend Across Kubernetes and LLM Provider BillsTimothy Urista2026-09-21下载Organizations pay for AI through disconnected ledgers: Kubernetes allocations for self-hosted inference, gateway logs, and per-token bills from API providers.
SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled ArchitectureGabriele Tombesi, William Baisi, Je Yang, Elisavet Lydia Alvanaki, Kevin Lee, Michael Lippe, Biruk Seyoum, Luca P. Carloni2026-09-21下载LLM inference on edge devices is constrained by computational and memory resources, making efficient autoregressive decoding challenging. Speculative decoding alleviates this bottleneck by generating ...
Tiga: Compiling Graph Message Passing at ScaleMingyuan Chi2026-09-21下载Graph message passing offers a common way to express learning algorithms, physical simulations, and numerical solvers. Efficient execution depends on interaction structure and data movement, which can...
Mitigating Front-Running Attacks through Fair and Resilient Transaction DisseminationWassim Yahyaoui, Joachim Bruneau-Queyreix, Jérémie Decouchant, Marcus Völp2026-09-21下载In modern blockchains, efficient, fair, and faulttolerant information dissemination is critical for performance and security. Several stages of the transaction lifecycle are affected, from the creatio...
Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI InferenceMingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, Wenqi Cui2026-09-21下载Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption.
Bridging the Vendor Gap: Enabling AMD GPU Support for Awkward Array via ROCm/HIP for the HL-LHC EraIanna Osborne, Maxym Naumchyk, Tai Sakuma, Andres Rios-Tascon, Peter Elmer2026-09-21下载The High-Luminosity LHC (HL-LHC) will demand order-of-magnitude gains in analysis throughput, and increasingly those gains must come from GPUs that are not made by a single vendor.
Conduit: An Experience Data Plane for Distributed Reinforcement LearningSitong Zhang, Tuo Shi, Mario Di Francesco, Zeke Wang, Bo Zhao2026-09-21下载Distributed reinforcement learning (RL) scales training by parallelizing actors and learners around an Experience Buffer. As RL workloads grow, however, the buffer becomes more than a replay queue: it...
Categorical Message Passing Language (CaMPL): Syntax and SemanticsRobin Cockett, Daniel Kiyoshi Hashimoto, Alexanna Little Berg, Priyaa Varshinee Srinivasan2026-09-21下载We introduce a novel functional-style concurrent programming language called Categorical Message Passing Language (CaMPL) which is designed using the mathematics of linear actegories.
SkelOT: Reusing AOT Compilation Across EVM Contract FamiliesSipeng Xie, Qianhong Wu, Minghang Li, Qin Wang, Zhipeng Wang, Bo Qin2026-09-21下载Ahead-of-time (AOT) compilers (e.g., revmc, evmone, and DTVM) for the Ethereum Virtual Machine (EVM) reuse compilation artifacts at contract-code-hash granularity.
Toward GPU-Resident Climate Models: A Feasibility Study on Lossy Compression for the Spherical Harmonic Transform's Communication BottleneckLorenzo Breschi, Flavio Vella2026-09-21下载Operational pseudospectral atmospheric models such as the ECMWF Integrated Forecasting System (IFS) run today almost exclusively on CPUs; GPU ports are under active development but not yet used in pro...
Dissecting How Die Scaling Breaks GPU Fine-grained SchedulingXiaoze Fan, Jianhao Wang, Weihao Cui, Han Zhao, Zhuobin Huang, Yangjie Zhou, Yuxian Qiu, Shixuan Sun, Bingsheng He, Quan Chen, Minyi Guo2026-09-21下载Modern GPUs are no longer physically symmetric. Die scaling leads to both manufacturing-driven floorsweeping and cache and memory partitioning.
A principled approach for energy-efficient training via phase-aware GPU frequency tuningMiguel Braga, Júlio Pinto, Rahma Nouaji, Olivier Michaud, Bettina Kemme, Oana Balmau, Cláudia Brito, Ricardo Macedo2026-09-21下载Modern AI model training imposes unprecedented computational demands, making it a key contributor to datacenter energy consumption. Yet a significant fraction of the energy consumed during training do...
MCP-GRANITE Benchmark: GRANularity Interface TEsting for MCP-Based LLM AgentsDemetris Paschalides, Moysis Symeonides, George Pallis, Marios D. Dikaiakos2026-09-21下载As LLM agents increasingly interact with external tools through standardized protocols such as MCP, tool-interface design becomes a critical yet underexplored factor.
Byzantine Causal Reliable Broadcast (BCRB) with Constant-Size Message MetadataPurv Patel, Ajay D. Kshemkalyani2026-09-21下载Asynchronous Byzantine Reliable Broadcast (BRB) is a fundamental primitive that guarantees agreement and validity in distributed systems subject to Byzantine faults, but it lacks ordering guarantees.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
WeightBridge: An Efficient Weight Transfer Library for Reinforcement LearningXuanlin Jiang, Samuel Hsia, Michael Kuchnik, Zachary DeVito, Minlan Yu, Carole-Jean Wu2026-09-21下载Weight transfer - the propagation of updated parameters from trainers to rollout generators - is becoming an important performance bottleneck in reinforcement learning (RL) systems for LLMs.
Tetris: Circuit Scheduling for Rearrangeably Non-Blocking Photonic InterconnectsEliezer Amponsah, Deeksha P Rao, Vamsi Addanki2026-09-21下载Reconfigurable photonic interconnects are emerging as a promising communication architecture for next-generation distributed computing. Yet, most circuit schedulers are designed around an idealized vi...
Spreading Factor Assignment Strategy for Coverage and Capacity Flexible TradeoffLuiz Filho, Alvaro Medeiros, Jéssika Silva, Vicente Angelo de Sousa Junior, Níbia Bezerra2026-09-21下载LoRa is a physical layer technology with the ability to connect multiple devices in a wide area of coverage, with low power consumption and with interference robustness.
Objective Video Quality Assessment in FWA-Based Over-the-Top Content Delivery Across Open Source 5G NetworksNelson Ion, FlÁvio Silva, Paulo Silva, Ricardo Silva, Mathews Lima, Daniel Luna, Antonio Campos, Augusto Neto, Vicente Sousa2026-09-21下载This paper presents a comprehensive experimental video quality assessment in 5G-based Fixed Wireless Access (FWA) networks, leveraging a real open source 5G network testbed.
Case for Vehicle-Edge Collaborative Multi-Sensor Data Fusion for Autonomous Vehicle TeleoperationQixin Zhang, Ajay Kumar Gurumadaiah, Wei Ye, Eman Ramadan, Zhi-Li Zhang2026-09-21下载Teleoperation provides a critical safety fallback when autonomous vehicles (AVs) encounter scenarios that are outside their operational design domain.
Multi-Agent Video Prediction: Self-Correcting Conditional Frames for Dynamic Scene ForecastingQixin Zhang, Ajay Kumar, Zhi-Li Zhang2026-09-21下载Transmission latency significantly degrades user quality of experience in real-time interactive perception systems. In remote driving, maintaining reliable visual feedback is critical for safe operati...
Impact of Data Compression on Downstream AI Tasks: A Study using Teleoperated Driving over 5GQixin Zhang, Steven Sleder, Xinyue Hu, Faaiq Bilal, Wei Ye, Zhi-Li Zhang2026-09-21下载Teleoperation, such as remote driving, is considered as a key use case of 5G and Next-Generation (NextG) networks. In this context, robots, autonomous vehicles, or other autonomous agents transmit sen...
ALARM: Adaptive Layer-Aware Resource Management for Power-Efficient vRANsAli Srour, Farzad Veisi, Sami Taktak, Vania Conan2026-09-21下载The transition to virtualized Radio Access Networks (vRAN) enables dynamic power control through fine-grained CPU resource management. However, existing approaches treat gNB as a monolithic entity, fa...
Trust in Edge-Enabled IoT Security: Features, Challenges and Research DirectionsEsin Ece Aydın, Şerif Bahtiyar, Gürkan Gür2026-09-21下载Providing autonomous intelligence, pervasive connectivity and usability to human life and industry has led to the emergence of the Internet of Things (IoT).
5G-Shark: A Network Security Auditor for 5G Subscriber Privacy and Unauthenticated Signalling ResilienceOscar Lasierra, Gines Garcia-Aviles, Antonio Skarmeta, Xavier Costa-Pérez2026-09-21下载The fifth generation of mobile networks was standardised with an explicit mandate to close long-standing privacy and security gaps, mandating the concealment of the subscriber's permanent identity, re...
Joint Energy Efficiency and Fairness Optimization for D2D Communications in Aerial-Ground Integrated Heterogeneous NetworksChuan-Chi Lai, Ang-Hsun Tsai, Shang-Long Wu2026-09-21下载This study investigates an Aerial-Ground Integrated Heterogeneous network (AGIHN) architecture that combines terrestrial macro base stations and unmanned aerial vehicles (UAVs) serving as aerial base ...
rApp/xApp Attestation: A New Security Use Case for O-RANHamed Alimohammadi, Burcu Şahin, Arda Akman, Chuan Heng Foh, Periklis Chatzimisios, Mohammad Shojafar2026-09-21下载The disaggregation and softwarization introduced by the Open Radio Access Network (O-RAN) architecture enable multi-vendor innovation but also expose the RAN Intelligent Controller (RIC) ecosystem to ...
Learning to Maximize Energy Efficiency in 6G in-X SubnetworksRamoni Adeogun2026-09-21下载This paper investigates energy-efficient power control in 6G in-X subnetworks. We consider a graph neural network (GNN) framework that captures inter-subnetwork interference and the underlying wireles...
Zero-Knowledge Remote Adversarial Attack against Wi-Fi-based Human Activity Recognition for Privacy ProtectionByungjun Kim, Amogh Panchagatti, Peter Gerstoft, Xinyu Zhang, Minsung Kim2026-09-21下载The growing capability of Wi-Fi devices to identify human activities using channel state information (CSI) raises privacy concerns. To counter this threat, we propose GRAW, an adversary system, acting...
Explanation-Guided Federated Deep Reinforcement Learning for Joint Resource Allocation and Scheduling in 6G in-X SubnetworksRamoni Adeogun2026-09-21下载Sixth-generation (6G) wireless systems are envisioned as networks of networks, integrating diverse in-X subnetworks that provide localized, high-performance connectivity.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Who Pays for the KV Cache? Attributing Shared AI Inference Spend Across Kubernetes and LLM Provider BillsTimothy Urista2026-09-21下载Organizations pay for AI through disconnected ledgers: Kubernetes allocations for self-hosted inference, gateway logs, and per-token bills from API providers.
Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI InferenceMingyuan Yan, Haiyu Wang, Linxuan Biao, H. Jonathan Chao, Sai Qian Zhang, Wenqi Cui2026-09-21下载Power availability increasingly constrains the operation of AI inference fleets, creating a need for provisioning methods that jointly consider serving capacity and power consumption.

基于 VitePress 构建 · 使用本地搜索查找论文