Skip to content

2026-06-09 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUsWesley Pang, Gregory Hyegang Jun, Feiyang Liu, Deming Chen2026-06-09下载With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets.
Revisiting "Cooler is Better": ITD-Aware Per-CPU Thermal Optimization for Sustainable Data Center OperationJason Crop, Hayden Moore, Sudeep Pasricha2026-06-09下载As data center energy demand approaches grid-level constraints, optimizing conventional server infrastructure is essential for sustainable growth.
Defeat the Heap: Zero-Copy Data Movement in AXI4MLIRElam Cohavi, Nicolas Bohm Agostini, Jude Haris, Antonino Tumeo, David Kaeli, José Cano2026-06-09下载As custom hardware accelerators become increasingly central to machine learning workloads, efficient data transfer is critical for maximizing accelerator performance on linear algebra kernels.
Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDAVinamra Sharma, Xingjian Fu, Jude Haris, José Cano2026-06-09下载Designing FPGA-based accelerators for modern artificial intelligence workloads requires exploring a large and complex hardware design space that involves architectural parameters, data flow strategies...
Coset Ensemble Decoder for Quantum Error Correction with Algorithm-Hardware Co-DesignShuang Liang, Jubo Xu, Giulio Bassanino, Qianzhou Wang, Yidong Zhou, Yuncheng Lu, Zhiwen Mo, Paul H. J. Kelly, Bo Yuan, Wayne Luk, Hongxiang Fan2026-06-09下载Reliable large-scale quantum computation relies on fault-tolerant architectures, where quantum error correction (QEC) continuously extracts and decodes error syndromes in real time.
Arithmetic Packing on Wide Integer Datapaths in DSP Primitives of Modern FPGA DevicesTitus Bornträger, Shane Fleming, Philipp Holzinger, Dietmar Fey, Michaela Blott, Thomas B. Preußer2026-06-09下载Deep Neural Networks increasingly employ low-precision quantization to reduce computational requirements. While FPGAs are well suited for workloads with heterogeneous precisions, their dedicated digit...
A 185 TOPS/W/mm2 Bayesian Inference Engine with 640 aJ Write-Free FeFET GRNG for Uncertainty-Aware Aerial Search and RescueZephan M. Enciso, Xuezhong Niu, Xingtian Wang, Mohammad Mehdi Sharifi, Subhasish Mukherjee, Likai Pei, Halid Mulaosmanovic, Stefan Duenkel, Sven Beyer, Michael Niemier, Kai Ni, Ningyuan Cao2026-06-09下载Aerial search and rescue missions require fast and reliable victim detection under uncertain and rapidly changing environments. Deterministic deep learning models can produce overconfident false posit...
A Hybrid Edge-Cloud Architecture for Low-Latency Entitlement Verification in Resource-Constrained DevicesPravin Nagare, Aditya Sabbineni, Devendra Dahiphale, Faiz Gouri, Pratik Thantharate2026-06-09下载As digital media consumption shifts toward large-scale Over-the-Top (OTT) platforms, the efficiency of the control plane, specifically entitlement and identity verification, has become a critical fact...
Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-based AcceleratorsChenguang Zhang, Yuanpeng Zhang, Chenhao Xue, Yihan Yin, Chen Zhang, Guangyu Sun2026-06-09下载Level-4+ autonomous driving systems (ADS) must run dozens of heterogeneous deep neural networks (DNNs) as end-to-end (E2E) pipelines under a strict latency constraint (<=100 ms), even as execution ti...
LLM-Guided Neural Architecture Search for Robust Co-Design of Physical Neural NetworksTyler King, Timothee Leleu2026-06-09下载Deploying neural networks on unconventional hardware demands architectures that co-optimize task accuracy and platform-specific constraints such as energy cost, physical non-idealities, and numerical ...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
A Scalable PyTorch Abstraction for Multi-GPU Gaussian SplattingMatthew Cong, Francis Williams, Jonathan Swartz, Mark Harris, Sanja Fidler, Ken Museth2026-06-09下载Gaussian splatting methods have become increasingly popular for neural reconstruction of the real world. However, they are often limited in scale and resolution due to compute and memory constraints.
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUsWesley Pang, Gregory Hyegang Jun, Feiyang Liu, Deming Chen2026-06-09下载With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets.
An Ocean Model Ported by a Large Language Model: Experience and Lessons from FESOM2 (Fortran to C to C++/Kokkos)Nikolay V. Koldunov, Suvarchal K. Cheedela, Sergey Danilov, Dmitry Sidorenko, Sebastian Beyer, Thomas Jung2026-06-09下载Large language models (LLMs) can translate and modify source code, and have been shown to do so for codes of different complexity. Whether they can port a complete, production geophysical model to a d...
Piper: A Programmable Distributed Training SystemMegan Frisella, Shubham Tiwari, Andy Ruan, Yi Pan, Parker Gustafson, Mat Jacob, Gilbert Bernstein, Stephanie Wang2026-06-09下载Large-scale model training increasingly relies on composing multiple parallelism strategies, such as data, pipeline, and expert parallelism, together with memory-saving optimizations like ZeRO.
Revisiting "Cooler is Better": ITD-Aware Per-CPU Thermal Optimization for Sustainable Data Center OperationJason Crop, Hayden Moore, Sudeep Pasricha2026-06-09下载As data center energy demand approaches grid-level constraints, optimizing conventional server infrastructure is essential for sustainable growth.
A Neurosymbolic Prolog Skill for LLM-Driven Service PlacementJacopo Massa, Giuseppe Bisicchia, Patrizio Dazzi, Antonio Brogi2026-06-09下载Service placement in the cloud-edge continuum requires assigning application components to heterogeneous resources under multiple constraints, including latency, locality, and policy requirements.
FairWave : A Fairness-Aware Asynchronous DAG-BFT ConsensusSyariful Mujaddiq2026-06-09下载Combining asynchronous Byzantine Fault Tolerant (BFT) consensus with Proof-of-Stake (PoS) creates a trilemma between Sybil resistance, reward distribution fairness, and protection against persistent p...
Dynamic Software Updates using CRDTsSeppe Wyns, Jim Bauwens, Elisa Gonzalez Boix2026-06-09下载This paper investigates how Conflict-free Replicated Data Types (CRDTs) can be used for dynamic software updates of distributed applications. We propose to model application updates as a new App CRDT ...
Inverse Probability Weighting and Age-of-Information Aggregation for Decentralized Federated Learning under Partial ReceptionChanuka A. S. Hewa Kaluannakkage, Rajkumar Buyya2026-06-09下载Decentralized Federated Learning (DFL) over lossy wireless networks faces two key challenges: selection bias, where updates from poor-quality links are systematically underrepresented due to partial m...
Generalizing LCL Complexity Gaps to Unbounded Degree via Monadic Second-Order PropertiesChiara Piombi2026-06-09下载The last decade of research on the LOCAL model has seen tremendous progress in understanding locally checkable labeling (LCL) problems, culminating in an almost complete classification of the possible...
A Hybrid Edge-Cloud Architecture for Low-Latency Entitlement Verification in Resource-Constrained DevicesPravin Nagare, Aditya Sabbineni, Devendra Dahiphale, Faiz Gouri, Pratik Thantharate2026-06-09下载As digital media consumption shifts toward large-scale Over-the-Top (OTT) platforms, the efficiency of the control plane, specifically entitlement and identity verification, has become a critical fact...
Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid DesignWenxin Wang, Yule Hou, Yu Ji, Peng Qu, Youhui Zhang2026-06-09下载Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even under low-concurrency workloads.
ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure ModelingWilliam Won, Jinsun Yoo, Tuan Ta, Moumita Dey, Andy Balogh, Pradosh Datta, Furkan Eris, Conor Green, Winston Liu, Changhai Man, Kingshuk Mandal, Amos Rai, Vinay Ramakrishnaiah, Ruchi Shah, David Sidler, Harsh Sikhwal, Hanjiang Wu, Tushar Krishna, Bradford M. Beckmann2026-06-09下载Distributed machine learning (ML) is a key paradigm for today's large-scale artificial intelligence applications. As model inference arises as an important use case, faithful modeling of latency-sensi...
RATrain: A Resource-Aware Training Runtime for Large Language Models on Bandwidth-Constrained Heterogeneous Supercomputing PlatformsYao Lu, Shiqing Ma, Zhongzhi Luan, Gen Li, Jiaxing Qi, Bin Han, Hailong Yang, Depei Qian2026-06-09下载Production heterogeneous supercomputing platforms are increasingly used to host large language model (LLM) training workloads. However, existing GPU-oriented training runtimes typically rely on high-b...
Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-based AcceleratorsChenguang Zhang, Yuanpeng Zhang, Chenhao Xue, Yihan Yin, Chen Zhang, Guangyu Sun2026-06-09下载Level-4+ autonomous driving systems (ADS) must run dozens of heterogeneous deep neural networks (DNNs) as end-to-end (E2E) pipelines under a strict latency constraint (<=100 ms), even as execution ti...
Determination Provenance: From Ambiguity to AlgebraJoseph M. Hellerstein2026-06-09下载Many data systems admit multiple admissible outcomes for the same input: concurrent transactions may serialize in one of many orders; a logic program may have multiple stable models.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Predictive and Spatially Aware Scheduling in Flexible Duplexing for Deterministic CommunicationsSyed Morsleen Riaz, Baldomero Coll-Perales, M. Carmen Lucas-Estañ, Javier Gozalvez, Miguel Sepulcre2026-06-09下载Next generation wireless networks must sustain deterministic service levels for time-sensitive closed-loop applications. Flexible duplexing (FD) is an efficient solution to support these services, as ...
Internet Quality Barometer (IQB): A preliminary data-driven evaluation of the IQB frameworkPavlos Sermpezis, Zeynep Arslan2026-06-09下载The Internet Quality Barometer (IQB) framework was designed to transform raw Internet measurement data into actionable insights about Internet quality.
Generative Explainability for Next-Generation Networks: LLM-Augmented XAI with Mutual Feature InteractionsKiarash Rezaei, Omran Ayoub, Sebastian Troia, Francesco Lelli, Paolo Monti, Carlos Natalino2026-06-09下载As artificial intelligence and machine learning (AI/ML) models become integral to network operations, their lack of transparency poses a significant barrier to operator trust.
A Unified Siamese Learning Framework for Zero-Day Anomaly Detection and Classification in Optical NetworksCarlos Natalino, Flávia P. Monteiro, Paolo Monti2026-06-09下载A multi-similarity Siamese neural network unifies zero-day anomaly detection and one-shot classification in optical networks, achieving over 99% accuracy and instant adaptability across lightpaths and...
High-Speed Generation of Periodic Traffic Patterns on P4TG for DDoS and Burst-Load EvaluationFabian Ihle, Etienne Zink, Michael Menth2026-06-09下载Traffic generators are essential tools for evaluating the robustness and performance of networked systems. P4TG is an open-source, hardware-accelerated traffic generator implemented in P4 for the Inte...
CAMASA: A CAM-based Dataset from the MASA Living LabSalvatore Iandolo, Marco Savarese, Gaetano Orazio Cauchi, Antonio Solida, Martin Klapez, Maurizio Casoni, Angelo Porrello, Carlo Augusto Grazia2026-06-09下载Trajectory prediction is a key enabler of autonomous and cooperative driving systems. However, most existing benchmarks are either sensor-centric, geographically constrained, or based on synthetic mob...
From Stacks to Circuits: A Regenerative Socio-Technical Roadmap for AI Infrastructure within Planetary BoundariesHan-Teng Liao, Karen Ang2026-06-09下载Current scaling trajectories for Generative AI, typified by linear supply-side "stacks," prioritize performance density while externalizing significant thermodynamic and material costs.
A Deployment-Oriented Framework for Explainable AI-Assisted eBPF/XDP Mitigation at the IoT EdgeAbdurrahman Tolay2026-06-09下载Internet of Things (IoT) deployments combine heterogeneous, resource-constrained devices with weak security configurations, exposed services, limited logging, patching constraints, and long lifecycles...
ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure ModelingWilliam Won, Jinsun Yoo, Tuan Ta, Moumita Dey, Andy Balogh, Pradosh Datta, Furkan Eris, Conor Green, Winston Liu, Changhai Man, Kingshuk Mandal, Amos Rai, Vinay Ramakrishnaiah, Ruchi Shah, David Sidler, Harsh Sikhwal, Hanjiang Wu, Tushar Krishna, Bradford M. Beckmann2026-06-09下载Distributed machine learning (ML) is a key paradigm for today's large-scale artificial intelligence applications. As model inference arises as an important use case, faithful modeling of latency-sensi...

cs.PF - Performance ​

标题作者发布日期PDF摘要
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUsWesley Pang, Gregory Hyegang Jun, Feiyang Liu, Deming Chen2026-06-09下载With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets.
Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDAVinamra Sharma, Xingjian Fu, Jude Haris, José Cano2026-06-09下载Designing FPGA-based accelerators for modern artificial intelligence workloads requires exploring a large and complex hardware design space that involves architectural parameters, data flow strategies...
Flash-GMM: A Memory-Efficient Kernel for Scalable Soft ClusteringGal Bloch, Ariel Gera, Matan Orbach, Ohad Eytan, Assaf Toledo2026-06-09下载We present \textbf{Flash-GMM}, a fused Triton kernel for efficient computation of Gaussian Mixture Models (GMMs) over large-scale data in a single GPU pass.
Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X EliteZhiyuan Cheng, Longying Lai2026-06-09下载Retrieval-Augmented Generation (RAG) pipelines are compute-intensive, combining embedding, retrieval, reranking, and large language model (LLM) generation.

基于 VitePress 构建 · 使用本地搜索查找论文