2026-06-09
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs | Wesley Pang, Gregory Hyegang Jun, Feiyang Liu, Deming Chen | 2026-06-09 | 下载 | With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets. |
| Revisiting "Cooler is Better": ITD-Aware Per-CPU Thermal Optimization for Sustainable Data Center Operation | Jason Crop, Hayden Moore, Sudeep Pasricha | 2026-06-09 | 下载 | As data center energy demand approaches grid-level constraints, optimizing conventional server infrastructure is essential for sustainable growth. |
| Defeat the Heap: Zero-Copy Data Movement in AXI4MLIR | Elam Cohavi, Nicolas Bohm Agostini, Jude Haris, Antonino Tumeo, David Kaeli, José Cano | 2026-06-09 | 下载 | As custom hardware accelerators become increasingly central to machine learning workloads, efficient data transfer is critical for maximizing accelerator performance on linear algebra kernels. |
| Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA | Vinamra Sharma, Xingjian Fu, Jude Haris, José Cano | 2026-06-09 | 下载 | Designing FPGA-based accelerators for modern artificial intelligence workloads requires exploring a large and complex hardware design space that involves architectural parameters, data flow strategies... |
| Coset Ensemble Decoder for Quantum Error Correction with Algorithm-Hardware Co-Design | Shuang Liang, Jubo Xu, Giulio Bassanino, Qianzhou Wang, Yidong Zhou, Yuncheng Lu, Zhiwen Mo, Paul H. J. Kelly, Bo Yuan, Wayne Luk, Hongxiang Fan | 2026-06-09 | 下载 | Reliable large-scale quantum computation relies on fault-tolerant architectures, where quantum error correction (QEC) continuously extracts and decodes error syndromes in real time. |
| Arithmetic Packing on Wide Integer Datapaths in DSP Primitives of Modern FPGA Devices | Titus Bornträger, Shane Fleming, Philipp Holzinger, Dietmar Fey, Michaela Blott, Thomas B. Preußer | 2026-06-09 | 下载 | Deep Neural Networks increasingly employ low-precision quantization to reduce computational requirements. While FPGAs are well suited for workloads with heterogeneous precisions, their dedicated digit... |
| A 185 TOPS/W/mm2 Bayesian Inference Engine with 640 aJ Write-Free FeFET GRNG for Uncertainty-Aware Aerial Search and Rescue | Zephan M. Enciso, Xuezhong Niu, Xingtian Wang, Mohammad Mehdi Sharifi, Subhasish Mukherjee, Likai Pei, Halid Mulaosmanovic, Stefan Duenkel, Sven Beyer, Michael Niemier, Kai Ni, Ningyuan Cao | 2026-06-09 | 下载 | Aerial search and rescue missions require fast and reliable victim detection under uncertain and rapidly changing environments. Deterministic deep learning models can produce overconfident false posit... |
| A Hybrid Edge-Cloud Architecture for Low-Latency Entitlement Verification in Resource-Constrained Devices | Pravin Nagare, Aditya Sabbineni, Devendra Dahiphale, Faiz Gouri, Pratik Thantharate | 2026-06-09 | 下载 | As digital media consumption shifts toward large-scale Over-the-Top (OTT) platforms, the efficiency of the control plane, specifically entitlement and identity verification, has become a critical fact... |
| Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-based Accelerators | Chenguang Zhang, Yuanpeng Zhang, Chenhao Xue, Yihan Yin, Chen Zhang, Guangyu Sun | 2026-06-09 | 下载 | Level-4+ autonomous driving systems (ADS) must run dozens of heterogeneous deep neural networks (DNNs) as end-to-end (E2E) pipelines under a strict latency constraint (<=100 ms), even as execution ti... |
| LLM-Guided Neural Architecture Search for Robust Co-Design of Physical Neural Networks | Tyler King, Timothee Leleu | 2026-06-09 | 下载 | Deploying neural networks on unconventional hardware demands architectures that co-optimize task accuracy and platform-specific constraints such as energy cost, physical non-idealities, and numerical ... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A Scalable PyTorch Abstraction for Multi-GPU Gaussian Splatting | Matthew Cong, Francis Williams, Jonathan Swartz, Mark Harris, Sanja Fidler, Ken Museth | 2026-06-09 | 下载 | Gaussian splatting methods have become increasingly popular for neural reconstruction of the real world. However, they are often limited in scale and resolution due to compute and memory constraints. |
| TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs | Wesley Pang, Gregory Hyegang Jun, Feiyang Liu, Deming Chen | 2026-06-09 | 下载 | With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets. |
| An Ocean Model Ported by a Large Language Model: Experience and Lessons from FESOM2 (Fortran to C to C++/Kokkos) | Nikolay V. Koldunov, Suvarchal K. Cheedela, Sergey Danilov, Dmitry Sidorenko, Sebastian Beyer, Thomas Jung | 2026-06-09 | 下载 | Large language models (LLMs) can translate and modify source code, and have been shown to do so for codes of different complexity. Whether they can port a complete, production geophysical model to a d... |
| Piper: A Programmable Distributed Training System | Megan Frisella, Shubham Tiwari, Andy Ruan, Yi Pan, Parker Gustafson, Mat Jacob, Gilbert Bernstein, Stephanie Wang | 2026-06-09 | 下载 | Large-scale model training increasingly relies on composing multiple parallelism strategies, such as data, pipeline, and expert parallelism, together with memory-saving optimizations like ZeRO. |
| Revisiting "Cooler is Better": ITD-Aware Per-CPU Thermal Optimization for Sustainable Data Center Operation | Jason Crop, Hayden Moore, Sudeep Pasricha | 2026-06-09 | 下载 | As data center energy demand approaches grid-level constraints, optimizing conventional server infrastructure is essential for sustainable growth. |
| A Neurosymbolic Prolog Skill for LLM-Driven Service Placement | Jacopo Massa, Giuseppe Bisicchia, Patrizio Dazzi, Antonio Brogi | 2026-06-09 | 下载 | Service placement in the cloud-edge continuum requires assigning application components to heterogeneous resources under multiple constraints, including latency, locality, and policy requirements. |
| FairWave : A Fairness-Aware Asynchronous DAG-BFT Consensus | Syariful Mujaddiq | 2026-06-09 | 下载 | Combining asynchronous Byzantine Fault Tolerant (BFT) consensus with Proof-of-Stake (PoS) creates a trilemma between Sybil resistance, reward distribution fairness, and protection against persistent p... |
| Dynamic Software Updates using CRDTs | Seppe Wyns, Jim Bauwens, Elisa Gonzalez Boix | 2026-06-09 | 下载 | This paper investigates how Conflict-free Replicated Data Types (CRDTs) can be used for dynamic software updates of distributed applications. We propose to model application updates as a new App CRDT ... |
| Inverse Probability Weighting and Age-of-Information Aggregation for Decentralized Federated Learning under Partial Reception | Chanuka A. S. Hewa Kaluannakkage, Rajkumar Buyya | 2026-06-09 | 下载 | Decentralized Federated Learning (DFL) over lossy wireless networks faces two key challenges: selection bias, where updates from poor-quality links are systematically underrepresented due to partial m... |
| Generalizing LCL Complexity Gaps to Unbounded Degree via Monadic Second-Order Properties | Chiara Piombi | 2026-06-09 | 下载 | The last decade of research on the LOCAL model has seen tremendous progress in understanding locally checkable labeling (LCL) problems, culminating in an almost complete classification of the possible... |
| A Hybrid Edge-Cloud Architecture for Low-Latency Entitlement Verification in Resource-Constrained Devices | Pravin Nagare, Aditya Sabbineni, Devendra Dahiphale, Faiz Gouri, Pratik Thantharate | 2026-06-09 | 下载 | As digital media consumption shifts toward large-scale Over-the-Top (OTT) platforms, the efficiency of the control plane, specifically entitlement and identity verification, has become a critical fact... |
| Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design | Wenxin Wang, Yule Hou, Yu Ji, Peng Qu, Youhui Zhang | 2026-06-09 | 下载 | Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even under low-concurrency workloads. |
| ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure Modeling | William Won, Jinsun Yoo, Tuan Ta, Moumita Dey, Andy Balogh, Pradosh Datta, Furkan Eris, Conor Green, Winston Liu, Changhai Man, Kingshuk Mandal, Amos Rai, Vinay Ramakrishnaiah, Ruchi Shah, David Sidler, Harsh Sikhwal, Hanjiang Wu, Tushar Krishna, Bradford M. Beckmann | 2026-06-09 | 下载 | Distributed machine learning (ML) is a key paradigm for today's large-scale artificial intelligence applications. As model inference arises as an important use case, faithful modeling of latency-sensi... |
| RATrain: A Resource-Aware Training Runtime for Large Language Models on Bandwidth-Constrained Heterogeneous Supercomputing Platforms | Yao Lu, Shiqing Ma, Zhongzhi Luan, Gen Li, Jiaxing Qi, Bin Han, Hailong Yang, Depei Qian | 2026-06-09 | 下载 | Production heterogeneous supercomputing platforms are increasingly used to host large language model (LLM) training workloads. However, existing GPU-oriented training runtimes typically rely on high-b... |
| Isolation-aware Scheduling Framework for DNN-based End-to-End Autonomous Driving System on Tile-based Accelerators | Chenguang Zhang, Yuanpeng Zhang, Chenhao Xue, Yihan Yin, Chen Zhang, Guangyu Sun | 2026-06-09 | 下载 | Level-4+ autonomous driving systems (ADS) must run dozens of heterogeneous deep neural networks (DNNs) as end-to-end (E2E) pipelines under a strict latency constraint (<=100 ms), even as execution ti... |
| Determination Provenance: From Ambiguity to Algebra | Joseph M. Hellerstein | 2026-06-09 | 下载 | Many data systems admit multiple admissible outcomes for the same input: concurrent transactions may serialize in one of many orders; a logic program may have multiple stable models. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Predictive and Spatially Aware Scheduling in Flexible Duplexing for Deterministic Communications | Syed Morsleen Riaz, Baldomero Coll-Perales, M. Carmen Lucas-Estañ, Javier Gozalvez, Miguel Sepulcre | 2026-06-09 | 下载 | Next generation wireless networks must sustain deterministic service levels for time-sensitive closed-loop applications. Flexible duplexing (FD) is an efficient solution to support these services, as ... |
| Internet Quality Barometer (IQB): A preliminary data-driven evaluation of the IQB framework | Pavlos Sermpezis, Zeynep Arslan | 2026-06-09 | 下载 | The Internet Quality Barometer (IQB) framework was designed to transform raw Internet measurement data into actionable insights about Internet quality. |
| Generative Explainability for Next-Generation Networks: LLM-Augmented XAI with Mutual Feature Interactions | Kiarash Rezaei, Omran Ayoub, Sebastian Troia, Francesco Lelli, Paolo Monti, Carlos Natalino | 2026-06-09 | 下载 | As artificial intelligence and machine learning (AI/ML) models become integral to network operations, their lack of transparency poses a significant barrier to operator trust. |
| A Unified Siamese Learning Framework for Zero-Day Anomaly Detection and Classification in Optical Networks | Carlos Natalino, Flávia P. Monteiro, Paolo Monti | 2026-06-09 | 下载 | A multi-similarity Siamese neural network unifies zero-day anomaly detection and one-shot classification in optical networks, achieving over 99% accuracy and instant adaptability across lightpaths and... |
| High-Speed Generation of Periodic Traffic Patterns on P4TG for DDoS and Burst-Load Evaluation | Fabian Ihle, Etienne Zink, Michael Menth | 2026-06-09 | 下载 | Traffic generators are essential tools for evaluating the robustness and performance of networked systems. P4TG is an open-source, hardware-accelerated traffic generator implemented in P4 for the Inte... |
| CAMASA: A CAM-based Dataset from the MASA Living Lab | Salvatore Iandolo, Marco Savarese, Gaetano Orazio Cauchi, Antonio Solida, Martin Klapez, Maurizio Casoni, Angelo Porrello, Carlo Augusto Grazia | 2026-06-09 | 下载 | Trajectory prediction is a key enabler of autonomous and cooperative driving systems. However, most existing benchmarks are either sensor-centric, geographically constrained, or based on synthetic mob... |
| From Stacks to Circuits: A Regenerative Socio-Technical Roadmap for AI Infrastructure within Planetary Boundaries | Han-Teng Liao, Karen Ang | 2026-06-09 | 下载 | Current scaling trajectories for Generative AI, typified by linear supply-side "stacks," prioritize performance density while externalizing significant thermodynamic and material costs. |
| A Deployment-Oriented Framework for Explainable AI-Assisted eBPF/XDP Mitigation at the IoT Edge | Abdurrahman Tolay | 2026-06-09 | 下载 | Internet of Things (IoT) deployments combine heterogeneous, resource-constrained devices with weak security configurations, exposed services, limited logging, patching constraints, and long lifecycles... |
| ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure Modeling | William Won, Jinsun Yoo, Tuan Ta, Moumita Dey, Andy Balogh, Pradosh Datta, Furkan Eris, Conor Green, Winston Liu, Changhai Man, Kingshuk Mandal, Amos Rai, Vinay Ramakrishnaiah, Ruchi Shah, David Sidler, Harsh Sikhwal, Hanjiang Wu, Tushar Krishna, Bradford M. Beckmann | 2026-06-09 | 下载 | Distributed machine learning (ML) is a key paradigm for today's large-scale artificial intelligence applications. As model inference arises as an important use case, faithful modeling of latency-sensi... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs | Wesley Pang, Gregory Hyegang Jun, Feiyang Liu, Deming Chen | 2026-06-09 | 下载 | With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets. |
| Towards Autonomous Accelerator Design: FPGA Accelerator Generation with SECDA | Vinamra Sharma, Xingjian Fu, Jude Haris, José Cano | 2026-06-09 | 下载 | Designing FPGA-based accelerators for modern artificial intelligence workloads requires exploring a large and complex hardware design space that involves architectural parameters, data flow strategies... |
| Flash-GMM: A Memory-Efficient Kernel for Scalable Soft Clustering | Gal Bloch, Ariel Gera, Matan Orbach, Ohad Eytan, Assaf Toledo | 2026-06-09 | 下载 | We present \textbf{Flash-GMM}, a fused Triton kernel for efficient computation of Gaussian Mixture Models (GMMs) over large-scale data in a single GPU pass. |
| Energy-Efficient On-Device RAG on a Mobile NPU: System Design and Benchmark on Snapdragon X Elite | Zhiyuan Cheng, Longying Lai | 2026-06-09 | 下载 | Retrieval-Augmented Generation (RAG) pipelines are compute-intensive, combining embedding, retrieval, reranking, and large language model (LLM) generation. |