2026-09-08
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Benchmarking Agentic HLS Design Tasks With HLS-Eval | Stefan Abi-Karam, Callie Hao | 2026-09-08 | 下载 | Large language models (LLMs) and AI agents are increasingly explored for hardware design, including high-level digital design. While most work targets code generation and editing for hardware descript... |
| HLSFactory-Agent: Large-Scale Agentic HLS Dataset Construction from Academic and Open-Source Projects | Kaushik Chandana, Jay Imperatori, Tanmay Shukla, Justin Zhou, Stefan Abi-Karam, Callie Hao | 2026-09-08 | 下载 | Building large, diverse datasets of high-level synthesis (HLS) designs beyond common community benchmarks remains an open challenge. This challenge is made urgent by the rise of deep learning and LLMs... |
| FPGA Acceleration of Fully Homomorphic Encryption with Adaptive Key Switching | Zhihan Xu, Jayashree Adivarahan, Rajgopal Kannan, Viktor K. Prasanna | 2026-09-08 | 下载 | Fully Homomorphic Encryption (FHE) enables privacy-preserving cloud services but incurs substantial computation overhead, making hardware acceleration essential. |
| Academia x Industry: The Role of Fundamentals for Silicon in an AI Native Era | Vincent T. Lee, Armin Alaghi, Carole-Jean Wu, Sai Zhang, Brandon Reagen, Thierry Tambe, Jean Boufarhat, Matheus Trevisan Moreira | 2026-09-08 | 下载 | Agentic AI is set to become one of the most transformational technologies in generations and materially change how we approach silicon design and engineering. |
| Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores | Satoshi Matsuoka | 2026-09-08 | 下载 | FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibriu... |
| Towards Standardized Evaluation of GPU Memory Safety with GMSBench | Saurabh Singh, Jaewon Lee, Seonjin Na, Hyesoon Kim | 2026-09-08 | 下载 | As GPUs become increasingly integral to high-performance computing and machine learning, ensuring memory safety in GPU programs has become crucial for reliable and secure execution. |
| DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable Connectivity | Jiaqi Ye, Xinrui Gong, Jingcun Wang, Olga Kondrateva, Bing Li, Grace Li Zhang | 2026-09-08 | 下载 | Field-programmable gate arrays (FPGAs) enable efficient neural-network inference, but most deployment flows either accelerate multiply-accumulate operations or convert pretrained quantized models into... |
| HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing | Haochen Huang, Shuzhang Zhong, Shengxuan Qiu, Zhe Zhang, Shuangchen Li, Cong Li, Dimin Niu, Hongzhong Zheng, Guangyu Sun, Runsheng Wang, Meng Li | 2026-09-08 | 下载 | Mixture-of-Experts (MoE) architectures have become a key technique for scaling Large Language Models (LLMs), enabling high model capacity with reduced computational cost. |
| FlexSpIM: An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Hybrid Stationarity | Nicolas Chauvaux, Adrian Kneip, Charlotte Frenkel | 2026-09-08 | 下载 | Compute-in-memory (CIM) accelerators for spiking neural networks (SNNs) offer a promising solution for achieving μs-level inference latency and ultra-low energy in edge vision applications. |
| PENDA: An Efficient Processing Element via Norm-of-Difference for Deep Learning Accelerators | Kai-Chieh Hsu, Tian-Sheuan Chang | 2026-09-08 | 下载 | Inner product computation dominates the computational cost of deep learning models; thus, accelerating this primitive is key to improving hardware efficiency. |
| Routing Dense Layouts with History-Aware Offline Reinforcement Learning using LSTM | Afsara Khan, Austin Rovinski | 2026-09-08 | 下载 | Detailed routing remains a dominant runtime bottleneck in physical design due to increasing complexity of design rules. Modern routers can struggle to resolve persistent violations under dense operati... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management | Antonino Vaccarella, Lanpei Li, Vincenzo Lomonaco, Massimo Coppola | 2026-09-08 | 下载 | Managing resources across IoT, edge, and cloud layers calls for continuous, context-aware decisions under constraints that rarely stay fixed. Deep reinforcement learning (DRL) handles this class of pr... |
| Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs | Joseph K. L. Lee, Mehrdad Malekmohammadi, Hong-Sheng Zheng, Shuli Shu, Cheick Doumbia, Kalman Szenes, Mehran Zamani Abnili, Thomas Ainsworth, Matthew Seymour, Thomas Germain, Leonhard Neuhaus, Josh Izaac, Lee J. O'Riordan | 2026-09-08 | 下载 | Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. |
| Distributed Linear Programming on GPU Clusters at Extreme Scale | Arnaud Deza, Santanu Dey, Pascal Van Hentenryck | 2026-09-08 | 下载 | Large linear programs can exceed the memory of a single compute node. Although first-order methods replace sparse factorizations with GPU-suited matrix-vector products, other solver phases can reintro... |
| Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores | Satoshi Matsuoka | 2026-09-08 | 下载 | FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibriu... |
| Impossibility of One-Way One-Round Quantum 4-Coloring via Matrix-Space Stability | Tom Gur, Longcheng Li | 2026-09-08 | 下载 | We show that one-way one-round quantum LOCAL algorithms cannot -color directed cycles with high probability, even with unbounded local computation and quantum message length. |
| GraphFAS: A Distributed System for Automated Graph Feature Generation and Selection in Industrial Transaction Networks | Yice Luo, Yun Zhu, Xi Chen, Yongchao Liu, Xintan Zeng, Chengying Huan, Kai Zhang, Jinrui Zhang, Juelu Zhang, Jiajun Zheng | 2026-09-08 | 下载 | Industrial fraud detection often relies on costly expert-crafted features that overlook graph-structured relational signals, while GNNs often do not meet the interpretability and deployment requiremen... |
| ContinuumBench: Benchmarking Joint Autoscaling and Placement Across Evaluation Regimes in the Cloud-Edge Continuum | Lanpei Li, Antonino Vaccarella, Vincenzo Lomonaco, Massimo Coppola | 2026-09-08 | 下载 | Cloud-edge controllers coordinate service placement, replica scaling, and resource pre-warming to keep end-to-end latency within application deadlines. |
| Exploring the Genesis Platform Capabilities to Accelerate Scientific Discovery in OPAL | Daniel Rosendo, Renan Souza, Kelsey Carter, John Lagergren, Frédéric Suter, Shelaine L. Curd, David Weston, Rafael Ferreira da Silva | 2026-09-08 | 下载 | Autonomous, cross-facility science requires capabilities that no individual project should have to build for itself: managed execution for long-lived services, versioned distribution of models to remo... |
| Tools-CC-Bench: a Benchmark Suite for Collective Communication with Compression in HPC and AI Workloads | Haozhe Fan, Wei Wang, Xingchen Liu, Man Liu, Xingjian Tian, Haoquan Long, Zedong Liu, Daran Sun, Jinwu Yang, Bo Yang, Jie Liu, Yonggang Che, Hairui Zhao, Guangming Tan, Dingwen Tao | 2026-09-08 | 下载 | Distributed HPC and LLM workloads increasingly require efficient communication for scalability, yet growing data movement has become a major performance bottleneck. |
| Measuring Sustainability in Multi-Scale High-Performance Computing | Carlos J Barrios, Frédéric Le Mouël, Yves Denneulin | 2026-09-08 | 下载 | The transition from traditional High Performance Computing (HPC) to the Computing Continuum emphasizes efficient resource management and sustainable practices across Multi-Scale hybrid architectures. |
| Sample-Guided Exact Top-K Selection for Long-Context Sparse Attention | Siran Liu, Yang Xue, Theo Tang, Changxu Shao, Qian Cheng, Haimeng Ren, Donghua Jiang, Haipeng Ming, Lehua Ding, Zhonghan Lin, Shengying Wei, Wei Liu, Kai Liu, Jianchen Zhu | 2026-09-08 | 下载 | Sparse attention bounds downstream attention work by retaining a fixed-size subset of indexed tokens, but its standalone exact Top- stage must still process materialized score rows whose length gro... |
| A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware | Maysam Khatib, Moysis Symeonides, Demetris Trihinas, George Pallis, Marios D. Dikaiakos | 2026-09-08 | 下载 | Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy... |
| Transversal Fanout for Fault Tolerant Distributed Quantum Computing: Analysis and Application | Seng W. Loke | 2026-09-08 | 下载 | We study a resource-efficient approach for implementing logical fanout operations in fault-tolerant distributed quantum computing using transversal operations on quantum error-correcting code blocks. |
| SemBridge: Compiling Consumer Observations into Cross-Stack Communication Plans | Genlang Chen, Junyi Zhu, Yuanshan Lin | 2026-09-08 | 下载 | Distributed-tensor systems specify where values reside, while collective systems optimize how requested operations execute. At a boundary between vendor runtimes that cannot share a native communicato... |
| Generalized DBLog: A Verified Contract for Interleaving Database Rows with a Change Log | Andreas Andreakis | 2026-09-08 | 下载 | Change-data capture (CDC) feeds downstream systems like caches, search indexes, and data warehouses from a database's log of committed row changes. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Prototyping QoE-Aware Rate Adaptation in Cellular Networks with Commercial Applications | Szilveszter Nádas, Lars Ernström, Dan Druta, Igor Pruzhansky, David Lindero, Jonathan Lynam, Eric Petajan | 2026-09-08 | 下载 | Prior work has shown that QoE-aware resource sharing for real-time interactive video can support up to three times more simultaneous sessions at acceptable quality compared to rate-fair allocation. |
| Concept drift mitigation through community and spectral graph analysis for the detectionof cyberattacks in network traffic | Julien Michel, Abdul Qadir Khan, Majed Jaber, Pierre Parrend | 2026-09-08 | 下载 | In network traffic, legitimate behaviours and attack techniques evolve jointly - the phenomenon known as 'concept drift' [1]. Every detector is thereby left obsolete between two updates, and always on... |
| QPS-ToR: A Parallel Iterative Switching Algorithm for Reconfigurable Optical Datacenter Switching | Dongzhao Song, Qianru Yu, Jun Xu | 2026-09-08 | 下载 | Reconfigurable optical data center networks (RODCNs) have emerged as a promising solution for scaling DCN capacity, yet their scheduling mechanisms remain a performance bottleneck: traffic-oblivious s... |
| Efficient User Association and Wireless Scheduling with Shorter Time-Scale Rate Adaptation | Xiaoyi Wu, Huacheng Zeng, Bin Li | 2026-09-08 | 下载 | Rate adaptation is a crucial mechanism in IEEE 802.11 networks and next-generation cellular systems. Since the time scale for rate adaptation is typically much shorter than that for user association a... |
| Improving 5G AI-RAN MCS Selection by Predicting Retransmissions | Tamerlan Aghayev, Maxime Elkael, Michele Polese, Reshma Prasad, Salvatore D'Oro, Yunseong Lee, Koichiro Furueda, Tommaso Melodia | 2026-09-08 | 下载 | Link Adaptation (LA) in 5G NR is inherently reactive, relying on channel measurements and HARQ feedback that may become quickly obsolete when the channel changes quickly. |
| A New Backscattering Dual-Polarized Rectenna for Wireless Power Transfer and IoT Applications | Taki Eddine Djidjekh, Quentin Bernyer, Alexandru Takacs | 2026-09-08 | 下载 | This paper proposes an innovative dual-polarized backscattering rectenna that operates in two distinct modesenergy harvesting and backscattering modulation-driven by two-bit digital control signals. |
| Hybrid Continuous DoA Estimation with Shared-Radius Co-Prime Circular Arrays | Keyvan Aghababaiyan | 2026-09-08 | 下载 | This paper proposes a shared-radius co-prime circular array for high-resolution, continuous 2D Direction-of-Arrival (DoA) estimation in 3D space, jointly estimating azimuth and elevation angles. |
| CleanCity-BinSense: An IoT-Enabled Smart Waste Management System with Configurable Real-Time Fill Monitoring and Nearest-Neighbor Route Optimization | Mohammad Adnan Kabir, Intifad Muhammad Sayeed | 2026-09-08 | 下载 | CleanCity-BinSense addresses inefficiencies in urban waste management in developing cities, where fixed-schedule collection routes lead to overflowing bins and wasted fuel. |
| AI-Native Orchestration in the 6G Continuum: Evolving Operator Platforms with Agentic AI | Claudia Carballo González, Hatim Chergui, Sergio Giménez-Antón, Mohammadreza Mosahebfard, Juan Sebastián Camargo, Pouria Sayyad Khodashenas, Vasileios Theodorou, Christos Verikoukis | 2026-09-08 | 下载 | As Sixth-Generation (6G) networks evolve towards a seamless Cloud-Edge-Internet of Things (IoT) continuum, autonomous orchestration across distributed compute and network domains becomes critical. |
| Toward Fully Autonomous 6G Networks: AI-driven Operational Efficiency and Optimization | David Reiss, Oriol Sallent, Miguel Catalan-Cid, Daniel Camps-Mur | 2026-09-08 | 下载 | Mobile networks evolution is characterized by a substantial increase in system complexity, driven by the need to accommodate a growing number of heterogeneous services on top of the digital infrastruc... |
| QoS-Aware RACH Preamble Slicing via Quota-Projected Branching Deep Reinforcement Learning | Jiulin Guo, Jiahan Xu, Jiashuo Zhang, Heng Yang, Yizhen Sun, Yutong Xie, Shanshan Li, Zhenyu Liu, Lei Zhang | 2026-09-08 | 下载 | Quality-of-service (QoS)-aware random access requires adaptive allocation of a finite random access channel (RACH) preamble budget across heterogeneous traffic and access procedures. |
| Information-Entropy-Driven Fault Propagation Modeling for Probabilistic Network Performance Prediction | Lusha Mo, Fengxiao Tang, Xiaonan Wang, Ming Zhao | 2026-09-08 | 下载 | Network faults can trigger cascading effects that cause abrupt and nonstationary performance degradation. Existing learning-based performance predictors mainly focus on normal operation or treat fault... |
| 6SEVEN: System for EValuating IPv6 ENumeration algorithms | Chase Kanipe, Erik Rye, Dave Levin, Robert Beverly | 2026-09-08 | 下载 | The vast, sparsely populated, and often ephemeral IPv6 address space makes discovering active addresses challenging. In response, the community has developed over thirty different IPv6 Target Generati... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor Cores | Satoshi Matsuoka | 2026-09-08 | 下载 | FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibriu... |