Skip to content

2026-09-08 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Benchmarking Agentic HLS Design Tasks With HLS-EvalStefan Abi-Karam, Callie Hao2026-09-08下载Large language models (LLMs) and AI agents are increasingly explored for hardware design, including high-level digital design. While most work targets code generation and editing for hardware descript...
HLSFactory-Agent: Large-Scale Agentic HLS Dataset Construction from Academic and Open-Source ProjectsKaushik Chandana, Jay Imperatori, Tanmay Shukla, Justin Zhou, Stefan Abi-Karam, Callie Hao2026-09-08下载Building large, diverse datasets of high-level synthesis (HLS) designs beyond common community benchmarks remains an open challenge. This challenge is made urgent by the rise of deep learning and LLMs...
FPGA Acceleration of Fully Homomorphic Encryption with Adaptive Key SwitchingZhihan Xu, Jayashree Adivarahan, Rajgopal Kannan, Viktor K. Prasanna2026-09-08下载Fully Homomorphic Encryption (FHE) enables privacy-preserving cloud services but incurs substantial computation overhead, making hardware acceleration essential.
Academia x Industry: The Role of Fundamentals for Silicon in an AI Native EraVincent T. Lee, Armin Alaghi, Carole-Jean Wu, Sai Zhang, Brandon Reagen, Thierry Tambe, Jean Boufarhat, Matheus Trevisan Moreira2026-09-08下载Agentic AI is set to become one of the most transformational technologies in generations and materially change how we approach silicon design and engineering.
Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor CoresSatoshi Matsuoka2026-09-08下载FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibriu...
Towards Standardized Evaluation of GPU Memory Safety with GMSBenchSaurabh Singh, Jaewon Lee, Seonjin Na, Hyesoon Kim2026-09-08下载As GPUs become increasingly integral to high-performance computing and machine learning, ensuring memory safety in GPU programs has become crucial for reliable and secure execution.
DiffLUT-Net: Differentiable Training of FPGA LUT Networks with Learnable ConnectivityJiaqi Ye, Xinrui Gong, Jingcun Wang, Olga Kondrateva, Bing Li, Grace Li Zhang2026-09-08下载Field-programmable gate arrays (FPGAs) enable efficient neural-network inference, but most deployment flows either accelerate multiply-accumulate operations or convert pretrained quantized models into...
HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory ProcessingHaochen Huang, Shuzhang Zhong, Shengxuan Qiu, Zhe Zhang, Shuangchen Li, Cong Li, Dimin Niu, Hongzhong Zheng, Guangyu Sun, Runsheng Wang, Meng Li2026-09-08下载Mixture-of-Experts (MoE) architectures have become a key technique for scaling Large Language Models (LLMs), enabling high model capacity with reduced computational cost.
FlexSpIM: An Event-Based Digital Compute-In-Memory Accelerator with Flexible Operand Resolution and Layer-Wise Hybrid StationarityNicolas Chauvaux, Adrian Kneip, Charlotte Frenkel2026-09-08下载Compute-in-memory (CIM) accelerators for spiking neural networks (SNNs) offer a promising solution for achieving μs-level inference latency and ultra-low energy in edge vision applications.
PENDA: An Efficient Processing Element via Norm-of-Difference for Deep Learning AcceleratorsKai-Chieh Hsu, Tian-Sheuan Chang2026-09-08下载Inner product computation dominates the computational cost of deep learning models; thus, accelerating this primitive is key to improving hardware efficiency.
Routing Dense Layouts with History-Aware Offline Reinforcement Learning using LSTMAfsara Khan, Austin Rovinski2026-09-08下载Detailed routing remains a dominant runtime bottleneck in physical design due to increasing complexity of design rules. Modern routers can struggle to resolve persistent violations under dense operati...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource ManagementAntonino Vaccarella, Lanpei Li, Vincenzo Lomonaco, Massimo Coppola2026-09-08下载Managing resources across IoT, edge, and cloud layers calls for continuous, context-aware decisions under constraints that rarely stay fixed. Deep reinforcement learning (DRL) handles this class of pr...
Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAsJoseph K. L. Lee, Mehrdad Malekmohammadi, Hong-Sheng Zheng, Shuli Shu, Cheick Doumbia, Kalman Szenes, Mehran Zamani Abnili, Thomas Ainsworth, Matthew Seymour, Thomas Germain, Leonhard Neuhaus, Josh Izaac, Lee J. O'Riordan2026-09-08下载Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders.
Distributed Linear Programming on GPU Clusters at Extreme ScaleArnaud Deza, Santanu Dey, Pascal Van Hentenryck2026-09-08下载Large linear programs can exceed the memory of a single compute node. Although first-order methods replace sparse factorizations with GPU-suited matrix-vector products, other solver phases can reintro...
Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor CoresSatoshi Matsuoka2026-09-08下载FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibriu...
Impossibility of One-Way One-Round Quantum 4-Coloring via Matrix-Space StabilityTom Gur, Longcheng Li2026-09-08下载We show that one-way one-round quantum LOCAL algorithms cannot 44-color directed cycles with high probability, even with unbounded local computation and quantum message length.
GraphFAS: A Distributed System for Automated Graph Feature Generation and Selection in Industrial Transaction NetworksYice Luo, Yun Zhu, Xi Chen, Yongchao Liu, Xintan Zeng, Chengying Huan, Kai Zhang, Jinrui Zhang, Juelu Zhang, Jiajun Zheng2026-09-08下载Industrial fraud detection often relies on costly expert-crafted features that overlook graph-structured relational signals, while GNNs often do not meet the interpretability and deployment requiremen...
ContinuumBench: Benchmarking Joint Autoscaling and Placement Across Evaluation Regimes in the Cloud-Edge ContinuumLanpei Li, Antonino Vaccarella, Vincenzo Lomonaco, Massimo Coppola2026-09-08下载Cloud-edge controllers coordinate service placement, replica scaling, and resource pre-warming to keep end-to-end latency within application deadlines.
Exploring the Genesis Platform Capabilities to Accelerate Scientific Discovery in OPALDaniel Rosendo, Renan Souza, Kelsey Carter, John Lagergren, Frédéric Suter, Shelaine L. Curd, David Weston, Rafael Ferreira da Silva2026-09-08下载Autonomous, cross-facility science requires capabilities that no individual project should have to build for itself: managed execution for long-lived services, versioned distribution of models to remo...
Tools-CC-Bench: a Benchmark Suite for Collective Communication with Compression in HPC and AI WorkloadsHaozhe Fan, Wei Wang, Xingchen Liu, Man Liu, Xingjian Tian, Haoquan Long, Zedong Liu, Daran Sun, Jinwu Yang, Bo Yang, Jie Liu, Yonggang Che, Hairui Zhao, Guangming Tan, Dingwen Tao2026-09-08下载Distributed HPC and LLM workloads increasingly require efficient communication for scalability, yet growing data movement has become a major performance bottleneck.
Measuring Sustainability in Multi-Scale High-Performance ComputingCarlos J Barrios, Frédéric Le Mouël, Yves Denneulin2026-09-08下载The transition from traditional High Performance Computing (HPC) to the Computing Continuum emphasizes efficient resource management and sustainable practices across Multi-Scale hybrid architectures.
Sample-Guided Exact Top-K Selection for Long-Context Sparse AttentionSiran Liu, Yang Xue, Theo Tang, Changxu Shao, Qian Cheng, Haimeng Ren, Donghua Jiang, Haipeng Ming, Lehua Ding, Zhonghan Lin, Shengying Wei, Wei Liu, Kai Liu, Jianchen Zhu2026-09-08下载Sparse attention bounds downstream attention work by retaining a fixed-size subset of indexed tokens, but its standalone exact Top-KK stage must still process materialized score rows whose length gro...
A Measurement Study of LLM Inference Trade-offs Across Edge Continuum HardwareMaysam Khatib, Moysis Symeonides, Demetris Trihinas, George Pallis, Marios D. Dikaiakos2026-09-08下载Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy...
Transversal Fanout for Fault Tolerant Distributed Quantum Computing: Analysis and ApplicationSeng W. Loke2026-09-08下载We study a resource-efficient approach for implementing logical fanout operations in fault-tolerant distributed quantum computing using transversal operations on quantum error-correcting code blocks.
SemBridge: Compiling Consumer Observations into Cross-Stack Communication PlansGenlang Chen, Junyi Zhu, Yuanshan Lin2026-09-08下载Distributed-tensor systems specify where values reside, while collective systems optimize how requested operations execute. At a boundary between vendor runtimes that cannot share a native communicato...
Generalized DBLog: A Verified Contract for Interleaving Database Rows with a Change LogAndreas Andreakis2026-09-08下载Change-data capture (CDC) feeds downstream systems like caches, search indexes, and data warehouses from a database's log of committed row changes.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Prototyping QoE-Aware Rate Adaptation in Cellular Networks with Commercial ApplicationsSzilveszter Nádas, Lars Ernström, Dan Druta, Igor Pruzhansky, David Lindero, Jonathan Lynam, Eric Petajan2026-09-08下载Prior work has shown that QoE-aware resource sharing for real-time interactive video can support up to three times more simultaneous sessions at acceptable quality compared to rate-fair allocation.
Concept drift mitigation through community and spectral graph analysis for the detectionof cyberattacks in network trafficJulien Michel, Abdul Qadir Khan, Majed Jaber, Pierre Parrend2026-09-08下载In network traffic, legitimate behaviours and attack techniques evolve jointly - the phenomenon known as 'concept drift' [1]. Every detector is thereby left obsolete between two updates, and always on...
QPS-ToR: A Parallel Iterative Switching Algorithm for Reconfigurable Optical Datacenter SwitchingDongzhao Song, Qianru Yu, Jun Xu2026-09-08下载Reconfigurable optical data center networks (RODCNs) have emerged as a promising solution for scaling DCN capacity, yet their scheduling mechanisms remain a performance bottleneck: traffic-oblivious s...
Efficient User Association and Wireless Scheduling with Shorter Time-Scale Rate AdaptationXiaoyi Wu, Huacheng Zeng, Bin Li2026-09-08下载Rate adaptation is a crucial mechanism in IEEE 802.11 networks and next-generation cellular systems. Since the time scale for rate adaptation is typically much shorter than that for user association a...
Improving 5G AI-RAN MCS Selection by Predicting RetransmissionsTamerlan Aghayev, Maxime Elkael, Michele Polese, Reshma Prasad, Salvatore D'Oro, Yunseong Lee, Koichiro Furueda, Tommaso Melodia2026-09-08下载Link Adaptation (LA) in 5G NR is inherently reactive, relying on channel measurements and HARQ feedback that may become quickly obsolete when the channel changes quickly.
A New Backscattering Dual-Polarized Rectenna for Wireless Power Transfer and IoT ApplicationsTaki Eddine Djidjekh, Quentin Bernyer, Alexandru Takacs2026-09-08下载This paper proposes an innovative dual-polarized backscattering rectenna that operates in two distinct modesenergy harvesting and backscattering modulation-driven by two-bit digital control signals.
Hybrid Continuous DoA Estimation with Shared-Radius Co-Prime Circular ArraysKeyvan Aghababaiyan2026-09-08下载This paper proposes a shared-radius co-prime circular array for high-resolution, continuous 2D Direction-of-Arrival (DoA) estimation in 3D space, jointly estimating azimuth and elevation angles.
CleanCity-BinSense: An IoT-Enabled Smart Waste Management System with Configurable Real-Time Fill Monitoring and Nearest-Neighbor Route OptimizationMohammad Adnan Kabir, Intifad Muhammad Sayeed2026-09-08下载CleanCity-BinSense addresses inefficiencies in urban waste management in developing cities, where fixed-schedule collection routes lead to overflowing bins and wasted fuel.
AI-Native Orchestration in the 6G Continuum: Evolving Operator Platforms with Agentic AIClaudia Carballo González, Hatim Chergui, Sergio Giménez-Antón, Mohammadreza Mosahebfard, Juan Sebastián Camargo, Pouria Sayyad Khodashenas, Vasileios Theodorou, Christos Verikoukis2026-09-08下载As Sixth-Generation (6G) networks evolve towards a seamless Cloud-Edge-Internet of Things (IoT) continuum, autonomous orchestration across distributed compute and network domains becomes critical.
Toward Fully Autonomous 6G Networks: AI-driven Operational Efficiency and OptimizationDavid Reiss, Oriol Sallent, Miguel Catalan-Cid, Daniel Camps-Mur2026-09-08下载Mobile networks evolution is characterized by a substantial increase in system complexity, driven by the need to accommodate a growing number of heterogeneous services on top of the digital infrastruc...
QoS-Aware RACH Preamble Slicing via Quota-Projected Branching Deep Reinforcement LearningJiulin Guo, Jiahan Xu, Jiashuo Zhang, Heng Yang, Yizhen Sun, Yutong Xie, Shanshan Li, Zhenyu Liu, Lei Zhang2026-09-08下载Quality-of-service (QoS)-aware random access requires adaptive allocation of a finite random access channel (RACH) preamble budget across heterogeneous traffic and access procedures.
Information-Entropy-Driven Fault Propagation Modeling for Probabilistic Network Performance PredictionLusha Mo, Fengxiao Tang, Xiaonan Wang, Ming Zhao2026-09-08下载Network faults can trigger cascading effects that cause abrupt and nonstationary performance degradation. Existing learning-based performance predictors mainly focus on normal operation or treat fault...
6SEVEN: System for EValuating IPv6 ENumeration algorithmsChase Kanipe, Erik Rye, Dave Levin, Robert Beverly2026-09-08下载The vast, sparsely populated, and often ephemeral IPv6 address space makes discovering active addresses challenging. In response, the community has developed over thirty different IPv6 Target Generati...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Ozaki 2.5: Engineering the Deconstruction Path of fp64-Emulated Dense Matrix Multiplication on FP8 Tensor CoresSatoshi Matsuoka2026-09-08下载FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes (the deconstruction term in the Tensor-Memory Equilibriu...

基于 VitePress 构建 · 使用本地搜索查找论文