Skip to content

2026-08-21 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Model Compression and Hardware-Aware Acceleration for Deep Learning on FPGAs: A Co-Design Taxonomy and Comparative AnalysisPeter Forcha, H. Kajekusumadhar, Mbua Peter, Muhammed Kawser, Audrey Cyriell Mo, Christophe Bobda2026-08-21下载Deploying deep neural networks on Field-Programmable Gate Arrays (FPGAs) requires joint reasoning about model compression and hardware acceleration, however the most comprehensive existing cross-platf...
Programmable Compute-in-Transit using Integrated PhotonicsImon Kundu, Livi Hammond, Jamie Todd, Kriti Goel, Peter Simpson, Jaganath Rajendra, Florent Michel, Jack Crawford, Flavio Bergamaschi, Robert Todd, Nick New2026-08-21下载Modern hardware designs for AI and cryptography treat data transit and processing separately. At Optalysys we have demonstrated programmable Photonic hardware that computes mathematical functions on d...
AI with Authority, from Application to SiliconJason Hickey2026-08-21下载For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here we report that generative AI inverts this relationship: at AI speed, machine verif...
Assessing Triple Modular Redundancy for Wide-Link, Low-Latency NoC Routers: Reliability and Physical Design ChallengesChen Wu, Michael Rogenmoser, Luca Benini, Angelo Garofalo2026-08-21下载Protecting the Network-on-Chip (NoC) of physical-AI tile-based accelerators deployed in harsh environments against single-event effects (SEEs) is paramount for preventing NoC failures that can lead to...
SPICE: Speculative Prefetching with Low-Rank Expert Surrogates and Heterogeneous Orchestration for MoE Inference AccelerationYongxiang Lyu, Ning Li, Bonian Jia2026-08-21下载Mixture-of-Experts (MoE) models are increasingly used in LLMs because sparse activation decouples model capacity from compute cost. However, the large expert parameter footprint often exceeds GPU memo...
Event-triggered Implicit Perturbation for Zeroth-Order Fine-Tuning of Spiking TransformersTengteng Lei, Prabodh Katti, Rashi Dutt, Houssem Sifaou, Tan Peng, Osvaldo Simeone, Kai Xu, Bipin Rajendran2026-08-21下载Zeroth-order (ZO) optimization estimates gradients using only forward-pass evaluations, making it suitable for fine-tuning non-differentiable, event-driven spiking neural networks (SNNs).

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought ReasoningYujie Zhang, Bin Gao, Tulika Mitra2026-08-21下载Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature increases decoding latency and memory usage.
Thermo-FL: Thermal-Aware Robust Federated Fine-Tuning of Large Language Models for Edge AIShiva Shrestha, Kazi Shaharair Sharif, Zongxing Xie, Jiajing Huang, Anhao Xiang, Honghui Xu2026-08-21下载Federated fine-tuning enables large language models to adapt on edge devices without centralizing private data, but practical deployments must address hardware instability and adversarial update corru...
GrAND: GPU-based Dynamic Graph Indexes for Approximate Nearest Neighbour SearchKarthik Venkatasubba, Shivendra Deshpande, Shivram S, Jyothi Vedurada2026-08-21下载Modern Approximate Nearest Neighbour Search (ANNS) applications operate over continuously evolving vector collections and require graph indexes that sustain high-throughput searches while incorporatin...
HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel OptimizationJinghao Wang, Qiqi Gu, Chenpeng Wu, Jianguo Yao, Haibing Guan, Xijun Li2026-08-21下载High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardware evolves rapidly, developing efficient methods for automat...
Integrating a Python Dynamical core into ICONMauro Bianco, Till Ehrengruber, Enrique González Paredes, Andreas Jocksch, Christos Kotsalos, Ioannis Magkanaris, Philip Müller, Edoardo Paone, Mikael Simberg, Hannes Vogt, Jacopo Canton, Yilu Chen, Anurag Dipankar, Nicoletta Farabullini, Michael Jähn, Matthieu Leclair, Ong Chia Rui, Nathan Beech, Nicolas Gruber, Christoph Müller, Daniel Hupp, Xavier Lapillonne2026-08-21下载The transition of Earth-system models to exascale is often hindered by rigid, monolithic Fortran codebases and maintenance-heavy compiler directives.
BackDFL: A Unified Benchmark For Backdoor Attacks and Defenses In Decentralized Federated LearningMouhamed Amine Bouchiha, Gregory Blanc, Yufei Han2026-08-21下载Decentralized Federated Learning (DFL) promises trust-free collaborative learning by replacing the centralized parameter server with peer-to-peer model exchange.
AI Infrastructure in Space: How Far Can We Go?Qing Li, Qiyang Zhang, Daliang Xu, Tianze Huang, Dingge Zhang, Yihao Zhao, Xiaolong Huang, Jinfeng Wen, Xiameng Hu, Tao Qi, Mengwei Xu, Shangguang Wang, Xuanzhe Liu2026-08-21下载Satellites are becoming programmable computing platforms capable of running increasingly demanding AI workloads. This shift raises a systems problem: how can AI services remain deployable, manageable,...
TreeWY: Speculative Verification for Gated DeltaNet HybridsSneha Murthy Ghantasala2026-08-21下载Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state instead of a growing key-value (KV) cache.
PRICE: Pricing-based Resource Incentives for Quality-of-Result-aware Computing at the EdgeUwe Gropengießer, Sebastian Frenz, Max Mühlhäuser2026-08-21下载Edge nodes are capacity-constrained by design, yet many edge workloads can trade result quality for resource efficiency at runtime. Existing edge pricing mechanisms largely treat requests as fixed-con...
MEMPOWER: Efficient Power Management with Fine-grained Memory Analysis and Modeling for HPC WorkloadsNanda Velugoti, Joseph Manzano, Andres Marquez, Nathan Tallent, Kyle Hale2026-08-21下载Managing the energy consumption and power efficiency of parallel applications is a significant issue in both HPC environments and in the cloud.
Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor CoresXiang Fu, Jixiang Ma, Xinpeng Zhang, Peng Zhao, Shuai Lu, Xu Tony Liu2026-08-21下载Convolution is a principal computational bottleneck in deep neural networks, and its efficiency depends on tight integration between algorithms and GPU hardware.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Scalable Quantum Key Distribution via GHZ Entanglement and Qubit ReuseTasdiqul Islam, Rasman Mubtasim Swargo, Engin Arslan, Md Arifuzzaman2026-08-21下载Conventional Quantum Key Distribution (QKD) requires the transmission of qubits proportional to or exceeding the length of the key, as protocols such as BB84 transmit more qubits than the final key si...
Age-Optimal Target Wake Time: Provably Good Wake Schedules for Energy-Constrained Wi-Fi Status UpdatingHaoyu Wang, Bo Sheng, Xiaoqian Zhang2026-08-21下载Target Wake Time (TWT), introduced in IEEE 802.11ax, lets an access point schedule exactly when each station wakes, transmits, and dozes. Existing TWT schedulers optimize energy or throughput, treatin...
Z2Z^2-ACT: End-to-End Verifiable Agentic Intent Control for Open 6G RANSunder Ali Khowaja, Kapal Dev, George C. Alexandropoulos2026-08-21下载With the progression in open and disaggregated 6G radio access networks, it is expected that the system will be able to host multi-vendors. In order to host multi-vendors, it is essential that AI-assi...
Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-JudgeRishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu2026-08-21下载Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause explanations and recommended actions.
Tools for Reducing Service Time in Near-Term Quantum NetworksJake Smith, Thomas R. Beauchamp, Scarlett Gauthier, Oumayma Bouchmal, Stephanie Wehner2026-08-21下载Architectures have been proposed to control entanglement generation in multi-user quantum networks. To allow time for local operations and classical communication at end nodes, these architectures ins...
Orchra: Stateful-aware Cross-slice Workload Migrations in the 6G Control PlaneAnthony Kiggundu, Bin Han, Hans D. Schotten2026-08-21下载Network slicing is a foundational capability of Fifth Generation (5G)-Advanced and emerging Sixth Generation (6G) networks, yet practical support for seamless runtime slice transitions remains limited...
Explainable Adaptive Zero Trust Framework for AWS with Adversarial Robustness EvaluationOm Singh, Yagyaraj Pandey, Nandini Pathak2026-08-21下载Cloud environments built on Amazon Web Services face a structural security vulnerability: once a credential passes authentication, the resulting session is often treated as trusted for its entire dura...
Mitigating Proxy-Induced Traffic Drift in Website Fingerprinting via Model-Agnostic Traffic TailoringLinxiao Yu, Tianyu Cui, Xinhao Deng, Yuqi Qing, Jun Tao, Ke Xu, Qi Li2026-08-21下载Website fingerprinting (WF) based on deep learning can effectively identify websites from encrypted traffic. However, users often rely on proxy protocols to bypass censorship, and the diversity of the...
Fluid-Dynamic Interference Modeling for LEO Mega-Constellations: A Spatiotemporal Kinetic Field ApproachWen-Yu Dong, Weiwei Jiang, Song Zhao, Rui-Si Han, Qi Bi, Sheng Chen2026-08-21下载Low Earth orbit (LEO) mega-constellations create a highly non-stationary interference environment that cannot be accurately captured by static stochastic-geometry snapshots.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Portable to Efficient: Auto-Tuning Hardware-Agnostic GPU Kernels in JuliaFloris-Jan Willemsen, Evelyne Ringoot, Alan Edelman2026-08-21下载Traditionally, GPU kernels have been developed and optimized within vendor-specific programming models to achieve high performance, resulting in software that is difficult to optimize and adapt across...
TreeWY: Speculative Verification for Gated DeltaNet HybridsSneha Murthy Ghantasala2026-08-21下载Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state instead of a growing key-value (KV) cache.
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMsBakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús2026-08-21下载Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits.

基于 VitePress 构建 · 使用本地搜索查找论文