Skip to content

2026-09-14 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
FINNAS: FINN-Guided Hardware-Aware NAS and Pruning for FPGA Jet Substructure ClassificationEva Chauffour, Changhong Li, Georgios Floros, Shreejith Shanker2026-09-14下载FPGAs are well suited to deploying quantised neural networks (QNNs) under strict accuracy, latency, and resource constraints; however, identifying efficient model-accelerator combinations commonly req...
FSNIC: A Low-Latency Flow-Based Intrusion Detection Architecture for FPGA SmartNICsNise O'Cuill, Changhong Li, Georgios Floros, Shreejith Shanker2026-09-14下载Modern data centres require high-performance networking alongside effective real-time security. Traditional Intrusion Detection Systems (IDS) commonly rely on general-purpose processors and often stru...
EBL: Efficient Broad Learning for Distributed Adaptive Harmonic AnalysisChanghong Li, Georgios Floros, Biswajit Basu, Shreejith Shanker2026-09-14下载Renewable energy systems and electrified transport have found widespread adoption in recent years. The integration of these non-linear loads, dominated by electric vehicle (EV) charging, however, has ...
The World Model Hardware AcceleratorShashank Chaurasia2026-09-14下载Diffusion transformers invert the arithmetic that autoregressive decoding made familiar. There is no token-by-token recurrence: every denoising step is a full-sequence forward pass over static shapes,...
Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility TwinWeqaar Janjua, Eoin OConnell, Mihai Penica2026-09-14下载Modern AI and HPC systems integrate accelerators, high-speed networks, and management controllers at rack scale. Developing software for this infrastructure typically requires access to scarce, costly...
Trillion-Parameter MoE in a Box: Decoupling Memory Provisioning with High-Bandwidth FlashPengfei Xia, Tuo Hao, Shengwei Li, Jinjing Chen, Shiru Wei, Wenjun Zou, Rui Zhang, Hui Zang2026-09-14下载An MoE appliance for trillion-parameter models at low concurrency must host terabytes of weights on one node and serve prefill and decode with fixed resources.
DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance EngineeringSiming Huang2026-09-14下载We present the enablement, correctness recovery, and performance engineering of DeepSeek-V4-Flash inference on AMD Instinct MI250 GPUs using the gfx90a/CDNA2 architecture.
A Memristive Synapse for Online STDP Learning and Inference in SNNsElia Mateu-Barriendos, Álvaro Gómez-Pau, Daniel Arumí, Rosa Rodríguez-Montañés, Salvador Manich2026-09-14下载This work presents a fully analog memristive synaptic circuit for online spike-timing-dependent plasticity (STDP) learning in spiking neural networks (SNNs).
LLM-enabled Behavior Driven Development Workflow for Formally Verified Hardware DesignsLuca Müller, Qian Liu, Rolf Drechsler2026-09-14下载Recently, the use of Large Language Models (LLMs) for different tasks in the Electronic Design Automation (EDA) life-cycle has been studied extensively, but an integrated view is lacking.
FlashGPU-sim: Enabling GPU Modeling for Modern Architectures and AI WorkloadsSiying Yu, Yixun Hong, Guozhi Qiu, Jingci Liu, Feng Gu, Chenbo Geng, Zhengrong Wang, Chen Zhang, Bei Yu2026-09-14下载As AI becomes increasingly ubiquitous, modern AI systems are shaped by a tight software-hardware co-design loop. Later GPUs expose features such as asynchronous data movement, tensor core pipelines, a...
A 25-μs/inf Event-driven Graph Neural Network Processor with Spatiotemporal Caching and Spline Convolution for Ultra-low-latency AI at the EdgeAdrian Kneip, Martin Lefebvre, Daniel Gehrig, Victoria Catalán Pastor, Davide Scaramuzza, Marian Verhelst, Charlotte Frenkel2026-09-14下载Dynamic-vision-sensor (DVS) cameras generate events on a per-pixel basis with a μs-level temporal resolution, calling for new algorithm-hardware co-design approaches compared to standard frame-based...
Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive AcceleratorsYuyeong Shin2026-09-14下载Eight-bit integer (INT8) post-training quantization is the default recipe for edge deployment, under a widely held assumption: INT8 makes inference faster at a small, predictable accuracy cost, and a ...
FastPair: GPU-Optimized String DecodingJoseph Isaacs, Francesco Gargiulo, Peter Boncz, Robert Kruszewski, Nicholas Gates, Rossano Venturini, Will Manning, Martin Prammer2026-09-14下载Modern data systems compress data at rest and decompress it only when needed to preserve interconnect bandwidth. This design is often inefficient on GPU-based compute platforms because many convention...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed SystemsJun He, Deying Yu2026-09-14下载In agentic distributed systems, an agent may be authorized to mutate external infrastructure while lacking evidence that the mutation is ready to execute.
BOA: Beamwidth Online Adaptation for Filtered-ANNS on a GPUFarhana Akter Tumpa, Rajiv Gupta2026-09-14下载Filtered approximate nearest neighbor search, i.e. returning the top-k vectors nearest to a query vector among those satisfying one or more attribute predicates, has become a fundamental operation in ...
Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility TwinWeqaar Janjua, Eoin OConnell, Mihai Penica2026-09-14下载Modern AI and HPC systems integrate accelerators, high-speed networks, and management controllers at rack scale. Developing software for this infrastructure typically requires access to scarce, costly...
SWB-DM: A Calibrated Sliced-Wasserstein-Barycenter Aggregator with Delayed-Momentum Caching for Byzantine-Robust Federated Learning under Partial ParticipationSaranraj S, Saranya M S, Alex David S, Ajay Kumar A2026-09-14下载Robust aggregation methods for federated learning quietly rest on a fragile assumption: that whoever shows up in a given round is a fair sample of the full population. In practice, they rarely are.
Scalability and Performance Evaluation of Federated Learning Frameworks: A Comparative AnalysisBassel Soudan, Sohail Abbas, Ahmed Kubba, Manar Wasif Abu Talib, Qassim Nasir2026-09-14下载This paper presents a systematic examination and experimental comparison of the prominent Federated Learning (FL) frameworks FedML, Flower, Substra, and OpenFL.
CIDERS: Cloud-Edge LLM Collaborative Learning via Accelerating Personalized Bilevel OptimizationVictor H. Chen, Hairui Yu, Stella K. Chung, Hong Yan2026-09-14下载Amid the rapid advancement of physical-world intelligence, cloud-edge collaborative large language models (LLMs) have emerged as a promising roadmap for practical LLM deployment.
DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance EngineeringSiming Huang2026-09-14下载We present the enablement, correctness recovery, and performance engineering of DeepSeek-V4-Flash inference on AMD Instinct MI250 GPUs using the gfx90a/CDNA2 architecture.
Opportunistic ZGC: Leveraging Idle Cores for More Effective Concurrent Garbage CollectionJacob Malloy, Michael R. Jantz, Terry Jones2026-09-14下载Managed language runtimes often provide concurrent garbage collectors so that latency-critical applications with large working sets can keep running while most collection work proceeds in the backgrou...
When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool BoundaryArtem Trofimov, Boris Novikov2026-09-14下载AI agents increasingly execute long-running workflows that externalize effects through independently supplied tools. Under retries, speculative execution, concurrency, and partial failures, the result...
FlashGPU-sim: Enabling GPU Modeling for Modern Architectures and AI WorkloadsSiying Yu, Yixun Hong, Guozhi Qiu, Jingci Liu, Feng Gu, Chenbo Geng, Zhengrong Wang, Chen Zhang, Bei Yu2026-09-14下载As AI becomes increasingly ubiquitous, modern AI systems are shaped by a tight software-hardware co-design loop. Later GPUs expose features such as asynchronous data movement, tensor core pipelines, a...
ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI DatacentersRui Lu, Rui Ge, Huanghuang Liang, Xiaobo Zhou, Dan Wang2026-09-14下载Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and c...
Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive LawsTristan Chenaille, Francesca Cuteri, Rapha{ë}l Prat, Guillaume Latu, Thomas Helfer2026-09-14下载Many applications require solving large numbers of independent linear systems on GPUs. While this need is well addressed for small to large systems, tiny ones, understood here as systems of dimension ...
FastPair: GPU-Optimized String DecodingJoseph Isaacs, Francesco Gargiulo, Peter Boncz, Robert Kruszewski, Nicholas Gates, Rossano Venturini, Will Manning, Martin Prammer2026-09-14下载Modern data systems compress data at rest and decompress it only when needed to preserve interconnect bandwidth. This design is often inefficient on GPU-based compute platforms because many convention...
Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCacheFrank Li2026-09-14下载External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5.
Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance BoundariesFrank Li2026-09-14下载Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality.
Agentic Autoscaling through Worker-Pool Orchestration for LLM-driven Text Classification in Cloud Computing EnvironmentsBablu Kumar, Anshul Verma, Rajkumar Buyya2026-09-14下载The growing adoption of large language model (LLM)-based systems for large-scale text processing has created a critical need for dynamic autoscaling to manage high-latency, bursty, and computationally...
Stability-Aware Proactive Autoscaling Using a Double Deep Q-Network in Cloud Computing EnvironmentsBablu Kumar, Anshul Verma, Rajkumar Buyya2026-09-14下载Dynamic workloads and latency-sensitive applications require efficient autoscaling in cloud computing environments. However, most existing approaches rely on reactive mechanisms based on static thresh...
Fast Stencil Computations on a Single Arbitrarily Moving IntervalAaron Gregory2026-09-14下载A stencil computation repeatedly updates every cell of a grid from its neighbours' values at the previous timestep. Simulating T steps on N cells directly costs Theta(NT), and a line of work beginning...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge CachingFarnaz Niknia, Ping Wang2026-09-14下载Wireless edge caching networks typically consist of many independent Base Stations (BSs), each facing its own request rate and content popularity profile.
Utility-Based Path Selection and Configuration in Quantum Networks via Layered Shortest PathsLeonardo Bacciottini, Subhransu Maji, Don Towsley, Gayane Vardoyan2026-09-14下载A path in a quantum network is a chain of repeaters that distributes entanglement between two users. Selecting a path requires balancing the rate and quality (e.g.
Understanding the oversubscription behaviour of DragonFly+ networksVlad-Adrian Ulmeanu, Costin Raiciu, Iulian-Ilie Drăcea2026-09-14下载The Max-Host Dragonfly+ topology's original paper proves that there is a 2:1 worst-case oversubscription ratio in expectation for the permutation traffic pattern.
Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility TwinWeqaar Janjua, Eoin OConnell, Mihai Penica2026-09-14下载Modern AI and HPC systems integrate accelerators, high-speed networks, and management controllers at rack scale. Developing software for this infrastructure typically requires access to scarce, costly...
Private Information Retrieval With Arbitrary Privacy Requirements: Introduction and Capacity ResultsMohamed Nomeir, Shreya Meel, Sennur Ulukus2026-09-14下载In this paper, we introduce the problem of private information retrieval (PIR) under arbitrary privacy requirements, in a graph-based storage system.
Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge IntelligenceThai T. Vu, John Le, Tu N. Nguyen, Jun Shen, Quang Vinh Duong, Ha Nguyen2026-09-14下载This paper proposes FREDI (Fair Resource Allocation for Edge Dual-Threshold Inference), a secure wireless edge-intelligence framework for event-triggered inference in a cooperative user equipment (UE)...
AI-Native Open RAN: A Roadmap from xApps and rApps to Autonomous Network AgentsRyan Barker, Alireza Ebrahimi Dorcheh, Tolunay Seyfi, Mohammad Raihan Uddin, Alireza Mohammadhosseini, Julia Boone, Stephen Streit, Drew Schlesener, Fatemeh Afghah2026-09-14下载Open Radio Access Networks (O-RAN) have emerged as a transformative paradigm for future wireless systems by introducing openness, virtualization, disaggregation, and programmable intelligence through ...
Cloud Workflow Scheduling Based on Graph Attention-Driven Hierarchical Reinforcement LearningZongjin Li, Shaohan Feng, Chunxi Yang, Wenbo Wang2026-09-14下载Dynamic cloud workflow scheduling must balance deadline satisfaction, container utilization, and energy consumption while dealing with stochastic task-execution speeds, placement-dependent communicati...
Burst-mode timing recovery based on fourth-power phase detector for passive optical networksJi Zhou, Haide Wang, Xiaofeng Zhang, Zhiyang Liu, Miao Yu, Changyuan Yu, Liangchuan Li, Xiangjun Xin2026-09-14下载Driven by the ever-increasing capacity demands, 50G passive optical network (50G-PON) is ready for practical application. It is highly challenging to realize 50GHz burst-mode analog components; theref...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Dichoptic FoveationHenry Kam, Colin Groth, Jenna Kang, Pratham Saraf, Qi Sun, Kenneth Chen2026-09-14下载Interocular differences in visual perception can induce a variety of effects when fused by the brain. For example, prior works have found that carefully crafted binocular differences in local detail c...
The AI-Enabled Scientific FrontierGabriel Manso, Emma Fu, Neil Thompson2026-09-14下载As artificial intelligence's capabilities improve, it is increasingly viewed as a general scientific method. But how true are these claims? Does AI outperform all techniques, or only some, and how is ...
Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance RankingMd Arafat Hossain, Thomas Randall, Akash Dutta, Xingfu Wu, Rong Ge, Ali Jannesari2026-09-14下载As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide pro...
DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance EngineeringSiming Huang2026-09-14下载We present the enablement, correctness recovery, and performance engineering of DeepSeek-V4-Flash inference on AMD Instinct MI250 GPUs using the gfx90a/CDNA2 architecture.
Opportunistic ZGC: Leveraging Idle Cores for More Effective Concurrent Garbage CollectionJacob Malloy, Michael R. Jantz, Terry Jones2026-09-14下载Managed language runtimes often provide concurrent garbage collectors so that latency-critical applications with large working sets can keep running while most collection work proceeds in the backgrou...
Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice EvaluationM. Ali Bayram2026-09-14下载Adding answer options can lower multiple-choice scores without improving assessment validity. Turkish MMLU Pro examines this distinction using 12,000 Turkish-source questions across 58 sections.
ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI DatacentersRui Lu, Rui Ge, Huanghuang Liang, Xiaobo Zhou, Dan Wang2026-09-14下载Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and c...
Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive AcceleratorsYuyeong Shin2026-09-14下载Eight-bit integer (INT8) post-training quantization is the default recipe for edge deployment, under a widely held assumption: INT8 makes inference faster at a small, predictable accuracy cost, and a ...
Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCacheFrank Li2026-09-14下载External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5.
Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance BoundariesFrank Li2026-09-14下载Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality.
GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three SystemsXinyu Qiu, Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai2026-09-14下载We predict single-sequence model throughput from GGUF metadata using roofline-shaped predictors with quantization-specific scale factors fitted on reference models.

基于 VitePress 构建 · 使用本地搜索查找论文