2026-09-14
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| FINNAS: FINN-Guided Hardware-Aware NAS and Pruning for FPGA Jet Substructure Classification | Eva Chauffour, Changhong Li, Georgios Floros, Shreejith Shanker | 2026-09-14 | 下载 | FPGAs are well suited to deploying quantised neural networks (QNNs) under strict accuracy, latency, and resource constraints; however, identifying efficient model-accelerator combinations commonly req... |
| FSNIC: A Low-Latency Flow-Based Intrusion Detection Architecture for FPGA SmartNICs | Nise O'Cuill, Changhong Li, Georgios Floros, Shreejith Shanker | 2026-09-14 | 下载 | Modern data centres require high-performance networking alongside effective real-time security. Traditional Intrusion Detection Systems (IDS) commonly rely on general-purpose processors and often stru... |
| EBL: Efficient Broad Learning for Distributed Adaptive Harmonic Analysis | Changhong Li, Georgios Floros, Biswajit Basu, Shreejith Shanker | 2026-09-14 | 下载 | Renewable energy systems and electrified transport have found widespread adoption in recent years. The integration of these non-linear loads, dominated by electric vehicle (EV) charging, however, has ... |
| The World Model Hardware Accelerator | Shashank Chaurasia | 2026-09-14 | 下载 | Diffusion transformers invert the arithmetic that autoregressive decoding made familiar. There is no token-by-token recurrence: every denoising step is a full-sequence forward pass over static shapes,... |
| Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin | Weqaar Janjua, Eoin OConnell, Mihai Penica | 2026-09-14 | 下载 | Modern AI and HPC systems integrate accelerators, high-speed networks, and management controllers at rack scale. Developing software for this infrastructure typically requires access to scarce, costly... |
| Trillion-Parameter MoE in a Box: Decoupling Memory Provisioning with High-Bandwidth Flash | Pengfei Xia, Tuo Hao, Shengwei Li, Jinjing Chen, Shiru Wei, Wenjun Zou, Rui Zhang, Hui Zang | 2026-09-14 | 下载 | An MoE appliance for trillion-parameter models at low concurrency must host terabytes of weights on one node and serve prefill and decode with fixed resources. |
| DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering | Siming Huang | 2026-09-14 | 下载 | We present the enablement, correctness recovery, and performance engineering of DeepSeek-V4-Flash inference on AMD Instinct MI250 GPUs using the gfx90a/CDNA2 architecture. |
| A Memristive Synapse for Online STDP Learning and Inference in SNNs | Elia Mateu-Barriendos, Álvaro Gómez-Pau, Daniel Arumí, Rosa Rodríguez-Montañés, Salvador Manich | 2026-09-14 | 下载 | This work presents a fully analog memristive synaptic circuit for online spike-timing-dependent plasticity (STDP) learning in spiking neural networks (SNNs). |
| LLM-enabled Behavior Driven Development Workflow for Formally Verified Hardware Designs | Luca Müller, Qian Liu, Rolf Drechsler | 2026-09-14 | 下载 | Recently, the use of Large Language Models (LLMs) for different tasks in the Electronic Design Automation (EDA) life-cycle has been studied extensively, but an integrated view is lacking. |
| FlashGPU-sim: Enabling GPU Modeling for Modern Architectures and AI Workloads | Siying Yu, Yixun Hong, Guozhi Qiu, Jingci Liu, Feng Gu, Chenbo Geng, Zhengrong Wang, Chen Zhang, Bei Yu | 2026-09-14 | 下载 | As AI becomes increasingly ubiquitous, modern AI systems are shaped by a tight software-hardware co-design loop. Later GPUs expose features such as asynchronous data movement, tensor core pipelines, a... |
| A 25-μs/inf Event-driven Graph Neural Network Processor with Spatiotemporal Caching and Spline Convolution for Ultra-low-latency AI at the Edge | Adrian Kneip, Martin Lefebvre, Daniel Gehrig, Victoria Catalán Pastor, Davide Scaramuzza, Marian Verhelst, Charlotte Frenkel | 2026-09-14 | 下载 | Dynamic-vision-sensor (DVS) cameras generate events on a per-pixel basis with a μs-level temporal resolution, calling for new algorithm-hardware co-design approaches compared to standard frame-based... |
| Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive Accelerators | Yuyeong Shin | 2026-09-14 | 下载 | Eight-bit integer (INT8) post-training quantization is the default recipe for edge deployment, under a widely held assumption: INT8 makes inference faster at a small, predictable accuracy cost, and a ... |
| FastPair: GPU-Optimized String Decoding | Joseph Isaacs, Francesco Gargiulo, Peter Boncz, Robert Kruszewski, Nicholas Gates, Rossano Venturini, Will Manning, Martin Prammer | 2026-09-14 | 下载 | Modern data systems compress data at rest and decompress it only when needed to preserve interconnect bandwidth. This design is often inefficient on GPU-based compute platforms because many convention... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Cognitive Admission Control: Risk-Conditioned Assurance for Consequential Actions in Agentic Distributed Systems | Jun He, Deying Yu | 2026-09-14 | 下载 | In agentic distributed systems, an agent may be authorized to mutate external infrastructure while lacking evidence that the mutation is ready to execute. |
| BOA: Beamwidth Online Adaptation for Filtered-ANNS on a GPU | Farhana Akter Tumpa, Rajiv Gupta | 2026-09-14 | 下载 | Filtered approximate nearest neighbor search, i.e. returning the top-k vectors nearest to a query vector among those satisfying one or more attribute predicates, has become a fundamental operation in ... |
| Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin | Weqaar Janjua, Eoin OConnell, Mihai Penica | 2026-09-14 | 下载 | Modern AI and HPC systems integrate accelerators, high-speed networks, and management controllers at rack scale. Developing software for this infrastructure typically requires access to scarce, costly... |
| SWB-DM: A Calibrated Sliced-Wasserstein-Barycenter Aggregator with Delayed-Momentum Caching for Byzantine-Robust Federated Learning under Partial Participation | Saranraj S, Saranya M S, Alex David S, Ajay Kumar A | 2026-09-14 | 下载 | Robust aggregation methods for federated learning quietly rest on a fragile assumption: that whoever shows up in a given round is a fair sample of the full population. In practice, they rarely are. |
| Scalability and Performance Evaluation of Federated Learning Frameworks: A Comparative Analysis | Bassel Soudan, Sohail Abbas, Ahmed Kubba, Manar Wasif Abu Talib, Qassim Nasir | 2026-09-14 | 下载 | This paper presents a systematic examination and experimental comparison of the prominent Federated Learning (FL) frameworks FedML, Flower, Substra, and OpenFL. |
| CIDERS: Cloud-Edge LLM Collaborative Learning via Accelerating Personalized Bilevel Optimization | Victor H. Chen, Hairui Yu, Stella K. Chung, Hong Yan | 2026-09-14 | 下载 | Amid the rapid advancement of physical-world intelligence, cloud-edge collaborative large language models (LLMs) have emerged as a promising roadmap for practical LLM deployment. |
| DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering | Siming Huang | 2026-09-14 | 下载 | We present the enablement, correctness recovery, and performance engineering of DeepSeek-V4-Flash inference on AMD Instinct MI250 GPUs using the gfx90a/CDNA2 architecture. |
| Opportunistic ZGC: Leveraging Idle Cores for More Effective Concurrent Garbage Collection | Jacob Malloy, Michael R. Jantz, Terry Jones | 2026-09-14 | 下载 | Managed language runtimes often provide concurrent garbage collectors so that latency-critical applications with large working sets can keep running while most collection work proceeds in the backgrou... |
| When Tool Calls Succeed but Workflows Fail: Anomalies at the Agent-Tool Boundary | Artem Trofimov, Boris Novikov | 2026-09-14 | 下载 | AI agents increasingly execute long-running workflows that externalize effects through independently supplied tools. Under retries, speculative execution, concurrency, and partial failures, the result... |
| FlashGPU-sim: Enabling GPU Modeling for Modern Architectures and AI Workloads | Siying Yu, Yixun Hong, Guozhi Qiu, Jingci Liu, Feng Gu, Chenbo Geng, Zhengrong Wang, Chen Zhang, Bei Yu | 2026-09-14 | 下载 | As AI becomes increasingly ubiquitous, modern AI systems are shaped by a tight software-hardware co-design loop. Later GPUs expose features such as asynchronous data movement, tensor core pipelines, a... |
| ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters | Rui Lu, Rui Ge, Huanghuang Liang, Xiaobo Zhou, Dan Wang | 2026-09-14 | 下载 | Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and c... |
| Accelerating the Solving of Many Tiny General Linear Systems on GPUs: Application to Constitutive Laws | Tristan Chenaille, Francesca Cuteri, Rapha{ë}l Prat, Guillaume Latu, Thomas Helfer | 2026-09-14 | 下载 | Many applications require solving large numbers of independent linear systems on GPUs. While this need is well addressed for small to large systems, tiny ones, understood here as systems of dimension ... |
| FastPair: GPU-Optimized String Decoding | Joseph Isaacs, Francesco Gargiulo, Peter Boncz, Robert Kruszewski, Nicholas Gates, Rossano Venturini, Will Manning, Martin Prammer | 2026-09-14 | 下载 | Modern data systems compress data at rest and decompress it only when needed to preserve interconnect bandwidth. This design is often inefficient on GPU-based compute platforms because many convention... |
| Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache | Frank Li | 2026-09-14 | 下载 | External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5. |
| Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries | Frank Li | 2026-09-14 | 下载 | Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. |
| Agentic Autoscaling through Worker-Pool Orchestration for LLM-driven Text Classification in Cloud Computing Environments | Bablu Kumar, Anshul Verma, Rajkumar Buyya | 2026-09-14 | 下载 | The growing adoption of large language model (LLM)-based systems for large-scale text processing has created a critical need for dynamic autoscaling to manage high-latency, bursty, and computationally... |
| Stability-Aware Proactive Autoscaling Using a Double Deep Q-Network in Cloud Computing Environments | Bablu Kumar, Anshul Verma, Rajkumar Buyya | 2026-09-14 | 下载 | Dynamic workloads and latency-sensitive applications require efficient autoscaling in cloud computing environments. However, most existing approaches rely on reactive mechanisms based on static thresh... |
| Fast Stencil Computations on a Single Arbitrarily Moving Interval | Aaron Gregory | 2026-09-14 | 下载 | A stencil computation repeatedly updates every cell of a grid from its neighbours' values at the previous timestep. Simulating T steps on N cells directly costs Theta(NT), and a line of work beginning... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Fast-Convergent Meta-RL via Gradient-Clustered BS Sampling for Edge Caching | Farnaz Niknia, Ping Wang | 2026-09-14 | 下载 | Wireless edge caching networks typically consist of many independent Base Stations (BSs), each facing its own request rate and content popularity profile. |
| Utility-Based Path Selection and Configuration in Quantum Networks via Layered Shortest Paths | Leonardo Bacciottini, Subhransu Maji, Don Towsley, Gayane Vardoyan | 2026-09-14 | 下载 | A path in a quantum network is a chain of repeaters that distributes entanglement between two users. Selecting a path requires balancing the rate and quality (e.g. |
| Understanding the oversubscription behaviour of DragonFly+ networks | Vlad-Adrian Ulmeanu, Costin Raiciu, Iulian-Ilie Drăcea | 2026-09-14 | 下载 | The Max-Host Dragonfly+ topology's original paper proves that there is a 2:1 worst-case oversubscription ratio in expectation for the permutation traffic pattern. |
| Cnuas: A Software-Defined AI/HPC Rack-scale Emulation Platform and Hyperscale Data Center Facility Twin | Weqaar Janjua, Eoin OConnell, Mihai Penica | 2026-09-14 | 下载 | Modern AI and HPC systems integrate accelerators, high-speed networks, and management controllers at rack scale. Developing software for this infrastructure typically requires access to scarce, costly... |
| Private Information Retrieval With Arbitrary Privacy Requirements: Introduction and Capacity Results | Mohamed Nomeir, Shreya Meel, Sennur Ulukus | 2026-09-14 | 下载 | In this paper, we introduce the problem of private information retrieval (PIR) under arbitrary privacy requirements, in a graph-based storage system. |
| Proportional-Fair Resource Allocation and Dual-Threshold Early-Exit Inference for Secure Cooperative Multi-Layer Edge Intelligence | Thai T. Vu, John Le, Tu N. Nguyen, Jun Shen, Quang Vinh Duong, Ha Nguyen | 2026-09-14 | 下载 | This paper proposes FREDI (Fair Resource Allocation for Edge Dual-Threshold Inference), a secure wireless edge-intelligence framework for event-triggered inference in a cooperative user equipment (UE)... |
| AI-Native Open RAN: A Roadmap from xApps and rApps to Autonomous Network Agents | Ryan Barker, Alireza Ebrahimi Dorcheh, Tolunay Seyfi, Mohammad Raihan Uddin, Alireza Mohammadhosseini, Julia Boone, Stephen Streit, Drew Schlesener, Fatemeh Afghah | 2026-09-14 | 下载 | Open Radio Access Networks (O-RAN) have emerged as a transformative paradigm for future wireless systems by introducing openness, virtualization, disaggregation, and programmable intelligence through ... |
| Cloud Workflow Scheduling Based on Graph Attention-Driven Hierarchical Reinforcement Learning | Zongjin Li, Shaohan Feng, Chunxi Yang, Wenbo Wang | 2026-09-14 | 下载 | Dynamic cloud workflow scheduling must balance deadline satisfaction, container utilization, and energy consumption while dealing with stochastic task-execution speeds, placement-dependent communicati... |
| Burst-mode timing recovery based on fourth-power phase detector for passive optical networks | Ji Zhou, Haide Wang, Xiaofeng Zhang, Zhiyang Liu, Miao Yu, Changyuan Yu, Liangchuan Li, Xiangjun Xin | 2026-09-14 | 下载 | Driven by the ever-increasing capacity demands, 50G passive optical network (50G-PON) is ready for practical application. It is highly challenging to realize 50GHz burst-mode analog components; theref... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Dichoptic Foveation | Henry Kam, Colin Groth, Jenna Kang, Pratham Saraf, Qi Sun, Kenneth Chen | 2026-09-14 | 下载 | Interocular differences in visual perception can induce a variety of effects when fused by the brain. For example, prior works have found that carefully crafted binocular differences in local detail c... |
| The AI-Enabled Scientific Frontier | Gabriel Manso, Emma Fu, Neil Thompson | 2026-09-14 | 下载 | As artificial intelligence's capabilities improve, it is increasingly viewed as a general scientific method. But how true are these claims? Does AI outperform all techniques, or only some, and how is ... |
| Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking | Md Arafat Hossain, Thomas Randall, Akash Dutta, Xingfu Wu, Rong Ge, Ali Jannesari | 2026-09-14 | 下载 | As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide pro... |
| DeepSeek-V4-Flash on AMD gfx90a: Correctness Recovery and Inference Performance Engineering | Siming Huang | 2026-09-14 | 下载 | We present the enablement, correctness recovery, and performance engineering of DeepSeek-V4-Flash inference on AMD Instinct MI250 GPUs using the gfx90a/CDNA2 architecture. |
| Opportunistic ZGC: Leveraging Idle Cores for More Effective Concurrent Garbage Collection | Jacob Malloy, Michael R. Jantz, Terry Jones | 2026-09-14 | 下载 | Managed language runtimes often provide concurrent garbage collectors so that latency-critical applications with large working sets can keep running while most collection work proceeds in the backgrou... |
| Turkish MMLU Pro: Traceable Option Augmentation and Its Validity Limits in Turkish Multiple-Choice Evaluation | M. Ali Bayram | 2026-09-14 | 下载 | Adding answer options can lower multiple-choice scores without improving assessment validity. Turkish MMLU Pro examines this distinction using 12,000 Turkish-source questions across 58 sections. |
| ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters | Rui Lu, Rui Ge, Huanghuang Liang, Xiaobo Zhou, Dan Wang | 2026-09-14 | 下载 | Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and c... |
| Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive Accelerators | Yuyeong Shin | 2026-09-14 | 下载 | Eight-bit integer (INT8) post-training quantization is the default recipe for edge deployment, under a widely held assumption: INT8 makes inference faster at a small, predictable accuracy cost, and a ... |
| Validating Hybrid-State Cache Recovery for GLM-5.3-Flash with vLLM and LMCache | Frank Li | 2026-09-14 | 下载 | External cache transfers can succeed while a hybrid language model resumes from an inconsistent state. We examine the full 45-layer GLM-5.3-Flash model, using the RedHatAI/ GLM-5. |
| Shared KV Caching for Replicated 27B Inference: Correctness Failures and Performance Boundaries | Frank Li | 2026-09-14 | 下载 | Shared host-memory caching can avoid repeated prefill when a request moves between inference replicas. Its usefulness depends on both correct state transfer and lost prefix locality. |
| GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems | Xinyu Qiu, Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai | 2026-09-14 | 下载 | We predict single-sequence model throughput from GGUF metadata using roofline-shaped predictors with quantization-specific scale factors fitted on reference models. |