2026-08-04
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A Centralized Performance Monitoring Architecture for Heterogeneous Multicore SoCs | Mohammed Sajjad Jafri, Abdur Rahman, Emon Sarkar, Mahdi Hassen, Gopishankar Thayyil, Ashwin Krishna Mani, Rodolfo Pellizzoni | 2026-08-04 | 下载 | Hardware Performance Counters (HPCs) are widely used to enable event-driven software mechanisms such as profile guided optimization, performance analysis, and dynamic resource management in real-time ... |
| On Design Principles for Efficient Heterogeneous DRAM-PIM-GPU Systems | Corey Lammie, Hadjer Benmeziane, William Andrew Simon, Irem Boybat | 2026-08-04 | 下载 | Heterogeneous DRAM-based processing-in-memory (PIM)-GPU systems promise significant efficiency gains for decode-phase large language model (LLM) inference, particularly in long-output generation, yet ... |
| Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM Inference | Junyi Luo, Xinting Jiang, Tai-Hao Wen, Ruichen Qi, Minxing Chu, Hongyi Wu, Gregory Kielian, Ben Laurie, Qirui Zhang, Quan Cheng, Dennis Sylvester, Mehdi Saligane | 2026-08-04 | 下载 | Microscaling (MX) is now the standard for low-bit large language model (LLM) inference. Its 4-bit form MXFP4 still loses substantial accuracy, because existing MX formats fix either the element format... |
| DiffPower: GPU-Accelerated Differentiable Switching Power Analysis and Optimization | Isaac Jacobson, Zheng Zhao, Rashmi Mehrotra, Guanglei Zhou, Vineet Rashingkar, Yiran Chen | 2026-08-04 | 下载 | Accurate and scalable switching power analysis remains a critical bottleneck in modern physical design, often forcing a trade-off between computational speed and modeling fidelity. |
| Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse Attention | Hyungkyu Ham, Junhyeong Bae, Seungheon Lee, Myeongjae Jeon, Gwangsun Kim | 2026-08-04 | 下载 | This paper presents a heterogeneous decode-phase serving system that relocates the KV cache out of GPU memory, motivated by the retrieval-based sparse attention that recent frontier LLMs adopt to serv... |
| ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU Architectures | Di Mu, Tengyuan Jin, Zhenkun Wang, Jialin Yang, Yusen Li, Mian Huo, Shusong Guo, Gang Wang, Xiaoguang Liu | 2026-08-04 | 下载 | Modern deep learning workloads increasingly comprise heterogeneous computation graphs that combine compute-intensive operators with memory-intensive subgraphs. |
| Beyond Peak TOPS/W: A System-Level Perspective on Hybrid Digital, Analogue and Neuromorphic Computing | Eiman Kanjo, Varuna De Silva | 2026-08-04 | 下载 | The digital revolution, which progressively replaced analogue methods with digital circuits, has entered a new phase as AI expands across cloud infrastructure, mobile networks, wearables and physical ... |
| Fovea: Physical-Implication-Aware Wafer-Scale DSE with Decision-Domain-Guided Cross-Fidelity Refinement | Jinxi Li, Huizheng Wang, Jinyi Deng, Yang Hu, Shouyi Yin | 2026-08-04 | 下载 | Modern pre-silicon design-space exploration (DSE) follows a coarse-to-fine workflow: low-cost evaluators screen candidate spaces, while detailed evaluation is reserved for a shortlist. |
| Unified Lookup-Table Inference with Signed-Digit K/V Caches for Ternary LLMs | Ziang Duan, Jiajun Wu, Zetian Chen, Hao Song, Yanwen Deng, Zixuan Shen, Nuobei Xie, Simo Wu, Bolun Wang, Peng Zhou, Chao Wang | 2026-08-04 | 下载 | Ternary LLMs make their weight-dominated projections compact and efficient, but attention remains a mismatch: its K/V cache is created online and is typically processed by a separate higher-precision ... |
| Interpolation of Non-Linear Functions for LLMs using Partial Reconfiguration in FPGAs | Roger Morales-Monge, Nazareth Jimenez-Chacon, Jose Gabriel Villalobos-Alvarado, Luis G. Leon-Vega, Jorge Castro-Godinez | 2026-08-04 | 下载 | Non-linear functions such as exponential and sigmoid are essential in AI and LLM acceleration, although implementing them efficiently on FPGAs is still costly. |
| CAMTA: A Reconfigurable Multi-Region Activation Unit for Nonlinear Function Approximation | Carlos Soto-Porras, Jose Fonseca-Cruz, Pablo Ramirez-Morera, Erick Obregon-Fonseca, Luis G. Leon-Vega, Jorge Castro-Godinez | 2026-08-04 | 下载 | Nonlinear activation functions are widely used in machine learning workloads, but their direct hardware implementation is often costly, function-specific, or difficult to reuse across different models... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A Distributed Quantum Approximate Optimization Algorithm For Unit Commitment | Ali Rajabi, Milad Hasanzadeh, Amin Kargarian | 2026-08-04 | 下载 | This paper presents a distributed quantum approximate optimization algorithm (DQAOA)-enabled three-block alternating direction method of multipliers (ADMM) framework for unit commitment (UC). |
| Real-time decoding of quantum error correction codes using high-performance computing | Lingling Lao, Qiang Wang, Yuanqi Liu, Yantong Liu, Haowen Wang, Yitao Chen, Yankang Zhao, Zhenwei Wu, Wei Zhang, Yong Dong, Yingwen Liu, Mingche Lai, Junjie Wu | 2026-08-04 | 下载 | Quantum error correction (QEC) is indispensable for building scalable fault-tolerant quantum computers. Effective QEC demands stringent real-time decoding: the decoder must process syndrome measuremen... |
| Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training | Niklas Enskat, Philipp Wiesner | 2026-08-04 | 下载 | High-fidelity performance simulators are essential for designing and configuring efficient AI systems, yet today's tools lack the ability to predict power consumption. |
| Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers | Sajjad Khan | 2026-08-04 | 下载 | A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired. |
| DiffPower: GPU-Accelerated Differentiable Switching Power Analysis and Optimization | Isaac Jacobson, Zheng Zhao, Rashmi Mehrotra, Guanglei Zhou, Vineet Rashingkar, Yiran Chen | 2026-08-04 | 下载 | Accurate and scalable switching power analysis remains a critical bottleneck in modern physical design, often forcing a trade-off between computational speed and modeling fidelity. |
| When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference | Przemyslaw Forys, Haoran Wu, Can Xiao, Jiayi Nie, Tony Liu, Rika Antonova, Timothy Jones, Robert Mullins, Wayne Luk, Aaron Zhao, George A. Constantinides | 2026-08-04 | 下载 | Agentic inference now dominates the LLM inference landscape, requiring LLMs to actively engage in multi-turn interactions with tool-calling capabilities. |
| Accelerating Dynamic Graph Clustering on GPU Architectures with cuGraph | Nelson Aloysio Reis de Almeida Passos, Emanuele Carlini, Salvatore Trani | 2026-08-04 | 下载 | This work addresses community detection in temporal networks through GPU-accelerated extensions of spectral clustering and modularity-based algorithms originally designed for static graphs. |
| TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE Training | Lingyun Zhang, Henghua Zhang, Shilei Gu, Kai Mo, Shuai Han, Shiyong Li, Yanpeng Wang, Dou Shen | 2026-08-04 | 下载 | Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs), yet its dynamic routing causes severe load imbalance in expert-parallel training. |
| ReputationChain: Robust Trust Updating for Blockchain-Enabled Supply Chains | Adnan Iftekhar, Chengliang Zheng, Xiaohui Cui, Mir Hassan | 2026-08-04 | 下载 | Blockchain can preserve supply-chain records, but ledger integrity alone does not show whether a participant should be trusted in a future risk-sensitive transaction. |
| FedCARE: A Multi-Objective Personalised Federated Learning Framework for Smart Healthcare | Rojalini Tripathy, Padmalochan Bera, Shreya Ghosh, Rajkumar Buyya | 2026-08-04 | 下载 | Federated Learning (FL) enables collaborative model training across distributed healthcare institutions without centralising sensitive patient data. |
| Certified Split Points for Parallel Lexing: Exact and Modulo Discarded Tokens | Nicklas Nidhögg | 2026-08-04 | 下载 | Table-driven DFA lexing is sequential: each transition depends on the previous byte's state. Scanning one input in parallel needs each chunk's entry state, which existing methods recover by simulation... |
| FedRings: A Scalable and Topology-Aware Federated Learning Framework for LEO Satellite Constellations | Ziwu Liu, Inês Pinto Gouveia, Rehana Yasmin, Paulo Esteves-Verissimo, Ali Shoker | 2026-08-04 | 下载 | Federated learning over low Earth orbit (LEO) satellite networks is limited by frequent link changes, short contact times, and a highly dynamic topology, making centralized or synchronized training in... |
| Flying over The Uncertain Nature (FORTUNE): Intelligent and Humanistic 3D Path Planning for Low-Altitude Collaboration | Minghui Liwang, Wenhan Jia, Xinlei Yi, Wenbo Zhu, Yuhan Su, Xianbin Wang | 2026-08-04 | 下载 | The proliferation of low-altitude intelligent agents is increasing the demand for timely and socially responsible collaborative sensing in dynamic urban environments. |
| LPV Control for Dynamic Power Capping in High-Performance Computing under Mixed Workloads | Mohamed Abdeldjalil Maziz, Kouds Halitim, Bogdan Robu, Sophie Cerf | 2026-08-04 | 下载 | Balancing energy consumption and performance remains a critical challenge in High Performance Computing (HPC) systems. While static power capping mechanisms such as Intel's Running Average Power Limit... |
| Trust-Aware Topology Learning for Dynamic Decentralized Federated Learning under Adversaries | Shubham Vaishnav, Murtaza Rangwala, Ali Beikmohammadi, Sindri Magnússon, Rajkumar Buyya | 2026-08-04 | 下载 | In dynamic mobile decentralized federated learning (DFL), adversaries can poison both model updates and the topology information devices use to choose collaborators. |
| CUDA MPC: A GPU-Native Solver for Model Predictive Control | Babak Akbari, Melissa Greeff | 2026-08-04 | 下载 | Model Predictive Control (MPC) delivers constraint-aware control, but its reliance on online optimization limits its use on systems with fast dynamics, high-dimensional models, or long horizons. |
| Pruning-Aware Multi-Cluster Co-Inference for Large AI Models in AI-RANs | Xiaowen Cao, Zhonghao Lyu, Shicheng Chu, Zezhong Zhang, Dingzhu Wen, Guangxu Zhu, Kaibin Huang, Shuguang Cui, Jie Xu | 2026-08-04 | 下载 | The increasing scale and computational demands of large artificial intelligence models (LAIMs) present significant challenges for efficient inference in resource-constrained distributed environments. |
| AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative Decoding | Shuang Liang, Hao, Chen, Zhiwen Mo, Qianzhou Wang, Guoyu Li, Lingxiao Ma, Wayne Luk | 2026-08-04 | 下载 | Speculative decoding verifies a tree of draft tokens in one target-model forward pass. For a mixture-of-experts (MoE) target, however, parallel verification can activate the union of the experts selec... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Securing Load Balancing over QUIC | Garegin Grigoryan, Dagim Mindaye, Shireen Maini, Minseok Kwon | 2026-08-04 | 下载 | In-network load balancing outperforms traditional software load balancing while costing less. For instance, programmable switch ASICs can use hashing to select the backend server for the initial packe... |
| Data-Driven Online Slice Admission Control and Resource Allocation in NextG Mobile Networks | Muhammad Sulaiman, Bo Sun, Mohammad Ali Salahuddin, Xiaoqi Tan, Raouf Boutaba | 2026-08-04 | 下载 | Virtualization in 5G and beyond networks enables the creation of virtual networks (i.e., network slices) tailored to the needs of different applications. |
| FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs | Amin Farajzadeh, Melike Erol-Kantarci | 2026-08-04 | 下载 | This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-leve... |
| AP Association for RHS-Enabled Cell-Free Uplink MIMO in Industrial Indoor UAV Networks | Liangshun Wu, Wen Chen, Zhendong Li, Qiong Wu, Ying Wang | 2026-08-04 | 下载 | Indoor industrial UAV uplink networks face serious blockage and shadowing from shelves, metal equipment, and production facilities. UAVs are also often clustered and fly along similar straight inspect... |
| FM4WiFi: Flow Matching for Multi-AP Coordination in Dense Deployments of Beyond Wi-Fi 8 Networks | Maksymilian Wojnar, Krzysztof Rusek, Katarzyna Kosek-Szott, Szymon Szott | 2026-08-04 | 下载 | Wi-Fi networks are moving beyond random channel access toward tightly coordinated operation across access points (APs), a shift reflected in Wi-Fi 8's multi-AP coordination (MAPC). |
| ProCAVE: A Self-Adaptive, Full-Lifecycle Edge Caching Framework for Video Streaming via Predictive Bandwidth Estimation and Preference-Aware Deep Reinforcement Learning | Yeganeh Chatri, Behzad Akbari, Foad Ghaderi, Pejman Goudarzi | 2026-08-04 | 下载 | The growing demand for mobile video streaming requires edge delivery systems that adapt efficiently to rapid network fluctuations and diverse user preferences. |
| PECR: A Reproducible Specification and Synthetic Stress Test of Telemetry-Informed Vulnerability Prioritization for SD-WAN | Saeed Alam | 2026-08-04 | 下载 | Software-defined wide-area networking concentrates operational authority in controllers, orchestrators, and Internet-facing edges, but severity-only remediation queues do not represent current exposur... |
| The Frontier LLM Trap in Network Automation | Minhao Jin, Sean Wang, Aarti Gupta, Maria Apostolaki | 2026-08-04 | 下载 | Large LLMs are powerful tools for network automation, but they are expensive, slow to serve, hard to audit, poorly tailored to individual networks, and create long-term dependencies on a small number ... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM Training | Niklas Enskat, Philipp Wiesner | 2026-08-04 | 下载 | High-fidelity performance simulators are essential for designing and configuring efficient AI systems, yet today's tools lack the ability to predict power consumption. |
| SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG | Kaysarul Anas Apurba, Md. Hasibul Hasan, Rofiqul Alam Shehab, Asab Azad | 2026-08-04 | 下载 | We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific... |