Skip to content

2026-08-04 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
A Centralized Performance Monitoring Architecture for Heterogeneous Multicore SoCsMohammed Sajjad Jafri, Abdur Rahman, Emon Sarkar, Mahdi Hassen, Gopishankar Thayyil, Ashwin Krishna Mani, Rodolfo Pellizzoni2026-08-04下载Hardware Performance Counters (HPCs) are widely used to enable event-driven software mechanisms such as profile guided optimization, performance analysis, and dynamic resource management in real-time ...
On Design Principles for Efficient Heterogeneous DRAM-PIM-GPU SystemsCorey Lammie, Hadjer Benmeziane, William Andrew Simon, Irem Boybat2026-08-04下载Heterogeneous DRAM-based processing-in-memory (PIM)-GPU systems promise significant efficiency gains for decode-phase large language model (LLM) inference, particularly in long-output generation, yet ...
Heterogeneity-Aware Microscaling for Efficient Low-Bit LLM InferenceJunyi Luo, Xinting Jiang, Tai-Hao Wen, Ruichen Qi, Minxing Chu, Hongyi Wu, Gregory Kielian, Ben Laurie, Qirui Zhang, Quan Cheng, Dennis Sylvester, Mehdi Saligane2026-08-04下载Microscaling (MX) is now the standard for low-bit large language model (LLM) inference. Its 4-bit form MXFP4 still loses substantial accuracy, because existing MX formats fix either the element format...
DiffPower: GPU-Accelerated Differentiable Switching Power Analysis and OptimizationIsaac Jacobson, Zheng Zhao, Rashmi Mehrotra, Guanglei Zhou, Vineet Rashingkar, Yiran Chen2026-08-04下载Accurate and scalable switching power analysis remains a critical bottleneck in modern physical design, often forcing a trade-off between computational speed and modeling fidelity.
Heterogeneous LLM Serving with General-Purpose Processing-Near-Memory for Retrieval-Based Sparse AttentionHyungkyu Ham, Junhyeong Bae, Seungheon Lee, Myeongjae Jeon, Gwangsun Kim2026-08-04下载This paper presents a heterogeneous decode-phase serving system that relocates the KV cache out of GPU memory, motivated by the retrieval-based sparse attention that recent frontier LLMs adopt to serv...
ComFuse: Fusing Complex Memory-Intensive Subgraphs with Compute-Intensive Kernels For Modern GPU ArchitecturesDi Mu, Tengyuan Jin, Zhenkun Wang, Jialin Yang, Yusen Li, Mian Huo, Shusong Guo, Gang Wang, Xiaoguang Liu2026-08-04下载Modern deep learning workloads increasingly comprise heterogeneous computation graphs that combine compute-intensive operators with memory-intensive subgraphs.
Beyond Peak TOPS/W: A System-Level Perspective on Hybrid Digital, Analogue and Neuromorphic ComputingEiman Kanjo, Varuna De Silva2026-08-04下载The digital revolution, which progressively replaced analogue methods with digital circuits, has entered a new phase as AI expands across cloud infrastructure, mobile networks, wearables and physical ...
Fovea: Physical-Implication-Aware Wafer-Scale DSE with Decision-Domain-Guided Cross-Fidelity RefinementJinxi Li, Huizheng Wang, Jinyi Deng, Yang Hu, Shouyi Yin2026-08-04下载Modern pre-silicon design-space exploration (DSE) follows a coarse-to-fine workflow: low-cost evaluators screen candidate spaces, while detailed evaluation is reserved for a shortlist.
Unified Lookup-Table Inference with Signed-Digit K/V Caches for Ternary LLMsZiang Duan, Jiajun Wu, Zetian Chen, Hao Song, Yanwen Deng, Zixuan Shen, Nuobei Xie, Simo Wu, Bolun Wang, Peng Zhou, Chao Wang2026-08-04下载Ternary LLMs make their weight-dominated projections compact and efficient, but attention remains a mismatch: its K/V cache is created online and is typically processed by a separate higher-precision ...
Interpolation of Non-Linear Functions for LLMs using Partial Reconfiguration in FPGAsRoger Morales-Monge, Nazareth Jimenez-Chacon, Jose Gabriel Villalobos-Alvarado, Luis G. Leon-Vega, Jorge Castro-Godinez2026-08-04下载Non-linear functions such as exponential and sigmoid are essential in AI and LLM acceleration, although implementing them efficiently on FPGAs is still costly.
CAMTA: A Reconfigurable Multi-Region Activation Unit for Nonlinear Function ApproximationCarlos Soto-Porras, Jose Fonseca-Cruz, Pablo Ramirez-Morera, Erick Obregon-Fonseca, Luis G. Leon-Vega, Jorge Castro-Godinez2026-08-04下载Nonlinear activation functions are widely used in machine learning workloads, but their direct hardware implementation is often costly, function-specific, or difficult to reuse across different models...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
A Distributed Quantum Approximate Optimization Algorithm For Unit CommitmentAli Rajabi, Milad Hasanzadeh, Amin Kargarian2026-08-04下载This paper presents a distributed quantum approximate optimization algorithm (DQAOA)-enabled three-block alternating direction method of multipliers (ADMM) framework for unit commitment (UC).
Real-time decoding of quantum error correction codes using high-performance computingLingling Lao, Qiang Wang, Yuanqi Liu, Yantong Liu, Haowen Wang, Yitao Chen, Yankang Zhao, Zhenwei Wu, Wei Zhang, Yong Dong, Yingwen Liu, Mingche Lai, Junjie Wu2026-08-04下载Quantum error correction (QEC) is indispensable for building scalable fault-tolerant quantum computers. Effective QEC demands stringent real-time decoding: the decoder must process syndrome measuremen...
Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM TrainingNiklas Enskat, Philipp Wiesner2026-08-04下载High-fidelity performance simulators are essential for designing and configuring efficient AI systems, yet today's tools lack the ability to predict power consumption.
Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence LayersSajjad Khan2026-08-04下载A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired.
DiffPower: GPU-Accelerated Differentiable Switching Power Analysis and OptimizationIsaac Jacobson, Zheng Zhao, Rashmi Mehrotra, Guanglei Zhou, Vineet Rashingkar, Yiran Chen2026-08-04下载Accurate and scalable switching power analysis remains a critical bottleneck in modern physical design, often forcing a trade-off between computational speed and modeling fidelity.
When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM InferencePrzemyslaw Forys, Haoran Wu, Can Xiao, Jiayi Nie, Tony Liu, Rika Antonova, Timothy Jones, Robert Mullins, Wayne Luk, Aaron Zhao, George A. Constantinides2026-08-04下载Agentic inference now dominates the LLM inference landscape, requiring LLMs to actively engage in multi-turn interactions with tool-calling capabilities.
Accelerating Dynamic Graph Clustering on GPU Architectures with cuGraphNelson Aloysio Reis de Almeida Passos, Emanuele Carlini, Salvatore Trani2026-08-04下载This work addresses community detection in temporal networks through GPU-accelerated extensions of spectral clustering and modularity-based algorithms originally designed for static graphs.
TAOT: Topology-Aware Optimal Transport for Dynamic Expert Replica Placement in MoE TrainingLingyun Zhang, Henghua Zhang, Shilei Gu, Kai Mo, Shuai Han, Shiyong Li, Yanpeng Wang, Dou Shen2026-08-04下载Mixture-of-Experts (MoE) has become a key architecture for scaling large language models (LLMs), yet its dynamic routing causes severe load imbalance in expert-parallel training.
ReputationChain: Robust Trust Updating for Blockchain-Enabled Supply ChainsAdnan Iftekhar, Chengliang Zheng, Xiaohui Cui, Mir Hassan2026-08-04下载Blockchain can preserve supply-chain records, but ledger integrity alone does not show whether a participant should be trusted in a future risk-sensitive transaction.
FedCARE: A Multi-Objective Personalised Federated Learning Framework for Smart HealthcareRojalini Tripathy, Padmalochan Bera, Shreya Ghosh, Rajkumar Buyya2026-08-04下载Federated Learning (FL) enables collaborative model training across distributed healthcare institutions without centralising sensitive patient data.
Certified Split Points for Parallel Lexing: Exact and Modulo Discarded TokensNicklas Nidhögg2026-08-04下载Table-driven DFA lexing is sequential: each transition depends on the previous byte's state. Scanning one input in parallel needs each chunk's entry state, which existing methods recover by simulation...
FedRings: A Scalable and Topology-Aware Federated Learning Framework for LEO Satellite ConstellationsZiwu Liu, Inês Pinto Gouveia, Rehana Yasmin, Paulo Esteves-Verissimo, Ali Shoker2026-08-04下载Federated learning over low Earth orbit (LEO) satellite networks is limited by frequent link changes, short contact times, and a highly dynamic topology, making centralized or synchronized training in...
Flying over The Uncertain Nature (FORTUNE): Intelligent and Humanistic 3D Path Planning for Low-Altitude CollaborationMinghui Liwang, Wenhan Jia, Xinlei Yi, Wenbo Zhu, Yuhan Su, Xianbin Wang2026-08-04下载The proliferation of low-altitude intelligent agents is increasing the demand for timely and socially responsible collaborative sensing in dynamic urban environments.
LPV Control for Dynamic Power Capping in High-Performance Computing under Mixed WorkloadsMohamed Abdeldjalil Maziz, Kouds Halitim, Bogdan Robu, Sophie Cerf2026-08-04下载Balancing energy consumption and performance remains a critical challenge in High Performance Computing (HPC) systems. While static power capping mechanisms such as Intel's Running Average Power Limit...
Trust-Aware Topology Learning for Dynamic Decentralized Federated Learning under AdversariesShubham Vaishnav, Murtaza Rangwala, Ali Beikmohammadi, Sindri Magnússon, Rajkumar Buyya2026-08-04下载In dynamic mobile decentralized federated learning (DFL), adversaries can poison both model updates and the topology information devices use to choose collaborators.
CUDA MPC: A GPU-Native Solver for Model Predictive ControlBabak Akbari, Melissa Greeff2026-08-04下载Model Predictive Control (MPC) delivers constraint-aware control, but its reliance on online optimization limits its use on systems with fast dynamics, high-dimensional models, or long horizons.
Pruning-Aware Multi-Cluster Co-Inference for Large AI Models in AI-RANsXiaowen Cao, Zhonghao Lyu, Shicheng Chu, Zezhong Zhang, Dingzhu Wen, Guangxu Zhu, Kaibin Huang, Shuguang Cui, Jie Xu2026-08-04下载The increasing scale and computational demands of large artificial intelligence models (LAIMs) present significant challenges for efficient inference in resource-constrained distributed environments.
AcceptMoE: Commitment-Weighted Self-Sizing Verifier Expert Sets for Efficient MoE Speculative DecodingShuang Liang, Hao, Chen, Zhiwen Mo, Qianzhou Wang, Guoyu Li, Lingxiao Ma, Wayne Luk2026-08-04下载Speculative decoding verifies a tree of draft tokens in one target-model forward pass. For a mixture-of-experts (MoE) target, however, parallel verification can activate the union of the experts selec...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Securing Load Balancing over QUICGaregin Grigoryan, Dagim Mindaye, Shireen Maini, Minseok Kwon2026-08-04下载In-network load balancing outperforms traditional software load balancing while costing less. For instance, programmable switch ASICs can use hashing to select the backend server for the initial packe...
Data-Driven Online Slice Admission Control and Resource Allocation in NextG Mobile NetworksMuhammad Sulaiman, Bo Sun, Mohammad Ali Salahuddin, Xiaoqi Tan, Raouf Boutaba2026-08-04下载Virtualization in 5G and beyond networks enables the creation of virtual networks (i.e., network slices) tailored to the needs of different applications.
FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANsAmin Farajzadeh, Melike Erol-Kantarci2026-08-04下载This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-leve...
AP Association for RHS-Enabled Cell-Free Uplink MIMO in Industrial Indoor UAV NetworksLiangshun Wu, Wen Chen, Zhendong Li, Qiong Wu, Ying Wang2026-08-04下载Indoor industrial UAV uplink networks face serious blockage and shadowing from shelves, metal equipment, and production facilities. UAVs are also often clustered and fly along similar straight inspect...
FM4WiFi: Flow Matching for Multi-AP Coordination in Dense Deployments of Beyond Wi-Fi 8 NetworksMaksymilian Wojnar, Krzysztof Rusek, Katarzyna Kosek-Szott, Szymon Szott2026-08-04下载Wi-Fi networks are moving beyond random channel access toward tightly coordinated operation across access points (APs), a shift reflected in Wi-Fi 8's multi-AP coordination (MAPC).
ProCAVE: A Self-Adaptive, Full-Lifecycle Edge Caching Framework for Video Streaming via Predictive Bandwidth Estimation and Preference-Aware Deep Reinforcement LearningYeganeh Chatri, Behzad Akbari, Foad Ghaderi, Pejman Goudarzi2026-08-04下载The growing demand for mobile video streaming requires edge delivery systems that adapt efficiently to rapid network fluctuations and diverse user preferences.
PECR: A Reproducible Specification and Synthetic Stress Test of Telemetry-Informed Vulnerability Prioritization for SD-WANSaeed Alam2026-08-04下载Software-defined wide-area networking concentrates operational authority in controllers, orchestrators, and Internet-facing edges, but severity-only remediation queues do not represent current exposur...
The Frontier LLM Trap in Network AutomationMinhao Jin, Sean Wang, Aarti Gupta, Maria Apostolaki2026-08-04下载Large LLMs are powerful tools for network automation, but they are expensive, slow to serve, hard to audit, poorly tailored to individual networks, and create long-term dependencies on a small number ...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Evaluating MFU as a Proxy for GPU Power for Energy-Aware Simulation of LLM TrainingNiklas Enskat, Philipp Wiesner2026-08-04下载High-fidelity performance simulators are essential for designing and configuring efficient AI systems, yet today's tools lack the ability to predict power consumption.
SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAGKaysarul Anas Apurba, Md. Hasibul Hasan, Rofiqul Alam Shehab, Asab Azad2026-08-04下载We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rather than proposing a new model, we evaluate a fixed scientific...

基于 VitePress 构建 · 使用本地搜索查找论文