2026-09-04
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM Decoding | Jiahao Zheng, Yifan Qin, Xiaobo Sharon Hu, Yiyu Shi | 2026-09-04 | 下载 | The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so decoding is memory-bandwidth bound. |
| RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems | Zlatan Feric, Amir Taherin, Bin Ren, Yanzhi Wang, Jennifer Dy, David Kaeli | 2026-09-04 | 下载 | We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU environments. |
| LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision Diagrams | Yoonju Sim, Federico Berto, Chuanbo Hua, Jinkyoo Park, Changhyun Kwon | 2026-09-04 | 下载 | Quantum circuits are central to implementing quantum algorithms on quantum devices, where quantum gates must be reversible. Many quantum algorithms rely on Boolean functions, which must therefore be i... |
| Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoC | Saad Memon, Rafal Graczyk, Jan Swakoń, Leszek Grzanka, Sebastian Kusyk, Mike Papadakis | 2026-09-04 | 下载 | As spaceborne computing systems increasingly rely on neural network (NN) accelerators, the opacity of commercial, black-box architectures severely restricts the development of verifiable radiation mit... |
| TETRIS-Q: Tiling-based Effective Transient-fault Reduction on Interleaved Superconducting Qubits | Marzio Vallero, Gioele Casagranda, Flavio Vella, Paolo Rech | 2026-09-04 | 下载 | The struggle of the hour in quantum computing research is achieving effective suppression of the error mechanisms induced by the interaction of external radiation with superconducting quantum devices. |
| APEX-RBD: Mixed-Precision Exploration Framework for Hardware-Efficient Robot Dynamics Accelerator Design | Xingyu Liu, Hanwei Fan, Chaofang Ma, Jiawei Liang, Guangyu Hu, Jiang Xu, Wei Zhang | 2026-09-04 | 下载 | Rigid Body Dynamics (RBD) forms the computational core of real-time robotic control, but its immense computational complexity creates a performance bottleneck that necessitates dedicated hardware acce... |
| TreeFI: Value-Aware Statistical Fault Injection for Deep Neural Networks | Noam Bires, Marcello Traiola, Angeliki Kritikakou, Elisa Fromont | 2026-09-04 | 下载 | Reliability evaluation of deep neural networks under hardware faults commonly relies on fault injection, but exhaustive campaigns are intractable for modern models and datasets. |
| A Piecewise-Linear Approximation-based Energy-Efficient Error-Optimized Unsigned Square Rooter for Accuracy-Critical Applications | Prateek Goyal, Sujit Kumar Sahoo | 2026-09-04 | 下载 | Approximate computing improves energy efficiency in error-resilient applications, but square root units remain challenging due to the trade-off between hardware cost and computational accuracy. |
| FlexPosit: Tunable Fractional Precision for LLM Inference Accelerators | Yimin Gao, Liangtao Dai, Jun Yin, Xinfei Guo, Mircea Stan | 2026-09-04 | 下载 | Large language models (LLMs) offer remarkable capabilities but impose prohibitive compute and energy costs. Quantization governs the trade-offs between accuracy and hardware efficiency across granular... |
| Sustainable Edge Vision via Empirically Calibrated DVFS: Eliminating Thermal Throttling on Passively Cooled Hardware | Aayush Marasini, Zhaoxian Zhou | 2026-09-04 | 下载 | Passive cooling eliminates the energy overhead and mechanical failure modes of fans, making it attractive for edge deployment, yet sustained Deep Neural Network (DNN) inference on passively cooled edg... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| DejaVu: Unifying Memory Allocations to Eliminate Redundant Copies on Unified-Memory SoCs | Yuheng Zhu, Yanbo Zhao, Jiajia Li, Man-Ki Yoon | 2026-09-04 | 下载 | GPU applications on unified-memory (UMA) edge platforms often inherit a discrete-GPU memory abstraction in which they allocate one buffer for the CPU, another for the GPU, and copy data between them b... |
| Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load Relief | Meiyi Li | 2026-09-04 | 下载 | Grid studies often represent data-center flexibility as a fixed percentage of load, although no public production trace has shown how much eligible load persists across event durations or co-moves acr... |
| From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned Baseline | Andrew James Amos | 2026-09-04 | 下载 | Comparisons between GPU implementations are usually asymmetric: one side is tuned by its author, the other is run as found. I report a programme that tuned both a novel SOM algorithm (SparseBin) and t... |
| Sharpedo: Dual-Mode Uncertified DAG-Based Consensus Protocol | Zeno de Angeli | 2026-09-04 | 下载 | Mysticeti and Mahi-Mahi represent leading approaches to consensus protocols, leveraging a novel uncertified Directed Acyclic Graph data structure to achieve substantial performance benefits compared t... |
| Solution-space heterogeneity shapes federated learning dynamics across partial differential equations | Ping Luo, Jiahuan Wang, Ziqing Wen, Tao Sun, Dongsheng Li | 2026-09-04 | 下载 | Federated scientific machine learning enables institutions to train neural surrogates without centralizing local physical data, yet studies of partial differential equations (PDEs) lack a transferable... |
| GreenPipe: Power Modeling for Containerized DNN Inference on Kubernetes Edge Nodes | Mengxue Wang, Peini Liu, Amir Taherkordi, Jordi Guitart | 2026-09-04 | 下载 | Distributed DNN inference is increasingly deployed in containerized edge-cloud environments, where workloads run on-device or are exposed to remote clients over the network. |
| Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated Learning | Ming Xiang, Stratis Ioannidis, Edmund Yeh, Carlee Joe-Wong, Lili Su | 2026-09-04 | 下载 | Due to resource constraints or external and internal uncertainties, clients in real-world federated learning systems are often intermittently available edge devices. |
| Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM Serving | Aditi Patodiya | 2026-09-04 | 下载 | Prefix caching, in which a serving engine reuses the key and value tensors of a shared prompt prefix across requests, is enabled by default in the major open-source stacks and treated as a transparent... |
| BF16 Component-Product Emulation of FP32 and FP64 GEMM on Intel AMX | Bing Cui, Yu Liu | 2026-09-04 | 下载 | Modern CPUs increasingly integrate high-throughput matrix engines optimized for low-precision AI workloads, while many scientific computing applications still rely on FP32 and FP64 GEMM to meet their ... |
| CIERA: Cross-Iteration Exponent Reuse for Lossless Allgather in Sharded MoE Training | Ali Zafar Sadiq, Haiying Shen, Masahiro Tanaka | 2026-09-04 | 下载 | In training Mixture-of-Experts (MoE) models, sharded data parallelism partitions each expert's parameters across GPUs, requiring an Allgather operation to reconstruct the full weight matrix before eac... |
| Improving Progressive Compression with Adaptive Interpolation and Coefficient Decomposition | Wenbo Li, Xuan Wu, Qian Gong, Pu Jiao, Jieyang Chen, Qing Liu, Norbert Podhorszki, Scott Klasky, Xin Liang | 2026-09-04 | 下载 | Exascale simulations generate data far faster than it can be stored or analyzed, making efficient data reduction essential. Error-controlled lossy compression offers high compression ratios under user... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| SDN-Orchestrated Dual-Path 5G/SATCOM Maritime Communications for Carrier Strike Groups | Avinash Srinivasan, Dalibor Španjević, Kevin H. Nguyen, Bannon Ireton, Christopher B. Landis | 2026-09-04 | 下载 | Carrier Strike Group (CSG) communications must sustain mission traffic across heterogeneous links whose quality varies with distance, weather, fading, and intermittent outages. |
| WIP: Energy-Efficient LLM-Based Serving Cluster Formulation in Cell-Free Massive MIMO | Marcin Hoffmann, Paweł Kryszkiewicz | 2026-09-04 | 下载 | One way to increase the Energy Efficiency (EE) of 6G wireless networks is to utilize existing network infrastructure more efficiently. This can be achieved by introducing User-Centric Cell-Free Massiv... |
| XAI-SDN: An Explainable Entropy-Guided Machine Learning Framework for Real-Time DDoS Detection in Software Defined Networks | Adeel Ahmad, Ali Akarma, Ahmad Ali, Hammad Muneer, Toqeer Ali Syed | 2026-09-04 | 下载 | One of the biggest risks faced by Software Defined Networks (SDN) is the Distributed Denial of Service (DDoS) attack in which a compromised controller can make an entire network unusable. |
| Towards Federated, Green, and Resilient 6G Non-Terrestrial Networks | Sarath Babu, Victor Baños-Gonzalez, Mario Cordina, Debabrata Dalai, Tomaso de Cola, Franco Davoli, Etienne Victor Depasquale, Ashutosh Dutta, Hesham ElBakoury, Michael A. Enright, Giovanni Giambene, Sumit Goswami, Ramesh Gupta, Wael Jaafar, Eman Hammad, B. S. Manoj, Tony Li, Manuel M. H. Roth, Paresh Saxena, Pat Scanlan, Zhili Sun, Daniele Tarchi, Saviour Zammit | 2026-09-04 | 下载 | This study focuses on future Non-Terrestrial Networks (NTN) integrated with Terrestrial Networks (TN) for future 5G/6G systems. NTN envisions a 3D architecture, where Low Earth Orbit (LEO) satellite n... |
| Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless Networks | Abdessamed Qchohi, Jessica Moysen Cortes, Matteo Zecchin | 2026-09-04 | 下载 | Conformal counterfactual inference enables network operators to use logged telemetry to reliably answer 'what-if' questions about network operation. |
| On the Delay-Constrained Maximum Concurrent Flow Problem | Walid Ben-Ameur, Guillaume Beraud-Sudreau, Hervé Kerivin, Sebastien Martin | 2026-09-04 | 下载 | Real-time services, such as VoIP and large-scale neural network training, require strict transmission delay guarantees. While routing under hop constraints is tractable, real-world delays increase sha... |
| Performance Evaluation of HAPS-enabled Coverage Enhancement in Hard-to-Reach Areas | Hao Lin, Mustafa A. Kishk, Mohamed-Slim Alouini | 2026-09-04 | 下载 | High altitude platform stations (HAPSs) are becoming a key component of future non-terrestrial networks (NTNs). HAPSs can serve a larger area than uncrewed aerial vehicles (UAVs) and offer lower propa... |
| QUASAR: Quantum Satellite Architecture and Routing Simulator | Yaliang Shi, Zi Wang, Bangguo Yuan, Gaojie Wu, Zhiwei Zhao | 2026-09-04 | 下载 | The deployment of Low Earth Orbit (LEO) satellite constellations is an important step toward global-scale quantum networking. However, evaluating satellite quantum network protocols under spatiotempor... |
| A Wavelength Borrowing Architecture for Optical Data Center Networks - Extended Version | Andrea Detti, Chiara Lodovisi, Silvello Betti | 2026-09-04 | 下载 | The growth of east-west traffic, along with the cost and power consumption of electronic switching, is motivating the integration of a low-power, high-rate, all-optical layer within the data center ne... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| DejaVu: Unifying Memory Allocations to Eliminate Redundant Copies on Unified-Memory SoCs | Yuheng Zhu, Yanbo Zhao, Jiajia Li, Man-Ki Yoon | 2026-09-04 | 下载 | GPU applications on unified-memory (UMA) edge platforms often inherit a discrete-GPU memory abstraction in which they allocate one buffer for the CPU, another for the GPU, and copy data between them b... |
| Adaptive Context Parallelism for Production LLM Serving | Jiarui Guo, Rongle Wang, Peijun Huang, Zongwei Lv, Ziqing Wang, Kan Liu, Tao Lan, Lin Qu, Xiaolin Wang, Tong Yang | 2026-09-04 | 下载 | As LLM context windows expand and input sequences grow longer, serving systems face increasing computational and memory demands. Context parallelism (CP), which partitions the input sequence across mu... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation Systems | Zlatan Feric, Amir Taherin, Bin Ren, Yanzhi Wang, Jennifer Dy, David Kaeli | 2026-09-04 | 下载 | We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU environments. |
| GreenPipe: Power Modeling for Containerized DNN Inference on Kubernetes Edge Nodes | Mengxue Wang, Peini Liu, Amir Taherkordi, Jordi Guitart | 2026-09-04 | 下载 | Distributed DNN inference is increasingly deployed in containerized edge-cloud environments, where workloads run on-device or are exposed to remote clients over the network. |