Skip to content

2026-09-04 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Interface-Aware KV Cache Quantization for Dense On-Chip NVM in Long-Context LLM DecodingJiahao Zheng, Yifan Qin, Xiaobo Sharon Hu, Yiyu Shi2026-09-04下载The key-value (KV) cache is the dominant memory bottleneck in long-context large language model (LLM) decoding: every step reads it entirely, so decoding is memory-bandwidth bound.
RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation SystemsZlatan Feric, Amir Taherin, Bin Ren, Yanzhi Wang, Jennifer Dy, David Kaeli2026-09-04下载We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU environments.
LLM-Driven Algorithm Design for Quantum Circuit Synthesis based on Binary Decision DiagramsYoonju Sim, Federico Berto, Chuanbo Hua, Jinkyoo Park, Changhyun Kwon2026-09-04下载Quantum circuits are central to implementing quantum algorithms on quantum devices, where quantum gates must be reversible. Many quantum algorithms rely on Boolean functions, which must therefore be i...
Proton Irradiation Characterization of an Open-Source ML Accelerator on a Zynq UltraScale+ MPSoCSaad Memon, Rafal Graczyk, Jan Swakoń, Leszek Grzanka, Sebastian Kusyk, Mike Papadakis2026-09-04下载As spaceborne computing systems increasingly rely on neural network (NN) accelerators, the opacity of commercial, black-box architectures severely restricts the development of verifiable radiation mit...
TETRIS-Q: Tiling-based Effective Transient-fault Reduction on Interleaved Superconducting QubitsMarzio Vallero, Gioele Casagranda, Flavio Vella, Paolo Rech2026-09-04下载The struggle of the hour in quantum computing research is achieving effective suppression of the error mechanisms induced by the interaction of external radiation with superconducting quantum devices.
APEX-RBD: Mixed-Precision Exploration Framework for Hardware-Efficient Robot Dynamics Accelerator DesignXingyu Liu, Hanwei Fan, Chaofang Ma, Jiawei Liang, Guangyu Hu, Jiang Xu, Wei Zhang2026-09-04下载Rigid Body Dynamics (RBD) forms the computational core of real-time robotic control, but its immense computational complexity creates a performance bottleneck that necessitates dedicated hardware acce...
TreeFI: Value-Aware Statistical Fault Injection for Deep Neural NetworksNoam Bires, Marcello Traiola, Angeliki Kritikakou, Elisa Fromont2026-09-04下载Reliability evaluation of deep neural networks under hardware faults commonly relies on fault injection, but exhaustive campaigns are intractable for modern models and datasets.
A Piecewise-Linear Approximation-based Energy-Efficient Error-Optimized Unsigned Square Rooter for Accuracy-Critical ApplicationsPrateek Goyal, Sujit Kumar Sahoo2026-09-04下载Approximate computing improves energy efficiency in error-resilient applications, but square root units remain challenging due to the trade-off between hardware cost and computational accuracy.
FlexPosit: Tunable Fractional Precision for LLM Inference AcceleratorsYimin Gao, Liangtao Dai, Jun Yin, Xinfei Guo, Mircea Stan2026-09-04下载Large language models (LLMs) offer remarkable capabilities but impose prohibitive compute and energy costs. Quantization governs the trade-offs between accuracy and hardware efficiency across granular...
Sustainable Edge Vision via Empirically Calibrated DVFS: Eliminating Thermal Throttling on Passively Cooled HardwareAayush Marasini, Zhaoxian Zhou2026-09-04下载Passive cooling eliminates the energy overhead and mechanical failure modes of fans, making it attractive for edge deployment, yet sustained Deep Neural Network (DNN) inference on passively cooled edg...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
DejaVu: Unifying Memory Allocations to Eliminate Redundant Copies on Unified-Memory SoCsYuheng Zhu, Yanbo Zhao, Jiajia Li, Man-Ki Yoon2026-09-04下载GPU applications on unified-memory (UMA) edge platforms often inherit a discrete-GPU memory abstraction in which they allocate one buffer for the CPU, another for the GPU, and copy data between them b...
Beyond Scalar Flexibility: From Eligible AI Workloads to Dependable Load ReliefMeiyi Li2026-09-04下载Grid studies often represent data-center flexibility as a fixed percentage of load, although no public production trace has shown how much eligible load persists across event durations or co-moves acr...
From 80x to 385x: A Best-Matching-Unit Search at the L2 Roof, Measured Against a Symmetrically Tuned BaselineAndrew James Amos2026-09-04下载Comparisons between GPU implementations are usually asymmetric: one side is tuned by its author, the other is run as found. I report a programme that tuned both a novel SOM algorithm (SparseBin) and t...
Sharpedo: Dual-Mode Uncertified DAG-Based Consensus ProtocolZeno de Angeli2026-09-04下载Mysticeti and Mahi-Mahi represent leading approaches to consensus protocols, leveraging a novel uncertified Directed Acyclic Graph data structure to achieve substantial performance benefits compared t...
Solution-space heterogeneity shapes federated learning dynamics across partial differential equationsPing Luo, Jiahuan Wang, Ziqing Wen, Tao Sun, Dongsheng Li2026-09-04下载Federated scientific machine learning enables institutions to train neural surrogates without centralizing local physical data, yet studies of partial differential equations (PDEs) lack a transferable...
GreenPipe: Power Modeling for Containerized DNN Inference on Kubernetes Edge NodesMengxue Wang, Peini Liu, Amir Taherkordi, Jordi Guitart2026-09-04下载Distributed DNN inference is increasingly deployed in containerized edge-cloud environments, where workloads run on-device or are exposed to remote clients over the network.
Resilience Beyond Stationary Client Unavailability: Unlocking Efficient and Unbiased Federated LearningMing Xiang, Stratis Ioannidis, Edmund Yeh, Carlee Joe-Wong, Lili Su2026-09-04下载Due to resource constraints or external and internal uncertainties, clients in real-world federated learning systems are often intermittently available edge devices.
Same Request, Different Answer: Quantization Amplifies Cache-Induced Divergence in LLM ServingAditi Patodiya2026-09-04下载Prefix caching, in which a serving engine reuses the key and value tensors of a shared prompt prefix across requests, is enabled by default in the major open-source stacks and treated as a transparent...
BF16 Component-Product Emulation of FP32 and FP64 GEMM on Intel AMXBing Cui, Yu Liu2026-09-04下载Modern CPUs increasingly integrate high-throughput matrix engines optimized for low-precision AI workloads, while many scientific computing applications still rely on FP32 and FP64 GEMM to meet their ...
CIERA: Cross-Iteration Exponent Reuse for Lossless Allgather in Sharded MoE TrainingAli Zafar Sadiq, Haiying Shen, Masahiro Tanaka2026-09-04下载In training Mixture-of-Experts (MoE) models, sharded data parallelism partitions each expert's parameters across GPUs, requiring an Allgather operation to reconstruct the full weight matrix before eac...
Improving Progressive Compression with Adaptive Interpolation and Coefficient DecompositionWenbo Li, Xuan Wu, Qian Gong, Pu Jiao, Jieyang Chen, Qing Liu, Norbert Podhorszki, Scott Klasky, Xin Liang2026-09-04下载Exascale simulations generate data far faster than it can be stored or analyzed, making efficient data reduction essential. Error-controlled lossy compression offers high compression ratios under user...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
SDN-Orchestrated Dual-Path 5G/SATCOM Maritime Communications for Carrier Strike GroupsAvinash Srinivasan, Dalibor Španjević, Kevin H. Nguyen, Bannon Ireton, Christopher B. Landis2026-09-04下载Carrier Strike Group (CSG) communications must sustain mission traffic across heterogeneous links whose quality varies with distance, weather, fading, and intermittent outages.
WIP: Energy-Efficient LLM-Based Serving Cluster Formulation in Cell-Free Massive MIMOMarcin Hoffmann, Paweł Kryszkiewicz2026-09-04下载One way to increase the Energy Efficiency (EE) of 6G wireless networks is to utilize existing network infrastructure more efficiently. This can be achieved by introducing User-Centric Cell-Free Massiv...
XAI-SDN: An Explainable Entropy-Guided Machine Learning Framework for Real-Time DDoS Detection in Software Defined NetworksAdeel Ahmad, Ali Akarma, Ahmad Ali, Hammad Muneer, Toqeer Ali Syed2026-09-04下载One of the biggest risks faced by Software Defined Networks (SDN) is the Distributed Denial of Service (DDoS) attack in which a compromised controller can make an entire network unusable.
Towards Federated, Green, and Resilient 6G Non-Terrestrial NetworksSarath Babu, Victor Baños-Gonzalez, Mario Cordina, Debabrata Dalai, Tomaso de Cola, Franco Davoli, Etienne Victor Depasquale, Ashutosh Dutta, Hesham ElBakoury, Michael A. Enright, Giovanni Giambene, Sumit Goswami, Ramesh Gupta, Wael Jaafar, Eman Hammad, B. S. Manoj, Tony Li, Manuel M. H. Roth, Paresh Saxena, Pat Scanlan, Zhili Sun, Daniele Tarchi, Saviour Zammit2026-09-04下载This study focuses on future Non-Terrestrial Networks (NTN) integrated with Terrestrial Networks (TN) for future 5G/6G systems. NTN envisions a 3D architecture, where Low Earth Orbit (LEO) satellite n...
Confounding-Valid Conformal Inference for Counterfactual KPIs in Wireless NetworksAbdessamed Qchohi, Jessica Moysen Cortes, Matteo Zecchin2026-09-04下载Conformal counterfactual inference enables network operators to use logged telemetry to reliably answer 'what-if' questions about network operation.
On the Delay-Constrained Maximum Concurrent Flow ProblemWalid Ben-Ameur, Guillaume Beraud-Sudreau, Hervé Kerivin, Sebastien Martin2026-09-04下载Real-time services, such as VoIP and large-scale neural network training, require strict transmission delay guarantees. While routing under hop constraints is tractable, real-world delays increase sha...
Performance Evaluation of HAPS-enabled Coverage Enhancement in Hard-to-Reach AreasHao Lin, Mustafa A. Kishk, Mohamed-Slim Alouini2026-09-04下载High altitude platform stations (HAPSs) are becoming a key component of future non-terrestrial networks (NTNs). HAPSs can serve a larger area than uncrewed aerial vehicles (UAVs) and offer lower propa...
QUASAR: Quantum Satellite Architecture and Routing SimulatorYaliang Shi, Zi Wang, Bangguo Yuan, Gaojie Wu, Zhiwei Zhao2026-09-04下载The deployment of Low Earth Orbit (LEO) satellite constellations is an important step toward global-scale quantum networking. However, evaluating satellite quantum network protocols under spatiotempor...
A Wavelength Borrowing Architecture for Optical Data Center Networks - Extended VersionAndrea Detti, Chiara Lodovisi, Silvello Betti2026-09-04下载The growth of east-west traffic, along with the cost and power consumption of electronic switching, is motivating the integration of a low-power, high-rate, all-optical layer within the data center ne...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
DejaVu: Unifying Memory Allocations to Eliminate Redundant Copies on Unified-Memory SoCsYuheng Zhu, Yanbo Zhao, Jiajia Li, Man-Ki Yoon2026-09-04下载GPU applications on unified-memory (UMA) edge platforms often inherit a discrete-GPU memory abstraction in which they allocate one buffer for the CPU, another for the GPU, and copy data between them b...
Adaptive Context Parallelism for Production LLM ServingJiarui Guo, Rongle Wang, Peijun Huang, Zongwei Lv, Ziqing Wang, Kan Liu, Tao Lan, Lin Qu, Xiaolin Wang, Tong Yang2026-09-04下载As LLM context windows expand and input sequences grow longer, serving systems face increasing computational and memory demands. Context parallelism (CP), which partitions the input sequence across mu...

cs.PF - Performance ​

标题作者发布日期PDF摘要
RAGMark: A Comprehensive Framework for Benchmarking Retrieval-Augmented Generation SystemsZlatan Feric, Amir Taherin, Bin Ren, Yanzhi Wang, Jennifer Dy, David Kaeli2026-09-04下载We present RAGMark, a modular benchmarking framework for advanced Retrieval-Augmented Generation (RAG) systems targeting small-scale multi-GPU environments.
GreenPipe: Power Modeling for Containerized DNN Inference on Kubernetes Edge NodesMengxue Wang, Peini Liu, Amir Taherkordi, Jordi Guitart2026-09-04下载Distributed DNN inference is increasingly deployed in containerized edge-cloud environments, where workloads run on-device or are exposed to remote clients over the network.

基于 VitePress 构建 · 使用本地搜索查找论文