Skip to content

2026-08-14 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM PathsSeeyeon Kim, Juhyeong Jin, Joo-Young Kim2026-08-14下载Modern mixture-of-experts (MoE) language models increasingly strain the capacity and cost efficiency of high-bandwidth memory (HBM), as rapidly growing expert weights must be provisioned close to GPUs...
Qu-Trefoil: Large-Scale Quantum Circuit Simulator Working on FPGA With SATA StoragesKaijie Wei, Hideharu Amano, Ryohei Niwase, Yoshiki Yamaguchi, Takefumi Miyoshi2026-08-14下载Quantum circuits are fundamental components of quantum computing, and state-vector-based quantum circuit simulation is a widely used technique for tracking qubit behavior throughout circuit evolution.
The Quartic Hessian Conjecture in Dimension FourZixiang Ni2026-08-14下载The Hessian conjecture asks whether a polynomial with nonzero constant Hessian determinant has a polynomial gradient inverse. It is known in dimensions at most three, false in dimensions at least five...
Experimental Study on System-Level Performance Impact of Read Disturbance in Modern SSDsYonggon Park, Hyunuk Cho, Onur Mutlu, Sungjin Lee, Jisung Park2026-08-14下载This work investigates the system-level performance impact of read disturbance in modern NAND flash-based SSDs, aiming to provide new insights that can help develop better storage architectures and op...
MoE Expert Execution in Disaggregated LLM Serving with a High-Bandwidth ReRAM Near-Memory ArchitectureKunming Shao, Ming Zeng, Xin Yuan, Binbin Liao, Yangming Zhang, Wei Wang, Tim Kwang-Ting Cheng, Chi-Ying Tsui2026-08-14下载Attention-FFN disaggregation maps LLM modules to specialized pools, creating an opening to keep Mixture-of-Experts (MoE) weights resident in a high-bandwidth FFN pool.
Characterizing the Variance Envelope: A Multi-Dimensional Analysis of Spectre Telemetry Across Architectures and WorkloadsJaya Keshava Chandra Kotha, Jean-Luc Gaudiot2026-08-14下载Hardware attacks like Spectre exploit built-in processor vulnerabilities, leaving anomalous footprints in Hardware Performance Counter (HPC) metrics.
Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and ChallengesDowon Son, Yonggon Park, Hyunuk Cho, Hyungkyu Ham, Onur Mutlu, Sungjin Lee, Gwangsun Kim, Jisung Park2026-08-14下载This work investigates the potential benefits and technical challenges of using high-bandwidth flash (HBF) for large language model (LLM) inference.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Evaluating Agentic Code Repair Capabilities in Distributed SystemsYibo Yan, Huijuan Wang, Junzhou He, Yizhuo Liang, Shaoyu Wang, Huanchen Sun, Seo Jin Park2026-08-14下载LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified.
Enabling Hybrid HPCQC Workflows with a Heterogeneous Software StackMuhammad Nufail Farooqi, Minh Chung, Burak Mete, Eric Mansfield, Bernd Hoffmann, Teemu Mattsson, Laura Schulz, Jorge Echavarria2026-08-14下载In this work, we demonstrate hybrid High Performance Computing-Quantum Computing (HPCQC) workflows on a production petascale system. The demonstration combines three components: the SuperMUC-NG superc...
Porting and Benchmarking Chapel on Emerging RISC-V Hardware: an HPC Viability StudyIan Henriksen, Chris Taylor, Patrick Diehl, Jade Abraham, Palmer Cox, Bradford L. Chamberlain, Stephen L. Olivier2026-08-14下载The Chapel programming language recently added support for the RISC-V architecture. Here we discuss what changes were needed for Chapel to work on RISC-V as well as lessons learned from the porting pr...
Validating LLM-Modernized Scientific Software Through Differential Fault InjectionEvan Coleman, Yuzhong Shen, Masha Sosonkina, Peng Xu2026-08-14下载Large language model (LLM) agents are increasingly used to modernize the legacy Fortran underlying production scientific software, but validation of these transformations emphasizes nominal executions...
Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-TrainingHanfeng Lu, Tianyu Feng, Suyi Li, Yuheng Zhao, Wei Gao, Shaopan Xiong, Ju Huang, Siran Yang, Jiamang Wang, Lin Qu, Wei Wang2026-08-14下载Vision-language models (VLMs) enable embodied agents to reason and act from visual observations and language instructions. Reinforcement learning (RL) post-training enhances these capabilities using t...
Large-scale workflow placement in serverless computing using integer nonlinear programmingJoshua Adamek, Natalie Carl, Trever Schirmer, Moritz Heinlein, David Bermbach, Sergio Lucia2026-08-14下载Serverless edge computing has become a powerful cloud framework that enables the execution of large workflows without the need for the user to manage the underlying servers and edge devices.
Could Model Partitioning Make Federated Learning More Sustainable?Tobias Frohlich, Tiffany Vlaar, Lauritz Thamsen2026-08-14下载As federated learning (FL) extends from distributed machine learning between low-power devices to cross-silo scenarios involving edge servers and data centres, its carbon footprint has become a growin...
Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal LearningChun-Hua Lin, Samuel Yen-Chi Chen, Yu-Chao Hsu, Kuo-Chung Peng, Jiun-Cheng Jiang, Chi-Sheng Chen, Tai-Yue Li, Nan-Yow Chen, En-Jui Kuo, Hsi-Sheng Goan2026-08-14下载Electrocardiogram (ECG) recordings are sensitive biomedical data, limiting the ability of hospitals and wearable devices to share raw signals for centralized model training.
Federated Prompt Learning: A Unified Framework, Empirical Analysis, and Future DirectionsQinglin Yang, Chen Qiu, Hongyuan Zhang, Pengdeng Li, Yuan Liu, Zhihong Tian2026-08-14下载Large language models (LLMs) have become core components of cloud-based intelligent services in academia and industry, yet their training and deployment are hindered by high computational costs, data ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Deep Reinforcement Learning for 6G AI-RAN: A Comprehensive SurveyJie Lu, Peihao Yan, Qijun Wang, Ruxin Lin, Huacheng Zeng2026-08-14下载The evolution toward sixth-generation (6G) networks is transforming the radio access network (RAN) into a programmable and intelligent control platform that must continuously adapt to heterogeneous se...
Scaling 5G-TSN Bridges: Operating Regimes, Scheduling, and Time Synchronisation Under Heterogeneous Industrial TrafficMohamed Seliem, Utz Roedig, Cormac Sreenan, Dirk Pesch2026-08-14下载3GPP Release 16 enables a 5G system to operate as a transparent IEEE 802.1 TSN bridge, but its scalability under heterogeneous industrial workloads remains insufficiently characterised.
Robust Constraint-Aware Bayesian Tuning of BBRv2 for QUIC under Tactile Internet ConstraintsMuhammad Hanif Lashari, Shakil Ahmed, Wafa Batayneh, Ashfaq Khokhar2026-08-14下载Tactile Internet applications place strict require- ments on latency, jitter, loss, and responsiveness, which makes transport configuration a critical design factor.
CipherSight: Robust Website Fingerprinting via Record-Resource Semantic Supervision under Distribution ShiftsRunhan Song, Qiqi Liu, Chuanzhou Pan, Zhenquan Ding, Youquan Xian, Chongru Fan, Lei Cui, Wei Wang, Zhiyu Hao2026-08-14下载HTTPS website fingerprinting (WF) aims to identify visited websites from metadata observable in encrypted traffic. However, real-world deployments introduce a significant out-of-distribution (OOD) pro...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
CoRun: Padding is Simple and Efficient for Deterministic LLM InferenceShiju Zhao, Jiacheng Yang, Qihang Chen, Junhao Hu, Jiaqi Zheng, Guihai Chen, Xusheng Chen2026-08-14下载Despite fixed sampling parameters and random seeds, Large Language Model (LLM) inference exhibits output inconsistency, which undermines downstream tasks such as model evaluation and reinforcement lea...

cs.PF - Performance ​

标题作者发布日期PDF摘要
CoRun: Padding is Simple and Efficient for Deterministic LLM InferenceShiju Zhao, Jiacheng Yang, Qihang Chen, Junhao Hu, Jiaqi Zheng, Guihai Chen, Xusheng Chen2026-08-14下载Despite fixed sampling parameters and random seeds, Large Language Model (LLM) inference exhibits output inconsistency, which undermines downstream tasks such as model evaluation and reinforcement lea...

基于 VitePress 构建 · 使用本地搜索查找论文