Skip to content

2026-09-30 ​

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
MANTA: Machine Learning Augmented Tiering AdvisorJohannes Freischuetz, Kiet Pham, Sujay Yadalam, Konstantinos Kanellis, Michael Swift, Shivaram Venkataraman2026-09-30下载Memory tiering has been used to expand memory capacity, particularly in datacenters, by combining fast DRAM with slower tiers, including CXL-attached memory.
Ditto: Generalized Reconfigurable Linearizable ReadsAleksey Panas2026-09-30下载Linearizability gives developers the illusion that operations on a distributed datastore execute sequentially on a single machine. Providing it is expensive, so numerous specialized reads algo- rithms...
Progressive-Resolution Secure Aggregation for Federated LearningSeyed Mohammad Azimi-Abarghouyi2026-09-30下载Secure aggregation lets a server recover an aggregate of client updates without observing any individual update, but conventional protocols fix the aggregate precision when clients upload.
Leto: Fast In-Place Recovery for LLM Training on Surviving HardwareGeon-Woo Kim, Joon Ha Kim, Daehyeok Kim2026-09-30下载Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement.
MegaFlux: Skew-Resilient MoE Megakernels via Pipelined Expert ReplicationJianzhu Yao, Siva Kumar Sastry Hari, Vignesh Balaji, Sana Damani, Insu Jang, Pramod Viswanath, Christos Kozyrakis2026-09-30下载Mixture-of-experts (MoE) megakernels fuse expert-parallel communication with expert computation. However, under fixed expert placement, routing skew creates GPU stragglers: overloaded GPUs determine l...
Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular OptimizationZheng Lin, Shaoke Fang, Yuxin Zhang, Jinfeng Xu, Zihan Fang, Zhe Chen, Wei Ni, Jun Luo, Symeon Chatzinotas2026-09-30下载While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements.
Local Relaxation Hierarchies for Quantum Ground State Energies: Convergence Guarantees and Message Passing AlgorithmsSheng-Ku Lin, Ricardo Rivera Cardoso, Roberto Bondesan2026-09-30下载Convex relaxation hierarchies provide lower bounds to the ground state energy of quantum many-body systems that can be computed in polynomial time on a classical computer, at any fixed hierarchy level...
Reinforcement Learning-Guided Graph Transformations for SpTRSV OptimizationBuse Yılmaz2026-09-30下载Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit ...
Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUsJaehwan Lee, Sangmin Lee, Chaewon Kim, Junsik Shin, Jaejin Lee2026-09-30下载Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer...
LatencyLab: A DPDK-Based P4 Pipeline Latency Measurement Framework for FPGA SmartNICsPavani Kuppili, Zhaoyang Han, Yicheng Qian, Suranga Handagala, Michael Zink, Miriam Leeser, Robert Ricci2026-09-30下载P4-programmable FPGA SmartNICs place packet processing directly on the wire, but open FPGA P4 toolflows do not expose timestamping at the pipeline boundary, so the latency a P4 program adds on the tar...
EPR Count for Runtime Prediction in Distributed Quantum ComputingFatih E. Bilgen, Ozgur B. Akan2026-09-30下载EPR-pair consumption is commonly used as a communication-cost objective in distributed quantum computing, but minimizing EPR cost does not necessarily minimize distributed execution time.
From Pilots to Production: Lessons in Cross-Institutional Federated Training and Artificial Intelligence for ScienceOlivera Kotevska, Max Carlson, Yan Gao, Francis Jeanson, Yijiang Li, William Lindskog, Mohammad Naseri, Minseok Ryu, Sahil Tyagi, Jerry Watkins, Feiyi Wang, Ravi Madduri, Kibaek Kim2026-09-30下载Many of the most valuable scientific datasets cannot be centralized: they are proprietary, export-controlled, classified, or bound by data-sovereignty restrictions.
Darpan: A Digital Twin Framework for the Next-Generation Computing ContinuumZhiyu Wang, Rajkumar Buyya2026-09-30下载Computing-continuum applications distribute work across devices, edge systems, fog resources, and clouds. While a placement, scheduling, or recovery decision is being made, resource availability, netw...
Exploring Adaptive Byzantine Quorum Systems to Improve Latency in the WANLinus Gnan, Rüdiger Kapitza, Christian Berger2026-09-30下载Quorum systems enforce strict consistency in Byzantine fault-tolerant (BFT) state machine replication: Before a value is decided, a subset of replicas (called quorum) must exchange votes for the value...
Kirin: Cloud-native WebAssembly Service OrchestrationJoshua Bauer, Sebastian Werner, Maria C. Borges2026-09-30下载Modern cloud computing infrastructure relies heavily on virtualization to provide workload isolation and resource efficiency. While containers have become the dominant deployment primitive due to thei...
About the Influence of Workflow Topology on Task Intensity Prediction through Graph LearningMax Otto, Haci Ismail Aslan, Joel Witzke, Jonathan Bader, Odej Kao2026-09-30下载Efficient resource provisioning for large-scale workflows on cloud infrastructures is a critical performance engineering challenge. These workflows are often structured as directed acyclic graphs (DAG...
FissionReady: Joint Workload and Power Scheduling for Data Centers Powered by Small Modular ReactorsRaghavendra Kanakagiri, Rohan Basu Roy, Yankai Jiang, Pranathi Wuppuluru, Devesh Tiwari2026-09-30下载Data centers are increasingly exploring small modular nuclear reactors (SMRs) as a carbon-free power source, but variable datacenter demand and negative grid prices require the SMR plant to load-follo...
Communication-Efficient (1+\varepsilon)Δ-Edge Coloring and Lovász Local LemmaYi-Jun Chang, Nima Dolatabadi, Hung Thuan Nguyen2026-09-30下载We study edge coloring in the two-party edge-partition model, where Alice and Bob each know part of the edge set and must jointly produce a proper coloring with little communication.
Robustifying Asynchronous SGD via Soft ThrottlingKaoru Otsuka, Maxime Meyer, Yuki Takezawa, Makoto Yamada, Anastasia Koloskova2026-09-30下载Asynchronous SGD is a popular algorithm for distributed learning where each client's gradient update is applied on arrival. This leads to a speed-up, but also an increased vulnerability to attacks, as...
HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models TrainingMengyuan Fan, Peizhuang Cong, Zixiao Huang, Si Xu, Tong Qiao, Yanghao Li, Jing Yang, Tong Yang, Quanlu Zhang, Yu Wang2026-09-30下载As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving supe...
Taming Speculative Search for Test-Time Scaling in LLM ServingJinwoo Jeong, Woohyung Choi, Myeongjae Jeon, Jeongseob Ahn2026-09-30下载Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, substantially enhancing accuracy on challenging tasks s...
XIM: The XDC Interledger Messaging ProtocolAtul Khekade, Ritesh Kakkad, Wanwiset Peerapatanapokin, Behnam Mohammadkhani2026-09-30下载Distributed ledgers, privacy-preserving institutional networks, and conventional payment systems increasingly need to exchange authenticated messages and settle assets across heterogeneous trust domai...
An Island-Based Parallel Biased Random-Key Genetic Algorithm for the Three-Dimensional Trailer Loading ProblemA. del Río, L. Díaz, L. C. de Vicente, J. Cameselle, B. Fernández2026-09-30下载The Three-Dimensional Trailer Loading Problem (3D-TLP) involves determining the optimal placement and orientation of heterogeneous items within the confined space of a trailer while maximizing volume ...
Client and Training Data Selection for Computationally Efficient Synchronized Federated LearningMuzaffer Citir, Hiroki Nishikawa, Sangyoung Park2026-09-30下载Federated learning (FL) is a promising paradigm of machine learning, which preserves user privacy by enabling learning without sharing raw data with a cloud server.
Characterizing High Bandwidth Flash for LLM ServingZack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami2026-09-30下载Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become...
Argus: A Real-EKS Study of When Predicting Spot Interruptions Beats Simple CheckpointingAngshuman Chakravertty, MD Rayyan2026-09-30下载Elastic Compute Cloud (EC2) Spot is 60% to 90% cheaper than On-Demand but can be reclaimed on just a 2-minute notice; for expensive multi-node training this loss can be severe, with one reclaim costin...
Vosti: Specifying, Implementing, and Verifying Deterministic LLM InferenceJianxing Qin, Alexander Du, Danfeng Zhang, Matthew Lentz, Danyang Zhuo2026-09-30下载LLM inference systems may vary batch composition, prompt chunking, prefill/decode execution, and KV-cache reuse, eviction, or recomputation. These optimizations should not affect system outputs.
Towards Efficient HPC Systems for Agents: Challenges and OpportunitiesYunjia Zheng, Bintang Dwi Marthen, Zachary Pan, Minghao Li, Raminder Singh, Manasvita Joshi, Minlan Yu, Juncheng Yang2026-09-30下载Coding agents have become real users of high-performance computing (HPC) systems, yet today's HPC abstractions, interfaces, and policies remain designed for human-driven workflows.
Preserving Provenance in Shared KV Caches for LLM ServingWei Song, Yuxin Cao, Xi Zheng, Leo Zhang, Xiao Cheng2026-09-30下载Production LLM serving stacks combine an inference engine's local prefix cache with a shared KV-cache tier for fleet-wide reuse. The local cache distinguishes requests by adapter, weight configuration...
Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI InfrastructureMatias Parij, Pawan Paudel, Tate Berenbaum, Muthaiah Venkatachalam2026-09-30下载We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
ARCTAN: Arbitrary RF Containment Using Tactical Aerial Networks and Differentiable Ray TracingSamuel Rivera, Zhihui Gao, Yiming Li, Tingjun Chen2026-09-30下载Aerial base stations (ABSs) can rapidly establish connectivity in ad hoc, infrastructure-deprived environments, but their broadcast, line-of-sight transmissions leak far beyond the intended service ar...
Resource-Efficient Semantic Communication for Heterogeneous Agentic TeamsFarhad Rezazadeh, Hatim Chergui, Lingjia Liu, Merouane Debbah2026-09-30下载Teams of autonomous agents, including large language model (LLM) agents, must coordinate over scarce and unreliable wireless links. We propose goal-oriented semantic communication (GOSC), a closed-loo...
Spatio-Temporal Wireless-Optical Planning for Multi-UAV NetworksBinglei Wang, Huiru Ao, Fan Yang, Zhenjie Zhou, Zhonghua Peng, Jialong Li2026-09-30下载Multi-unmanned aerial vehicle (UAV) networks in urban low-altitude environments couple UAV mobility, wireless access, and optical backhaul resources.
MoSE: Mode-Switching Expander for Mixed LLM Training and InferenceFan Yang, Ying Zhou, Binglei Wang, Zhenjie Zhou, Jialong Li2026-09-30下载AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode...
Can LLMs help find Ambiguities in Protocol Specifications?Ziyue Dang, Sixu Tan, Atharva Nevasekar, Zhaowei Tan, George Varghese, Songwu Lu2026-09-30下载Internet protocol specifications written in RFCs are subject to ambiguities and multiple interpretations that can cause interoperability failure.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
MANTA: Machine Learning Augmented Tiering AdvisorJohannes Freischuetz, Kiet Pham, Sujay Yadalam, Konstantinos Kanellis, Michael Swift, Shivaram Venkataraman2026-09-30下载Memory tiering has been used to expand memory capacity, particularly in datacenters, by combining fast DRAM with slower tiers, including CXL-attached memory.
Herschel: Continuous Optimization of Production LLM Inference through On-Demand ProfilingLuping Wang, Weigao Chen, Yifei Wu, Yonghe Zhang, Rui Zhang, Wenchao Wu, Jiyu Luo, Haoran Geng, Xin Yang, Chen Cao, Yuemin Wu, Cheng Huang, Guodong Yang, Liping Zhang2026-09-30下载Model-as-a-service platforms call for continuous optimization as complex serving conditions expose inefficiencies missed before deployment. Detailed always-on profiling can incur substantial overhead,...
Tide: Reclaiming Phased Memory in Agent MicroVMsYiyang Wu, Chengfan Liao, Jinyu Gu2026-09-30下载Cloud agents run each task in an isolated MicroVM. The trouble is the harness loop inside that guest: the harness is nearly idle while it waits on the model, then usage rises on a tool whose size is k...
Capture the lifecycle: KV Cache management in ReAct Agents with KVTetherKaihua Fu, Yukun Zhou, Chaokun Chang, Yinghao Yu, Luping Wang, Guodong Yang, Jiuchen Shi, Quan Chen, Wei Wang2026-09-30下载Efficient serving of long-context reasoning-and-acting (ReAct) agents relies on KV cache reuse to reduce large language model (LLM) prefill latency and monetary cost.
Extending eBPF observability to Non-standard execution environmentsPamenas Kariuk, André Martin, Christof Fetzer2026-09-30下载eBPF observability of non-standard execution environments (NEEs) like TEEs or LibOSes is hindered by their unconventional exception-handling and memory-access mechanisms that limit standard Linux tool...
Taming Speculative Search for Test-Time Scaling in LLM ServingJinwoo Jeong, Woohyung Choi, Myeongjae Jeon, Jeongseob Ahn2026-09-30下载Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, substantially enhancing accuracy on challenging tasks s...

cs.PF - Performance ​

标题作者发布日期PDF摘要
LLTA: A Simplicity-Oriented Open-Source WCET AnalyserNils Hölscher, Kay Heider, Jian-Jia Chen2026-09-30下载Deriving a safe upper bound on the worst-case execution time (WCET) of a real-time task is essential for hard real-time systems. Many WCET analysers exist, but they are either (i) closed source or (ii...
Distribution of Age of Information in the Erlang Loss SystemNail Akar, Sennur Ulukus2026-09-30下载In this paper, we study the exact distributions of the age of information (AoI) and peak AoI (PAoI) in a bufferless setting in which time-stamped updates, or processed tasks, generated by one or sever...
Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time CostSietse Schelpe2026-09-30下载A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out...
FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRsShivam Saini, Eric Bezzam, Georg Götz, Alessia Milo, Steinar Guðjónsson, Konstantinos Gkanos, Finnur Pind, Daniel Gert Nielsen2026-09-30下载Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech.

基于 VitePress 构建 · 使用本地搜索查找论文