2026-09-30
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| MANTA: Machine Learning Augmented Tiering Advisor | Johannes Freischuetz, Kiet Pham, Sujay Yadalam, Konstantinos Kanellis, Michael Swift, Shivaram Venkataraman | 2026-09-30 | 下载 | Memory tiering has been used to expand memory capacity, particularly in datacenters, by combining fast DRAM with slower tiers, including CXL-attached memory. |
| Ditto: Generalized Reconfigurable Linearizable Reads | Aleksey Panas | 2026-09-30 | 下载 | Linearizability gives developers the illusion that operations on a distributed datastore execute sequentially on a single machine. Providing it is expensive, so numerous specialized reads algo- rithms... |
| Progressive-Resolution Secure Aggregation for Federated Learning | Seyed Mohammad Azimi-Abarghouyi | 2026-09-30 | 下载 | Secure aggregation lets a server recover an aggregate of client updates without observing any individual update, but conventional protocols fix the aggregate precision when clients upload. |
| Leto: Fast In-Place Recovery for LLM Training on Surviving Hardware | Geon-Woo Kim, Joon Ha Kim, Daehyeok Kim | 2026-09-30 | 下载 | Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement. |
| MegaFlux: Skew-Resilient MoE Megakernels via Pipelined Expert Replication | Jianzhu Yao, Siva Kumar Sastry Hari, Vignesh Balaji, Sana Damani, Insu Jang, Pramod Viswanath, Christos Kozyrakis | 2026-09-30 | 下载 | Mixture-of-experts (MoE) megakernels fuse expert-parallel communication with expert computation. However, under fixed expert placement, routing skew creates GPU stragglers: overloaded GPUs determine l... |
| Redundancy Meets Synergy: Dependency-aware Expert Selection for MoE via Submodular Optimization | Zheng Lin, Shaoke Fang, Yuxin Zhang, Jinfeng Xu, Zihan Fang, Zhe Chen, Wei Ni, Jun Luo, Symeon Chatzinotas | 2026-09-30 | 下载 | While Mixture-of-Experts (MoE) models effectively scale model capacity through sparse activation, their deployment is often bottlenecked by prohibitive memory requirements. |
| Local Relaxation Hierarchies for Quantum Ground State Energies: Convergence Guarantees and Message Passing Algorithms | Sheng-Ku Lin, Ricardo Rivera Cardoso, Roberto Bondesan | 2026-09-30 | 下载 | Convex relaxation hierarchies provide lower bounds to the ground state energy of quantum many-body systems that can be computed in polynomial time on a classical computer, at any fixed hierarchy level... |
| Reinforcement Learning-Guided Graph Transformations for SpTRSV Optimization | Buse Yılmaz | 2026-09-30 | 下载 | Sparse triangular solve (SpTRSV) is a fundamental kernel in numerous scientific and engineering applications. However, the data dependencies inherent in sparse triangular matrices significantly limit ... |
| Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs | Jaehwan Lee, Sangmin Lee, Chaewon Kim, Junsik Shin, Jaejin Lee | 2026-09-30 | 下载 | Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer... |
| LatencyLab: A DPDK-Based P4 Pipeline Latency Measurement Framework for FPGA SmartNICs | Pavani Kuppili, Zhaoyang Han, Yicheng Qian, Suranga Handagala, Michael Zink, Miriam Leeser, Robert Ricci | 2026-09-30 | 下载 | P4-programmable FPGA SmartNICs place packet processing directly on the wire, but open FPGA P4 toolflows do not expose timestamping at the pipeline boundary, so the latency a P4 program adds on the tar... |
| EPR Count for Runtime Prediction in Distributed Quantum Computing | Fatih E. Bilgen, Ozgur B. Akan | 2026-09-30 | 下载 | EPR-pair consumption is commonly used as a communication-cost objective in distributed quantum computing, but minimizing EPR cost does not necessarily minimize distributed execution time. |
| From Pilots to Production: Lessons in Cross-Institutional Federated Training and Artificial Intelligence for Science | Olivera Kotevska, Max Carlson, Yan Gao, Francis Jeanson, Yijiang Li, William Lindskog, Mohammad Naseri, Minseok Ryu, Sahil Tyagi, Jerry Watkins, Feiyi Wang, Ravi Madduri, Kibaek Kim | 2026-09-30 | 下载 | Many of the most valuable scientific datasets cannot be centralized: they are proprietary, export-controlled, classified, or bound by data-sovereignty restrictions. |
| Darpan: A Digital Twin Framework for the Next-Generation Computing Continuum | Zhiyu Wang, Rajkumar Buyya | 2026-09-30 | 下载 | Computing-continuum applications distribute work across devices, edge systems, fog resources, and clouds. While a placement, scheduling, or recovery decision is being made, resource availability, netw... |
| Exploring Adaptive Byzantine Quorum Systems to Improve Latency in the WAN | Linus Gnan, Rüdiger Kapitza, Christian Berger | 2026-09-30 | 下载 | Quorum systems enforce strict consistency in Byzantine fault-tolerant (BFT) state machine replication: Before a value is decided, a subset of replicas (called quorum) must exchange votes for the value... |
| Kirin: Cloud-native WebAssembly Service Orchestration | Joshua Bauer, Sebastian Werner, Maria C. Borges | 2026-09-30 | 下载 | Modern cloud computing infrastructure relies heavily on virtualization to provide workload isolation and resource efficiency. While containers have become the dominant deployment primitive due to thei... |
| About the Influence of Workflow Topology on Task Intensity Prediction through Graph Learning | Max Otto, Haci Ismail Aslan, Joel Witzke, Jonathan Bader, Odej Kao | 2026-09-30 | 下载 | Efficient resource provisioning for large-scale workflows on cloud infrastructures is a critical performance engineering challenge. These workflows are often structured as directed acyclic graphs (DAG... |
| FissionReady: Joint Workload and Power Scheduling for Data Centers Powered by Small Modular Reactors | Raghavendra Kanakagiri, Rohan Basu Roy, Yankai Jiang, Pranathi Wuppuluru, Devesh Tiwari | 2026-09-30 | 下载 | Data centers are increasingly exploring small modular nuclear reactors (SMRs) as a carbon-free power source, but variable datacenter demand and negative grid prices require the SMR plant to load-follo... |
| Communication-Efficient (1+\varepsilon)Δ-Edge Coloring and Lovász Local Lemma | Yi-Jun Chang, Nima Dolatabadi, Hung Thuan Nguyen | 2026-09-30 | 下载 | We study edge coloring in the two-party edge-partition model, where Alice and Bob each know part of the edge set and must jointly produce a proper coloring with little communication. |
| Robustifying Asynchronous SGD via Soft Throttling | Kaoru Otsuka, Maxime Meyer, Yuki Takezawa, Makoto Yamada, Anastasia Koloskova | 2026-09-30 | 下载 | Asynchronous SGD is a popular algorithm for distributed learning where each client's gradient update is applied on arrival. This leads to a speed-up, but also an increased vulnerability to attacks, as... |
| HAPMoE: Heterogeneity-Aware Automatic Parallelism Planning for Mixture-of-Experts Models Training | Mengyuan Fan, Peizhuang Cong, Zixiao Huang, Si Xu, Tong Qiao, Yanghao Li, Jing Yang, Tong Yang, Quanlu Zhang, Yu Wang | 2026-09-30 | 下载 | As model sizes continue to scale, distributed training has become inevitable. Automatic parallelization techniques can derive efficient training parallelism strategies at low cost while achieving supe... |
| Taming Speculative Search for Test-Time Scaling in LLM Serving | Jinwoo Jeong, Woohyung Choi, Myeongjae Jeon, Jeongseob Ahn | 2026-09-30 | 下载 | Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, substantially enhancing accuracy on challenging tasks s... |
| XIM: The XDC Interledger Messaging Protocol | Atul Khekade, Ritesh Kakkad, Wanwiset Peerapatanapokin, Behnam Mohammadkhani | 2026-09-30 | 下载 | Distributed ledgers, privacy-preserving institutional networks, and conventional payment systems increasingly need to exchange authenticated messages and settle assets across heterogeneous trust domai... |
| An Island-Based Parallel Biased Random-Key Genetic Algorithm for the Three-Dimensional Trailer Loading Problem | A. del Río, L. Díaz, L. C. de Vicente, J. Cameselle, B. Fernández | 2026-09-30 | 下载 | The Three-Dimensional Trailer Loading Problem (3D-TLP) involves determining the optimal placement and orientation of heterogeneous items within the confined space of a trailer while maximizing volume ... |
| Client and Training Data Selection for Computationally Efficient Synchronized Federated Learning | Muzaffer Citir, Hiroki Nishikawa, Sangyoung Park | 2026-09-30 | 下载 | Federated learning (FL) is a promising paradigm of machine learning, which preserves user privacy by enabling learning without sharing raw data with a cloud server. |
| Characterizing High Bandwidth Flash for LLM Serving | Zack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami | 2026-09-30 | 下载 | Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become... |
| Argus: A Real-EKS Study of When Predicting Spot Interruptions Beats Simple Checkpointing | Angshuman Chakravertty, MD Rayyan | 2026-09-30 | 下载 | Elastic Compute Cloud (EC2) Spot is 60% to 90% cheaper than On-Demand but can be reclaimed on just a 2-minute notice; for expensive multi-node training this loss can be severe, with one reclaim costin... |
| Vosti: Specifying, Implementing, and Verifying Deterministic LLM Inference | Jianxing Qin, Alexander Du, Danfeng Zhang, Matthew Lentz, Danyang Zhuo | 2026-09-30 | 下载 | LLM inference systems may vary batch composition, prompt chunking, prefill/decode execution, and KV-cache reuse, eviction, or recomputation. These optimizations should not affect system outputs. |
| Towards Efficient HPC Systems for Agents: Challenges and Opportunities | Yunjia Zheng, Bintang Dwi Marthen, Zachary Pan, Minghao Li, Raminder Singh, Manasvita Joshi, Minlan Yu, Juncheng Yang | 2026-09-30 | 下载 | Coding agents have become real users of high-performance computing (HPC) systems, yet today's HPC abstractions, interfaces, and policies remain designed for human-driven workflows. |
| Preserving Provenance in Shared KV Caches for LLM Serving | Wei Song, Yuxin Cao, Xi Zheng, Leo Zhang, Xiao Cheng | 2026-09-30 | 下载 | Production LLM serving stacks combine an inference engine's local prefix cache with a shared KV-cache tier for fleet-wide reuse. The local cache distinguishes requests by adapter, weight configuration... |
| Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure | Matias Parij, Pawan Paudel, Tate Berenbaum, Muthaiah Venkatachalam | 2026-09-30 | 下载 | We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ARCTAN: Arbitrary RF Containment Using Tactical Aerial Networks and Differentiable Ray Tracing | Samuel Rivera, Zhihui Gao, Yiming Li, Tingjun Chen | 2026-09-30 | 下载 | Aerial base stations (ABSs) can rapidly establish connectivity in ad hoc, infrastructure-deprived environments, but their broadcast, line-of-sight transmissions leak far beyond the intended service ar... |
| Resource-Efficient Semantic Communication for Heterogeneous Agentic Teams | Farhad Rezazadeh, Hatim Chergui, Lingjia Liu, Merouane Debbah | 2026-09-30 | 下载 | Teams of autonomous agents, including large language model (LLM) agents, must coordinate over scarce and unreliable wireless links. We propose goal-oriented semantic communication (GOSC), a closed-loo... |
| Spatio-Temporal Wireless-Optical Planning for Multi-UAV Networks | Binglei Wang, Huiru Ao, Fan Yang, Zhenjie Zhou, Zhonghua Peng, Jialong Li | 2026-09-30 | 下载 | Multi-unmanned aerial vehicle (UAV) networks in urban low-altitude environments couple UAV mobility, wireless access, and optical backhaul resources. |
| MoSE: Mode-Switching Expander for Mixed LLM Training and Inference | Fan Yang, Ying Zhou, Binglei Wang, Zhenjie Zhou, Jialong Li | 2026-09-30 | 下载 | AI clusters increasingly run large language model (LLM) inference and training on the same fabric. Prefill-decode (P-D) disaggregation creates key-value (KV) cache transfers between prefill and decode... |
| Can LLMs help find Ambiguities in Protocol Specifications? | Ziyue Dang, Sixu Tan, Atharva Nevasekar, Zhaowei Tan, George Varghese, Songwu Lu | 2026-09-30 | 下载 | Internet protocol specifications written in RFCs are subject to ambiguities and multiple interpretations that can cause interoperability failure. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| MANTA: Machine Learning Augmented Tiering Advisor | Johannes Freischuetz, Kiet Pham, Sujay Yadalam, Konstantinos Kanellis, Michael Swift, Shivaram Venkataraman | 2026-09-30 | 下载 | Memory tiering has been used to expand memory capacity, particularly in datacenters, by combining fast DRAM with slower tiers, including CXL-attached memory. |
| Herschel: Continuous Optimization of Production LLM Inference through On-Demand Profiling | Luping Wang, Weigao Chen, Yifei Wu, Yonghe Zhang, Rui Zhang, Wenchao Wu, Jiyu Luo, Haoran Geng, Xin Yang, Chen Cao, Yuemin Wu, Cheng Huang, Guodong Yang, Liping Zhang | 2026-09-30 | 下载 | Model-as-a-service platforms call for continuous optimization as complex serving conditions expose inefficiencies missed before deployment. Detailed always-on profiling can incur substantial overhead,... |
| Tide: Reclaiming Phased Memory in Agent MicroVMs | Yiyang Wu, Chengfan Liao, Jinyu Gu | 2026-09-30 | 下载 | Cloud agents run each task in an isolated MicroVM. The trouble is the harness loop inside that guest: the harness is nearly idle while it waits on the model, then usage rises on a tool whose size is k... |
| Capture the lifecycle: KV Cache management in ReAct Agents with KVTether | Kaihua Fu, Yukun Zhou, Chaokun Chang, Yinghao Yu, Luping Wang, Guodong Yang, Jiuchen Shi, Quan Chen, Wei Wang | 2026-09-30 | 下载 | Efficient serving of long-context reasoning-and-acting (ReAct) agents relies on KV cache reuse to reduce large language model (LLM) prefill latency and monetary cost. |
| Extending eBPF observability to Non-standard execution environments | Pamenas Kariuk, André Martin, Christof Fetzer | 2026-09-30 | 下载 | eBPF observability of non-standard execution environments (NEEs) like TEEs or LibOSes is hindered by their unconventional exception-handling and memory-access mechanisms that limit standard Linux tool... |
| Taming Speculative Search for Test-Time Scaling in LLM Serving | Jinwoo Jeong, Woohyung Choi, Myeongjae Jeon, Jeongseob Ahn | 2026-09-30 | 下载 | Test-time scaling has recently emerged as a powerful approach for improving LLM reasoning by allocating additional computation during inference, substantially enhancing accuracy on challenging tasks s... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| LLTA: A Simplicity-Oriented Open-Source WCET Analyser | Nils Hölscher, Kay Heider, Jian-Jia Chen | 2026-09-30 | 下载 | Deriving a safe upper bound on the worst-case execution time (WCET) of a real-time task is essential for hard real-time systems. Many WCET analysers exist, but they are either (i) closed source or (ii... |
| Distribution of Age of Information in the Erlang Loss System | Nail Akar, Sennur Ulukus | 2026-09-30 | 下载 | In this paper, we study the exact distributions of the age of information (AoI) and peak AoI (PAoI) in a bufferless setting in which time-stamped updates, or processed tasks, generated by one or sever... |
| Working Around the Compute Ceiling: Byte-Exact Memory in Galahad Makes LLM Reading a One-Time Cost LLM Reading a One-Time Cost | Sietse Schelpe | 2026-09-30 | 下载 | A transformer language model performs a bounded amount of computation per token, and recent work by Vishal Sikka, former CEO of Infosys, argues that this bound limits which tasks a model can carry out... |
| FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs | Shivam Saini, Eric Bezzam, Georg Götz, Alessia Milo, Steinar Guðjónsson, Konstantinos Gkanos, Finnur Pind, Daniel Gert Nielsen | 2026-09-30 | 下载 | Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. |