2026-07-13
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Overcoming Orchestration Bottlenecks at Exascale: A Decentralized, Policy-Driven Approach for Sim-AI Ensembles | Harikrishna Tummalapalli, Christine M. Simpson, Riccardo Balin, Vitali A. Morozov, Thang D. Pham, Murat Keceli, Thomas D. Uram | 2026-07-13 | 下载 | Scientific computing is increasingly shifting from monolithic applications to coupled simulation-AI workflows composed of highly heterogeneous tasks with diverse hardware, scale, and runtime requireme... |
| Decentralized Gradient Descent: Bottleneck Regimes and Budget Complexity | Nicolò Michelusi | 2026-07-13 | 下载 | Decentralized gradient descent (DGD) is widely used for solving distributed optimization problems over networks of agents. While its convergence properties are well understood, less is known about the... |
| FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving | Yaqi Qiao, Ping He, Songrun Xie, Ayush Barik, Chensong Zhang, Zhengzhong Tu, Fan Lai | 2026-07-13 | 下载 | Diffusion models have become the central backbone for modern image, video, and audio generation, but their efficient service remains a challenge. |
| Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap | Rafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage, Laura Biven, Michael Bussmann, Kyle Chard, Ryan Coffee, Stephen DeWitt, Sagar Dolas, Carrie Eckert, David Elbert, Ian Foster, Tirthankar Ghosal, Anna Giannakou, Tom Gibbs, Leslie Hamilton, Glenn Lockwood, Theresa Mayer, Ben Mintz, Raffi Nazikian, Sal Nimer, Amanda Randles, Woong Shin, Sreenivas Rangan Sukumar, Frédéric Suter, Mitra Taheri, Michela Taufer, Draguna Vrabie | 2026-07-13 | 下载 | One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions. |
| Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning | Jing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu | 2026-07-13 | 下载 | Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust ... |
| PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs | Jing Liu, Kun Yang, Yan Wang, Dingkang Yang, Xiaoshuai Hao, Wei Zhang, Yang Liu, Wei Zhou | 2026-07-13 | 下载 | Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges. |
| HARP-ME: Closure-Driven Exact Induced Motif Enumeration on GPUs | Ashwina Kumar, Rupesh Nasre | 2026-07-13 | 下载 | Exact induced motif enumeration is a fundamental operation in graph mining, but it remains challenging on GPUs because candidate expansion is irregular, repeated set intersections dominate execution, ... |
| MemExchange: Utility-Driven Distributed Memory Reallocation for Multi-Tenant Datacenters | AmirHossein Seyri, Abhisek Pan, Balajee Vamanan | 2026-07-13 | 下载 | To handle unpredictable workloads, cloud providers typically over-provision memory to meet peak demand, resulting in substantial underutilization across datacenter clusters. |
| Time Is Money: Incentivized Causal Transaction Ordering | Hongyin Chen, Xu Zheng, Jichen Li, Ittay Eyal | 2026-07-13 | 下载 | Front-running is a subtle and persistent problem for blockchains. A blockchain is a stateful virtual machine executing instructions called transactions. |
| Stage-Level Executor Allocation in Apache Spark with Cost-Performance Trade-offs | Miriam Rateike, Isaac Waweru Wambugu, Celia Cintas, Michael Kaufmann, Ioana Giurgiu, Skyler Speakman | 2026-07-13 | 下载 | Allocating executors (i.e. compute resources) to distributed processing systems must balance resource costs of scaling-out unnecessarily against artificial, performance-limiting bottlenecks. |
| Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs | Weijia Han, Lisha Qu | 2026-07-13 | 下载 | Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four NVIDIA RTX A5000 GPU... |
| GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-Design | Yitong Ding, Jiawei Huang, Renyang Guan, Yangjie Zhou, Zihan Liu, Yu Feng, Shixuan Sun, Mingyi Guo, Jingwen Leng, Jian Weng | 2026-07-13 | 下载 | Modern LLM (large language model) workloads increasingly rely on optimized GPU kernels through hardware-software co-design. These kernels achieve high-performance through fine-grained dependency sched... |
| Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration | Xueze Kang, Guangyu Xiang, Suyi Li, Yuxin Wang, Shaohuai Shi, Lin Zhang, Xiaowen Chu | 2026-07-13 | 下载 | Diffusion models are increasingly deployed as production visual-generation services, where serving high-resolution image and long video generation is often limited by GPU memory. |
| Energy Calculus: A Compositional Algebra of Energy in Computational Systems | Mosharaf Chowdhury, Jae-Won Chung, Jeff J. Ma, Nishil Talati, Ruofan Wu | 2026-07-13 | 下载 | Energy is a binding constraint for AI scaling, yet it lacks the formal treatment that computation, communication, and learning have long enjoyed. |
| Domain Extension of Lock-Freedom and Wait-Freedom for Group Computations | Raaghav Ravishankar, Sandeep Kulkarni, Sathya Peri, Gokarna Sharma, Manaswini Piduguralla, Yashi Rastogi | 2026-07-13 | 下载 | A domain extension of a definition refers to broadening the scope of a definition so that it applies to a larger set of cases than originally specified. |
| [AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent Workflows | Arup Kumar Sarker, Alexander James Halpern, Mills Staylor, Aymen Alsaadi, Gregor von Laszewski, Yue Cheng, Shantenu Jha, Geoffrey Fox | 2026-07-13 | 下载 | Multi-agent LLM systems increasingly integrate retrieval, planning, and reasoning, but remain fundamentally text-centric, requiring agents to repeatedly recompute shared context through expensive pref... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| OSNR/GSNR Prediction in Brownfield Links via a DLM-Anchored Hybrid Physics/ML Model | Agastya Raj, Venkata Virajit Garbhapu, Hiroyuki Ishihara, Peyman Pahlevanzadeh, Hideki Nishizawa, Takeo Sasai, Daniel C. Kilper, Marco Ruffini | 2026-07-13 | 下载 | We present a DLM-anchored hybrid physics/ML framework for brownfield optical links that accurately predicts per-channel power, OSNR, and GSNR. |
| PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs | Jing Liu, Kun Yang, Yan Wang, Dingkang Yang, Xiaoshuai Hao, Wei Zhang, Yang Liu, Wei Zhou | 2026-07-13 | 下载 | Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges. |
| GNSS Spoofing Detection in TDD Networks: A 3GPP Standards-Based Security Framework | Ravi Kant Sharma, John Owens, Kevin Kiernan | 2026-07-13 | 下载 | Time Division Duplex (TDD) mobile networks require synchronization accuracy of $\pm$1.5 μs (3GPP TS 38.104), with GNSS-disciplined grandmaster clocks as the predominant timing source. |
| An Experimental Evaluation of Accurate Scheduling and Hardware Timestamping on NVIDIA ConnectX NICs | Takahiro Hirofuchi, Takaaki Fukai | 2026-07-13 | 下载 | High-precision packet transmission is becoming increasingly important in deterministic networking applications, including 5G fronthaul and Time-Sensitive Networking (TSN). |
| SAIL: Perceptual Quality-Aware Rate Control for Cloud Gaming | Houde Qian, Chenglei Wu, Jiaxing Zhang, Rui-Xiao Zhang, Jing Wang, Meijia Song, Sijia Chen, Xiaozhong Xu, Zhi Wang, Lifeng Sun, Honghao Liu | 2026-07-13 | 下载 | Cloud gaming streams cloud-rendered frames under strict motion-to-photon latency, yet its at-scale viability is increasingly constrained by bandwidth cost: in our study of the T cloud gaming platform,... |
| SkillComm: Skill-Driven Semantic Communication for Sequential Workflows via Incremental Token Transmission | Ziyang Meng, Lu Lu | 2026-07-13 | 下载 | As wireless visual intelligence evolves from isolated task inference to ordered skill workflows, the communication bottleneck shifts from transmitting a single semantic representation to coordinating ... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUs | Weijia Han, Lisha Qu | 2026-07-13 | 下载 | Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four NVIDIA RTX A5000 GPU... |
| Energy Calculus: A Compositional Algebra of Energy in Computational Systems | Mosharaf Chowdhury, Jae-Won Chung, Jeff J. Ma, Nishil Talati, Ruofan Wu | 2026-07-13 | 下载 | Energy is a binding constraint for AI scaling, yet it lacks the formal treatment that computation, communication, and learning have long enjoyed. |