Skip to content

2026-07-13 ​

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Overcoming Orchestration Bottlenecks at Exascale: A Decentralized, Policy-Driven Approach for Sim-AI EnsemblesHarikrishna Tummalapalli, Christine M. Simpson, Riccardo Balin, Vitali A. Morozov, Thang D. Pham, Murat Keceli, Thomas D. Uram2026-07-13下载Scientific computing is increasingly shifting from monolithic applications to coupled simulation-AI workflows composed of highly heterogeneous tasks with diverse hardware, scale, and runtime requireme...
Decentralized Gradient Descent: Bottleneck Regimes and Budget ComplexityNicolò Michelusi2026-07-13下载Decentralized gradient descent (DGD) is widely used for solving distributed optimization problems over networks of agents. While its convergence properties are well understood, less is known about the...
FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model ServingYaqi Qiao, Ping He, Songrun Xie, Ayush Barik, Chensong Zhang, Zhengzhong Tu, Fan Lai2026-07-13下载Diffusion models have become the central backbone for modern image, video, and audio generation, but their efficient service remains a challenge.
Toward Trustworthy Autonomous Science: A Two-Year Community RoadmapRafael Ferreira da Silva, Milad Abolhasani, Peter Beaucage, Laura Biven, Michael Bussmann, Kyle Chard, Ryan Coffee, Stephen DeWitt, Sagar Dolas, Carrie Eckert, David Elbert, Ian Foster, Tirthankar Ghosal, Anna Giannakou, Tom Gibbs, Leslie Hamilton, Glenn Lockwood, Theresa Mayer, Ben Mintz, Raffi Nazikian, Sal Nimer, Amanda Randles, Woong Shin, Sreenivas Rangan Sukumar, Frédéric Suter, Mitra Taheri, Michela Taufer, Draguna Vrabie2026-07-13下载One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organized around five critical dimensions.
Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-TuningJing Liu, Chenxuanyin Zou, Jiayang Ren, Gaoyun Fang, Chengfang Li, Yan Wang, Zhenchao Ma, Bo Hu2026-07-13下载Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolving data streams, yet a fundamental obstacle prevents robust ...
PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMsJing Liu, Kun Yang, Yan Wang, Dingkang Yang, Xiaoshuai Hao, Wei Zhang, Yang Liu, Wei Zhou2026-07-13下载Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges.
HARP-ME: Closure-Driven Exact Induced Motif Enumeration on GPUsAshwina Kumar, Rupesh Nasre2026-07-13下载Exact induced motif enumeration is a fundamental operation in graph mining, but it remains challenging on GPUs because candidate expansion is irregular, repeated set intersections dominate execution, ...
MemExchange: Utility-Driven Distributed Memory Reallocation for Multi-Tenant DatacentersAmirHossein Seyri, Abhisek Pan, Balajee Vamanan2026-07-13下载To handle unpredictable workloads, cloud providers typically over-provision memory to meet peak demand, resulting in substantial underutilization across datacenter clusters.
Time Is Money: Incentivized Causal Transaction OrderingHongyin Chen, Xu Zheng, Jichen Li, Ittay Eyal2026-07-13下载Front-running is a subtle and persistent problem for blockchains. A blockchain is a stateful virtual machine executing instructions called transactions.
Stage-Level Executor Allocation in Apache Spark with Cost-Performance Trade-offsMiriam Rateike, Isaac Waweru Wambugu, Celia Cintas, Michael Kaufmann, Ioana Giurgiu, Skyler Speakman2026-07-13下载Allocating executors (i.e. compute resources) to distributed processing systems must balance resource costs of scaling-out unnecessarily against artificial, performance-limiting bottlenecks.
Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUsWeijia Han, Lisha Qu2026-07-13下载Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four NVIDIA RTX A5000 GPU...
GPU-Tile-Sim: A Tile-Centric GPU Simulation Framework for LLM Hardware-Software Co-DesignYitong Ding, Jiawei Huang, Renyang Guan, Yangjie Zhou, Zihan Liu, Yu Feng, Shixuan Sun, Mingyi Guo, Jingwen Leng, Jian Weng2026-07-13下载Modern LLM (large language model) workloads increasingly rely on optimized GPU kernels through hardware-software co-design. These kernels achieve high-performance through fine-grained dependency sched...
Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-ConfigurationXueze Kang, Guangyu Xiang, Suyi Li, Yuxin Wang, Shaohuai Shi, Lin Zhang, Xiaowen Chu2026-07-13下载Diffusion models are increasingly deployed as production visual-generation services, where serving high-resolution image and long video generation is often limited by GPU memory.
Energy Calculus: A Compositional Algebra of Energy in Computational SystemsMosharaf Chowdhury, Jae-Won Chung, Jeff J. Ma, Nishil Talati, Ruofan Wu2026-07-13下载Energy is a binding constraint for AI scaling, yet it lacks the formal treatment that computation, communication, and learning have long enjoyed.
Domain Extension of Lock-Freedom and Wait-Freedom for Group ComputationsRaaghav Ravishankar, Sandeep Kulkarni, Sathya Peri, Gokarna Sharma, Manaswini Piduguralla, Yashi Rastogi2026-07-13下载A domain extension of a definition refers to broadening the scope of a definition so that it applies to a larger set of cases than originally specified.
[AAFLOW+] Stateful Operator Abstraction with Zero-Copy Distributed KV Cache Orchestration for Multi-Agent WorkflowsArup Kumar Sarker, Alexander James Halpern, Mills Staylor, Aymen Alsaadi, Gregor von Laszewski, Yue Cheng, Shantenu Jha, Geoffrey Fox2026-07-13下载Multi-agent LLM systems increasingly integrate retrieval, planning, and reasoning, but remain fundamentally text-centric, requiring agents to repeatedly recompute shared context through expensive pref...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
OSNR/GSNR Prediction in Brownfield Links via a DLM-Anchored Hybrid Physics/ML ModelAgastya Raj, Venkata Virajit Garbhapu, Hiroyuki Ishihara, Peyman Pahlevanzadeh, Hideki Nishizawa, Takeo Sasai, Daniel C. Kilper, Marco Ruffini2026-07-13下载We present a DLM-anchored hybrid physics/ML framework for brownfield optical links that accurately predicts per-channel power, OSNR, and GSNR.
PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMsJing Liu, Kun Yang, Yan Wang, Dingkang Yang, Xiaoshuai Hao, Wei Zhang, Yang Liu, Wei Zhou2026-07-13下载Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning while maintaining data privacy at network edges.
GNSS Spoofing Detection in TDD Networks: A 3GPP Standards-Based Security FrameworkRavi Kant Sharma, John Owens, Kevin Kiernan2026-07-13下载Time Division Duplex (TDD) mobile networks require synchronization accuracy of $\pm$1.5 μs (3GPP TS 38.104), with GNSS-disciplined grandmaster clocks as the predominant timing source.
An Experimental Evaluation of Accurate Scheduling and Hardware Timestamping on NVIDIA ConnectX NICsTakahiro Hirofuchi, Takaaki Fukai2026-07-13下载High-precision packet transmission is becoming increasingly important in deterministic networking applications, including 5G fronthaul and Time-Sensitive Networking (TSN).
SAIL: Perceptual Quality-Aware Rate Control for Cloud GamingHoude Qian, Chenglei Wu, Jiaxing Zhang, Rui-Xiao Zhang, Jing Wang, Meijia Song, Sijia Chen, Xiaozhong Xu, Zhi Wang, Lifeng Sun, Honghao Liu2026-07-13下载Cloud gaming streams cloud-rendered frames under strict motion-to-photon latency, yet its at-scale viability is increasingly constrained by bandwidth cost: in our study of the T cloud gaming platform,...
SkillComm: Skill-Driven Semantic Communication for Sequential Workflows via Incremental Token TransmissionZiyang Meng, Lu Lu2026-07-13下载As wireless visual intelligence evolves from isolated task inference to ordered skill workflows, the communication bottleneck shifts from transmitting a single semantic representation to coordinating ...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Decomposing Runtime, Kernel, and Quantization Speedups via a Matched FP16 Intermediate: A Hardware-Conditioned Case Study on Four NVIDIA RTX A5000 GPUsWeijia Han, Lisha Qu2026-07-13下载Reported serving speedups from quantized kernels typically bundle the weight format, the kernel, and the inference runtime into one number. We present an attribution study on four NVIDIA RTX A5000 GPU...
Energy Calculus: A Compositional Algebra of Energy in Computational SystemsMosharaf Chowdhury, Jae-Won Chung, Jeff J. Ma, Nishil Talati, Ruofan Wu2026-07-13下载Energy is a binding constraint for AI scaling, yet it lacks the formal treatment that computation, communication, and learning have long enjoyed.

基于 VitePress 构建 · 使用本地搜索查找论文