Skip to content

2026-08-16 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Concurrency Response of Plain Global Loads on the NVIDIA H100Somashekar Manjunath, Rahul Ramachandra M2026-08-16下载The bandwidth a memory-bound GPU kernel sustains is set by how many bytes it keeps in flight. We use Little's Law here as throughput accounting, not as a measured hardware pool.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
KV-Pipe: On the Relation Between KV Sharing and Pipeline Parallel Efficiency in LLMsMaryam Dialameh, Hossein Rajabzadeh, Harish Krishnamoorthy Murali, Walid Ahmed, Weiwei Zhang, Hyock Ju Kwon2026-08-16下载Pipeline parallelism (PP) is widely used to scale large language model (LLM) training, but its efficiency is often limited by stage imbalance and pipeline bubbles.
Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model ServingWeinan Liu, Zeyuan Ding, Dian Ding, Chengcheng Wan, Lu Tang, Guangtao Xue, Jiwu Shu, Yiming Zhang2026-08-16下载Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited. Reallocating containers is too heavyweight to utilize such fine-grained opportunities under SLA ...
Adaptive Heterogeneous Compression for Resource-Efficient Federated Knowledge DistillationChenwang Liu, Yijun Liu, Chang Liu, Xu Zhang, Pengchao Han2026-08-16下载Federated learning (FL) enables privacy-preserving distributed model training but faces challenges from heterogeneous model architectures and limited communication resources at the network edge.
When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency EstimationWenhao Yuan, Chenchen Lin, Wenhao Hu, Jian Chen, Jinfeng Xu, Shujie Li, Edith Cheuk Han Ngai2026-08-16下载\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, under client heterogeneity, the conventional static split stra...
DeltaLog: Deferred Materialization of Recurrent States for Linear Attention DecodingJunqing Lin, Jingwei Sun, Guangzhong Sun2026-08-16下载Linear attention models eliminate the quadratic prefix computation and context-growing KV cache of softmax attention by replacing pairwise token interactions with recurrent state updates.
FlashQuant: Sparse-Dense Fusion for Memory-Efficient Outlier-Aware LLM InferenceJunqing Lin, Jingwei Sun, Zhengding Hu, Guangzhong Sun2026-08-16下载Low-bit quantization reduces the memory footprint and computational cost of large language model (LLM) inference. However, high-magnitude outlier weights can induce substantial quantization errors and...
Q-First: Attention and Feed-Forward Concurrency at the Smallest Change to the BlockWenJie Fan2026-08-16下载Disaggregated LLM serving puts the KV-cache sweep on memory-optimised hardware and the projections and feed-forward on compute-optimised hardware, then inherits from the decoder block a dependency nei...
eAVID: Asynchronous Verifiable Information Dispersal with Post-Dissemination PruningRithwik Kerur, Dahlia Malkhi, Michael K. Reiter2026-08-16下载Asynchronous verifiable information dispersal (AVID) lets a sender spread a message across N=3F+1N=3F+1 nodes such that it remains recoverable despite up to FF Byzantine failures.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Scaling the Lightning Network with Practical Set ReconciliationXingyu Chen, Anish Sinha, David Starobinski, Ari Trachtenberg2026-08-16下载The Lightning Network (LN) utilizes gossip to share network topology, channel announcements and updates, and node announcements among its local constituents.
WiFiSpectralJam: A Large-Scale Open Wi-Fi Spectral Scan Dataset with Controlled RF JammingDania Herzalla, Govind Singh, Willian T. Lunardi, Martin Andreoni2026-08-16下载WiFiSpectralJam is a Wi-Fi spectral-scan dataset comprising 14.52 GB, 96,090 CSV files, and 522,771,130 ordered spectral observations using commodity Wi-Fi sensing hardware.
Maintaining IoT Device Identification under Concept Drift via Budget-Aware Traffic LabelingShayan Azizi, Norihiro Okui, Masataka Nakahara, Ayumu Kubota, Gustavo Batista, Hassan Habibi Gharakaheili2026-08-16下载Identification of IoT device types from passive traffic is increasingly used for security management in enterprise and ISP networks. However, the performance of machine learning-based classifiers grad...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model ServingWeinan Liu, Zeyuan Ding, Dian Ding, Chengcheng Wan, Lu Tang, Guangtao Xue, Jiwu Shu, Yiming Zhang2026-08-16下载Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited. Reallocating containers is too heavyweight to utilize such fine-grained opportunities under SLA ...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Concurrency Response of Plain Global Loads on the NVIDIA H100Somashekar Manjunath, Rahul Ramachandra M2026-08-16下载The bandwidth a memory-bound GPU kernel sustains is set by how many bytes it keeps in flight. We use Little's Law here as throughput accounting, not as a measured hardware pool.

基于 VitePress 构建 · 使用本地搜索查找论文