2026-08-16
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Concurrency Response of Plain Global Loads on the NVIDIA H100 | Somashekar Manjunath, Rahul Ramachandra M | 2026-08-16 | 下载 | The bandwidth a memory-bound GPU kernel sustains is set by how many bytes it keeps in flight. We use Little's Law here as throughput accounting, not as a measured hardware pool. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| KV-Pipe: On the Relation Between KV Sharing and Pipeline Parallel Efficiency in LLMs | Maryam Dialameh, Hossein Rajabzadeh, Harish Krishnamoorthy Murali, Walid Ahmed, Weiwei Zhang, Hyock Ju Kwon | 2026-08-16 | 下载 | Pipeline parallelism (PP) is widely used to scale large language model (LLM) training, but its efficiency is often limited by stage imbalance and pipeline bubbles. |
| Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model Serving | Weinan Liu, Zeyuan Ding, Dian Ding, Chengcheng Wan, Lu Tang, Guangtao Xue, Jiwu Shu, Yiming Zhang | 2026-08-16 | 下载 | Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited. Reallocating containers is too heavyweight to utilize such fine-grained opportunities under SLA ... |
| Adaptive Heterogeneous Compression for Resource-Efficient Federated Knowledge Distillation | Chenwang Liu, Yijun Liu, Chang Liu, Xu Zhang, Pengchao Han | 2026-08-16 | 下载 | Federated learning (FL) enables privacy-preserving distributed model training but faces challenges from heterogeneous model architectures and limited communication resources at the network edge. |
| When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation | Wenhao Yuan, Chenchen Lin, Wenhao Hu, Jian Chen, Jinfeng Xu, Shujie Li, Edith Cheuk Han Ngai | 2026-08-16 | 下载 | \textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, under client heterogeneity, the conventional static split stra... |
| DeltaLog: Deferred Materialization of Recurrent States for Linear Attention Decoding | Junqing Lin, Jingwei Sun, Guangzhong Sun | 2026-08-16 | 下载 | Linear attention models eliminate the quadratic prefix computation and context-growing KV cache of softmax attention by replacing pairwise token interactions with recurrent state updates. |
| FlashQuant: Sparse-Dense Fusion for Memory-Efficient Outlier-Aware LLM Inference | Junqing Lin, Jingwei Sun, Zhengding Hu, Guangzhong Sun | 2026-08-16 | 下载 | Low-bit quantization reduces the memory footprint and computational cost of large language model (LLM) inference. However, high-magnitude outlier weights can induce substantial quantization errors and... |
| Q-First: Attention and Feed-Forward Concurrency at the Smallest Change to the Block | WenJie Fan | 2026-08-16 | 下载 | Disaggregated LLM serving puts the KV-cache sweep on memory-optimised hardware and the projections and feed-forward on compute-optimised hardware, then inherits from the decoder block a dependency nei... |
| eAVID: Asynchronous Verifiable Information Dispersal with Post-Dissemination Pruning | Rithwik Kerur, Dahlia Malkhi, Michael K. Reiter | 2026-08-16 | 下载 | Asynchronous verifiable information dispersal (AVID) lets a sender spread a message across nodes such that it remains recoverable despite up to Byzantine failures. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Scaling the Lightning Network with Practical Set Reconciliation | Xingyu Chen, Anish Sinha, David Starobinski, Ari Trachtenberg | 2026-08-16 | 下载 | The Lightning Network (LN) utilizes gossip to share network topology, channel announcements and updates, and node announcements among its local constituents. |
| WiFiSpectralJam: A Large-Scale Open Wi-Fi Spectral Scan Dataset with Controlled RF Jamming | Dania Herzalla, Govind Singh, Willian T. Lunardi, Martin Andreoni | 2026-08-16 | 下载 | WiFiSpectralJam is a Wi-Fi spectral-scan dataset comprising 14.52 GB, 96,090 CSV files, and 522,771,130 ordered spectral observations using commodity Wi-Fi sensing hardware. |
| Maintaining IoT Device Identification under Concept Drift via Budget-Aware Traffic Labeling | Shayan Azizi, Norihiro Okui, Masataka Nakahara, Ayumu Kubota, Gustavo Batista, Hassan Habibi Gharakaheili | 2026-08-16 | 下载 | Identification of IoT device types from passive traffic is increasingly used for security management in enterprise and ISP networks. However, the performance of machine learning-based classifiers grad... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Global Simulation-Guided Dynamic Operator Scheduling for Efficient Multi-Tenant Model Serving | Weinan Liu, Zeyuan Ding, Dian Ding, Chengcheng Wan, Lu Tang, Guangtao Xue, Jiwu Shu, Yiming Zhang | 2026-08-16 | 下载 | Container-granularity scheduling leaves abundant short-lived idle slices within containers unexploited. Reallocating containers is too heavyweight to utilize such fine-grained opportunities under SLA ... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Concurrency Response of Plain Global Loads on the NVIDIA H100 | Somashekar Manjunath, Rahul Ramachandra M | 2026-08-16 | 下载 | The bandwidth a memory-bound GPU kernel sustains is set by how many bytes it keeps in flight. We use Little's Law here as throughput accounting, not as a measured hardware pool. |