2026-08-09
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| IDRAAK: From Multi-Agent NLP to Few-Shot Prompting for Semantic Drift Detection in Technical Requirements | Shiva Ahir | 2026-08-09 | 下载 | Translating technical requirements across languages can introduce semantic drift, altering numerical constraints, polarities, modalities, or other specification-critical meaning. |
| Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence | Jatin Chopra | 2026-08-09 | 下载 | In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor manufacturing and operation has reached a critical threshold. |
| C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing Systems | Jiayi Li, Di Wu, Qingxu Li, Hongxiao Zhao, Jiaqi Yang, Anjunyi Fan, Wenbin Zhang, Boqiang Wu, Shuting Liu, Shifeng Fang, Jianbo Dong, Dimin Niu, Bonan Yan | 2026-08-09 | 下载 | The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communication. |
| ReVolt: Power Delivery Network-Aware Voltage Droop Control for 2.5D PIM Chiplet Architectures | Vibhanshu Sharma, Alish Kanani, Miao Sun, Janardhan Rao Doppa, Umit Y. Ogras, Partha Pratim Pande | 2026-08-09 | 下载 | Processing-in-memory (PIM)-based 2.5D multi-chiplet platforms are enablers for machine learning (ML) workloads. However, their performance is affected by the power delivery network (PDN), where varyin... |
| ARMOR: Accelerating RTL Simulation by Mitigating the Front-End Bottleneck Using Node Compression | Jiaping Tang, Jianan Mu, Zhiteng Chao, Jingzhong Wen, Jing Ye, Huawei Li | 2026-08-09 | 下载 | RTL simulation is indispensable in chip design. High-performance simulators typically lower each node in the RTL graph into an instruction sequence. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Randomized Tree-Intersection Leader Election | Yuval Emek, Shay Kutten, Ido Rafael, Gadi Taubenfeld | 2026-08-09 | 下载 | We present a randomized leader election algorithm for synchronous complete -node graphs in the \textsf{CONGEST} model that introduces a highly tunable trade-off between time complexity and the per-... |
| Measuring and Reducing WebGPU Dispatch Overhead for LLM Inference | Jędrzej Maczan | 2026-08-09 | 下载 | Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platform standard. |
| C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing Systems | Jiayi Li, Di Wu, Qingxu Li, Hongxiao Zhao, Jiaqi Yang, Anjunyi Fan, Wenbin Zhang, Boqiang Wu, Shuting Liu, Shifeng Fang, Jianbo Dong, Dimin Niu, Bonan Yan | 2026-08-09 | 下载 | The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communication. |
| Robust Reputation-Driven Crowdsourced Federated Learning | Mouhamed Amine Bouchiha, Gregory Blanc | 2026-08-09 | 下载 | Crowdsourced Federated Learning (CrowdFL) extends traditional federated learning by enabling open and heterogeneous participation through a crowdsourcing paradigm. |
| FlashBoot: Sub-Second Weight Loading for Large Models at Rack Scale | Issac Zhu, Hscos Zhang, Keith Jiang, Jack Li, Hugh Yin, Jason Zhao | 2026-08-09 | 下载 | Flagship Mixture-of-Experts (MoE) models are growing fast along two axes at once: total parameter count and the number of experts. In elastic deployment scenarios, many GPUs across many nodes must bec... |
| PSP: Low-Overhead Packet-Level Load Balancing for Stale-State and Bandwidth-Asymmetric Networks | Jiaqi Liu, Chunyang Zhang, Heng Pan, Yanbiao Li | 2026-08-09 | 下载 | With the rapid growth of large language model training and generative artificial intelligence services, data center networks face severe micro-burst traffic and high concurrency. |
| Empirical Analysis of Cloud-Edge Infrastructure Complexity: Practitioner Pain Points and Architectural Directions | Pawissanutt Lertpongrujikorn, Hai Duc Nguyen, Mohsen Amini Salehi | 2026-08-09 | 下载 | The proliferation of cloud, edge, and Internet of Things (IoT) computing has created unprecedented opportunities for distributed applications. |
| LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving | Shuowei Jin, Xueshen Liu, Jiaxin Shan, Le Xu, Tieying Zhang, Liguang Xie, Z. Morley Mao | 2026-08-09 | 下载 | As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractional sharing of the inference engine requires a r... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Carrier Leakage Suppression and Power Difference Equalization for TFDMA Coherent PON | Ji Zhou, Haide Wang, Liangchuan Li, Changyuan Yu, Xiangjun Xin | 2026-08-09 | 下载 | In this invited talk, we demonstrate the carrier leakage suppression and power difference equalization based on semiconductor optical amplifier (SOA) for TFDMA-based coherent PON, and propose a nonlin... |
| PSP: Low-Overhead Packet-Level Load Balancing for Stale-State and Bandwidth-Asymmetric Networks | Jiaqi Liu, Chunyang Zhang, Heng Pan, Yanbiao Li | 2026-08-09 | 下载 | With the rapid growth of large language model training and generative artificial intelligence services, data center networks face severe micro-burst traffic and high concurrency. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference | Asaad Althoubi | 2026-08-09 | 下载 | Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequence length, creating a severe memory bottleneck for... |
| Measuring and Reducing WebGPU Dispatch Overhead for LLM Inference | Jędrzej Maczan | 2026-08-09 | 下载 | Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platform standard. |
| Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs | Mohanad Odema, Gabrielle De Micheli, Dayin Gou, Nilesh Malpeddi, Prathamesh Vaste, Jacob Song | 2026-08-09 | 下载 | Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy. |