Skip to content

2026-08-09 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
IDRAAK: From Multi-Agent NLP to Few-Shot Prompting for Semantic Drift Detection in Technical RequirementsShiva Ahir2026-08-09下载Translating technical requirements across languages can introduce semantic drift, altering numerical constraints, polarities, modalities, or other specification-critical meaning.
Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial IntelligenceJatin Chopra2026-08-09下载In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor manufacturing and operation has reached a critical threshold.
C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing SystemsJiayi Li, Di Wu, Qingxu Li, Hongxiao Zhao, Jiaqi Yang, Anjunyi Fan, Wenbin Zhang, Boqiang Wu, Shuting Liu, Shifeng Fang, Jianbo Dong, Dimin Niu, Bonan Yan2026-08-09下载The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communication.
ReVolt: Power Delivery Network-Aware Voltage Droop Control for 2.5D PIM Chiplet ArchitecturesVibhanshu Sharma, Alish Kanani, Miao Sun, Janardhan Rao Doppa, Umit Y. Ogras, Partha Pratim Pande2026-08-09下载Processing-in-memory (PIM)-based 2.5D multi-chiplet platforms are enablers for machine learning (ML) workloads. However, their performance is affected by the power delivery network (PDN), where varyin...
ARMOR: Accelerating RTL Simulation by Mitigating the Front-End Bottleneck Using Node CompressionJiaping Tang, Jianan Mu, Zhiteng Chao, Jingzhong Wen, Jing Ye, Huawei Li2026-08-09下载RTL simulation is indispensable in chip design. High-performance simulators typically lower each node in the RTL graph into an instruction sequence.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Randomized Tree-Intersection Leader ElectionYuval Emek, Shay Kutten, Ido Rafael, Gadi Taubenfeld2026-08-09下载We present a randomized leader election algorithm for synchronous complete nn-node graphs in the \textsf{CONGEST} model that introduces a highly tunable trade-off between time complexity and the per-...
Measuring and Reducing WebGPU Dispatch Overhead for LLM InferenceJędrzej Maczan2026-08-09下载Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platform standard.
C2C-Explorer: An Exploration Framework for Chip-to-Chip Interconnect Architectures in LLM Cloud Computing SystemsJiayi Li, Di Wu, Qingxu Li, Hongxiao Zhao, Jiaqi Yang, Anjunyi Fan, Wenbin Zhang, Boqiang Wu, Shuting Liu, Shifeng Fang, Jianbo Dong, Dimin Niu, Bonan Yan2026-08-09下载The scaling-up of large language models (LLMs) necessitates computing systems to have multi-processor-chip architectures, elevating the importance of chip-to-chip (C2C) communication.
Robust Reputation-Driven Crowdsourced Federated LearningMouhamed Amine Bouchiha, Gregory Blanc2026-08-09下载Crowdsourced Federated Learning (CrowdFL) extends traditional federated learning by enabling open and heterogeneous participation through a crowdsourcing paradigm.
FlashBoot: Sub-Second Weight Loading for Large Models at Rack ScaleIssac Zhu, Hscos Zhang, Keith Jiang, Jack Li, Hugh Yin, Jason Zhao2026-08-09下载Flagship Mixture-of-Experts (MoE) models are growing fast along two axes at once: total parameter count and the number of experts. In elastic deployment scenarios, many GPUs across many nodes must bec...
PSP: Low-Overhead Packet-Level Load Balancing for Stale-State and Bandwidth-Asymmetric NetworksJiaqi Liu, Chunyang Zhang, Heng Pan, Yanbiao Li2026-08-09下载With the rapid growth of large language model training and generative artificial intelligence services, data center networks face severe micro-burst traffic and high concurrency.
Empirical Analysis of Cloud-Edge Infrastructure Complexity: Practitioner Pain Points and Architectural DirectionsPawissanutt Lertpongrujikorn, Hai Duc Nguyen, Mohsen Amini Salehi2026-08-09下载The proliferation of cloud, edge, and Internet of Things (IoT) computing has created unprecedented opportunities for distributed applications.
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM ServingShuowei Jin, Xueshen Liu, Jiaxin Shan, Le Xu, Tieying Zhang, Liguang Xie, Z. Morley Mao2026-08-09下载As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractional sharing of the inference engine requires a r...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Carrier Leakage Suppression and Power Difference Equalization for TFDMA Coherent PONJi Zhou, Haide Wang, Liangchuan Li, Changyuan Yu, Xiangjun Xin2026-08-09下载In this invited talk, we demonstrate the carrier leakage suppression and power difference equalization based on semiconductor optical amplifier (SOA) for TFDMA-based coherent PON, and propose a nonlin...
PSP: Low-Overhead Packet-Level Load Balancing for Stale-State and Bandwidth-Asymmetric NetworksJiaqi Liu, Chunyang Zhang, Heng Pan, Yanbiao Li2026-08-09下载With the rapid growth of large language model training and generative artificial intelligence services, data center networks face severe micro-burst traffic and high concurrency.

cs.PF - Performance ​

标题作者发布日期PDF摘要
DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM InferenceAsaad Althoubi2026-08-09下载Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequence length, creating a severe memory bottleneck for...
Measuring and Reducing WebGPU Dispatch Overhead for LLM InferenceJędrzej Maczan2026-08-09下载Large Language Models are deployed to multiple types of environments, from internet browsers to edge devices, and WebGPU serves as a modern cross-platform standard.
Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMsMohanad Odema, Gabrielle De Micheli, Dayin Gou, Nilesh Malpeddi, Prathamesh Vaste, Jacob Song2026-08-09下载Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model parameter count while maintaining task-level accuracy.

基于 VitePress 构建 · 使用本地搜索查找论文