2026-08-01
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Rethinking Agentic Kernel Generation for Emerging Accelerators | Ruijie Gao, Jirong Yang, Barry Lyu, Haoran Jin, Nathan Bleier | 2026-08-01 | 下载 | Emerging accelerators often lack mature compiler backends, motivating neural agents that generate and repair kernels from architectural documentation and simulator feedback. |
| NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems | Conor James Green, William Won, Tuan Ta, Bradford M. Beckmann | 2026-08-01 | 下载 | Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly... |
| CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAs | Oliver Cassidy, Marta Andronic, George A. Constantinides | 2026-08-01 | 下载 | Mapping neural networks to FPGAs enables low-latency, energy-efficient inference, particularly for lookup table (LUT)-based models that eliminate multipliers and map directly to reconfigurable fabric. |
| A Time-Multiplexed Spiking Neural Network Accelerator with Pipelined Readout for FPGA Inference | Reza Ansari, Maciej Wielgosz | 2026-08-01 | 下载 | Spiking Neural Networks (SNNs) provide a power-efficient neuromorphic alternative to traditional artificial neural networks by processing information through discrete temporal events. |
| A Journey in Shared Memory Land | Ran Ginosar | 2026-08-01 | 下载 | I have greatly enjoyed spending many years in studying parallel computing. My journey goes thorough MP-C, PLURAL, Async Plural, HAL, RC64 and more. |
| C2P-Cache: Scalable GPU L1 Cache Sharing via Concurrent Candidate Pruning | Hanqing Li, Lizhou Wu, Tiejun Li, Sheng Ma, Hanzhi Xun, Jianmin Zhang, Yuhan Tang, Jixuan Tang, Xuchao Xie | 2026-08-01 | 下载 | Modern GPUs rely on private per-SM L1 caches and a shared L2 cache, but this organization obscures cross-SM reuse: an L1 miss is typically forwarded to L2 even when the requested line already resides ... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Shiftfly: Scaling the Accelerator Interconnect Past the Pod with a Shift-Routed Optical Tier | Eylon E. Krause | 2026-08-01 | 下载 | Google's TPU interconnect spent nine generations as a -ary -cube, whose diameter grows as Θ(N^{1/n}), before TPU 8i replaced it with Boardfly: a three-tier hierarchy in which every tier is a c... |
| Execution Timing Control for Deterministic Task Offloading in the IoT-Edge-Cloud Continuum | Keyvan Aghababaiyan, Baldomero Coll-Perales, Javier Gozalvez | 2026-08-01 | 下载 | Latency-critical IoT applications, such as autonomous mobility and industrial automation, require deterministic guarantees to ensure that tasks are completed within strict deadlines. |
| NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up Systems | Conor James Green, William Won, Tuan Ta, Bradford M. Beckmann | 2026-08-01 | 下载 | Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly... |
| Multi-tenant Kubernetes Use Cases for AI, Secure Computing and Data Services, and More | Jake Watson, Sadaf R Alam, Christopher Woods, Abdelwahab Kawafi, Thomas Green, Ian Johnson, Ellis Pires, Jessica R. Jones, Utz-Uwe Haus | 2026-08-01 | 下载 | Kubernetes, as a container orchestration engine, has been widely used in cloud-native ecosystems for several years. In supercomputing ecosystems, especially where bare-metal performance for compute an... |
| Machine-Checked Dual-Write Recovery from a Committed Log | Andreas Andreakis | 2026-08-01 | 下载 | After a crash, a delivery process faces a question its own database cannot answer: did the other side already receive the effect? Transactional outboxes and change data capture remove the application'... |
| Error-bounded Point Cloud Compression Using Truncated Octahedron Quantization | Youyuan Liu, Longtao Zhang, Ruoyu Li, Bo Jiang, Taolue Yang, Kai Zhao, Sheng Di, Eduard Dragut, Sian Jin | 2026-08-01 | 下载 | With the rapid advancement of large-scale scientific simulations, the massive volume of point cloud data generated has increasingly become a critical bottleneck for scientific storage systems and data... |
| Collaborative Orbital Edge Intelligence: A Decentralized Paradigm for Energy-Efficient Computing in Space | Yuvraj Sahni, Jiannong Cao, Fu Xiao | 2026-08-01 | 下载 | In recent years, Low Earth Orbit (LEO) satellites have been increasingly deployed to enable connectivity in remote and disaster-prone areas. Researchers have proposed Orbital Edge Computing, which add... |
| AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement Learning | Yingqi Peng, Jiawei Zhang, Wenhao Zhou, Ruida Xu, Ran Yan, Wei Dong, Yi Gao, Zhiqiang Ding, Tongkai Yang, Binhang Yuan | 2026-08-01 | 下载 | Online agentic reinforcement learning implemented with micro-services separates policy training from rollout generation, improving scalability and modularity while potentially making frequent policy-w... |
| Cache-Consistent Dynamic Load Balancing for Kubernetes Controllers | Shoma Ansai, Yasuo Okabe, Daisuke Kotani | 2026-08-01 | 下载 | As Kubernetes clusters grow, the scalability of controllers can become a bottleneck for the performance of the system. Distributing the load dynamically across multiple controller instances, however, ... |
| Design and Implementation of Schwarz Information Criterion-Aided Intelligent Decentralized Resource Allocation in Dynamic LoRa Networks | Aohan Li, Ryota Ariyoshi, Mikio Hasegawa, Miao Pan, Tomoaki Ohtsuki, Zhu Han | 2026-08-01 | 下载 | This paper proposes a lightweight distributed learning method for selecting transmission parameters in Long-Range (LoRa) networks that adapts to dynamically changing communication environments. |
| HCCL: Collective Communication for Meta Training and Inference Accelerators | Wesley Bland, Tiago Antunes, Lars Paul Huse, Chidambaram Muthu, Adel Abouchaev, Rabib Alam, Abdullah Alperen, Alexey Andronov, Jose Anto Akkara, Vineet Badhwar, Pavan Balaji, Daniel Berkovitch, Bartosz Bogdanski, Shmeelok Chakraborty, Sungjun Cho, John Choi, James Custer, Rodrigo De Castro, Nguyen Dinh Pham, Matthew Edwards, Kristian Evensen, Evan Ezell, Alex Finestead, Seth Goldstein, Prankur Gupta, Ranwei Hu, Adam Incera, Anand Jayaraman, Prashanth Kannan, Soumil Kanwal, Martin Karp, Sameer Kumar, Naina Kuruballi Mahesh, Wei Lin Guay, Cristian Lumezanu, Cory Modlin, Dag Georg Moxnes, Hoang Nam Nguyen, Ashay Narsale, Jaden Padua, Kirtesh Patil, Minh Pham, Amin Qassoud, Ashwin Ramachandran, David Ramon Prados, Pallavi Shurpali, Gregory R. Steinbrecher, John Sundharam, Vangelis Tasoulas, Fuhou Tian, Srinivas Vaidyanathan, Vimal Vasudevan, Nicolaas Viljoen, Daniel Winkelman, Yijing Zeng, Zhaoqi Zhu, Stig Arne Olsen, Gilad Goldfarb, Rajiv Krishnamurthy, Rajeev Nair, Jonas Olsson, Joseph Provine, Sreeram Ravinoothala, Shivayogi Ugaji, Hongyi Zeng, Nairan Zhang | 2026-08-01 | 下载 | We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Shiftfly: Scaling the Accelerator Interconnect Past the Pod with a Shift-Routed Optical Tier | Eylon E. Krause | 2026-08-01 | 下载 | Google's TPU interconnect spent nine generations as a -ary -cube, whose diameter grows as Θ(N^{1/n}), before TPU 8i replaced it with Boardfly: a three-tier hierarchy in which every tier is a c... |
| Execution Timing Control for Deterministic Task Offloading in the IoT-Edge-Cloud Continuum | Keyvan Aghababaiyan, Baldomero Coll-Perales, Javier Gozalvez | 2026-08-01 | 下载 | Latency-critical IoT applications, such as autonomous mobility and industrial automation, require deterministic guarantees to ensure that tasks are completed within strict deadlines. |
| LLM-Assisted Coalition Formation for Cooperative Perception in Autonomous Driving | Ahmad Sarlak, Hao Wang, Rahul Amin, Abolfazl Razi | 2026-08-01 | 下载 | Cooperative perception (CP) enables connected autonomous vehicles (CAVs) to share complementary observations for safer navigation, but practical deployment is limited by bandwidth constraints, unrelia... |
| Domain Decoupling Attack: Exploiting the Validation Gap Between Protective DNS and Shared Edge Routing | Weizhe Wang, Minhong Dong, Jinhao Li, Yao Zhang, Hao Liu, Qiang Hu, Tao Luo, Guangquan Xu, Bin Wu | 2026-08-01 | 下载 | Network attackers often conceal malicious communication within legitimate Internet traffic. Existing CDN-based evasion techniques rely on SNI--Host inconsistency, insufficient domain ownership verific... |
| HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference | Xin Yuan, Ning Li, Wenchao Xu, Athanasios V. Vasilakos, Song Guo, Haijun Zhang | 2026-08-01 | 下载 | Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challenging. |
| TrimMoE A communication aware and adaptive depth framework for distributed edge inference | Ning Li, Shuting Bai, Xin Yuan, Wenchao Xu, Athanasios V. Vasilakos, Song Guo, Haijun Zhang | 2026-08-01 | 下载 | Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission. |
| Channel-Agnostic Semantic Compression for Bandwidth-Limited Visual Communication | Xuanhao Luo, Ruichen Gao, Zhizhen Li, Mingzhe Chen, Yuchen Liu | 2026-08-01 | 下载 | Bandwidth-limited visual communication systems require efficient transmission of high-dimensional data under dynamic wireless conditions. Existing approaches either rely on joint source-channel coding... |
| HCCL: Collective Communication for Meta Training and Inference Accelerators | Wesley Bland, Tiago Antunes, Lars Paul Huse, Chidambaram Muthu, Adel Abouchaev, Rabib Alam, Abdullah Alperen, Alexey Andronov, Jose Anto Akkara, Vineet Badhwar, Pavan Balaji, Daniel Berkovitch, Bartosz Bogdanski, Shmeelok Chakraborty, Sungjun Cho, John Choi, James Custer, Rodrigo De Castro, Nguyen Dinh Pham, Matthew Edwards, Kristian Evensen, Evan Ezell, Alex Finestead, Seth Goldstein, Prankur Gupta, Ranwei Hu, Adam Incera, Anand Jayaraman, Prashanth Kannan, Soumil Kanwal, Martin Karp, Sameer Kumar, Naina Kuruballi Mahesh, Wei Lin Guay, Cristian Lumezanu, Cory Modlin, Dag Georg Moxnes, Hoang Nam Nguyen, Ashay Narsale, Jaden Padua, Kirtesh Patil, Minh Pham, Amin Qassoud, Ashwin Ramachandran, David Ramon Prados, Pallavi Shurpali, Gregory R. Steinbrecher, John Sundharam, Vangelis Tasoulas, Fuhou Tian, Srinivas Vaidyanathan, Vimal Vasudevan, Nicolaas Viljoen, Daniel Winkelman, Yijing Zeng, Zhaoqi Zhu, Stig Arne Olsen, Gilad Goldfarb, Rajiv Krishnamurthy, Rajeev Nair, Jonas Olsson, Joseph Provine, Sreeram Ravinoothala, Shivayogi Ugaji, Hongyi Zeng, Nairan Zhang | 2026-08-01 | 下载 | We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Less Is More: Tuning Configurable Systems with Imperfect Fidelity | Yulong Ye, Miqing Li, Tao Chen | 2026-08-01 | 下载 | Configuration tuning is essential for optimizing the performance of highly configurable systems, e.g., throughput or runtime, under a given environment. |
| Multi-tenant Kubernetes Use Cases for AI, Secure Computing and Data Services, and More | Jake Watson, Sadaf R Alam, Christopher Woods, Abdelwahab Kawafi, Thomas Green, Ian Johnson, Ellis Pires, Jessica R. Jones, Utz-Uwe Haus | 2026-08-01 | 下载 | Kubernetes, as a container orchestration engine, has been widely used in cloud-native ecosystems for several years. In supercomputing ecosystems, especially where bare-metal performance for compute an... |
| An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed Models | Enzo Nicolas Spotorno, Josafat Leal Filho | 2026-08-01 | 下载 | Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the... |
| Error-bounded Point Cloud Compression Using Truncated Octahedron Quantization | Youyuan Liu, Longtao Zhang, Ruoyu Li, Bo Jiang, Taolue Yang, Kai Zhao, Sheng Di, Eduard Dragut, Sian Jin | 2026-08-01 | 下载 | With the rapid advancement of large-scale scientific simulations, the massive volume of point cloud data generated has increasingly become a critical bottleneck for scientific storage systems and data... |