Skip to content

2026-08-01 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Rethinking Agentic Kernel Generation for Emerging AcceleratorsRuijie Gao, Jirong Yang, Barry Lyu, Haoran Jin, Nathan Bleier2026-08-01下载Emerging accelerators often lack mature compiler backends, motivating neural agents that generate and repair kernels from architectural documentation and simulator feedback.
NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up SystemsConor James Green, William Won, Tuan Ta, Bradford M. Beckmann2026-08-01下载Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly...
CascadeLUT: Information-Ordered Streaming Inference for Bandwidth-Constrained FPGAsOliver Cassidy, Marta Andronic, George A. Constantinides2026-08-01下载Mapping neural networks to FPGAs enables low-latency, energy-efficient inference, particularly for lookup table (LUT)-based models that eliminate multipliers and map directly to reconfigurable fabric.
A Time-Multiplexed Spiking Neural Network Accelerator with Pipelined Readout for FPGA InferenceReza Ansari, Maciej Wielgosz2026-08-01下载Spiking Neural Networks (SNNs) provide a power-efficient neuromorphic alternative to traditional artificial neural networks by processing information through discrete temporal events.
A Journey in Shared Memory LandRan Ginosar2026-08-01下载I have greatly enjoyed spending many years in studying parallel computing. My journey goes thorough MP-C, PLURAL, Async Plural, HAL, RC64 and more.
C2P-Cache: Scalable GPU L1 Cache Sharing via Concurrent Candidate PruningHanqing Li, Lizhou Wu, Tiejun Li, Sheng Ma, Hanzhi Xun, Jianmin Zhang, Yuhan Tang, Jixuan Tang, Xuchao Xie2026-08-01下载Modern GPUs rely on private per-SM L1 caches and a shared L2 cache, but this organization obscures cross-SM reuse: an L1 miss is typically forwarded to L2 even when the requested line already resides ...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Shiftfly: Scaling the Accelerator Interconnect Past the Pod with a Shift-Routed Optical TierEylon E. Krause2026-08-01下载Google's TPU interconnect spent nine generations as a kk-ary nn-cube, whose diameter grows as Θ(N^{1/n}), before TPU 8i replaced it with Boardfly: a three-tier hierarchy in which every tier is a c...
Execution Timing Control for Deterministic Task Offloading in the IoT-Edge-Cloud ContinuumKeyvan Aghababaiyan, Baldomero Coll-Perales, Javier Gozalvez2026-08-01下载Latency-critical IoT applications, such as autonomous mobility and industrial automation, require deterministic guarantees to ensure that tasks are completed within strict deadlines.
NUNA: Characterizing and Mitigating Non-Uniform Network Access in Multi-Die GPU Scale-Up SystemsConor James Green, William Won, Tuan Ta, Bradford M. Beckmann2026-08-01下载Graphics processing unit (GPU) architectures are growing in size to meet the increasing compute and memory requirements. As GPU sizes increase, intra-socket wire transfer delay increases significantly...
Multi-tenant Kubernetes Use Cases for AI, Secure Computing and Data Services, and MoreJake Watson, Sadaf R Alam, Christopher Woods, Abdelwahab Kawafi, Thomas Green, Ian Johnson, Ellis Pires, Jessica R. Jones, Utz-Uwe Haus2026-08-01下载Kubernetes, as a container orchestration engine, has been widely used in cloud-native ecosystems for several years. In supercomputing ecosystems, especially where bare-metal performance for compute an...
Machine-Checked Dual-Write Recovery from a Committed LogAndreas Andreakis2026-08-01下载After a crash, a delivery process faces a question its own database cannot answer: did the other side already receive the effect? Transactional outboxes and change data capture remove the application'...
Error-bounded Point Cloud Compression Using Truncated Octahedron QuantizationYouyuan Liu, Longtao Zhang, Ruoyu Li, Bo Jiang, Taolue Yang, Kai Zhao, Sheng Di, Eduard Dragut, Sian Jin2026-08-01下载With the rapid advancement of large-scale scientific simulations, the massive volume of point cloud data generated has increasingly become a critical bottleneck for scientific storage systems and data...
Collaborative Orbital Edge Intelligence: A Decentralized Paradigm for Energy-Efficient Computing in SpaceYuvraj Sahni, Jiannong Cao, Fu Xiao2026-08-01下载In recent years, Low Earth Orbit (LEO) satellites have been increasingly deployed to enable connectivity in remote and disaster-prone areas. Researchers have proposed Orbital Edge Computing, which add...
AReaL-DTE: Sparse Policy-Weight Transfer for Online Agentic Reinforcement LearningYingqi Peng, Jiawei Zhang, Wenhao Zhou, Ruida Xu, Ran Yan, Wei Dong, Yi Gao, Zhiqiang Ding, Tongkai Yang, Binhang Yuan2026-08-01下载Online agentic reinforcement learning implemented with micro-services separates policy training from rollout generation, improving scalability and modularity while potentially making frequent policy-w...
Cache-Consistent Dynamic Load Balancing for Kubernetes ControllersShoma Ansai, Yasuo Okabe, Daisuke Kotani2026-08-01下载As Kubernetes clusters grow, the scalability of controllers can become a bottleneck for the performance of the system. Distributing the load dynamically across multiple controller instances, however, ...
Design and Implementation of Schwarz Information Criterion-Aided Intelligent Decentralized Resource Allocation in Dynamic LoRa NetworksAohan Li, Ryota Ariyoshi, Mikio Hasegawa, Miao Pan, Tomoaki Ohtsuki, Zhu Han2026-08-01下载This paper proposes a lightweight distributed learning method for selecting transmission parameters in Long-Range (LoRa) networks that adapts to dynamically changing communication environments.
HCCL: Collective Communication for Meta Training and Inference AcceleratorsWesley Bland, Tiago Antunes, Lars Paul Huse, Chidambaram Muthu, Adel Abouchaev, Rabib Alam, Abdullah Alperen, Alexey Andronov, Jose Anto Akkara, Vineet Badhwar, Pavan Balaji, Daniel Berkovitch, Bartosz Bogdanski, Shmeelok Chakraborty, Sungjun Cho, John Choi, James Custer, Rodrigo De Castro, Nguyen Dinh Pham, Matthew Edwards, Kristian Evensen, Evan Ezell, Alex Finestead, Seth Goldstein, Prankur Gupta, Ranwei Hu, Adam Incera, Anand Jayaraman, Prashanth Kannan, Soumil Kanwal, Martin Karp, Sameer Kumar, Naina Kuruballi Mahesh, Wei Lin Guay, Cristian Lumezanu, Cory Modlin, Dag Georg Moxnes, Hoang Nam Nguyen, Ashay Narsale, Jaden Padua, Kirtesh Patil, Minh Pham, Amin Qassoud, Ashwin Ramachandran, David Ramon Prados, Pallavi Shurpali, Gregory R. Steinbrecher, John Sundharam, Vangelis Tasoulas, Fuhou Tian, Srinivas Vaidyanathan, Vimal Vasudevan, Nicolaas Viljoen, Daniel Winkelman, Yijing Zeng, Zhaoqi Zhu, Stig Arne Olsen, Gilad Goldfarb, Rajiv Krishnamurthy, Rajeev Nair, Jonas Olsson, Joseph Provine, Sreeram Ravinoothala, Shivayogi Ugaji, Hongyi Zeng, Nairan Zhang2026-08-01下载We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Shiftfly: Scaling the Accelerator Interconnect Past the Pod with a Shift-Routed Optical TierEylon E. Krause2026-08-01下载Google's TPU interconnect spent nine generations as a kk-ary nn-cube, whose diameter grows as Θ(N^{1/n}), before TPU 8i replaced it with Boardfly: a three-tier hierarchy in which every tier is a c...
Execution Timing Control for Deterministic Task Offloading in the IoT-Edge-Cloud ContinuumKeyvan Aghababaiyan, Baldomero Coll-Perales, Javier Gozalvez2026-08-01下载Latency-critical IoT applications, such as autonomous mobility and industrial automation, require deterministic guarantees to ensure that tasks are completed within strict deadlines.
LLM-Assisted Coalition Formation for Cooperative Perception in Autonomous DrivingAhmad Sarlak, Hao Wang, Rahul Amin, Abolfazl Razi2026-08-01下载Cooperative perception (CP) enables connected autonomous vehicles (CAVs) to share complementary observations for safer navigation, but practical deployment is limited by bandwidth constraints, unrelia...
Domain Decoupling Attack: Exploiting the Validation Gap Between Protective DNS and Shared Edge RoutingWeizhe Wang, Minhong Dong, Jinhao Li, Yao Zhang, Hao Liu, Qiang Hu, Tao Luo, Guangquan Xu, Bin Wu2026-08-01下载Network attackers often conceal malicious communication within legitimate Internet traffic. Existing CDN-based evasion techniques rely on SNI--Host inconsistency, insufficient domain ownership verific...
HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE InferenceXin Yuan, Ning Li, Wenchao Xu, Athanasios V. Vasilakos, Song Guo, Haijun Zhang2026-08-01下载Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challenging.
TrimMoE A communication aware and adaptive depth framework for distributed edge inferenceNing Li, Shuting Bai, Xin Yuan, Wenchao Xu, Athanasios V. Vasilakos, Song Guo, Haijun Zhang2026-08-01下载Serving Mixture-of-Experts (MoE) large language models across distributed edge servers is bottlenecked by the cross-server expert transmission.
Channel-Agnostic Semantic Compression for Bandwidth-Limited Visual CommunicationXuanhao Luo, Ruichen Gao, Zhizhen Li, Mingzhe Chen, Yuchen Liu2026-08-01下载Bandwidth-limited visual communication systems require efficient transmission of high-dimensional data under dynamic wireless conditions. Existing approaches either rely on joint source-channel coding...
HCCL: Collective Communication for Meta Training and Inference AcceleratorsWesley Bland, Tiago Antunes, Lars Paul Huse, Chidambaram Muthu, Adel Abouchaev, Rabib Alam, Abdullah Alperen, Alexey Andronov, Jose Anto Akkara, Vineet Badhwar, Pavan Balaji, Daniel Berkovitch, Bartosz Bogdanski, Shmeelok Chakraborty, Sungjun Cho, John Choi, James Custer, Rodrigo De Castro, Nguyen Dinh Pham, Matthew Edwards, Kristian Evensen, Evan Ezell, Alex Finestead, Seth Goldstein, Prankur Gupta, Ranwei Hu, Adam Incera, Anand Jayaraman, Prashanth Kannan, Soumil Kanwal, Martin Karp, Sameer Kumar, Naina Kuruballi Mahesh, Wei Lin Guay, Cristian Lumezanu, Cory Modlin, Dag Georg Moxnes, Hoang Nam Nguyen, Ashay Narsale, Jaden Padua, Kirtesh Patil, Minh Pham, Amin Qassoud, Ashwin Ramachandran, David Ramon Prados, Pallavi Shurpali, Gregory R. Steinbrecher, John Sundharam, Vangelis Tasoulas, Fuhou Tian, Srinivas Vaidyanathan, Vimal Vasudevan, Nicolaas Viljoen, Daniel Winkelman, Yijing Zeng, Zhaoqi Zhu, Stig Arne Olsen, Gilad Goldfarb, Rajiv Krishnamurthy, Rajeev Nair, Jonas Olsson, Joseph Provine, Sreeram Ravinoothala, Shivayogi Ugaji, Hongyi Zeng, Nairan Zhang2026-08-01下载We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Less Is More: Tuning Configurable Systems with Imperfect FidelityYulong Ye, Miqing Li, Tao Chen2026-08-01下载Configuration tuning is essential for optimizing the performance of highly configurable systems, e.g., throughput or runtime, under a given environment.
Multi-tenant Kubernetes Use Cases for AI, Secure Computing and Data Services, and MoreJake Watson, Sadaf R Alam, Christopher Woods, Abdelwahab Kawafi, Thomas Green, Ian Johnson, Ellis Pires, Jessica R. Jones, Utz-Uwe Haus2026-08-01下载Kubernetes, as a container orchestration engine, has been widely used in cloud-native ecosystems for several years. In supercomputing ecosystems, especially where bare-metal performance for compute an...
An Embedded RISC-V Evaluation of Kolmogorov--Arnold Networks in Hard-Constrained Recurrent Physics-Informed ModelsEnzo Nicolas Spotorno, Josafat Leal Filho2026-08-01下载Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the...
Error-bounded Point Cloud Compression Using Truncated Octahedron QuantizationYouyuan Liu, Longtao Zhang, Ruoyu Li, Bo Jiang, Taolue Yang, Kai Zhao, Sheng Di, Eduard Dragut, Sian Jin2026-08-01下载With the rapid advancement of large-scale scientific simulations, the massive volume of point cloud data generated has increasingly become a critical bottleneck for scientific storage systems and data...

基于 VitePress 构建 · 使用本地搜索查找论文