2026-06-03
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| SET: Stream-Event-Triggered Scheduling for Efficient CUDA Graph Pipelines | Zhengxiong Li, Tsung-Wei Huang, Umit Ogras | 2026-06-03 | 下载 | Achieving peak GPU performance remains a significant challenge as the system throughput is constrained by host-device synchronization delays and kernel scheduling overheads, even with aggressive kerne... |
| MOSAIC: A Workload-Driven Simulation and Design-Space Exploration Framework for Heterogeneous NPUs | Arghadip Das, Hoseok Kim, Soomin Lee, Arnab Raha, Deepak A Mathaikutty, Vijay Raghunathan | 2026-06-03 | 下载 | AI model architectures are diversifying rapidly. Although dense matrix multiplication underlies today's CNNs and transformers, emerging architectures (state-space models, long convolutions via the fas... |
| BIDENT: Heterogeneous Operator-level Mapping for Efficient Edge Inference | Hoseok Kim, Arghadip Das, Soumendu Ghosh, Arnab Raha, Vijay Raghunathan | 2026-06-03 | 下载 | Modern edge System-on-Chips (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, yet current inference stacks map entire models to a single PU, leaving significant perfo... |
| GoldenFloat: A Phi-Derived Static-Split Floating-Point Family from GF4 to GF256 with a Lucas-Exact Integer Identity | Dmitrii Vasiliev | 2026-06-03 | 下载 | We present a hardware-oriented description of GoldenFloat (GF), a static-split floating-point family generated by a single closed rule, and three concrete artefacts: (i) an open multi-width RTL genera... |
| Uncertainty-Aware End-to-End Co-Design of Neural Network Processors: From Training and Mapping to Fabrication | Yuyang Du, Yujun Huang, Gioele Zardini | 2026-06-03 | 下载 | Designing a neural network processor is an end-to-end co-design problem: network architecture and training budget determine the inference workload; hardware mapping decisions determine chip area, late... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Latent Reasoning Guidance for Parallel Code Translation | Tomer Bitan, Erel Kaplan, Roee Bar-Yadin, Lian Ghrayeb, Le Chen, Samyak Jhaveri, Niranjan Hasabnis, Gal Oren | 2026-06-03 | 下载 | Tackling complex coding tasks often requires autonomous agents and iterative repair pipelines. These increasingly rely on large amounts of test-time computation, often spending many decoding and repai... |
| Bitcoin After Block Rewards | Junhyuk Lee | 2026-06-03 | 下载 | Bitcoin's block reward is scheduled to decline to zero, raising concerns about whether the network can remain secure once miners rely solely on transaction fees. |
| SET: Stream-Event-Triggered Scheduling for Efficient CUDA Graph Pipelines | Zhengxiong Li, Tsung-Wei Huang, Umit Ogras | 2026-06-03 | 下载 | Achieving peak GPU performance remains a significant challenge as the system throughput is constrained by host-device synchronization delays and kernel scheduling overheads, even with aggressive kerne... |
| Graph Traversal on Tensor Cores: A BFS Framework for Modern GPUs | Deniz Elbek, Kamer Kaya | 2026-06-03 | 下载 | Modern GPUs have Tensor Cores (TCs) capable of extremely high-throughput matrix operations, yet graph algorithms remain difficult to accelerate because of their irregular and data-dependent execution ... |
| The local complexity of certifying parity | Nicolas Bousquet, Laurent Feuilloley, Jorge Valenzuela, Sébastien Zeitoun | 2026-06-03 | 下载 | In this paper, we consider the problem of locally certifying that the size of a network is even, or more generally, congruent to some fixed number. |
| The Usefulness Gap in Proof-of-Useful-Work: An Empirical Study of Pearl's cuPOW Protocol | Abhinaba Basu | 2026-06-03 | 下载 | Pearl, a Layer-1 blockchain with high-profile AI industry endorsements, markets its Proof-of-Useful-Work (PoUW) protocol as simultaneously securing the network and performing AI inference. |
| Clownfish: Scaling DAG-based BFT Consensus via Sparse Edges | Feifan Wang, Jingfan Yu, Zixi Cai, Zhixuan Fang | 2026-06-03 | 下载 | Directed Acyclic Graph (DAG) based BFT protocols have demonstrated the capability to achieve significantly high throughput in practice. Recent advancements focused on minimizing the good-case latency ... |
| Rectangular Matrix Multiplication in the Low-Bandwidth Model | Chetan Gupta, Jukka Suomela, Hossein Vahidi | 2026-06-03 | 下载 | We study rectangular matrix multiplication in the low-bandwidth model of distributed computing. There are computers; initially the input matrices are distributed evenly between computers, and in e... |
| Ekka: Automated Diagnosis of Silent Errors in LLM Inference | Yile Gu, Zhen Zhang, Shaowei Zhu, Xinwei Fu, Jun Wu, Yida Wang, Baris Kasikci | 2026-06-03 | 下载 | LLM serving frameworks are quickly evolving with a complex software stack and a vast number of optimizations. The rapid development process can introduce silent errors where output quality silently de... |
| Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge | Haotian Zheng, Zhanwei Wang, Mingyao Cui, Chang Cai, Hongyang Du, Kaibin Huang | 2026-06-03 | 下载 | Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment to enable cooperative... |
| D^2SD: Accelerating Speculative Decoding with Dual Diffusion Draft Models | Liyuan Zhang, Jiarui Zhang, Jinwei Yao, Ran Yan, Yuchen Yang, Jiahao Zhang, Tongkai Yang, Yi Wu, Binhang Yuan | 2026-06-03 | 下载 | Speculative decoding accelerates autoregressive large language model inference by drafting multiple tokens and verifying them in a single target-model forward pass. |
| FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-location | Jiongjiong Gu, Jianfeng Wang, Zidong Han, Yongqiao Wang, Pengfei Xia, Mingjie Zhang, Hong Liu, Yuanyi Xia, Jiajia Chu, Yifeng Tang, Hui Zang, Xin Yao, Qijie Qiu, Yuzhao Wang, Chuanfei Xu, Lin Zhang, Zhuonan Lai, Hongming Huang, Jiawei Qiu, Gong Zhang, Zhong Ming, Weipeng Cao | 2026-06-03 | 下载 | Modern AI serving increasingly relies on NPUs for conventional inference and large language model serving. However, current NPU deployments commonly expose physical devices directly to applications, w... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| RAMC: Remote Access Memory Channels over HPE Slingshot | Whit Schonbein, Matthew G. F. Dosanjh, Scott Levy | 2026-06-03 | 下载 | In this paper, we present Remote Access Memory Channels (RAMC), an explicit one-sided communication library designed to leverage the capabilities of HPE Cray Slingshot network hardware. |
| Bridging High-Level Intent and Network Execution: Detecting Violations and Intent Drift Through Low-Level Traffic Analysis | Tonia Haikal, Shereen Ismail, Eman Hammad | 2026-06-03 | 下载 | Intent-Based Networking (IBN) structures a core management pillar for autonomous 6G networks by translating high-level administrative goals into autonomous configurations, yet a critical validation ga... |
| A Practical AI-Driven Strategy for Cell On/Off Switching under Adaptable QoS Constraints | David Reiss, Miguel Catalan-Cid, Daniel Camps, Oriol Sallent | 2026-06-03 | 下载 | The rapid expansion of 5G networks has intensified concerns over their sustainability, as denser Radio Access Network (RAN) deployments have increased overall power consumption. |
| COSMO: O-RAN-Based Service Management and Orchestration for Cross-Technology Multi-Tenant Radio Access Networks | M. Catalan-Cid, J. J. Aleixendri, J. Pueyo, P. Tomas, D. Camps-Mur | 2026-06-03 | 下载 | The evolution toward 6G networks envisions a heterogeneous Radio Access Network (RAN) comprising diverse access technologies, such as private 5G, public 4G/5G, and Wi-Fi, managed by multiple stakehold... |
| Demo: BeGREEN Intelligence Plane for AI-driven Energy Efficient O-RAN management | M. Catalan-Cid, D. Reiss, G. Castellanos, J. Armstrong | 2026-06-03 | 下载 | Cellular networks management is being enhanced by O-RAN architecture and AI/ML solutions, enabling automated intelligent control loops for RAN optimization across various use cases. |
| From Network Experience to Subscriber Retention: An Explainable AI Framework for Mobile Operators | Faris B. Mismar, Abdol Saleh, Ivan Maxmillian Putra Pasaribu, Suhelmy Syaifuddin | 2026-06-03 | 下载 | This article presents a framework for the prediction of subscriber churn in mobile operators also known as telecommunication operators (or telcos). |
| Contrastive Learning and Correlation Clustering for Sequences of Network Telescope Data | Jannik Presberger, Alexander Männel, Maynard Koch, Thomas C. Schmidt, Matthias Wählisch, Bjoern Andres | 2026-06-03 | 下载 | Understanding activities of Internet scanners is challenging; it often requires identifying relationships between sources, a task for which semantic annotations are scarce. |
| Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge | Haotian Zheng, Zhanwei Wang, Mingyao Cui, Chang Cai, Hongyang Du, Kaibin Huang | 2026-06-03 | 下载 | Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment to enable cooperative... |
| Treat Traffic Like Trees: A Semantic-Preserving Hierarchical Graph-Based Expert Framework for Encrypted Traffic Analysis | Yuantu Luo, Jun Tao, Linxiao Yu, Guang Cheng | 2026-06-03 | 下载 | Graph-based deep learning methods have been widely employed in encrypted traffic analysis to exploit latent correlations across different granularities. |
| Generalizable Multi-Task Learning for Wireless Networks Using Prompt Decision Transformers | Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci | 2026-06-03 | 下载 | Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conventional rule-based and optimization-driven radio ... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| GNStor: Design of GPU-Native High-Performance Remote All-Flash Array | Shushu Yi, Wenbo Wu, Guoci Chen, Junrong Zhu, Shengwen Liang, Mao Bo, Chenying Huan, Chen Tian, Jie Zhang | 2026-06-03 | 下载 | GPU has become the leading computing device for a wide range of data-intensive applications, which tightly collaborates with remote all-flash array (AFA) to accommodate ever-expanding datasets, facili... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Look Before You Leap: Checking in on Type Tag Checking | Stephen M. Watt | 2026-06-03 | 下载 | Tagging of generic dynamic values is important in symbolic-computation and dynamic-language systems, but the trade-offs change as machine architectures and workloads evolve. |