Skip to content

2026-06-03 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
SET: Stream-Event-Triggered Scheduling for Efficient CUDA Graph PipelinesZhengxiong Li, Tsung-Wei Huang, Umit Ogras2026-06-03下载Achieving peak GPU performance remains a significant challenge as the system throughput is constrained by host-device synchronization delays and kernel scheduling overheads, even with aggressive kerne...
MOSAIC: A Workload-Driven Simulation and Design-Space Exploration Framework for Heterogeneous NPUsArghadip Das, Hoseok Kim, Soomin Lee, Arnab Raha, Deepak A Mathaikutty, Vijay Raghunathan2026-06-03下载AI model architectures are diversifying rapidly. Although dense matrix multiplication underlies today's CNNs and transformers, emerging architectures (state-space models, long convolutions via the fas...
BIDENT: Heterogeneous Operator-level Mapping for Efficient Edge InferenceHoseok Kim, Arghadip Das, Soumendu Ghosh, Arnab Raha, Vijay Raghunathan2026-06-03下载Modern edge System-on-Chips (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, yet current inference stacks map entire models to a single PU, leaving significant perfo...
GoldenFloat: A Phi-Derived Static-Split Floating-Point Family from GF4 to GF256 with a Lucas-Exact Integer IdentityDmitrii Vasiliev2026-06-03下载We present a hardware-oriented description of GoldenFloat (GF), a static-split floating-point family generated by a single closed rule, and three concrete artefacts: (i) an open multi-width RTL genera...
Uncertainty-Aware End-to-End Co-Design of Neural Network Processors: From Training and Mapping to FabricationYuyang Du, Yujun Huang, Gioele Zardini2026-06-03下载Designing a neural network processor is an end-to-end co-design problem: network architecture and training budget determine the inference workload; hardware mapping decisions determine chip area, late...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Latent Reasoning Guidance for Parallel Code TranslationTomer Bitan, Erel Kaplan, Roee Bar-Yadin, Lian Ghrayeb, Le Chen, Samyak Jhaveri, Niranjan Hasabnis, Gal Oren2026-06-03下载Tackling complex coding tasks often requires autonomous agents and iterative repair pipelines. These increasingly rely on large amounts of test-time computation, often spending many decoding and repai...
Bitcoin After Block RewardsJunhyuk Lee2026-06-03下载Bitcoin's block reward is scheduled to decline to zero, raising concerns about whether the network can remain secure once miners rely solely on transaction fees.
SET: Stream-Event-Triggered Scheduling for Efficient CUDA Graph PipelinesZhengxiong Li, Tsung-Wei Huang, Umit Ogras2026-06-03下载Achieving peak GPU performance remains a significant challenge as the system throughput is constrained by host-device synchronization delays and kernel scheduling overheads, even with aggressive kerne...
Graph Traversal on Tensor Cores: A BFS Framework for Modern GPUsDeniz Elbek, Kamer Kaya2026-06-03下载Modern GPUs have Tensor Cores (TCs) capable of extremely high-throughput matrix operations, yet graph algorithms remain difficult to accelerate because of their irregular and data-dependent execution ...
The local complexity of certifying parityNicolas Bousquet, Laurent Feuilloley, Jorge Valenzuela, Sébastien Zeitoun2026-06-03下载In this paper, we consider the problem of locally certifying that the size of a network is even, or more generally, congruent to some fixed number.
The Usefulness Gap in Proof-of-Useful-Work: An Empirical Study of Pearl's cuPOW ProtocolAbhinaba Basu2026-06-03下载Pearl, a Layer-1 blockchain with high-profile AI industry endorsements, markets its Proof-of-Useful-Work (PoUW) protocol as simultaneously securing the network and performing AI inference.
Clownfish: Scaling DAG-based BFT Consensus via Sparse EdgesFeifan Wang, Jingfan Yu, Zixi Cai, Zhixuan Fang2026-06-03下载Directed Acyclic Graph (DAG) based BFT protocols have demonstrated the capability to achieve significantly high throughput in practice. Recent advancements focused on minimizing the good-case latency ...
Rectangular Matrix Multiplication in the Low-Bandwidth ModelChetan Gupta, Jukka Suomela, Hossein Vahidi2026-06-03下载We study rectangular matrix multiplication in the low-bandwidth model of distributed computing. There are nn computers; initially the input matrices are distributed evenly between computers, and in e...
Ekka: Automated Diagnosis of Silent Errors in LLM InferenceYile Gu, Zhen Zhang, Shaowei Zhu, Xinwei Fu, Jun Wu, Yida Wang, Baris Kasikci2026-06-03下载LLM serving frameworks are quickly evolving with a complex software stack and a vast number of optimizations. The rapid development process can introduce silent errors where output quality silently de...
Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the EdgeHaotian Zheng, Zhanwei Wang, Mingyao Cui, Chang Cai, Hongyang Du, Kaibin Huang2026-06-03下载Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment to enable cooperative...
D^2SD: Accelerating Speculative Decoding with Dual Diffusion Draft ModelsLiyuan Zhang, Jiarui Zhang, Jinwei Yao, Ran Yan, Yuchen Yang, Jiahao Zhang, Tongkai Yang, Yi Wu, Binhang Yuan2026-06-03下载Speculative decoding accelerates autoregressive large language model inference by drafting multiple tokens and verifying them in a single target-model forward pass.
FlexNPU: Transparent NPU Virtualization for Dynamic LLM Prefill-Decode Co-locationJiongjiong Gu, Jianfeng Wang, Zidong Han, Yongqiao Wang, Pengfei Xia, Mingjie Zhang, Hong Liu, Yuanyi Xia, Jiajia Chu, Yifeng Tang, Hui Zang, Xin Yao, Qijie Qiu, Yuzhao Wang, Chuanfei Xu, Lin Zhang, Zhuonan Lai, Hongming Huang, Jiawei Qiu, Gong Zhang, Zhong Ming, Weipeng Cao2026-06-03下载Modern AI serving increasingly relies on NPUs for conventional inference and large language model serving. However, current NPU deployments commonly expose physical devices directly to applications, w...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
RAMC: Remote Access Memory Channels over HPE SlingshotWhit Schonbein, Matthew G. F. Dosanjh, Scott Levy2026-06-03下载In this paper, we present Remote Access Memory Channels (RAMC), an explicit one-sided communication library designed to leverage the capabilities of HPE Cray Slingshot network hardware.
Bridging High-Level Intent and Network Execution: Detecting Violations and Intent Drift Through Low-Level Traffic AnalysisTonia Haikal, Shereen Ismail, Eman Hammad2026-06-03下载Intent-Based Networking (IBN) structures a core management pillar for autonomous 6G networks by translating high-level administrative goals into autonomous configurations, yet a critical validation ga...
A Practical AI-Driven Strategy for Cell On/Off Switching under Adaptable QoS ConstraintsDavid Reiss, Miguel Catalan-Cid, Daniel Camps, Oriol Sallent2026-06-03下载The rapid expansion of 5G networks has intensified concerns over their sustainability, as denser Radio Access Network (RAN) deployments have increased overall power consumption.
COSMO: O-RAN-Based Service Management and Orchestration for Cross-Technology Multi-Tenant Radio Access NetworksM. Catalan-Cid, J. J. Aleixendri, J. Pueyo, P. Tomas, D. Camps-Mur2026-06-03下载The evolution toward 6G networks envisions a heterogeneous Radio Access Network (RAN) comprising diverse access technologies, such as private 5G, public 4G/5G, and Wi-Fi, managed by multiple stakehold...
Demo: BeGREEN Intelligence Plane for AI-driven Energy Efficient O-RAN managementM. Catalan-Cid, D. Reiss, G. Castellanos, J. Armstrong2026-06-03下载Cellular networks management is being enhanced by O-RAN architecture and AI/ML solutions, enabling automated intelligent control loops for RAN optimization across various use cases.
From Network Experience to Subscriber Retention: An Explainable AI Framework for Mobile OperatorsFaris B. Mismar, Abdol Saleh, Ivan Maxmillian Putra Pasaribu, Suhelmy Syaifuddin2026-06-03下载This article presents a framework for the prediction of subscriber churn in mobile operators also known as telecommunication operators (or telcos).
Contrastive Learning and Correlation Clustering for Sequences of Network Telescope DataJannik Presberger, Alexander Männel, Maynard Koch, Thomas C. Schmidt, Matthias Wählisch, Bjoern Andres2026-06-03下载Understanding activities of Internet scanners is challenging; it often requires identifying relationships between sources, a task for which semantic annotations are scarce.
Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the EdgeHaotian Zheng, Zhanwei Wang, Mingyao Cui, Chang Cai, Hongyang Du, Kaibin Huang2026-06-03下载Speculative inference (SPIN) was originally developed as an efficient architecture to accelerate Large Language Models (LLMs). In this work, we propose its distributed deployment to enable cooperative...
Treat Traffic Like Trees: A Semantic-Preserving Hierarchical Graph-Based Expert Framework for Encrypted Traffic AnalysisYuantu Luo, Jun Tao, Linxiao Yu, Guang Cheng2026-06-03下载Graph-based deep learning methods have been widely employed in encrypted traffic analysis to exploit latent correlations across different granularities.
Generalizable Multi-Task Learning for Wireless Networks Using Prompt Decision TransformersFatih Temiz, Shavbo Salehi, Melike Erol-Kantarci2026-06-03下载Future wireless networks demand rapid adaptation to highly heterogeneous environments and dynamic task configurations, necessitating a shift from conventional rule-based and optimization-driven radio ...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
GNStor: Design of GPU-Native High-Performance Remote All-Flash ArrayShushu Yi, Wenbo Wu, Guoci Chen, Junrong Zhu, Shengwen Liang, Mao Bo, Chenying Huan, Chen Tian, Jie Zhang2026-06-03下载GPU has become the leading computing device for a wide range of data-intensive applications, which tightly collaborates with remote all-flash array (AFA) to accommodate ever-expanding datasets, facili...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Look Before You Leap: Checking in on Type Tag CheckingStephen M. Watt2026-06-03下载Tagging of generic dynamic values is important in symbolic-computation and dynamic-language systems, but the trade-offs change as machine architectures and workloads evolve.

基于 VitePress 构建 · 使用本地搜索查找论文