2026-09-23
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Xtrace: High-Fidelity GPU Intra-Kernel Tracing via Binary-Level Instruction Splicing | Zhuobin Huang, Kai Zhang, Weihao Cui, Hongshi Tan, Liang Luo, Christopher Dewan, Shen Li, Bingsheng He | 2026-09-23 | 下载 | Modern GPU kernels fuse increasingly more work into a single kernel, and intra-kernel tracing has become the mainstream method to profile them. |
| HeteroReason: Heterogeneous FPGA-GPU Acceleration for Disaggregated Speculative Reasoning | Zehuan Zhang, Quan Deng, Zibo Ren, Hao Mark Chen, Guoyu Li, Xuchun Hu, Jose G. F. Coutinho, Ce Guo, Wayne Luk, Zhiqiang Que, Hongxiang Fan | 2026-09-23 | 下载 | Large Reasoning Models (LRMs) have achieved state-of-the-art performance in reasoning tasks by utilizing Chain-of-Thought (CoT) reasoning. To achieve fast execution speed, speculative reasoning techni... |
| MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and Inference | Romain Facq, Sami Ben Ali, Olivier Sentieys | 2026-09-23 | 下载 | Microscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full precision accuracy. |
| MVP: A Motion-Predictive Speculative Vision Pipeline with Non-Blocking Drift Correction | Raul Taranco, Antonio González | 2026-09-23 | 下载 | Continuous Vision (CV) systems underpin real-time applications such as autonomous driving and augmented reality, where latency, throughput, and energy are tightly constrained on mobile platforms. |
| Precision and resource scaling of real-time flux distortion compensation for superconducting quantum control | Qi Zhou, Zi-Hao Mei, Peng Duan, Peng Wang, Liang-Liang Guo, Hao-Ran Tao, Wei-Cheng Kong, Hui Yang, Guo-Ping Guo, Zhao-Yun Chen | 2026-09-23 | 下载 | Real-time waveform generation supports dynamic quantum circuits without pre-storing complete waveforms for every execution path. However, long-lived distortions in flux-control lines degrade gate fide... |
| Implementation and Evaluation of BitNet Inference on a CGLA by Signed-Int4 Instructions | Takuto Ando, Yasuhiko Nakashima | 2026-09-23 | 下载 | Large language model (LLM) inference transfers model weights and activations for every generated token, making memory traffic and its energy cost part of the decode path. BitNet b1. |
| Energy-Oriented CGLA Mapping of a Memory-Polynomial Digital Predistortion Kernel | Takuto Ando, Yasuhiko Nakashima | 2026-09-23 | 下载 | Memory-polynomial digital predistortion (DPD) evaluates a small fixed coefficient set over a sliding input history, so its reduction step is a complex-MAC workload with local reuse. |
| Mamba-Family State-Space Model Kernels on a Programmable CGLA | Takuto Ando, Yasuhiko Nakashima | 2026-09-23 | 下载 | Edge and embedded inference is constrained by power and data movement. Mamba-family state-space models replace attention with sequence-linear recurrence, but their inference path combines dense projec... |
| Exploiting Decompression Latency for Covert Channels in Inter-Line-Compressed LLCs | David K. Oh, Hiroshi Sasaki | 2026-09-23 | 下载 | The recently proposed XOR cache is an inter-line-compressed last-level cache (LLC) that leverages the data-inclusion relationship between the private caches and the LLC, compressing two cache lines in... |
| SoK: You Find What You Seek: Rethinking Oracles, Guidance, and Input Generation in Hardware Fuzzing | G Abarajithan, Zhenghua Ma, Cristian Tirelli, Andres Meza, Francesco Restuccia, Cynthia Sturton, Ryan Kastner | 2026-09-23 | 下载 | Hardware fuzzing is an active area in security verification research, yet its industrial adoption remains in its early stages. This SoK examines which lessons from software fuzzing carry over to hardw... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Distributed Service Orchestration in Edge-Cloud Continuum for Digital Healthcare | Johirul Islam, Hafiz Faheem Shahid, Ijaz Ahmad, Tanesh Kumar, Ayan Mondal, Erkki Harjula | 2026-09-23 | 下载 | Today's digital healthcare services rely on various applications and functions that must be continuously accessible. Cloud computing enables global access to these services through public networks, wh... |
| Where Does Streaming State Cost Go? A Reproducible Comparison of Flink and Kafka Streams on Kafka | Kiran N Kumar, Santhosh Kumar Saminathan | 2026-09-23 | 下载 | This paper presents a controlled comparison of exactly-once Kafka pipelines implemented with Apache Flink and Kafka Streams, two engines with different state-management architectures. |
| Flamingo: On Load Balancing in DAG-based Consensus Protocols | Zhen Ping Khor, Garvit Gupta, Mohammad Javad Amiri, Boon Thau Loo | 2026-09-23 | 下载 | Distributed data management systems deployed in untrusted environments rely on Byzantine Fault-Tolerant (BFT) consensus protocols to tolerate malicious failures. |
| GLASS: Architecture-Tuned, Composable, Device-Side Linear Algebra for Edge Robotics and Beyond | Brian Plancher | 2026-09-23 | 下载 | GPU robotics lacks the reusable numerical infrastructure of mature CPU stacks, instead relying on compiler frameworks that introduce overhead or repeatedly reimplementing numerical libraries. |
| The KV Cache Working Set: Online Capacity Planning for LLM Inference Systems | Luchang Li, Shuaishuai Wang, Zhao Ruan, Dongfang Li, Bozhao Gong | 2026-09-23 | 下载 | Prefix caching is critical for efficient large language model (LLM) serving, particularly for agentic workloads that repeatedly invoke the model with a growing conversation and tool-use history. |
| Sovereign Grassroots Currencies: A CBDC Architecture for Credit and Monetary Policy (Full Version) | Ehud Shapiro | 2026-09-23 | 下载 | A Central Bank Digital Currency (CBDC) is central-bank money in digital form, held by the public. Leading designs have two limitations: conversion from bank deposits into CBDC can accelerate deposit f... |
| EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management | Liang Mi, Weijun Wang, Bowen Gao, Tianze Yu, Zixu Hao, Han Xiao, Xin Ding, Mingzhe Huang, Xin He, Lu Shi, Hao Wu, Haipeng Dai, Guihai Chen, Yunxin Liu, Ting Cao | 2026-09-23 | 下载 | Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates. |
| Backstitch: Restoring Request Causality Across a Production Microservice Fleet | Ziyue Dang, Qiuyu Wu, Haoyun Xu, Tongjue Wang, Yongqing Ling, Weihao Chen, Guangming Luo | 2026-09-23 | 下载 | A major video platform runs on thousands of microservices, each request propagating a context so downstream work can be traced and governed. At handoffs outside instrumented paths, e.g. |
| Distributed Stochastic Approximation Algorithms and Heavy-Tailed Age of Information | Adrian Redder, Arunselvan Ramaswamy, Holger Karl | 2026-09-23 | 下载 | Algorithms in multi-agent systems such as federated learning, mobile robotic swarming, and consensus control can be designed and analyzed as distributed stochastic approximation algorithms. |
| CerebroSim: Scalable Whole-Brain Simulator at 100-Trillion-Synapse Scale on the LineShine Supercomputer | Guangnan Feng, Tianxiang Lyu, Hao Huang, Honghui Liang, Jingjing Li, Zhiguang Chen, Yutong Lu | 2026-09-23 | 下载 | Building executable brain models is essential for moving neuroscience from description to mechanism and prediction. Human-brain-scale spiking simulation is constrained by highly irregular communicatio... |
| Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud Orchestration | An N. H. Phan, Dang Van Huynh, Muhammad Usman, Hoa T. Nguyen | 2026-09-23 | 下载 | Quantum cloud computing, delivered through the quantum-as-a-service (QaaS) model, provides access to quantum computing resources. However, applying uniform time-based pricing across fundamentally hete... |
| From PyTorch to the NPU: LLM-Agent-Driven Model Conversion Across Heterogeneous Inference Runtimes | Jianhao Su, Zhanwei Wu, Chia-Heng Tu, ShengTing Huang | 2026-09-23 | 下载 | Edge AI model deployment is a multi-stage engineering process involving model conversion, operator compatibility handling, runtime integration, and precision verification. |
| LayerCheck: Adaptive Layer-wise Checkpointing for Large Language Model Post-training | Minqiu Sun, Xin Huang, Luanzheng Guo, Nathan R. Tallent, Kento Sato, Dong Dai | 2026-09-23 | 下载 | With the rising computational and monetary costs of training large language models (LLMs), checkpointing---periodically storing model states for recovery---becomes essential for fault tolerance. |
| ZOCheck: CPU-Shadow Checkpointing for Zeroth-Order LLM Fine-Tuning | Minqiu Sun, Xin Huang, Luanzheng Guo, Nathan R. Tallent, Kento Sato, Dong Dai | 2026-09-23 | 下载 | Zeroth-order (ZO) optimization is an attractive option for memory-efficient LLM fine-tuning, but its fault tolerance remains underexplored. Unlike first-order training, ZO progress can be represented ... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Has The Physical Layer Matured? | Mansoor Shafi, Changlong Xu, Xingqin Lin, Gilwon Lee, Feifei Sun, Eko Onggosanusi, Oskari Tervo, Joonyoung Cho | 2026-09-23 | 下载 | The wireless physical (PHY) layer has enabled successive generations of cellular systems through advances in modulation, coding, waveforms, and multiple-input multiple-output (MIMO) transmission. |
| Unmasking Shortcut Learning in IoT Intrusion Detection: A Forensic, Multi-Paradigm Evaluation of Feature Dependence and Data Leakage | Uday Shankar Roy, Mahbuba Jahan Minu | 2026-09-23 | 下载 | Machine learning-based Network Intrusion Detection Systems often report near-perfect performance on IoT benchmarks. However, whether these models learn generalizable attack behavior or exploit spuriou... |
| Load Balancing with Partial Queue Information - Threshold Optimality and Indexability | Sathwik Chadaga, Eytan Modiano | 2026-09-23 | 下载 | We consider the problem of load balancing in a system with one dispatcher and parallel servers. The dispatcher must select one server to dispatch new jobs at every time-step and each server buffer... |
| A DRL-Driven Optimization of RAN Slice Resource Partitioning for V2X SLA Compliance in 5G Networks | M. Martínez, I. de-la-Bandera, D. E. García, P. Vera, S. Fortes, M. L. Luque, A. Mendo, J. Ramiro, R. Barco | 2026-09-23 | 下载 | Vehicle-to-Everything (V2X) communications impose very demanding requirements in terms of latency and reliability, which must be met in scenarios where multiple services with diverse performance targe... |
| Knowledge Distillation for Intelligent Softwarized Networks: Advances and Open Challenges | Mohamed Ali Zormati, Ghada Jaber, Hicham Lakhlef | 2026-09-23 | 下载 | The increasing adoption of software defined networking and network function virtualization, combined with rapid advances in Machine Learning (ML), is driving the evolution toward intelligent network s... |
| Distributed Stochastic Approximation Algorithms and Heavy-Tailed Age of Information | Adrian Redder, Arunselvan Ramaswamy, Holger Karl | 2026-09-23 | 下载 | Algorithms in multi-agent systems such as federated learning, mobile robotic swarming, and consensus control can be designed and analyzed as distributed stochastic approximation algorithms. |
| From Intents to Algorithms: Verified Algorithm Discovery for Transport Networks | Behnam Ojaghi, Ricard Vilalta, Raul Muñoz | 2026-09-23 | 下载 | Intent-based networking decouples desired outcomes from device-level configuration, but most systems still map intents to parameters of an algorithm selected in advance. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| xTier: Intelligent Tiering for CXL-Enabled Memory | Sriranga Ramaswamy, Yueqi Chen | 2026-09-23 | 下载 | CXL-enabled memory expands server memory capacity, but introduces a page-placement problem: the operating system must decide which pages should reside in DRAM and which should reside on slower CXL mem... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| The Canonical Parallel Form as a Substrate for Parallelizing Compilers and Agentic Optimizers | Yakup Koray Budanaz, Pratyai Mazumder, Alexandru Calotoiu, Torsten Hoefler | 2026-09-23 | 下载 | Imperative code fixes an execution order the computation does not require, and a parallelizing compiler must prove which parts of that order it can remove. |
| RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision Models | David Población-Criado, Dario Garcia-Gasulla, Eduardo Quinones | 2026-09-23 | 下载 | Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy. |