Skip to content

2026-09-23 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Xtrace: High-Fidelity GPU Intra-Kernel Tracing via Binary-Level Instruction SplicingZhuobin Huang, Kai Zhang, Weihao Cui, Hongshi Tan, Liang Luo, Christopher Dewan, Shen Li, Bingsheng He2026-09-23下载Modern GPU kernels fuse increasingly more work into a single kernel, and intra-kernel tracing has become the mainstream method to profile them.
HeteroReason: Heterogeneous FPGA-GPU Acceleration for Disaggregated Speculative ReasoningZehuan Zhang, Quan Deng, Zibo Ren, Hao Mark Chen, Guoyu Li, Xuchun Hu, Jose G. F. Coutinho, Ce Guo, Wayne Luk, Zhiqiang Que, Hongxiang Fan2026-09-23下载Large Reasoning Models (LRMs) have achieved state-of-the-art performance in reasoning tasks by utilizing Chain-of-Thought (CoT) reasoning. To achieve fast execution speed, speculative reasoning techni...
MicroQonv: Reshaping Convolution Tensors for Efficient Microscaling in Training and InferenceRomain Facq, Sami Ben Ali, Olivier Sentieys2026-09-23下载Microscaling quantization techniques are increasingly used to represent neural network parameters with 8 bits or fewer while preserving near-full precision accuracy.
MVP: A Motion-Predictive Speculative Vision Pipeline with Non-Blocking Drift CorrectionRaul Taranco, Antonio González2026-09-23下载Continuous Vision (CV) systems underpin real-time applications such as autonomous driving and augmented reality, where latency, throughput, and energy are tightly constrained on mobile platforms.
Precision and resource scaling of real-time flux distortion compensation for superconducting quantum controlQi Zhou, Zi-Hao Mei, Peng Duan, Peng Wang, Liang-Liang Guo, Hao-Ran Tao, Wei-Cheng Kong, Hui Yang, Guo-Ping Guo, Zhao-Yun Chen2026-09-23下载Real-time waveform generation supports dynamic quantum circuits without pre-storing complete waveforms for every execution path. However, long-lived distortions in flux-control lines degrade gate fide...
Implementation and Evaluation of BitNet Inference on a CGLA by Signed-Int4 InstructionsTakuto Ando, Yasuhiko Nakashima2026-09-23下载Large language model (LLM) inference transfers model weights and activations for every generated token, making memory traffic and its energy cost part of the decode path. BitNet b1.
Energy-Oriented CGLA Mapping of a Memory-Polynomial Digital Predistortion KernelTakuto Ando, Yasuhiko Nakashima2026-09-23下载Memory-polynomial digital predistortion (DPD) evaluates a small fixed coefficient set over a sliding input history, so its reduction step is a complex-MAC workload with local reuse.
Mamba-Family State-Space Model Kernels on a Programmable CGLATakuto Ando, Yasuhiko Nakashima2026-09-23下载Edge and embedded inference is constrained by power and data movement. Mamba-family state-space models replace attention with sequence-linear recurrence, but their inference path combines dense projec...
Exploiting Decompression Latency for Covert Channels in Inter-Line-Compressed LLCsDavid K. Oh, Hiroshi Sasaki2026-09-23下载The recently proposed XOR cache is an inter-line-compressed last-level cache (LLC) that leverages the data-inclusion relationship between the private caches and the LLC, compressing two cache lines in...
SoK: You Find What You Seek: Rethinking Oracles, Guidance, and Input Generation in Hardware FuzzingG Abarajithan, Zhenghua Ma, Cristian Tirelli, Andres Meza, Francesco Restuccia, Cynthia Sturton, Ryan Kastner2026-09-23下载Hardware fuzzing is an active area in security verification research, yet its industrial adoption remains in its early stages. This SoK examines which lessons from software fuzzing carry over to hardw...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Distributed Service Orchestration in Edge-Cloud Continuum for Digital HealthcareJohirul Islam, Hafiz Faheem Shahid, Ijaz Ahmad, Tanesh Kumar, Ayan Mondal, Erkki Harjula2026-09-23下载Today's digital healthcare services rely on various applications and functions that must be continuously accessible. Cloud computing enables global access to these services through public networks, wh...
Where Does Streaming State Cost Go? A Reproducible Comparison of Flink and Kafka Streams on KafkaKiran N Kumar, Santhosh Kumar Saminathan2026-09-23下载This paper presents a controlled comparison of exactly-once Kafka pipelines implemented with Apache Flink and Kafka Streams, two engines with different state-management architectures.
Flamingo: On Load Balancing in DAG-based Consensus ProtocolsZhen Ping Khor, Garvit Gupta, Mohammad Javad Amiri, Boon Thau Loo2026-09-23下载Distributed data management systems deployed in untrusted environments rely on Byzantine Fault-Tolerant (BFT) consensus protocols to tolerate malicious failures.
GLASS: Architecture-Tuned, Composable, Device-Side Linear Algebra for Edge Robotics and BeyondBrian Plancher2026-09-23下载GPU robotics lacks the reusable numerical infrastructure of mature CPU stacks, instead relying on compiler frameworks that introduce overhead or repeatedly reimplementing numerical libraries.
The KV Cache Working Set: Online Capacity Planning for LLM Inference SystemsLuchang Li, Shuaishuai Wang, Zhao Ruan, Dongfang Li, Bozhao Gong2026-09-23下载Prefix caching is critical for efficient large language model (LLM) serving, particularly for agentic workloads that repeatedly invoke the model with a growing conversation and tool-use history.
Sovereign Grassroots Currencies: A CBDC Architecture for Credit and Monetary Policy (Full Version)Ehud Shapiro2026-09-23下载A Central Bank Digital Currency (CBDC) is central-bank money in digital form, held by the public. Leading designs have two limitations: conversion from bank deposits into CBDC can accelerate deposit f...
EBRL: Asynchronous Embodied RL by Multi-Grained Resource ManagementLiang Mi, Weijun Wang, Bowen Gao, Tianze Yu, Zixu Hao, Han Xiao, Xin Ding, Mingzhe Huang, Xin He, Lu Shi, Hao Wu, Haipeng Dai, Guihai Chen, Yunxin Liu, Ting Cao2026-09-23下载Embodied reinforcement learning (RL) improves model capabilities with a pipeline of environment simulation, action generation, and model updates.
Backstitch: Restoring Request Causality Across a Production Microservice FleetZiyue Dang, Qiuyu Wu, Haoyun Xu, Tongjue Wang, Yongqing Ling, Weihao Chen, Guangming Luo2026-09-23下载A major video platform runs on thousands of microservices, each request propagating a context so downstream work can be traced and governed. At handoffs outside instrumented paths, e.g.
Distributed Stochastic Approximation Algorithms and Heavy-Tailed Age of InformationAdrian Redder, Arunselvan Ramaswamy, Holger Karl2026-09-23下载Algorithms in multi-agent systems such as federated learning, mobile robotic swarming, and consensus control can be designed and analyzed as distributed stochastic approximation algorithms.
CerebroSim: Scalable Whole-Brain Simulator at 100-Trillion-Synapse Scale on the LineShine SupercomputerGuangnan Feng, Tianxiang Lyu, Hao Huang, Honghui Liang, Jingjing Li, Zhiguang Chen, Yutong Lu2026-09-23下载Building executable brain models is essential for moving neuroscience from description to mechanism and prediction. Human-brain-scale spiking simulation is constrained by highly irregular communicatio...
Quantum Reinforcement Learning for Cost and Delay Tradeoffs in Quantum Cloud OrchestrationAn N. H. Phan, Dang Van Huynh, Muhammad Usman, Hoa T. Nguyen2026-09-23下载Quantum cloud computing, delivered through the quantum-as-a-service (QaaS) model, provides access to quantum computing resources. However, applying uniform time-based pricing across fundamentally hete...
From PyTorch to the NPU: LLM-Agent-Driven Model Conversion Across Heterogeneous Inference RuntimesJianhao Su, Zhanwei Wu, Chia-Heng Tu, ShengTing Huang2026-09-23下载Edge AI model deployment is a multi-stage engineering process involving model conversion, operator compatibility handling, runtime integration, and precision verification.
LayerCheck: Adaptive Layer-wise Checkpointing for Large Language Model Post-trainingMinqiu Sun, Xin Huang, Luanzheng Guo, Nathan R. Tallent, Kento Sato, Dong Dai2026-09-23下载With the rising computational and monetary costs of training large language models (LLMs), checkpointing---periodically storing model states for recovery---becomes essential for fault tolerance.
ZOCheck: CPU-Shadow Checkpointing for Zeroth-Order LLM Fine-TuningMinqiu Sun, Xin Huang, Luanzheng Guo, Nathan R. Tallent, Kento Sato, Dong Dai2026-09-23下载Zeroth-order (ZO) optimization is an attractive option for memory-efficient LLM fine-tuning, but its fault tolerance remains underexplored. Unlike first-order training, ZO progress can be represented ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Has The Physical Layer Matured?Mansoor Shafi, Changlong Xu, Xingqin Lin, Gilwon Lee, Feifei Sun, Eko Onggosanusi, Oskari Tervo, Joonyoung Cho2026-09-23下载The wireless physical (PHY) layer has enabled successive generations of cellular systems through advances in modulation, coding, waveforms, and multiple-input multiple-output (MIMO) transmission.
Unmasking Shortcut Learning in IoT Intrusion Detection: A Forensic, Multi-Paradigm Evaluation of Feature Dependence and Data LeakageUday Shankar Roy, Mahbuba Jahan Minu2026-09-23下载Machine learning-based Network Intrusion Detection Systems often report near-perfect performance on IoT benchmarks. However, whether these models learn generalizable attack behavior or exploit spuriou...
Load Balancing with Partial Queue Information - Threshold Optimality and IndexabilitySathwik Chadaga, Eytan Modiano2026-09-23下载We consider the problem of load balancing in a system with one dispatcher and NN parallel servers. The dispatcher must select one server to dispatch new jobs at every time-step and each server buffer...
A DRL-Driven Optimization of RAN Slice Resource Partitioning for V2X SLA Compliance in 5G NetworksM. Martínez, I. de-la-Bandera, D. E. García, P. Vera, S. Fortes, M. L. Luque, A. Mendo, J. Ramiro, R. Barco2026-09-23下载Vehicle-to-Everything (V2X) communications impose very demanding requirements in terms of latency and reliability, which must be met in scenarios where multiple services with diverse performance targe...
Knowledge Distillation for Intelligent Softwarized Networks: Advances and Open ChallengesMohamed Ali Zormati, Ghada Jaber, Hicham Lakhlef2026-09-23下载The increasing adoption of software defined networking and network function virtualization, combined with rapid advances in Machine Learning (ML), is driving the evolution toward intelligent network s...
Distributed Stochastic Approximation Algorithms and Heavy-Tailed Age of InformationAdrian Redder, Arunselvan Ramaswamy, Holger Karl2026-09-23下载Algorithms in multi-agent systems such as federated learning, mobile robotic swarming, and consensus control can be designed and analyzed as distributed stochastic approximation algorithms.
From Intents to Algorithms: Verified Algorithm Discovery for Transport NetworksBehnam Ojaghi, Ricard Vilalta, Raul Muñoz2026-09-23下载Intent-based networking decouples desired outcomes from device-level configuration, but most systems still map intents to parameters of an algorithm selected in advance.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
xTier: Intelligent Tiering for CXL-Enabled MemorySriranga Ramaswamy, Yueqi Chen2026-09-23下载CXL-enabled memory expands server memory capacity, but introduces a page-placement problem: the operating system must decide which pages should reside in DRAM and which should reside on slower CXL mem...

cs.PF - Performance ​

标题作者发布日期PDF摘要
The Canonical Parallel Form as a Substrate for Parallelizing Compilers and Agentic OptimizersYakup Koray Budanaz, Pratyai Mazumder, Alexandru Calotoiu, Torsten Hoefler2026-09-23下载Imperative code fixes an execution order the computation does not require, and a parallelizing compiler must prove which parts of that order it can remove.
RAMP: Robust Adaptive Mixed-Precision Quantization for Edge CPU Vision ModelsDavid Población-Criado, Dario Garcia-Gasulla, Eduardo Quinones2026-09-23下载Deploying deep learning models on edge CPUs is bottlenecked by computational and memory constraints. Mixed-precision quantization promises to reduce inference latency while preserving accuracy.

基于 VitePress 构建 · 使用本地搜索查找论文