2026-09-07
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A 28nm 27,648-Spin Multichip Digital Ising Accelerator with Pegasus Connectivity | Tong Wu, Atharva Raut, Ian Khor, Ziyad Alswaidan, Ting-Yu Lee, Hongchang Kuang, Navid Anjum Aadit, Anand Raju, Dhruv Chinmay, Siddharth Das, Ken Mai, Kerem Camsari, Tathagata Srimani | 2026-09-07 | 下载 | We present a 28nm digital Ising accelerator with 27,648 spins across four chips. A time-multiplexed spin-update array with local SRAM and scheduled interchip transfers delivers 41.5G updates/s at 1. |
| QROB: Quantifying Realization Overhead in Quantum Compilation via Reverse Construction | Jintao Li, Kaiqi Li, Rui Wang, Yilun Zhao, Kaixuan Huang, Ying Wang, Jialin Zhang, Zheng-An Wang, Xiaoming Sun, Heng Fan | 2026-09-07 | 下载 | Quantum compilation reconciles a program's idealized interaction topology with hardware locality constraints, yet evaluations at scale lack calibrated references for realization overhead. |
| TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning | Avik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann | 2026-09-07 | 下载 | The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly ... |
| Enabling High-Bandwidth Flash for Generative Recommendation Serving with Write-Aware KV Cache Policy | Danni Peng, Kai Wu, Tianyu Zuo, Pengfei Xia, Hui Zang | 2026-09-07 | 下载 | Generative recommendation (GR) systems increasingly leverage user-level KV cache reuse to avoid recomputing long user histories. However, the growing KV cache capacity and bandwidth requirements intro... |
| NOVA-CIM: Noise- and Correlation-Tolerant Stochastic Interfaces for Analog Compute-in-Memory | Jiachen Ren, Wenshuai Yao, Haobo Liu, Xincheng Feng, Chenxi Hu, Zhengwu Liu, Kechao Tang, Wenyong Zhou, Ngai Wong | 2026-09-07 | 下载 | Analog compute-in-memory (CIM) enables energy-efficient model acceleration, but its reliance on ADC-based readout, which directly quantizes noisy column currents, makes inference accuracy highly sensi... |
| Capability-Gated Conformance Testing of Quantum Error-Correction Decoder Libraries | Jiachen Shen, Hui Zhong | 2026-09-07 | 下载 | A quantum error correction decoder is a library other people's results depend on, judged in one dominant way. Sample errors, decode, and count wrong logical observables. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| MicroIntent: Intent-Based Placement Strategy for Microservice Application in the Compute Continuum Using LLMs | Koushikur Islam, Guilherme Da Cunha Rodrigues, Bahman Javadi, Rodrigo N. Calheiros | 2026-09-07 | 下载 | The placement of microservices in the compute continuum plays a vital role in delivering services that comply with customers' needs, such as reduced latency, storage requirements, quality of service a... |
| Interactive Debugger for Performance Portable Python HPC Kernels | Ivan Grigorik, Gabriel Kosmacher, George Biros, Milos Gligoric | 2026-09-07 | 下载 | We propose PKDB, the first interactive debugger for GPU and multithreaded low-level kernels written in Python. Python is widely used in high performance computing (HPC), with frameworks such as PyKokk... |
| Real-Time dApps for AI-RAN: Measured Interface Requirements for Inline PHY and Slot-Level Control | Timothy O'Shea, Matthew Pennybacker, Andriy Kharchenko | 2026-09-07 | 下载 | Distributed applications (dApps) bring AI to the microsecond-to-millisecond band beside the 5G distributed unit (DU), but every public dApp framework realizes them the same way: an external process th... |
| Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems | Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso | 2026-09-07 | 下载 | Kolmogorov-Arnold Networks (KANs) replace the fixed activation functions and linear weights of Multi-Layer Perceptrons (MLPs) with learnable univariate functions on network edges, offering improved in... |
| Attestream: Usage-Aware Intermittent Data Distribution with Verifiable Lifecycle Provenance for Machine-Learning Data Streams | Kentaro Oda | 2026-09-07 | 下载 | Providers of continuously produced, commercially valuable data -- sensor streams, telemetry, and other feeds sold as machine-learning training material -- cannot observe whether delivered data is actu... |
| Fast Multidimensional Approximate Agreement with Optimal Resilience Using Ball Validity | Tijana Milentijević, Stefan Schmid | 2026-09-07 | 下载 | Multidimensional approximate agreement requires processes with inputs in to output vectors close to each other, despite up to Byzantine faults. |
| GPU-Accelerated Hypergraph Partitioning and Placement to Map SNNs on Neuromorphic Hardware | Marco Ronzani, Cristina Silvano | 2026-09-07 | 下载 | SNNs running on neuromorphic hardware use spikes to achieve sparse and energy-efficient communication over a mesh of cores. In turn, system performance heavily depends on the assignment of neurons to ... |
| Analytical Resource Management for Fine-grained MoE Computation-Communication Overlap | Hongyu Liu, Minyu Cui, Miquel Pericas | 2026-09-07 | 下载 | Fine-grained computation--communication overlap in distributed Mixture-of-Experts (MoE) inference allows communication to begin as partial compute results become ready. |
| From Bracha to Coded MBRB: Benchmarking Byzantine Reliable Broadcast Implementations | Yenan Wang, Jesper Kullberg, Fabian Paglianno Persson, Elad Michael Schiller, Timothé Albouy | 2026-09-07 | 下载 | Byzantine Reliable Broadcast (BRB) and Message-Adversary-Tolerant Byzantine Reliable Broadcast (MBRB) are reliable-dissemination abstractions for fault-tolerant distributed systems. |
| PLATOS: A Power and Latency-Aware Task-Oriented Scheduling Strategy for Healthcare IoT in Fog Computing | Mohammed Alaa Ala'anzy, Zulfiqar Ahmad, Zhanar Mukash | 2026-09-07 | 下载 | Healthcare Internet of Things (HIoT) technology is revolutionising the healthcare industry by enabling real-time data collection and analysis for personalised patient care. |
| Robust Decentralized Personalized Federated Learning via Prediction-Constrained Neighborhood Collaboration | Xiao Ma, Hong Shen, Hui Tian, Wenqi Lyu, Wei Ke | 2026-09-07 | 下载 | This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine attacks via robust neighborhood direction estimation an... |
| The Maximum Mutual Visibility Set on a Cactus Graph and the Self-stabilizing Constructions | Yonghwan Kim, Yuichi Sudo | 2026-09-07 | 下载 | Given a graph , let () be a set of vertices. Two vertices are \emph{mutually visible} if there exists a shortest path in between them that does not contain any other vert... |
| An Efficient Out-of-Core Tomographic Imaging Framework for Edge Devices | Xuetao Chen, Cong Ma, Xiangyu Meng, Du Wu, Zhengyang Bai, Tao Luo, Zhaorui Zhang, Emmanuel Jeannot, Edgar Josafat Martinez Noriega, Xun Wang, Peng Chen, Amelie Chi Zhou, Mohamed Wahib | 2026-09-07 | 下载 | Computed Tomography (CT) is an essential 3D imaging technology widely used in medical diagnostics and scientific research. However, performing CT imaging on edge devices is challenging due to limitati... |
| Parallelism Strategy Chaining for Fast Training Convergence | Minchul Kang, Changyong Shin, Younghun Go, Hyunho Lee, Jinwoo Jeong, Chuck Yoo, Gyeongsik Yang | 2026-09-07 | 下载 | Selecting a parallelism strategy - the configuration of data, tensor, and pipeline parallelism degrees together with micro- and global-batch sizes - largely determines the training efficiency of large... |
| Robust Decentralized Federated Distillation via Multi-Modality Knowledge Collaboration | Xiao Ma, Hong Shen, Hui Tian, Wei Ke, Wenqi Lyu | 2026-09-07 | 下载 | This paper propose a robust decentralized federated distillation method that enables clients with heterogeneous models to collaborate through predictions on shared unlabeled public data. |
| Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training | Zili Wang, Zhaopeng Qiu, Yuekai Zhang, Shuang Yu, Junjie Lai | 2026-09-07 | 下载 | Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft's accuracy, yielding great... |
| TreeRedux: Separating Concerns in Spark's Distributed Tree Aggregation | David A. G. Harrison, Ivan Cao | 2026-09-07 | 下载 | By default, Apache Spark's tree aggregation primitives place the tree root on the driver, requiring the driver to participate in the same aggregation computation over intermediate aggregation state as... |
| Unified AI Gateway: A Framework for Joint Model Routing and KV Cache Management | Jiaxun Lu, Xiang Zhang, Yunfeng Shao | 2026-09-07 | 下载 | Large language model (LLM) inference increasingly spans models that differ in size, capability, price, and provider. This shift creates two costs for developers. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Movable Antennas Enabled Wireless Powered Networks: Principles and Technologies | Zhendong Li, Yiran Zheng, Tianyu Li, Zhou Su, Wen Chen, Ying Wang | 2026-09-07 | 下载 | As an emerging framework, movable antenna (MA)-enabled wireless powered networks (WPNs) have attracted growing attention. WPNs integrate wireless communication and energy transfer. |
| The OCUDU dApp Platform: An Open Runtime and E3 Interface for Real-Time AI-RAN | Timothy O'Shea, Matthew Pennybacker, Andriy Kharchenko | 2026-09-07 | 下载 | Machine learning has shown its largest gains in the band below 10 ms inside a 3GPP new radio (NR) 5G distributed unit (DU): link adaptation, per-slot scheduling, channel estimation, and the receiver i... |
| Real-Time dApps for AI-RAN: Measured Interface Requirements for Inline PHY and Slot-Level Control | Timothy O'Shea, Matthew Pennybacker, Andriy Kharchenko | 2026-09-07 | 下载 | Distributed applications (dApps) bring AI to the microsecond-to-millisecond band beside the 5G distributed unit (DU), but every public dApp framework realizes them the same way: an external process th... |
| How Bitcoin Forms Its Network: Peer-Table Sampling and Structural Properties | Taki E. M. Abedesselam, Antonio Cruciani, Fabio Giacomelli, Lucianna Kiffer, Francesco Pasquale | 2026-09-07 | 下载 | The global structure of P2P networks underlying modern cryptocurrencies is hidden by design: each node only knows its neighbors and maintains a local \textit{peer table} of IP addresses. |
| No-Regret Mixing of LRU and LFU with Optimal Switching Cost | Younes Ben Mazziane, Xinying Zou | 2026-09-07 | 下载 | Caching systems often rely on simple eviction policies such as Least Recently Used (LRU) and Least Frequently Used (LFU), which perform well in complementary request regimes. |
| Perception-Aware Joint Power and Sub-Band Allocation for 6G In-Body Subnetworks | Samira Abdelrahman, Hossam Farag | 2026-09-07 | 下载 | In-body subnetworks (IBSs) are expected to become a key enabler of immersive eXtended Reality (XR) services in sixth-generation (6G) networks by providing ultra-short-range, low-latency wireless conne... |
| Blockchain-based Proportional Fair Scheduling for Multi-Operator O-RAN | Kun Huang, Xintong Ling, Meining Wu, Jiaheng Wang, Zhi Ding, Xiqi Gao | 2026-09-07 | 下载 | The openness and disaggregation of Open radio access network (O-RAN) facilitate resource sharing and coordination across networks, creating new demands for efficient and trustworthy cross-operator sch... |
| Det-5G: Closing the Determinism Gap in 5G-Advanced for Industrial Closed-Loop Control | Adnan Aijaz | 2026-09-07 | 下载 | The ultra-reliable low-latency communication (uRLLC) capability of 5G has created significant opportunities for industrial wireless connectivity, yet widespread use of cellular networks for closed-loo... |
| Ollama in the Wild: A Longitudinal Measurement of Exposed Ollama LLM Endpoints at Internet Scale | Zuyao Xu, Xiang Li, Yuqi Qiu, Lu Sun | 2026-09-07 | 下载 | Self-hosted large language model (LLM) serving is emerging as a distinct category of Internet service, but we still know little about how these deployments appear and change on the public Internet. |
| ASTRA: Low-Overhead Runtime Architecture for STReam Adaptation in Video Analytics | Mahshid Ghasemi, Zoran Kostic, Javad Ghaderi, Gil Zussman | 2026-09-07 | 下载 | Real-time video analytics is crucial for smart city applications and cloud-connected vehicle control. To improve analytics accuracy, it is desirable to process the video at the highest resolution and ... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing Systems | Guangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso | 2026-09-07 | 下载 | Kolmogorov-Arnold Networks (KANs) replace the fixed activation functions and linear weights of Multi-Layer Perceptrons (MLPs) with learnable univariate functions on network edges, offering improved in... |
| TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge Learning | Avik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann | 2026-09-07 | 下载 | The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly ... |
| Mathematical Modeling of a Cognitive Continuum Digital Shadow for Large-Scale, Cross-Facility Workflows | Mark Asch, Marius Garénaux Gruau, François Bodin | 2026-09-07 | 下载 | We present the mathematical foundations of a \emph{Cognitive Continuum Digital Shadow} (CCDS), a decision-support layer between users and the cross-facility infrastructure---instruments, networks, dat... |