Skip to content

2026-09-07 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
A 28nm 27,648-Spin Multichip Digital Ising Accelerator with Pegasus ConnectivityTong Wu, Atharva Raut, Ian Khor, Ziyad Alswaidan, Ting-Yu Lee, Hongchang Kuang, Navid Anjum Aadit, Anand Raju, Dhruv Chinmay, Siddharth Das, Ken Mai, Kerem Camsari, Tathagata Srimani2026-09-07下载We present a 28nm digital Ising accelerator with 27,648 spins across four chips. A time-multiplexed spin-update array with local SRAM and scheduled interchip transfers delivers 41.5G updates/s at 1.
QROB: Quantifying Realization Overhead in Quantum Compilation via Reverse ConstructionJintao Li, Kaiqi Li, Rui Wang, Yilun Zhao, Kaixuan Huang, Ying Wang, Jialin Zhang, Zheng-An Wang, Xiaoming Sun, Heng Fan2026-09-07下载Quantum compilation reconciles a program's idealized interaction topology with hardware locality constraints, yet evaluations at scale lack calibrated references for realization overhead.
TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge LearningAvik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann2026-09-07下载The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly ...
Enabling High-Bandwidth Flash for Generative Recommendation Serving with Write-Aware KV Cache PolicyDanni Peng, Kai Wu, Tianyu Zuo, Pengfei Xia, Hui Zang2026-09-07下载Generative recommendation (GR) systems increasingly leverage user-level KV cache reuse to avoid recomputing long user histories. However, the growing KV cache capacity and bandwidth requirements intro...
NOVA-CIM: Noise- and Correlation-Tolerant Stochastic Interfaces for Analog Compute-in-MemoryJiachen Ren, Wenshuai Yao, Haobo Liu, Xincheng Feng, Chenxi Hu, Zhengwu Liu, Kechao Tang, Wenyong Zhou, Ngai Wong2026-09-07下载Analog compute-in-memory (CIM) enables energy-efficient model acceleration, but its reliance on ADC-based readout, which directly quantizes noisy column currents, makes inference accuracy highly sensi...
Capability-Gated Conformance Testing of Quantum Error-Correction Decoder LibrariesJiachen Shen, Hui Zhong2026-09-07下载A quantum error correction decoder is a library other people's results depend on, judged in one dominant way. Sample errors, decode, and count wrong logical observables.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
MicroIntent: Intent-Based Placement Strategy for Microservice Application in the Compute Continuum Using LLMsKoushikur Islam, Guilherme Da Cunha Rodrigues, Bahman Javadi, Rodrigo N. Calheiros2026-09-07下载The placement of microservices in the compute continuum plays a vital role in delivering services that comply with customers' needs, such as reduced latency, storage requirements, quality of service a...
Interactive Debugger for Performance Portable Python HPC KernelsIvan Grigorik, Gabriel Kosmacher, George Biros, Milos Gligoric2026-09-07下载We propose PKDB, the first interactive debugger for GPU and multithreaded low-level kernels written in Python. Python is widely used in high performance computing (HPC), with frameworks such as PyKokk...
Real-Time dApps for AI-RAN: Measured Interface Requirements for Inline PHY and Slot-Level ControlTimothy O'Shea, Matthew Pennybacker, Andriy Kharchenko2026-09-07下载Distributed applications (dApps) bring AI to the microsecond-to-millisecond band beside the 5G distributed unit (DU), but every public dApp framework realizes them the same way: an external process th...
Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing SystemsGuangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso2026-09-07下载Kolmogorov-Arnold Networks (KANs) replace the fixed activation functions and linear weights of Multi-Layer Perceptrons (MLPs) with learnable univariate functions on network edges, offering improved in...
Attestream: Usage-Aware Intermittent Data Distribution with Verifiable Lifecycle Provenance for Machine-Learning Data StreamsKentaro Oda2026-09-07下载Providers of continuously produced, commercially valuable data -- sensor streams, telemetry, and other feeds sold as machine-learning training material -- cannot observe whether delivered data is actu...
Fast Multidimensional Approximate Agreement with Optimal Resilience Using Ball ValidityTijana Milentijević, Stefan Schmid2026-09-07下载Multidimensional approximate agreement requires nn processes with inputs in Rd\mathbb{R}^d to output vectors close to each other, despite up to tt Byzantine faults.
GPU-Accelerated Hypergraph Partitioning and Placement to Map SNNs on Neuromorphic HardwareMarco Ronzani, Cristina Silvano2026-09-07下载SNNs running on neuromorphic hardware use spikes to achieve sparse and energy-efficient communication over a mesh of cores. In turn, system performance heavily depends on the assignment of neurons to ...
Analytical Resource Management for Fine-grained MoE Computation-Communication OverlapHongyu Liu, Minyu Cui, Miquel Pericas2026-09-07下载Fine-grained computation--communication overlap in distributed Mixture-of-Experts (MoE) inference allows communication to begin as partial compute results become ready.
From Bracha to Coded MBRB: Benchmarking Byzantine Reliable Broadcast ImplementationsYenan Wang, Jesper Kullberg, Fabian Paglianno Persson, Elad Michael Schiller, Timothé Albouy2026-09-07下载Byzantine Reliable Broadcast (BRB) and Message-Adversary-Tolerant Byzantine Reliable Broadcast (MBRB) are reliable-dissemination abstractions for fault-tolerant distributed systems.
PLATOS: A Power and Latency-Aware Task-Oriented Scheduling Strategy for Healthcare IoT in Fog ComputingMohammed Alaa Ala'anzy, Zulfiqar Ahmad, Zhanar Mukash2026-09-07下载Healthcare Internet of Things (HIoT) technology is revolutionising the healthcare industry by enabling real-time data collection and analysis for personalised patient care.
Robust Decentralized Personalized Federated Learning via Prediction-Constrained Neighborhood CollaborationXiao Ma, Hong Shen, Hui Tian, Wenqi Lyu, Wei Ke2026-09-07下载This paper proposes a robust decentralized personalized federated learning method R-DPFL, that enables clients to reduce the impact of Byzantine attacks via robust neighborhood direction estimation an...
The Maximum Mutual Visibility Set on a Cactus Graph and the Self-stabilizing ConstructionsYonghwan Kim, Yuichi Sudo2026-09-07下载Given a graph G=(V,E)G=(V,E), let SS (⊆V\subseteq V) be a set of vertices. Two vertices are \emph{mutually visible} if there exists a shortest path in GG between them that does not contain any other vert...
An Efficient Out-of-Core Tomographic Imaging Framework for Edge DevicesXuetao Chen, Cong Ma, Xiangyu Meng, Du Wu, Zhengyang Bai, Tao Luo, Zhaorui Zhang, Emmanuel Jeannot, Edgar Josafat Martinez Noriega, Xun Wang, Peng Chen, Amelie Chi Zhou, Mohamed Wahib2026-09-07下载Computed Tomography (CT) is an essential 3D imaging technology widely used in medical diagnostics and scientific research. However, performing CT imaging on edge devices is challenging due to limitati...
Parallelism Strategy Chaining for Fast Training ConvergenceMinchul Kang, Changyong Shin, Younghun Go, Hyunho Lee, Jinwoo Jeong, Chuck Yoo, Gyeongsik Yang2026-09-07下载Selecting a parallelism strategy - the configuration of data, tensor, and pipeline parallelism degrees together with micro- and global-batch sizes - largely determines the training efficiency of large...
Robust Decentralized Federated Distillation via Multi-Modality Knowledge CollaborationXiao Ma, Hong Shen, Hui Tian, Wei Ke, Wenqi Lyu2026-09-07下载This paper propose a robust decentralized federated distillation method that enables clients with heterogeneous models to collaborate through predictions on shared unlabeled public data.
Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-TrainingZili Wang, Zhaopeng Qiu, Yuekai Zhang, Shuang Yu, Junjie Lai2026-09-07下载Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft's accuracy, yielding great...
TreeRedux: Separating Concerns in Spark's Distributed Tree AggregationDavid A. G. Harrison, Ivan Cao2026-09-07下载By default, Apache Spark's tree aggregation primitives place the tree root on the driver, requiring the driver to participate in the same aggregation computation over intermediate aggregation state as...
Unified AI Gateway: A Framework for Joint Model Routing and KV Cache ManagementJiaxun Lu, Xiang Zhang, Yunfeng Shao2026-09-07下载Large language model (LLM) inference increasingly spans models that differ in size, capability, price, and provider. This shift creates two costs for developers.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Movable Antennas Enabled Wireless Powered Networks: Principles and TechnologiesZhendong Li, Yiran Zheng, Tianyu Li, Zhou Su, Wen Chen, Ying Wang2026-09-07下载As an emerging framework, movable antenna (MA)-enabled wireless powered networks (WPNs) have attracted growing attention. WPNs integrate wireless communication and energy transfer.
The OCUDU dApp Platform: An Open Runtime and E3 Interface for Real-Time AI-RANTimothy O'Shea, Matthew Pennybacker, Andriy Kharchenko2026-09-07下载Machine learning has shown its largest gains in the band below 10 ms inside a 3GPP new radio (NR) 5G distributed unit (DU): link adaptation, per-slot scheduling, channel estimation, and the receiver i...
Real-Time dApps for AI-RAN: Measured Interface Requirements for Inline PHY and Slot-Level ControlTimothy O'Shea, Matthew Pennybacker, Andriy Kharchenko2026-09-07下载Distributed applications (dApps) bring AI to the microsecond-to-millisecond band beside the 5G distributed unit (DU), but every public dApp framework realizes them the same way: an external process th...
How Bitcoin Forms Its Network: Peer-Table Sampling and Structural PropertiesTaki E. M. Abedesselam, Antonio Cruciani, Fabio Giacomelli, Lucianna Kiffer, Francesco Pasquale2026-09-07下载The global structure of P2P networks underlying modern cryptocurrencies is hidden by design: each node only knows its neighbors and maintains a local \textit{peer table} of IP addresses.
No-Regret Mixing of LRU and LFU with Optimal Switching CostYounes Ben Mazziane, Xinying Zou2026-09-07下载Caching systems often rely on simple eviction policies such as Least Recently Used (LRU) and Least Frequently Used (LFU), which perform well in complementary request regimes.
Perception-Aware Joint Power and Sub-Band Allocation for 6G In-Body SubnetworksSamira Abdelrahman, Hossam Farag2026-09-07下载In-body subnetworks (IBSs) are expected to become a key enabler of immersive eXtended Reality (XR) services in sixth-generation (6G) networks by providing ultra-short-range, low-latency wireless conne...
Blockchain-based Proportional Fair Scheduling for Multi-Operator O-RANKun Huang, Xintong Ling, Meining Wu, Jiaheng Wang, Zhi Ding, Xiqi Gao2026-09-07下载The openness and disaggregation of Open radio access network (O-RAN) facilitate resource sharing and coordination across networks, creating new demands for efficient and trustworthy cross-operator sch...
Det-5G: Closing the Determinism Gap in 5G-Advanced for Industrial Closed-Loop ControlAdnan Aijaz2026-09-07下载The ultra-reliable low-latency communication (uRLLC) capability of 5G has created significant opportunities for industrial wireless connectivity, yet widespread use of cellular networks for closed-loo...
Ollama in the Wild: A Longitudinal Measurement of Exposed Ollama LLM Endpoints at Internet ScaleZuyao Xu, Xiang Li, Yuqi Qiu, Lu Sun2026-09-07下载Self-hosted large language model (LLM) serving is emerging as a distinct category of Internet service, but we still know little about how these deployments appear and change on the public Internet.
ASTRA: Low-Overhead Runtime Architecture for STReam Adaptation in Video AnalyticsMahshid Ghasemi, Zoran Kostic, Javad Ghaderi, Gil Zussman2026-09-07下载Real-time video analytics is crucial for smart city applications and cloud-connected vehicle control. To improve analytics accuracy, it is desirable to process the video at the highest resolution and ...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Scalability Analysis of Distributed Kolmogorov-Arnold Network Training on High-Performance Computing SystemsGuangneng Chen, David Garcia Selfa, Pablo Quesada Barriuso2026-09-07下载Kolmogorov-Arnold Networks (KANs) replace the fixed activation functions and linear weights of Multi-Layer Perceptrons (MLPs) with learnable univariate functions on network edges, offering improved in...
TASTE: Throughput-Aware Batch Size Tuning for On-Device Edge LearningAvik Bhatnagar, Federico Nicolas Peccia, Oliver Bringmann2026-09-07下载The rise of privacy-preserving artificial intelligence (AI) has shifted the focus of model adaptation and personalization towards on-device learning, where deep learning models are finetuned directly ...
Mathematical Modeling of a Cognitive Continuum Digital Shadow for Large-Scale, Cross-Facility WorkflowsMark Asch, Marius Garénaux Gruau, François Bodin2026-09-07下载We present the mathematical foundations of a \emph{Cognitive Continuum Digital Shadow} (CCDS), a decision-support layer between users and the cross-facility infrastructure---instruments, networks, dat...

基于 VitePress 构建 · 使用本地搜索查找论文