Skip to content

2026-04-28 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
RAG-Enhanced Kernel-Based Heuristic Synthesis (RKHS): A Structured Methodology Using Large Language Models for Hardware DesignShiva Ahir, Alex Doboli2026-04-28下载Heuristic design upholds modern electronic design automation (EDA) tools, yet crafting effective placement, routing, and scheduling strategies entails substantial expertise.
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention ServingZhongkai Yu, Haotian Ye, Chenyang Zhou, Ohm Rishabh Venkatachalam, Zaifeng Pan, Zhengding Hu, Junsung Kim, Won Woo Ro, Po-An Tsai, Shuyi Pei, Yangwook Kang, Yufei Ding2026-04-28下载All current LLM serving systems place the GPU at the center, from production-level attention-FFN disaggregation to NVIDIA's Rubin GPU-LPU heterogeneous platform.
At the Edge of the Heart: ULP FPGA-Based CNN for On-Device Cardiac Feature Extraction in Smart Health Sensors for AstronautsKazi Mohammad Abidur Rahman, Davis Rakhshan, Philipp Lütke, Laura Harms, Ulf Kulau2026-04-28下载The convergence of accelerating human spaceflight ambitions and critical terrestrial health monitoring demands is driving unprecedented requirements for reliable, real-time feature extraction on extre...
NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM InferenceMingbo Hao, Changwei Yan, Haoyu Cui, Zhihao Yan, Yizhi Ding, Zhangrui Qian, Weiwei Shan2026-04-28下载The rapid growth of LLMs demands high-throughput, memory-capacity-intensive inference on resource-constrained edge devices, where single-batch decoding remains fundamentally memory-bound.
No Tile Left Behind: Multiprogramming for Surface-Code ArchitecturesArchisman Ghosh, Avimita Chatterjee, Swaroop Ghosh2026-04-28下载Fault-tolerant quantum computing (FTQC) is emerging as the architectural regime in which practical large-scale quantum workloads will execute.
TetrisG-SDK: Efficient Convolutional Layer Mapping with Adaptive Windows and Grouped Convolutions for Fast In-Memory ComputingKe Dong, Kejie Huang, Tao Luo, Bo Wang2026-04-28下载Shifted-and-Duplicated-Kernel (SDK) mapping has emerged as an effective strategy to accelerate convolutional layers on compute-in-memory (CIM) hardware. However, existing SDK variants (e.g.
RecFlash: Fast Recommendation System on In-Storage Computing with Frequency-Based Data MappingJangho Baik, Sunghyun Kim, Gisan Ji, Wonbo Shim, Sungju Ryu2026-04-28下载Recommendation system has gained a large popularity for a variety of personalized suggestion tasks, but the ever-increasing number of user data makes real-time processing of recommendation systems dif...
AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile DevicesMa Zirui, Fan Zhihua, Li Wenxing, Wu Haibin, Zhang Fulin, Ye Xiaochun, Li Wenming2026-04-28下载Speculative decoding enhances the inference efficiency of large language models (LLMs) by generating drafts using a small draft language model (DLM) and verifying them in batches with a large target l...
FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory ArchitectureZihao Xuan, Jia Chen, Yewen Li, Wei Xuan, Hegan Chen, Xiao Huo, Fengbin Tu2026-04-28下载In this paper, we propose FusionCIM, an operator-fusion-driven compute-in-memory (CIM) accelerator architecture for efficient and scalable LLM inference, with three key innovations: (1) a hybrid CIM p...
How Can Reinforcement Learning Achieve Expert-level Placement?Ruo-Tong Chen, Ke Xue, Chengrui Gao, Yunqi Shi, Tian Xu, Peng Xie, Siyuan Xu, Mingxuan Yuan, Chao Qian, Zhi-Hua Zhou2026-04-28下载Chip placement is a critical step in physical design. While reinforcement learning (RL)-based methods have recently emerged, their training primarily focuses on wirelength optimization, and therefore ...
Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM InferenceRobin Geens, Joran Heldens, Joren Dumoulin, Marian Verhelst2026-04-28下载Ternary weight quantization (e.g., BitNet b1.58) offers a promising path to mitigate the memory bandwidth bottleneck in Large Language Model (LLM) inference.
Agentic Architect: An Agentic AI Framework for Architecture Design Exploration and OptimizationAlexander Blasberg, Vasilis Kypriotis, Dimitrios Skarlatos2026-04-28下载Rapid advances in Large Language Models (LLMs) create new opportunities by enabling efficient exploration of broad, complex design spaces. This is particularly valuable in computer architecture, where...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention ServingZhongkai Yu, Haotian Ye, Chenyang Zhou, Ohm Rishabh Venkatachalam, Zaifeng Pan, Zhengding Hu, Junsung Kim, Won Woo Ro, Po-An Tsai, Shuyi Pei, Yangwook Kang, Yufei Ding2026-04-28下载All current LLM serving systems place the GPU at the center, from production-level attention-FFN disaggregation to NVIDIA's Rubin GPU-LPU heterogeneous platform.
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM InferenceShouxu Lin, Zhiyuan Guo, Jiaxin Lin2026-04-28下载LLM inference is constrained by GPU memory capacity and bandwidth. Tiered memory architectures mitigate this by allowing the GPU to offload memory to the remote tier.
RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-ExpertsVyom Sharma, Debajyoti Datta2026-04-28下载The optimal kernel configuration for Mixture-of-Experts (MoE) inference depends on both batch size and the expert routing distribution, yet production systems dispatch from batch size alone, leaving 1...
Pythia: Toward Predictability-Driven Agent-Native LLM ServingShan Yu, Junyi Shu, Yuanjiang Ni, Kun Qian, Xue Li, Yang Wang, Jinyuan Zhang, Ziyi Xu, Shuo Yang, Lingjun Zhu, Ennan Zhai, Qingda Lu, Jiarong Xing, Youyou Lu, Xin Jin, Xuanzhe Liu, Harry Xu2026-04-28下载As LLM applications grow more complex, developers are increasingly adopting multi-agent architectures to decompose workflows into specialized, collaborative components, introducing structure that cons...
SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed TransmissionCe Zheng, Xinghan Wang, Jiahong Ning, Yuxuan Shi, Ning Huang, Tingting Yang2026-04-28下载Federated inference enhances LLM performance in edge computing through weighted averaging of distributed model predictions. However, autoregressive LLM inference requires frequent full-model forward p...
Two Efficient Message-passing Exclusive Scan AlgorithmsJesper Larsson Träff2026-04-28下载Parallel scan primitives compute element-wise inclusive or exclusive prefix sums of input vectors contributed by pp consecutively ranked processors under an associative, possibly expensive, binary op...
Volitional Multiagent Atomic Transactions: Describing People and their MachinesAndy Lewis-Pye, Ehud Shapiro2026-04-28下载Formal models for concurrent and distributed systems describe machines; the people who operate them are either ignored or treated as external environment.
Economical and ecological impact of sector coupling applied to computing clustersP. Bechtle, O. Freyermuth, M. Geffers, M. Giffels, M. Hübner, F. Kirfel, J. Kreutz, S. Krieg, S. Matberg, M. Schnepf2026-04-28下载The rising share of abundant renewable energy inevitably increases volatility in the electricity production. The concept of sector coupling means that the volatility of electricity production to a lar...
CUDA Kernel Optimization and Counter-Free Performance Analysis for Depthwise Convolution in Cloud EnvironmentsHuriyeh Babak, Melanie Schaller2026-04-28下载Efficient GPU execution of convolution operators is governed by memory-access efficiency, on-chip data reuse, and execution mapping rather than arithmetic throughput alone.
Adaptive Management of Microservices in Dynamic Computing Environments: A Taxonomy and Future DirectionsMing Chen, Muhammed Tawfiqul Islam, Maria Rodriguez Read, Rajkumar Buyya2026-04-28下载Microservice-based cloud applications face changing workloads, evolving request paths, variable network conditions, interference, and failures.
CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache RestorationSean Nian, Jiahao Fang, Qilong Feng, Zhiyu Wu, Fan Lai2026-04-28下载KV cache restoration has emerged as a dominant bottleneck in serving long-context LLM workloads, including multi-turn conversations, retrieval-augmented generation, and agentic pipelines.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Application-Aware Twin-in-the-Loop Planning for Federated Split Learning over Wireless Edge NetworksZihao Ding, Beining Wu, Jun Huang, Shiwen Mao2026-04-28下载We investigate task-success-oriented resource allocation for federated split learning (FSL) at the wireless edge. In this setting, the server must jointly determine bandwidth, transmit power, split-la...
On the Role of Time Series Clustering in Traffic Matrix PredictionMartha Cash, Charlotte Fowler, Alexander M. Wyglinski2026-04-28下载This paper analyzes the role of time-series clustering in traffic matrix (TM) prediction. Traffic flows within a TM often exhibit heterogeneous behavior, which can reduce the effectiveness of global f...
NeuralEmu: in situ Measurement-Driven, ML-based, High-Fidelity 5G Network EmulationHaoran Wan, Yaxiong Xie, Kyle Jamieson2026-04-28下载Current and future applications demand ultra-low latency and consistent throughput, yet frequently traverse 5G cellular networks, so cope with volatile packet dynamics, as 5G base station schedulers d...
Decoding Delay Guarantees of Space Regulated Multiple Access Random Wireless Networks using Successive Interference CancellationKevin Zagalo, Jean-Marie Gorce, François Baccelli2026-04-28下载This paper is focused on decoding delay guarantees in wireless networks, where messages have a given signal-to-interference-plus-noise ratio threshold η_0 to meet in order to be successfully decoded...
Slice Agent: Identifying and Isolating Slices in Shared Open Radio UnitFelipe Arnholda, Flavio Rocha, Lucio Prade, Cristiano Bonato Both2026-04-28下载Network Slice as a Service (NSaaS) is a key enabler of Beyond Fifth Generation (5G) and Sixth Generation (6G) networks, supporting next-generation applications such as extended reality (XR), immersive...
EOS-Bench: A Comprehensive Benchmark for Earth Observation Satellite SchedulingQian Yin, Jiaxing Li, Jiaqi Cheng, Qizhang Luo, Annalisa Riccardi, Abhijit Chatterjee, Rafael Vazquez, Carlo Novara, Michalis Mavrovouniotis, Ponnuthurai Nagaratnam Suganthan, Shengzhou Bai, Xiaoxuan Hu, Lining Xing, Ming Xu, Shuang Li, Zixuan Zheng, Xin Shen, Xiaoyu Chen, Yi Gu, Yanjie Song, Witold Pedrycz, Evan L. Kramer, Laio Oriel Seman, Cletah Shoko, Guohua Wu, Xinwei Wang2026-04-28下载Earth observation satellite imaging scheduling is a challenging NP-hard combinatorial optimisation problem central to space mission operations.
Chorusing Synchronization Signals for Ambient 5G BackscatterYunyun Feng, Chenhong Cao, Si Chen, Wei Gong2026-04-28下载5G backscatter communication presents an emerging energy-efficient IoT connectivity solution with enhanced availability and data rate advantages over traditional wireless networks.
Design Insights into Partition Placement and Routing for DNN Inference in Multi-Hop Edge NetworksJinkun Zhang, Poonam Yadav2026-04-28下载Partitioned DNN inference is a promising approach for latency-sensitive intelligent services in edge networks, since it allows different parts of a model to be executed across end devices, edge server...
Assistants, Not Architects: The Role of LLMs in Networked Systems DesignPratyush Sahu, Rahul Bothra, Venkat Arun, Brighten Godfrey, Akshay Narayan, Ahmed Saeed2026-04-28下载Designing the architecture of modern networked systems requires navigating a large, combinatorial space of hardware, systems, and configuration choices with complex cross-layer interactions.
Probing for Better Age of Information in Energy-Harvesting Random Access NetworksZiyi Li, Fangming Zhao, Howard H. Yang2026-04-28下载In this paper, we investigate the impact of channel probing and reservation on the Age of Information (AoI) in energy-harvesting (EH) random access networks, where each source relies solely on harvest...
Digital Twin-assisted belief-state reinforcement learning for latency-robust ISAC in 6G networksHimanshu Tiwari, Binayak Kar, Priyanshu Tiwari2026-04-28下载Integrated Sensing and Communication (ISAC) enables joint data transmission and environmental perception for sixth-generation (6G) networks, but centralized and virtualized RAN control loops introduce...
Optimization of Model Splitting, Placement, and Chaining for Multi-hop Split Learning and InferenceTakanori Hara, Masahiro Sasabe2026-04-28下载Service Function Chaining (SFC) establishes efficient communication paths by ensuring that traffic traverses a predefined sequence of network functions in a specified order to meet particular service ...
Quantum-enhanced Network TomographyYufei Zheng, Zihao Gong, Saikat Guha, Don Towsley2026-04-28下载Network tomography refers to the use of inference techniques for inferring internal network states from end-to-end probes. Quantum probes, implemented by sending blocks of nn coherent-state pulses au...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Embedded Rust or C Firmware? Lessons from an Industrial Microcontroller Use Case with Ariel OSBipin Thapa, Daniele Alfonso, Lorenzo Bini, Licio Mapelli, Kaspar Schleiser, Romain Fouquet, Emmanuel Baccelli2026-04-28下载As Rust gains traction for developing safer systems software, a reality check for the microcontroller hardware segment becomes necessary. How ready is the Rust ecosystem for this segment? Can Rust com...

基于 VitePress 构建 · 使用本地搜索查找论文