2026-04-28
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| RAG-Enhanced Kernel-Based Heuristic Synthesis (RKHS): A Structured Methodology Using Large Language Models for Hardware Design | Shiva Ahir, Alex Doboli | 2026-04-28 | 下载 | Heuristic design upholds modern electronic design automation (EDA) tools, yet crafting effective placement, routing, and scheduling strategies entails substantial expertise. |
| AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving | Zhongkai Yu, Haotian Ye, Chenyang Zhou, Ohm Rishabh Venkatachalam, Zaifeng Pan, Zhengding Hu, Junsung Kim, Won Woo Ro, Po-An Tsai, Shuyi Pei, Yangwook Kang, Yufei Ding | 2026-04-28 | 下载 | All current LLM serving systems place the GPU at the center, from production-level attention-FFN disaggregation to NVIDIA's Rubin GPU-LPU heterogeneous platform. |
| At the Edge of the Heart: ULP FPGA-Based CNN for On-Device Cardiac Feature Extraction in Smart Health Sensors for Astronauts | Kazi Mohammad Abidur Rahman, Davis Rakhshan, Philipp Lütke, Laura Harms, Ulf Kulau | 2026-04-28 | 下载 | The convergence of accelerating human spaceflight ambitions and critical terrestrial health monitoring demands is driving unprecedented requirements for reliable, real-time feature extraction on extre... |
| NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference | Mingbo Hao, Changwei Yan, Haoyu Cui, Zhihao Yan, Yizhi Ding, Zhangrui Qian, Weiwei Shan | 2026-04-28 | 下载 | The rapid growth of LLMs demands high-throughput, memory-capacity-intensive inference on resource-constrained edge devices, where single-batch decoding remains fundamentally memory-bound. |
| No Tile Left Behind: Multiprogramming for Surface-Code Architectures | Archisman Ghosh, Avimita Chatterjee, Swaroop Ghosh | 2026-04-28 | 下载 | Fault-tolerant quantum computing (FTQC) is emerging as the architectural regime in which practical large-scale quantum workloads will execute. |
| TetrisG-SDK: Efficient Convolutional Layer Mapping with Adaptive Windows and Grouped Convolutions for Fast In-Memory Computing | Ke Dong, Kejie Huang, Tao Luo, Bo Wang | 2026-04-28 | 下载 | Shifted-and-Duplicated-Kernel (SDK) mapping has emerged as an effective strategy to accelerate convolutional layers on compute-in-memory (CIM) hardware. However, existing SDK variants (e.g. |
| RecFlash: Fast Recommendation System on In-Storage Computing with Frequency-Based Data Mapping | Jangho Baik, Sunghyun Kim, Gisan Ji, Wonbo Shim, Sungju Ryu | 2026-04-28 | 下载 | Recommendation system has gained a large popularity for a variety of personalized suggestion tasks, but the ever-increasing number of user data makes real-time processing of recommendation systems dif... |
| AHASD: Asynchronous Heterogeneous Architecture for LLM Adaptive Drafting Speculative Decoding on Mobile Devices | Ma Zirui, Fan Zhihua, Li Wenxing, Wu Haibin, Zhang Fulin, Ye Xiaochun, Li Wenming | 2026-04-28 | 下载 | Speculative decoding enhances the inference efficiency of large language models (LLMs) by generating drafts using a small draft language model (DLM) and verifying them in batches with a large target l... |
| FusionCIM: Accelerating LLM Inference with Fusion-Driven Computing-in-Memory Architecture | Zihao Xuan, Jia Chen, Yewen Li, Wei Xuan, Hegan Chen, Xiao Huo, Fengbin Tu | 2026-04-28 | 下载 | In this paper, we propose FusionCIM, an operator-fusion-driven compute-in-memory (CIM) accelerator architecture for efficient and scalable LLM inference, with three key innovations: (1) a hybrid CIM p... |
| How Can Reinforcement Learning Achieve Expert-level Placement? | Ruo-Tong Chen, Ke Xue, Chengrui Gao, Yunqi Shi, Tian Xu, Peng Xie, Siyuan Xu, Mingxuan Yuan, Chao Qian, Zhi-Hua Zhou | 2026-04-28 | 下载 | Chip placement is a critical step in physical design. While reinforcement learning (RL)-based methods have recently emerged, their training primarily focuses on wirelength optimization, and therefore ... |
| Hardware Generation and Exploration of Lookup Table-Based Accelerators for 1.58-bit LLM Inference | Robin Geens, Joran Heldens, Joren Dumoulin, Marian Verhelst | 2026-04-28 | 下载 | Ternary weight quantization (e.g., BitNet b1.58) offers a promising path to mitigate the memory bandwidth bottleneck in Large Language Model (LLM) inference. |
| Agentic Architect: An Agentic AI Framework for Architecture Design Exploration and Optimization | Alexander Blasberg, Vasilis Kypriotis, Dimitrios Skarlatos | 2026-04-28 | 下载 | Rapid advances in Large Language Models (LLMs) create new opportunities by enabling efficient exploration of broad, complex design spaces. This is particularly valuable in computer architecture, where... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AMMA: A Multi-Chiplet Memory-Centric Architecture for Low-Latency 1M Context Attention Serving | Zhongkai Yu, Haotian Ye, Chenyang Zhou, Ohm Rishabh Venkatachalam, Zaifeng Pan, Zhengding Hu, Junsung Kim, Won Woo Ro, Po-An Tsai, Shuyi Pei, Yangwook Kang, Yufei Ding | 2026-04-28 | 下载 | All current LLM serving systems place the GPU at the center, from production-level attention-FFN disaggregation to NVIDIA's Rubin GPU-LPU heterogeneous platform. |
| DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference | Shouxu Lin, Zhiyuan Guo, Jiaxin Lin | 2026-04-28 | 下载 | LLM inference is constrained by GPU memory capacity and bandwidth. Tiered memory architectures mitigate this by allowing the GPU to offload memory to the remote tier. |
| RaMP: Runtime-Aware Megakernel Polymorphism for Mixture-of-Experts | Vyom Sharma, Debajyoti Datta | 2026-04-28 | 下载 | The optimal kernel configuration for Mixture-of-Experts (MoE) inference depends on both batch size and the expert routing distribution, yet production systems dispatch from batch size alone, leaving 1... |
| Pythia: Toward Predictability-Driven Agent-Native LLM Serving | Shan Yu, Junyi Shu, Yuanjiang Ni, Kun Qian, Xue Li, Yang Wang, Jinyuan Zhang, Ziyi Xu, Shuo Yang, Lingjun Zhu, Ennan Zhai, Qingda Lu, Jiarong Xing, Youyou Lu, Xin Jin, Xuanzhe Liu, Harry Xu | 2026-04-28 | 下载 | As LLM applications grow more complex, developers are increasingly adopting multi-agent architectures to decompose workflows into specialized, collaborative components, introducing structure that cons... |
| SpecFed: Accelerating Federated LLM Inference with Speculative Decoding and Compressed Transmission | Ce Zheng, Xinghan Wang, Jiahong Ning, Yuxuan Shi, Ning Huang, Tingting Yang | 2026-04-28 | 下载 | Federated inference enhances LLM performance in edge computing through weighted averaging of distributed model predictions. However, autoregressive LLM inference requires frequent full-model forward p... |
| Two Efficient Message-passing Exclusive Scan Algorithms | Jesper Larsson Träff | 2026-04-28 | 下载 | Parallel scan primitives compute element-wise inclusive or exclusive prefix sums of input vectors contributed by consecutively ranked processors under an associative, possibly expensive, binary op... |
| Volitional Multiagent Atomic Transactions: Describing People and their Machines | Andy Lewis-Pye, Ehud Shapiro | 2026-04-28 | 下载 | Formal models for concurrent and distributed systems describe machines; the people who operate them are either ignored or treated as external environment. |
| Economical and ecological impact of sector coupling applied to computing clusters | P. Bechtle, O. Freyermuth, M. Geffers, M. Giffels, M. Hübner, F. Kirfel, J. Kreutz, S. Krieg, S. Matberg, M. Schnepf | 2026-04-28 | 下载 | The rising share of abundant renewable energy inevitably increases volatility in the electricity production. The concept of sector coupling means that the volatility of electricity production to a lar... |
| CUDA Kernel Optimization and Counter-Free Performance Analysis for Depthwise Convolution in Cloud Environments | Huriyeh Babak, Melanie Schaller | 2026-04-28 | 下载 | Efficient GPU execution of convolution operators is governed by memory-access efficiency, on-chip data reuse, and execution mapping rather than arithmetic throughput alone. |
| Adaptive Management of Microservices in Dynamic Computing Environments: A Taxonomy and Future Directions | Ming Chen, Muhammed Tawfiqul Islam, Maria Rodriguez Read, Rajkumar Buyya | 2026-04-28 | 下载 | Microservice-based cloud applications face changing workloads, evolving request paths, variable network conditions, interference, and failures. |
| CacheFlow: Efficient LLM Serving with 3D-Parallel KV Cache Restoration | Sean Nian, Jiahao Fang, Qilong Feng, Zhiyu Wu, Fan Lai | 2026-04-28 | 下载 | KV cache restoration has emerged as a dominant bottleneck in serving long-context LLM workloads, including multi-turn conversations, retrieval-augmented generation, and agentic pipelines. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Application-Aware Twin-in-the-Loop Planning for Federated Split Learning over Wireless Edge Networks | Zihao Ding, Beining Wu, Jun Huang, Shiwen Mao | 2026-04-28 | 下载 | We investigate task-success-oriented resource allocation for federated split learning (FSL) at the wireless edge. In this setting, the server must jointly determine bandwidth, transmit power, split-la... |
| On the Role of Time Series Clustering in Traffic Matrix Prediction | Martha Cash, Charlotte Fowler, Alexander M. Wyglinski | 2026-04-28 | 下载 | This paper analyzes the role of time-series clustering in traffic matrix (TM) prediction. Traffic flows within a TM often exhibit heterogeneous behavior, which can reduce the effectiveness of global f... |
| NeuralEmu: in situ Measurement-Driven, ML-based, High-Fidelity 5G Network Emulation | Haoran Wan, Yaxiong Xie, Kyle Jamieson | 2026-04-28 | 下载 | Current and future applications demand ultra-low latency and consistent throughput, yet frequently traverse 5G cellular networks, so cope with volatile packet dynamics, as 5G base station schedulers d... |
| Decoding Delay Guarantees of Space Regulated Multiple Access Random Wireless Networks using Successive Interference Cancellation | Kevin Zagalo, Jean-Marie Gorce, François Baccelli | 2026-04-28 | 下载 | This paper is focused on decoding delay guarantees in wireless networks, where messages have a given signal-to-interference-plus-noise ratio threshold η_0 to meet in order to be successfully decoded... |
| Slice Agent: Identifying and Isolating Slices in Shared Open Radio Unit | Felipe Arnholda, Flavio Rocha, Lucio Prade, Cristiano Bonato Both | 2026-04-28 | 下载 | Network Slice as a Service (NSaaS) is a key enabler of Beyond Fifth Generation (5G) and Sixth Generation (6G) networks, supporting next-generation applications such as extended reality (XR), immersive... |
| EOS-Bench: A Comprehensive Benchmark for Earth Observation Satellite Scheduling | Qian Yin, Jiaxing Li, Jiaqi Cheng, Qizhang Luo, Annalisa Riccardi, Abhijit Chatterjee, Rafael Vazquez, Carlo Novara, Michalis Mavrovouniotis, Ponnuthurai Nagaratnam Suganthan, Shengzhou Bai, Xiaoxuan Hu, Lining Xing, Ming Xu, Shuang Li, Zixuan Zheng, Xin Shen, Xiaoyu Chen, Yi Gu, Yanjie Song, Witold Pedrycz, Evan L. Kramer, Laio Oriel Seman, Cletah Shoko, Guohua Wu, Xinwei Wang | 2026-04-28 | 下载 | Earth observation satellite imaging scheduling is a challenging NP-hard combinatorial optimisation problem central to space mission operations. |
| Chorusing Synchronization Signals for Ambient 5G Backscatter | Yunyun Feng, Chenhong Cao, Si Chen, Wei Gong | 2026-04-28 | 下载 | 5G backscatter communication presents an emerging energy-efficient IoT connectivity solution with enhanced availability and data rate advantages over traditional wireless networks. |
| Design Insights into Partition Placement and Routing for DNN Inference in Multi-Hop Edge Networks | Jinkun Zhang, Poonam Yadav | 2026-04-28 | 下载 | Partitioned DNN inference is a promising approach for latency-sensitive intelligent services in edge networks, since it allows different parts of a model to be executed across end devices, edge server... |
| Assistants, Not Architects: The Role of LLMs in Networked Systems Design | Pratyush Sahu, Rahul Bothra, Venkat Arun, Brighten Godfrey, Akshay Narayan, Ahmed Saeed | 2026-04-28 | 下载 | Designing the architecture of modern networked systems requires navigating a large, combinatorial space of hardware, systems, and configuration choices with complex cross-layer interactions. |
| Probing for Better Age of Information in Energy-Harvesting Random Access Networks | Ziyi Li, Fangming Zhao, Howard H. Yang | 2026-04-28 | 下载 | In this paper, we investigate the impact of channel probing and reservation on the Age of Information (AoI) in energy-harvesting (EH) random access networks, where each source relies solely on harvest... |
| Digital Twin-assisted belief-state reinforcement learning for latency-robust ISAC in 6G networks | Himanshu Tiwari, Binayak Kar, Priyanshu Tiwari | 2026-04-28 | 下载 | Integrated Sensing and Communication (ISAC) enables joint data transmission and environmental perception for sixth-generation (6G) networks, but centralized and virtualized RAN control loops introduce... |
| Optimization of Model Splitting, Placement, and Chaining for Multi-hop Split Learning and Inference | Takanori Hara, Masahiro Sasabe | 2026-04-28 | 下载 | Service Function Chaining (SFC) establishes efficient communication paths by ensuring that traffic traverses a predefined sequence of network functions in a specified order to meet particular service ... |
| Quantum-enhanced Network Tomography | Yufei Zheng, Zihao Gong, Saikat Guha, Don Towsley | 2026-04-28 | 下载 | Network tomography refers to the use of inference techniques for inferring internal network states from end-to-end probes. Quantum probes, implemented by sending blocks of coherent-state pulses au... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Embedded Rust or C Firmware? Lessons from an Industrial Microcontroller Use Case with Ariel OS | Bipin Thapa, Daniele Alfonso, Lorenzo Bini, Licio Mapelli, Kaspar Schleiser, Romain Fouquet, Emmanuel Baccelli | 2026-04-28 | 下载 | As Rust gains traction for developing safer systems software, a reality check for the microcontroller hardware segment becomes necessary. How ready is the Rust ecosystem for this segment? Can Rust com... |