2026-04-29
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code Generation | Mahshid Rezakhani, Nowfel Mashnoor, Kimia Azar, Hadi Kamali | 2026-04-29 | 下载 | As large language models (LLMs) are increasingly fine-tuned for hardware tasks like RTL code generation, the scarcity of high-quality datasets often leads to the use of rapidly assembled or generated ... |
| Recent Advances in mm-Wave and Sub-THz/THz Oscillators for FutureG Technologies | Baktash Behmanesh, Ahmad Rezvanitabar | 2026-04-29 | 下载 | This paper provides a concise yet comprehensive review of recent advancements in millimeter-wave (mm-wave) oscillators below 100 GHz and sub-terahertz (sub-THz/THz) oscillators above 100 GHz for next-... |
| Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel | Yiqi Liu, Noelle Crawford, Michael Wang, Jilong Xue, Jian Huang | 2026-04-29 | 下载 | To overcome the well-known memory bottleneck of AI chips, 3D stacked architectures that employ advanced packaging technology with high-density through-silicon vias (TSVs) pins have proven to be a prom... |
| Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication Accelerators | Hyunsung Yoon, Sungju Ryu, Jae-Joon Kim | 2026-04-29 | 下载 | As the size of Deep Neural Networks (DNNs) increases dramatically to achieve high accuracy, the DNNs require a large amount of computations and memory footprint. |
| Verification and Validation (V&V)-in-the-Loop for RISC-V Design: The Holistic Vision of BZL | Sajjad Ahmed, Alexander Kropotov, Roberto Ignacio Genovese, Bernat Homs, Eloi Merino, Francesco Urbani, Henrique Yano, Iván Díaz, Joan Gracia Fernandez, Matteo Toselli, Muhammad Imran, Muhammad Abu Bakar Umar Haider Iqbal, Nadeem Yaseen, Quswar Abid, Shaista Cheema, Samuel Sanchez, Daniel Garcia, Joan Cabré, Mostafa Elyasi, Fernando Ayats, Miquel Moreto, Teresa Cervero, Oscar Palomar, Behzad Salami | 2026-04-29 | 下载 | The Barcelona Zetascale Lab (BZL) project aims to strengthening Europe's capacity in the design and manufacture of RISC-V based high-performance computing chips. |
| EMiX: Emulating Beyond Single-FPGA Limits | Alexander Kropotov, Miquel Moreto, Behzad Salami | 2026-04-29 | 下载 | FPGA-level emulation is a key step in pre-silicon chip design validation. However, emulating large-scale multi-core systems increasingly exceed the hardware resource capacity of a single FPGA, limitin... |
| Efficient, VRAM-Constrained xLM Inference on Clients | Aditya Ukarande, Deep Shekhar, Marc Blackstein, Ram Rangan | 2026-04-29 | 下载 | To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language models (LLMs) and vision language models (VLMs), joi... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Real-Time GPU-Accelerated Monte Carlo Evaluation of Safety-Critical AEB Systems Under Uncertainty | Akshay Karjol, Shadi Alawneh | 2026-04-29 | 下载 | Automatic Emergency Braking (AEB) systems represent a safety-critical national interest, with the National Highway Traffic Safety Administration (NHTSA) Federal Motor Vehicle Safety Standard (FMVSS No... |
| End-to-End and Phase-Level Performance Optimization for Hyperledger Fabric | Pavan Sollu, Aniruddha Mukherjee, Divya Pulivarthi, S. R. Eshwar, Gugan Thoppe, Kshitij Pratihast, Tittu Varghese, Hrishikesh Nashikkar, Yogesh Simmhan | 2026-04-29 | 下载 | Hyperledger Fabric (HLF) is a modular, permissioned blockchain widely adopted in enterprise settings. Enhancing its throughput and latency remains challenging, as optimization decisions made in one ph... |
| AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism | Ahan Gupta, Zhihao Wang, Neel Dani, Masahiro Tanaka, Olatunji Ruwase, Minjia Zhang | 2026-04-29 | 下载 | Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of thousands of tokens. |
| Efficient Training on Multiple Consumer GPUs with RoundPipe | Yibin Luo, Shiwei Gao, Huichuan Zheng, Youyou Lu, Jiwu Shu | 2026-04-29 | 下载 | Fine-tuning Large Language Models (LLMs) on consumer-grade GPUs is highly cost-effective, yet constrained by limited GPU memory and slow PCIe interconnects. |
| Adaptive Self-Organization in Anonymous Dynamic Networks | Garrett Parzych, Joshua J. Daymude | 2026-04-29 | 下载 | We introduce the problem of adaptive self-organization in which the nodes of an anonymous, synchronous dynamic network must distributively change the collective distribution of their responses (or "co... |
| FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts Serving | Minghe Wang, Trever Schirmer, Mohammadreza Malekabbasi, David Bermbach | 2026-04-29 | 下载 | Mixture-of-Experts (MoE) models offer high capacity with efficient inference cost by activating a small subset of expert models per input. However, deploying MoE models requires all experts to reside ... |
| A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPC | Petter Sandås, Íñigo Aréjula-Aísa, Sergio Iserte, Antonio J. Peña | 2026-04-29 | 下载 | High-performance computing (HPC) systems are increasingly exploring dynamic resource management and malleable MPI applications to better adapt to heterogeneous architectures, fluctuating workloads, an... |
| Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel | Yiqi Liu, Noelle Crawford, Michael Wang, Jilong Xue, Jian Huang | 2026-04-29 | 下载 | To overcome the well-known memory bottleneck of AI chips, 3D stacked architectures that employ advanced packaging technology with high-density through-silicon vias (TSVs) pins have proven to be a prom... |
| A Semantic Quantum Circuit Cache for Scalable and Distributed Quantum-Classical Workflows | Mar Tejedor, Javier Conejero, Rosa M. Badia | 2026-04-29 | 下载 | Hybrid quantum--classical workflows often execute large ensembles of circuits that differ syntactically but implement identical operations, leading to substantial redundant computation. |
| COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model Training | Akhmed Sakip, Erland Hilman Fuadi, Omar Sayedelahl, Zonghang Li, Jianshu She, Alham Fikri Aji, Steve Liu, Eric Xing, Qirong Ho | 2026-04-29 | 下载 | Training large language models requires jointly configuring two interdependent aspects of the system: the global batch size, which governs statistical efficiency, and the 3D parallelism strategy, whic... |
| FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow | Sina Heidari, Dimitrios S. Nikolopoulos | 2026-04-29 | 下载 | Deep learning compilers and vendor libraries deliver strong baseline performance but are bounded by finite, engineer-curated catalogs. When these omit needed optimizations, practitioners substitute ha... |
| DMRlib: Easy-coding and Efficient Resource Management for Job Malleability | Sergio Iserte, Rafael Mayo, Enrique S. Quintana-Ortí, Antonio J. Peña | 2026-04-29 | 下载 | Process malleability has proved to have a highly positive impact on the resource utilization and global productivity in data centers compared with the conventional static resource allocation policy. |
| MPI Malleability Validation under Replayed Real-World HPC Conditions | S. Iserte, M. Madon, G. Da, J. Pierson, A. J. Peña | 2026-04-29 | 下载 | Dynamic Resource Management (DRM) techniques can be leveraged to maximize throughput and resource utilization in computational clusters. Although DRM has been extensively studied through analytical wo... |
| DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference | Bodon Jeong, Hongsu Byun, Youngjae Kim, Weikuan Yu, Kyungkeun Lee, Jihoon Yang, Sungyong Park | 2026-04-29 | 下载 | The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key challenge arises from Key-Value (KV) caches, which of... |
| FloatSOM: GPU-Accelerated, Distributed, Topology-Flexible Self-Organizing Maps | Tony Xu, Sarah Klamt, Katherine Turner, Anne Brustle, Felix Marsh-Wakefield, Givanna Putri | 2026-04-29 | 下载 | GPU-accelerated Self-Organizing Map (SOM) implementations are among the most competitive options for large-scale SOM analysis, but growing dataset sizes increasingly challenge their practical use beca... |
| Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language Models | Cyril Shih-Huan Hsu, Wig Yuan-Cheng Cheng, Chrysa Papagianni | 2026-04-29 | 下载 | Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed the capabilities of resource-constrained embedded pl... |
| SplitFT: An Adaptive Federated Split Learning System For LLMs Fine-Tuning | Yimeng Shan, Zhaorui Zhang, Sheng Di, Yu Liu, Xiaoyi Lu, Benben Liu | 2026-04-29 | 下载 | Federated Split Learning has been identified as an efficient approach to address the computational resource constraints of clients in classical federated learning, while guaranteeing data privacy for ... |
| Efficient, VRAM-Constrained xLM Inference on Clients | Aditya Ukarande, Deep Shekhar, Marc Blackstein, Ram Rangan | 2026-04-29 | 下载 | To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language models (LLMs) and vision language models (VLMs), joi... |
| Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference | Vasu Shyam, Anna Golubeva, Quentin Anthony | 2026-04-29 | 下载 | We present tensor and sequence parallelism (TSP), a parallel execution strategy that folds tensor parallelism and sequence parallelism onto a single device axis. |
| DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training | Tianhao Hu, Xiangcheng Liu, Youshao Xiao, Yang Zheng, Xuan Huang, Jinrui Ding, Yufei Zhang, Tao Liang, Hongyu Zang, Quan Chen, Yueqing Sun, Wenjie Shi, Chao Zhang, Wei Wang, Qi Gu, Yerui Sun, Yucheng Xie, Xunliang Cai | 2026-04-29 | 下载 | Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed generation: long-taile... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| BLINC: Context-Specific Causal Learning for Automated RAN Configuration | Reshma Prasad, Michele Polese, Tommaso Melodia | 2026-04-29 | 下载 | Radio Access Network (RAN) configuration has traditionally required significant manual effort due to indirect causal dependencies between observable Key Performance Indicators (KPIs), and context-depe... |
| A 3GPP Perspective on Spectrum Sharing for the 5G-to-6G Migration: From DSS to MRSS | Xingqin Lin | 2026-04-29 | 下载 | Dynamic spectrum sharing (DSS) played an important role in the 4G-to-5G transition by allowing 5G new radio (NR) to enter valuable legacy spectrum without immediate static refarming. |
| Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language Models | Cyril Shih-Huan Hsu, Wig Yuan-Cheng Cheng, Chrysa Papagianni | 2026-04-29 | 下载 | Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed the capabilities of resource-constrained embedded pl... |
| SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering Tasks | Jiao Chen, Jianhua Tang, Xiaotong Yang, Zuohong Lv | 2026-04-29 | 下载 | AI coding agents demonstrate strong performance on general-purpose software benchmarks. However, their ability to handle 5G network engineering tasks remains unexplored. |
| StreamGuard: Exploring a 5G Architecture for Efficient, Quality of Experience-Aware Video Conferencing | Xuyang Cao, Oliver Michel, Kyle Jamieson | 2026-04-29 | 下载 | Video conferencing over 5G is increasingly prevalent, yet its Quality of Experience (QoE) often degrades under limited radio resources. This has two causes: 5G networks must serve many users, while in... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent Operations | Timothy Flavin, Sandip Sen | 2026-04-29 | 下载 | Reinforcement Learning (RL) algorithms exhibit high sample complexity, particularly when applied to Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs). |
| AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism | Ahan Gupta, Zhihao Wang, Neel Dani, Masahiro Tanaka, Olatunji Ruwase, Minjia Zhang | 2026-04-29 | 下载 | Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of thousands of tokens. |
| Revealing NVIDIA Closed-Source Driver Command Streams for CPU-GPU Runtime Behavior Insight | Yuang Yan, Ian Karlin, Ryan Grant | 2026-04-29 | 下载 | For NVIDIA GPUs, CUDA is the primary interface through which applications orchestrate GPU execution, yet much of the logic that realizes CUDA operations resides in NVIDIA's closed-source userspace dri... |
| What Is the Cost of Energy Monitoring? An Empirical Study on the Overhead of RAPL-Based Tools | Jeremy Diamond, Vincenzo Stoico | 2026-04-29 | 下载 | The Running Average Power Limit (RAPL) interface is widely used to estimate software energy consumption via CPU and DRAM counters, but tool design differences and high-frequency polling can introduce ... |
| FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow | Sina Heidari, Dimitrios S. Nikolopoulos | 2026-04-29 | 下载 | Deep learning compilers and vendor libraries deliver strong baseline performance but are bounded by finite, engineer-curated catalogs. When these omit needed optimizations, practitioners substitute ha... |
| DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM Inference | Bodon Jeong, Hongsu Byun, Youngjae Kim, Weikuan Yu, Kyungkeun Lee, Jihoon Yang, Sungyong Park | 2026-04-29 | 下载 | The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key challenge arises from Key-Value (KV) caches, which of... |