Skip to content

2026-04-29 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
SafeTune: Mitigating Data Poisoning in LLM Fine-Tuning for RTL Code GenerationMahshid Rezakhani, Nowfel Mashnoor, Kimia Azar, Hadi Kamali2026-04-29下载As large language models (LLMs) are increasingly fine-tuned for hardware tasks like RTL code generation, the scarcity of high-quality datasets often leads to the use of rapidly assembled or generated ...
Recent Advances in mm-Wave and Sub-THz/THz Oscillators for FutureG TechnologiesBaktash Behmanesh, Ahmad Rezvanitabar2026-04-29下载This paper provides a concise yet comprehensive review of recent advancements in millimeter-wave (mm-wave) oscillators below 100 GHz and sub-terahertz (sub-THz/THz) oscillators above 100 GHz for next-...
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with VoxelYiqi Liu, Noelle Crawford, Michael Wang, Jilong Xue, Jian Huang2026-04-29下载To overcome the well-known memory bottleneck of AI chips, 3D stacked architectures that employ advanced packaging technology with high-density through-silicon vias (TSVs) pins have proven to be a prom...
Sparse-on-Dense: Area and Energy-Efficient Computing of Sparse Neural Networks on Dense Matrix Multiplication AcceleratorsHyunsung Yoon, Sungju Ryu, Jae-Joon Kim2026-04-29下载As the size of Deep Neural Networks (DNNs) increases dramatically to achieve high accuracy, the DNNs require a large amount of computations and memory footprint.
Verification and Validation (V&V)-in-the-Loop for RISC-V Design: The Holistic Vision of BZLSajjad Ahmed, Alexander Kropotov, Roberto Ignacio Genovese, Bernat Homs, Eloi Merino, Francesco Urbani, Henrique Yano, Iván Díaz, Joan Gracia Fernandez, Matteo Toselli, Muhammad Imran, Muhammad Abu Bakar Umar Haider Iqbal, Nadeem Yaseen, Quswar Abid, Shaista Cheema, Samuel Sanchez, Daniel Garcia, Joan Cabré, Mostafa Elyasi, Fernando Ayats, Miquel Moreto, Teresa Cervero, Oscar Palomar, Behzad Salami2026-04-29下载The Barcelona Zetascale Lab (BZL) project aims to strengthening Europe's capacity in the design and manufacture of RISC-V based high-performance computing chips.
EMiX: Emulating Beyond Single-FPGA LimitsAlexander Kropotov, Miquel Moreto, Behzad Salami2026-04-29下载FPGA-level emulation is a key step in pre-silicon chip design validation. However, emulating large-scale multi-core systems increasingly exceed the hardware resource capacity of a single FPGA, limitin...
Efficient, VRAM-Constrained xLM Inference on ClientsAditya Ukarande, Deep Shekhar, Marc Blackstein, Ram Rangan2026-04-29下载To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language models (LLMs) and vision language models (VLMs), joi...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Real-Time GPU-Accelerated Monte Carlo Evaluation of Safety-Critical AEB Systems Under UncertaintyAkshay Karjol, Shadi Alawneh2026-04-29下载Automatic Emergency Braking (AEB) systems represent a safety-critical national interest, with the National Highway Traffic Safety Administration (NHTSA) Federal Motor Vehicle Safety Standard (FMVSS No...
End-to-End and Phase-Level Performance Optimization for Hyperledger FabricPavan Sollu, Aniruddha Mukherjee, Divya Pulivarthi, S. R. Eshwar, Gugan Thoppe, Kshitij Pratihast, Tittu Varghese, Hrishikesh Nashikkar, Yogesh Simmhan2026-04-29下载Hyperledger Fabric (HLF) is a modular, permissioned blockchain widely adopted in enterprise settings. Enhancing its throughput and latency remains challenging, as optimization decisions made in one ph...
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence ParallelismAhan Gupta, Zhihao Wang, Neel Dani, Masahiro Tanaka, Olatunji Ruwase, Minjia Zhang2026-04-29下载Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of thousands of tokens.
Efficient Training on Multiple Consumer GPUs with RoundPipeYibin Luo, Shiwei Gao, Huichuan Zheng, Youyou Lu, Jiwu Shu2026-04-29下载Fine-tuning Large Language Models (LLMs) on consumer-grade GPUs is highly cost-effective, yet constrained by limited GPU memory and slow PCIe interconnects.
Adaptive Self-Organization in Anonymous Dynamic NetworksGarrett Parzych, Joshua J. Daymude2026-04-29下载We introduce the problem of adaptive self-organization in which the nodes of an anonymous, synchronous dynamic network must distributively change the collective distribution of their responses (or "co...
FaaSMoE: A Serverless Framework for Multi-Tenant Mixture-of-Experts ServingMinghe Wang, Trever Schirmer, Mohammadreza Malekabbasi, David Bermbach2026-04-29下载Mixture-of-Experts (MoE) models offer high capacity with efficient inference cost by activating a small subset of expert models per input. However, deploying MoE models requires all experts to reside ...
A Test Taxonomy and Continuous Integration Ecosystem for Dynamic Resource Management in HPCPetter Sandås, Íñigo Aréjula-Aísa, Sergio Iserte, Antonio J. Peña2026-04-29下载High-performance computing (HPC) systems are increasingly exploring dynamic resource management and malleable MPI applications to better adapt to heterogeneous architectures, fluctuating workloads, an...
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with VoxelYiqi Liu, Noelle Crawford, Michael Wang, Jilong Xue, Jian Huang2026-04-29下载To overcome the well-known memory bottleneck of AI chips, 3D stacked architectures that employ advanced packaging technology with high-density through-silicon vias (TSVs) pins have proven to be a prom...
A Semantic Quantum Circuit Cache for Scalable and Distributed Quantum-Classical WorkflowsMar Tejedor, Javier Conejero, Rosa M. Badia2026-04-29下载Hybrid quantum--classical workflows often execute large ensembles of circuits that differ syntactically but implement identical operations, leading to substantial redundant computation.
COPUS: Co-adaptive Parallelism and Batch Size Selection in Large Language Model TrainingAkhmed Sakip, Erland Hilman Fuadi, Omar Sayedelahl, Zonghang Li, Jianshu She, Alham Fikri Aji, Steve Liu, Eric Xing, Qirong Ho2026-04-29下载Training large language models requires jointly configuring two interdependent aspects of the system: the global batch size, which governs statistical efficiency, and the 3D parallelism strategy, whic...
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic WorkflowSina Heidari, Dimitrios S. Nikolopoulos2026-04-29下载Deep learning compilers and vendor libraries deliver strong baseline performance but are bounded by finite, engineer-curated catalogs. When these omit needed optimizations, practitioners substitute ha...
DMRlib: Easy-coding and Efficient Resource Management for Job MalleabilitySergio Iserte, Rafael Mayo, Enrique S. Quintana-Ortí, Antonio J. Peña2026-04-29下载Process malleability has proved to have a highly positive impact on the resource utilization and global productivity in data centers compared with the conventional static resource allocation policy.
MPI Malleability Validation under Replayed Real-World HPC ConditionsS. Iserte, M. Madon, G. Da, J. Pierson, A. J. Peña2026-04-29下载Dynamic Resource Management (DRM) techniques can be leveraged to maximize throughput and resource utilization in computational clusters. Although DRM has been extensively studied through analytical wo...
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM InferenceBodon Jeong, Hongsu Byun, Youngjae Kim, Weikuan Yu, Kyungkeun Lee, Jihoon Yang, Sungyong Park2026-04-29下载The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key challenge arises from Key-Value (KV) caches, which of...
FloatSOM: GPU-Accelerated, Distributed, Topology-Flexible Self-Organizing MapsTony Xu, Sarah Klamt, Katherine Turner, Anne Brustle, Felix Marsh-Wakefield, Givanna Putri2026-04-29下载GPU-accelerated Self-Organizing Map (SOM) implementations are among the most competitive options for large-scale SOM analysis, but growing dataset sizes increasingly challenge their practical use beca...
Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language ModelsCyril Shih-Huan Hsu, Wig Yuan-Cheng Cheng, Chrysa Papagianni2026-04-29下载Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed the capabilities of resource-constrained embedded pl...
SplitFT: An Adaptive Federated Split Learning System For LLMs Fine-TuningYimeng Shan, Zhaorui Zhang, Sheng Di, Yu Liu, Xiaoyi Lu, Benben Liu2026-04-29下载Federated Split Learning has been identified as an efficient approach to address the computational resource constraints of clients in classical federated learning, while guaranteeing data privacy for ...
Efficient, VRAM-Constrained xLM Inference on ClientsAditya Ukarande, Deep Shekhar, Marc Blackstein, Ram Rangan2026-04-29下载To usher in the next round of client AI innovation, there is an urgent need to enable efficient, lossless inference of high-accuracy large language models (LLMs) and vision language models (VLMs), joi...
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & InferenceVasu Shyam, Anna Golubeva, Quentin Anthony2026-04-29下载We present tensor and sequence parallelism (TSP), a parallel execution strategy that folds tensor parallelism and sequence parallelism onto a single device axis.
DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model TrainingTianhao Hu, Xiangcheng Liu, Youshao Xiao, Yang Zheng, Xuan Huang, Jinrui Ding, Yufei Zhang, Tao Liang, Hongyu Zang, Quan Chen, Yueqing Sun, Wenjie Shi, Chao Zhang, Wei Wang, Qi Gu, Yerui Sun, Yucheng Xie, Xunliang Cai2026-04-29下载Reinforcement learning (RL) has become a critical paradigm for LLM post-training, yet the rollout phase -- accounting for 50--80% of total step time -- is bottlenecked by skewed generation: long-taile...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
BLINC: Context-Specific Causal Learning for Automated RAN ConfigurationReshma Prasad, Michele Polese, Tommaso Melodia2026-04-29下载Radio Access Network (RAN) configuration has traditionally required significant manual effort due to indirect causal dependencies between observable Key Performance Indicators (KPIs), and context-depe...
A 3GPP Perspective on Spectrum Sharing for the 5G-to-6G Migration: From DSS to MRSSXingqin Lin2026-04-29下载Dynamic spectrum sharing (DSS) played an important role in the 4G-to-5G transition by allowing 5G new radio (NR) to enter valuable legacy spectrum without immediate static refarming.
Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language ModelsCyril Shih-Huan Hsu, Wig Yuan-Cheng Cheng, Chrysa Papagianni2026-04-29下载Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, which exceed the capabilities of resource-constrained embedded pl...
SWE-Bench 5G: Benchmarking AI Coding Agents on Telecom Network Engineering TasksJiao Chen, Jianhua Tang, Xiaotong Yang, Zuohong Lv2026-04-29下载AI coding agents demonstrate strong performance on general-purpose software benchmarks. However, their ability to handle 5G network engineering tasks remains unexplored.
StreamGuard: Exploring a 5G Architecture for Efficient, Quality of Experience-Aware Video ConferencingXuyang Cao, Oliver Michel, Kyle Jamieson2026-04-29下载Video conferencing over 5G is increasingly prevalent, yet its Quality of Experience (QoE) often degrades under limited radio resources. This has two causes: 5G networks must serve many users, while in...

cs.PF - Performance ​

标题作者发布日期PDF摘要
A High-Throughput Compute-Efficient POMDP Hide-And-Seek-Engine (HASE) for Multi-Agent OperationsTimothy Flavin, Sandip Sen2026-04-29下载Reinforcement Learning (RL) algorithms exhibit high sample complexity, particularly when applied to Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs).
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence ParallelismAhan Gupta, Zhihao Wang, Neel Dani, Masahiro Tanaka, Olatunji Ruwase, Minjia Zhang2026-04-29下载Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of thousands of tokens.
Revealing NVIDIA Closed-Source Driver Command Streams for CPU-GPU Runtime Behavior InsightYuang Yan, Ian Karlin, Ryan Grant2026-04-29下载For NVIDIA GPUs, CUDA is the primary interface through which applications orchestrate GPU execution, yet much of the logic that realizes CUDA operations resides in NVIDIA's closed-source userspace dri...
What Is the Cost of Energy Monitoring? An Empirical Study on the Overhead of RAPL-Based ToolsJeremy Diamond, Vincenzo Stoico2026-04-29下载The Running Average Power Limit (RAPL) interface is widely used to estimate software energy consumption via CPU and DRAM counters, but tool design differences and high-frequency polling can introduce ...
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic WorkflowSina Heidari, Dimitrios S. Nikolopoulos2026-04-29下载Deep learning compilers and vendor libraries deliver strong baseline performance but are bounded by finite, engineer-curated catalogs. When these omit needed optimizations, practitioners substitute ha...
DUAL-BLADE: Dual-Path NVMe-Direct KV-Cache Offloading for Edge LLM InferenceBodon Jeong, Hongsu Byun, Youngjae Kim, Weikuan Yu, Kyungkeun Lee, Jihoon Yang, Sungyong Park2026-04-29下载The increasing deployment of Large Language Model (LLM) inference on edge AI systems demands efficient execution under tight memory budgets. A key challenge arises from Key-Value (KV) caches, which of...

基于 VitePress 构建 · 使用本地搜索查找论文