2026-04-20
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ChipLight: Cross-Layer Optimization of Chiplet Design with Optical Interconnects for LLM Training | Kangbo Bai, Zhantong Zhu, Yifan Ding, Tianyu Jia | 2026-04-20 | 下载 | In large-scale distributed LLM training, communication between devices becomes the key performance bottleneck. Chiplet technology can integrate multiple dies into a package to scale-up node performanc... |
| A Comparative Analysis of ARM and x86-64 Laptop-Class Processors: Architecture, Assembly-Level Performance, and Energy Efficiency | Mustafa Mert Özyılmaz | 2026-04-20 | 下载 | ARM-based and x86-64 laptop processors differ not only in instruction-set design, but also in memory hierarchy, core organization, system integration, and power-management mechanisms. |
| A PPA-Driven 3D-IC Partitioning Selection Framework with Surrogate Models | Shang Wang, Shuai Liu, Owen Randall, Matthew E. Taylor | 2026-04-20 | 下载 | 3D-IC netlist partitioning is commonly optimized using proxy objectives, while final PPA is treated as a costly evaluation rather than an optimization signal. |
| CHICO-Agent: An LLM Agent for the Cross-layer Optimization of 2.5D and 3D Chiplet-based Systems | Qihang Wu, Aman Arora, Vidya A. Chhabria | 2026-04-20 | 下载 | The rapid growth of large language models (LLMs) and AI workloads has pushed monolithic silicon to its reticle and economic limits, accelerating the adoption of 2.5D/3D chiplet systems. |
| Optimizing Branch Predictor for Graph Applications | Upasna, Venkata Kalyan Tavva | 2026-04-20 | 下载 | Real-world graph applications are generally larger than the size of the cache itself. Due to this reason, the memory hierarchy was identified as a key bottleneck by the earlier works. |
| AutoPPA: Automated Circuit PPA Optimization via Contrastive Code-based Rule Library Learning | Chongxiao Li, Pengwei Jin, Di Huang, Guangrun Sun, Husheng Han, Jianan Mu, Xinyao Zheng, Jiaguo Zhu, Shuyi Xing, Hanjun Wei, Tianyun Ma, Shuyao Cheng, Rui Zhang, Ying Wang, Zidong Du, Qi Guo, Xing Hu | 2026-04-20 | 下载 | Performance, power, and area (PPA) optimization is a fundamental task in RTL design, requiring a precise understanding of circuit functionality and the relationship between circuit structures and PPA ... |
| Scattering-Matrix-Based Parametric Characterization of a Two-Port Bridged-T Network for Microstrip Filter Applications | Naser Khatti Dizabadi, Douglas Jussaume | 2026-04-20 | 下载 | The purpose of this study is to characterize a two-port Bridged-T network using transmission (T) and scattering (S) matrices. Using mathematical derivations, scattering parameters including S11, S12, ... |
| VerilogCL: A Contrastive Learning Framework for Robust LLM-Based Verilog Generation | Yan Tan, Tong Liu, Xiangchen Meng, Yangdi Lyu | 2026-04-20 | 下载 | Large Language Models (LLMs) have recently achieved strong performance in software code generation. However, applying them to hardware description languages (HDLs), such as Verilog, remains challengin... |
| AQPIM: Breaking the PIM Capacity Wall for LLMs with In-Memory Activation Quantization | Kosuke Matsushima, Yasuyuki Okoshi, Masato Motomura, Daichi Fujiki | 2026-04-20 | 下载 | Processing-in-Memory (PIM) architectures offer a promising solution to the memory bottlenecks in data-intensive machine learning, yet often overlook the growing challenge of activation memory footprin... |
| Proxics: an efficient programming model for far memory accelerators | Zikai Liu, Niels Pressel, Jasmin Schult, Roman Meier, Pengcheng Xu, Timothy Roscoe | 2026-04-20 | 下载 | The use of disaggregated or far memory systems such as CXL memory pools has renewed interest in Near-Data Processing (NDP): situating cores close to memory to reduce bandwidth requirements to and from... |
| M100: An Orchestrated Dataflow Architecture Powering General AI Computing | Yan Xie, Changkui Mao, Changsong Wu, Chao Lu, Chao Suo, Cheng Qian, Chun Yang, Danyang Zhu, Hengchang Xiong, Hongzhan Lu, Hongzhen Liu, Jiafu Liu, Jie Chen, Jie Dai, Junfeng Tang, Kai Liu, Kun Li, Lipeng Ge, Meng Sun, Min Luo, Peng Chen, Peng Wang, Shaodong Yang, Shibin Tang, Shibo Chen, Weikang Zhang, Xiao Ling, Xiaobo Du, Xin Wu, Yang Liu, Yi Jiang, Yihua Jin, Yin Huang, Yuli Zhang, Zhen Yuan, Zhiyuan Man, Zhongxiao Yao | 2026-04-20 | 下载 | As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based architectures offer versatility for diverse AI work... |
| Enabling AI ASICs for Zero Knowledge Proof | Jianming Tong, Jingtian Dang, Simon Langowski, Tianhao Huang, Asra Ali, Jeremy Kun, Jevin Jiang, Srinivas Devadas, Tushar Krishna | 2026-04-20 | 下载 | Zero-knowledge proof (ZKP) provers remain costly because multi-scalar multiplication (MSM) and number-theoretic transforms (NTTs) dominate runtime as they need significant computation. |
| AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator | Chenhao Xue, Yukun Wang, An Guo, Yuhui Shi, Jinwei Zhou, Xiping Dong, Yihan Yin, Yuanpeng Zhang, Tianyu Jia, Wei Gao, Qiang Wu, Xin Si, Jun Yang, Guangyu Sun | 2026-04-20 | 下载 | SRAM-based compute-in-memory (CIM) offers high computational density and energy efficiency for deep neural network (DNN) accelerators, but its limited capacity causes on/off-chip data movement overhea... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Preserving Clusters in Error-Bounded Lossy Compression of Particle Data | Congrong Ren, Sheng Di, Katrin Heitmann, Franck Cappello, Hanqi Guo | 2026-04-20 | 下载 | Lossy compression is widely used to reduce storage and I/O costs for large-scale particle datasets in scientific applications such as cosmology, molecular dynamics, and fluid dynamics, where clusterin... |
| HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing | Mao Lin, Xi Wang, Guilherme Cox, Dong Li, Hyeran Jeon | 2026-04-20 | 下载 | As modern LLMs support thousands to millions of tokens, KV caches grow to hundreds of gigabytes, stressing memory capacity and bandwidth. Existing solutions, such as KV cache pruning and offloading, a... |
| User Experiences with MPI RMA and ULFM in a Resilient Key-Value Store Implementation | Claudia Fohry, Rainer Fink | 2026-04-20 | 下载 | As hardware failures such as node losses become increasingly common, MPI programmers may want to save vulnerable data in a resilient store. While third-party storage solutions such as Redis or the Haz... |
| Trust, but Verify: ByzTwin-Range, a Digital Twin Cyber-Range for Byzantine Faults | Tadeu Freitas, João Soares, Rolando Martins | 2026-04-20 | 下载 | Critical infrastructures increasingly rely on interconnected and software-driven Cyber-Physical Systems (CPS), exposing operational processes to both accidental failures and sophisticated adversarial ... |
| Optimizing Memory Allocation in Distributed Clusters with Predictive Modeling | Jonathan Bader, Edgar Blumenthal, Marten Eckardt, Justus Krebs, Joel Witzke, Xemena Wysokinska, Haci Ismail Aslan, Odej Kao | 2026-04-20 | 下载 | In modern distributed systems, efficient resource allocation is a vital aspect to maintain scalability, reduce operational costs, and ensure fast execution even across heterogeneous workloads. |
| Toward Optimality: A Tighter Analysis of Message Complexity for Leader Election in Diameter-Two Networks | Abhijit Sadhukhan, Adri Bhattacharya, Anisur Rahaman Molla | 2026-04-20 | 下载 | We study the message complexity of leader election in synchronous networks of diameter two. Our main contribution is a refined analysis of the randomized algorithm proposed by Chatterjee et al. |
| Matrix-Free 3D SIMP Topology Optimization with Fused Gather-GEMM-Scatter Kernels | Shaoliang Yang, Jun Wang, Yunsheng Wang | 2026-04-20 | 下载 | The matrix-free gather-batched-GEMM-scatter pattern eliminates global stiffness assembly for three-dimensional SIMP topology optimization, but the conventional three-stage implementation forces avoida... |
| Unlocking the Edge deployment and ondevice acceleration of multi-LoRA enabled one-for-all foundational LLM | Sravanth Kodavanti, Sowmya Vajrala, Srinivas Miriyala, Utsav Tiwari, Uttam Kumar, Utkarsh Kumar Mahawar, Achal Pratap Singh, Arya D, Narendra Mutyala, Vikram Nelvoy Rajendiran, Sharan Kumar Allur, Euntaik Lee, Dohyoung Kim, HyeonSu Lee, Gyusung Cho, JungBae Kim | 2026-04-20 | 下载 | Deploying large language models (LLMs) on smartphones poses significant engineering challenges due to stringent constraints on memory, latency, and runtime flexibility. |
| GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion | Yiwei Yang, Xiangyu Gao, Yuan Zhou, Yuhang Gan, Yusheng Zheng, Andi Quinn | 2026-04-20 | 下载 | Modern deep learning workloads often consist of many small tensor operations, especially in inference, attention, and micro-batched training. In these settings, kernel launch overhead can become a maj... |
| AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures | Jie Liu, Huanzhi Pu, Zhiru Zhang | 2026-04-20 | 下载 | Sparse Matrix-Matrix Multiplication (SpMM) is a fundamental kernel across scientific computing and machine learning. While prior work accelerates SpMM using Tensor Cores, no existing sparse kernel exp... |
| DeInfer: Efficient Parallel Inferencing for Decomposed Large Language Models | You-Liang Huang, Xinhao Huang, Chengxi Liao, Zeyi Wen | 2026-04-20 | 下载 | Existing works on large language model (LLM) decomposition mainly focus on improving performance on downstream tasks, but they ignore the poor parallel inference performance when trying to scale up th... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Spectrum Configuration Framework for Throughput Maximization in Open Systems with Roll-Off-Based QoT Optimization | Peyman Pahlevanzadeh, Venkata Virajit Garbhapu, Agastya Raj, Dmitrii Briantcev, Dan Kilper, Marco Ruffini | 2026-04-20 | 下载 | We propose a spectrum-configuration framework for open and disaggregated optical systems that maximizes throughput while guaranteeing the quality of transmission (QoT) margins. |
| Sub-additive service curves in the Network Calculus analysis | Anne Bouillard | 2026-04-20 | 下载 | Network Calculus is a theoretical model that aims at providing upper bounds of worst-case performance (such as delay or buffer occupancy). This is a mathematical framework that handles both network mo... |
| Tabu Search for Tactical Wireless Network Design in Challenging Environments | Wisssem Ahmed Zaid, Alain Hertz | 2026-04-20 | 下载 | Tactical wireless networks play a vital role in ensuring reliable connectivity in scenarios where conventional telecommunications infrastructure is unavailable or damaged, such as areas impacted by na... |
| Dynamic Risk Assessment by Bayesian Attack Graphs and Process Mining | Francesco Vitale, Simone Guarino, Stefano Perone, Massimiliano Rak, Nicola Mazzocca | 2026-04-20 | 下载 | While attack graphs are useful for identifying major cybersecurity threats affecting a system, they do not provide operational support for determining the likelihood of having a known vulnerability ex... |
| Lagrange Index based Scheduling for Minimizing Age of Updates from Heterogeneous Sources | Aniket Mukherjee, Joy Kuri, Chandramani Singh | 2026-04-20 | 下载 | Modern sensing systems generate heterogeneous updates ranging from small status packets to large data objects. We study a single-hop wireless uplink network where sensors generate updates at will, eac... |
| Enhancing Anomaly-Based Intrusion Detection Systems with Process Mining | Francesco Vitale, Francesco Grimaldi, Massimiliano Rak, Nicola Mazzocca | 2026-04-20 | 下载 | Anomaly-based Intrusion Detection Systems (IDSs) ensure protection against malicious attacks on networked systems. While deep learning-based IDSs achieve effective performance, their limited trustwort... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AgenTEE: Confidential LLM Agent Execution on Edge Devices | Sina Abdollahi, Mohammad M Maheri, Javad Forough, Amir Al Sadi, Josh Millar, David Kotz, Marios Kogias, Hamed Haddadi | 2026-04-20 | 下载 | Large Language Model (LLM) agents provide powerful automation capabilities, but they also create a substantially broader attack surface than traditional applications due to their tight integration wit... |
| Proxics: an efficient programming model for far memory accelerators | Zikai Liu, Niels Pressel, Jasmin Schult, Roman Meier, Pengcheng Xu, Timothy Roscoe | 2026-04-20 | 下载 | The use of disaggregated or far memory systems such as CXL memory pools has renewed interest in Near-Data Processing (NDP): situating cores close to memory to reduce bandwidth requirements to and from... |
| GPUOS: A GPU Operating System Primitive for Transparent Operation Fusion | Yiwei Yang, Xiangyu Gao, Yuan Zhou, Yuhang Gan, Yusheng Zheng, Andi Quinn | 2026-04-20 | 下载 | Modern deep learning workloads often consist of many small tensor operations, especially in inference, attention, and micro-batched training. In these settings, kernel launch overhead can become a maj... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing | Mao Lin, Xi Wang, Guilherme Cox, Dong Li, Hyeran Jeon | 2026-04-20 | 下载 | As modern LLMs support thousands to millions of tokens, KV caches grow to hundreds of gigabytes, stressing memory capacity and bandwidth. Existing solutions, such as KV cache pruning and offloading, a... |
| Lagrange Index based Scheduling for Minimizing Age of Updates from Heterogeneous Sources | Aniket Mukherjee, Joy Kuri, Chandramani Singh | 2026-04-20 | 下载 | Modern sensing systems generate heterogeneous updates ranging from small status packets to large data objects. We study a single-hop wireless uplink network where sensors generate updates at will, eac... |