2026-05-07
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| EDA-Schema-V2: A Multimodal Schema, Open Datasets, and Benchmarks for Machine Learning in Digital Physical Design | Pratik Shrestha, Alec Aversa, Ioannis Savidis | 2026-05-07 | 下载 | The continuous scaling of CMOS technology has significantly increased the complexity of very large-scale integrated circuits, driving interest in applying machine learning (ML) to electronic design au... |
| Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing | Pengju Liu, Nuo Xu, Jinwei Tang, Yu Cao, Caiwen Ding | 2026-05-07 | 下载 | LLM-based agents are increasingly applied to the "last mile" of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC) violations and converging Power-Performance-Area... |
| CARMEN: CORDIC-Accelerated Resource-Efficient Multi-Precision Inference Engine for Deep Learning | Sonu Kumar, Mukul Lokhande, Santosh Kumar Vishvakarma, Adam Teman | 2026-05-07 | 下载 | This paper presents CARMEN, a runtime-adaptive, CORDIC-accelerated multi-precision vector engine for resource-efficient deep learning inference. |
| EULER-ADAS: Energy-Efficient & SIMD-Unified Logarithmic-Posit Engine for Precision-Reconfigurable Approximate ADAS Acceleration | Mukul Lokhande, Ratko Pilipovic, Omkar Kokane, Adam Teman, Santosh Kumar Vishvakarma | 2026-05-07 | 下载 | Advanced driver-assistance systems (ADAS) require neural compute engines that deliver low-latency inference under strict power and area constraints. |
| Development of embedded target detection system based on FPGA and YOLOv3-Tiny | Zihan Jiang, Fanghao Liu, Huawei Wang, Mamataziz Mattohti, Xiangquan Chen, Jingfu Guo, Xiaotian Wu, Yongjun Dong | 2026-05-07 | 下载 | Computational complexity and storage requirements are crucial factors influencing the performance and efficiency of convolutional neural networks (CNNs) in resource-constrained environments. |
| On-Orbit Real-Time Wildfire Detection Under On-Board Constraints | Matthias Rötzer, Veronika Pörtge, Martin Ickerott, Jayendra Praveen Kumar Chorapalli, Dimitri Scheftelowitsch, Max Bereczky, Dmitry Rashkovetsky, Sai Manoj Appalla, Julia Gottfriedsen | 2026-05-07 | 下载 | We present a deployed system for on-orbit wildfire detection aboard a nine-satellite commercial thermal infrared constellation, operating under demanding joint constraints: sub-megabyte model footprin... |
| PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs | Rappy Saha, Jude Haris, Nicolas Bohm Agostini, David Kaeli, José Cano | 2026-05-07 | 下载 | Power-of-two (PoT) quantization significantly reduces the size of deep neural networks (DNNs) and replaces multiplications with bit-shift operations for inference. |
| XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGA | Feng Yu, Hongshi Tan, Yao Chen, Weng-Fai Wong, Bingsheng He | 2026-05-07 | 下载 | The widespread adoption of mixed-precision quantization in large language models (LLMs) has created demand for hardware that can efficiently perform multiply-accumulate (MAC) operations across mixed d... |
| A virtually connected probabilistic computer as a solver for higher-order, densely connected, or reconfigurable combinatorial optimisation problems | Amy J. Searle, Harry Youel, Fredrik Hasselgren, Annika Möslein, Ramy Aboushelbaya, Marko von der Leyen | 2026-05-07 | 下载 | Recently, there has been growing interest in unconventional computing as an approach for solving NP-hard problems, by developing dedicated hardware to find solutions more efficiently than conventional... |
| LLM-Driven Design Space Exploration of FPGA-based Accelerators | Vinamra Sharma, Xingjian Fu, Jude Haris, José Cano | 2026-05-07 | 下载 | Designing field-programmable gate array (FPGA)-based accelerators for modern artificial intelligence workloads requires navigating a large and complex hardware design space encompassing architectural ... |
| MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU Systems | Zhuoshan Zhou, Chen Zhang, Shuyi Zhang, Qijun Zhang, Haibo Wang, Zhe Zhou, Zhipeng Tu, Guangyu Sun, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He, Minyi Guo | 2026-05-07 | 下载 | The Mixture-of-Experts (MoE) architecture is crucial for scaling large language models, but its scalability is severely limited by inter-GPU communication bottlenecks in multi-GPU systems. |
| TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM Inference | Zhuoran Li, Zhuohang Bian, Zihao Huang, Guangyu Sun, Yun Liang, Youwei Zhuo | 2026-05-07 | 下载 | Large language model (LLM) serving is now limited by the key-value (KV) cache. During decode, each new token rereads prior KV state, so attention becomes a bandwidth- and capacity-heavy memory task. |
| Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems | Chen Zhang, Qijun Zhang, Zhuoshan Zhou, Yijia Diao, Haibo Wang, Zhe Zhou, Zhipeng Tu, Zhiyao Li, Guangyu Sun, Zhuoran Song, Zhigang Ji, Jingwen Leng, Minyi Guo | 2026-05-07 | 下载 | Tensor parallelism (TP) in large-scale LLM inference and training introduces frequent collective operations that dominate inter-GPU communication. |
| Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs | Qijun Zhang, Chen Zhang, Zhuoshan Zhou, Haibo Wang, Zhe Zhou, Zhipeng Tu, Guangyu Sun, Zhiyao Xie, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He, Minyi Guo | 2026-05-07 | 下载 | Mixture-of-Experts (MoE) has been adopted by many leading large models to reduce computational requirements. However, frequent inter-GPU communication in MoE expert parallelism (EP) becomes a performa... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Pomegranate: A Lightweight Compartmentalization Architecture using Virtualization Extensions | Shriram Raja, Zhiyuan Ruan, Richard West | 2026-05-07 | 下载 | The monolithic nature of widely used commodity operating systems means that vulnerabilities in one software component potentially compromise the entire kernel. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations | Afif Boudaoud, Lisa Gaedke-Merzhäuser, Alexandros Nikolaos Ziogas, Vincent Maillou, Alexandru Calotoiu, Marcin Copik, Håvard Rue, Mathieu Luisier, Torsten Hoefler | 2026-05-07 | 下载 | Spatio-temporal Bayesian inference drives environmental and health sciences using latent Gaussian models. Integrated Nested Laplace Approximations (INLA) enable inference for these models at HPC scale... |
| PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNs | Rappy Saha, Jude Haris, Nicolas Bohm Agostini, David Kaeli, José Cano | 2026-05-07 | 下载 | Power-of-two (PoT) quantization significantly reduces the size of deep neural networks (DNNs) and replaces multiplications with bit-shift operations for inference. |
| LLM-Driven Design Space Exploration of FPGA-based Accelerators | Vinamra Sharma, Xingjian Fu, Jude Haris, José Cano | 2026-05-07 | 下载 | Designing field-programmable gate array (FPGA)-based accelerators for modern artificial intelligence workloads requires navigating a large and complex hardware design space encompassing architectural ... |
| When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple Silicon | Mohamed Amine Bergach | 2026-05-07 | 下载 | KV-cache quantization is framed as a quality--latency trade-off. We show it is \emph{inverted} on Apple Silicon's unified memory: a single fused Metal kernel (sign-randomized FFT per-channel λ $... |