Skip to content

2026-05-07 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
EDA-Schema-V2: A Multimodal Schema, Open Datasets, and Benchmarks for Machine Learning in Digital Physical DesignPratik Shrestha, Alec Aversa, Ioannis Savidis2026-05-07下载The continuous scaling of CMOS technology has significantly increased the complexity of very large-scale integrated circuits, driving interest in applying machine learning (ML) to electronic design au...
Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC FixingPengju Liu, Nuo Xu, Jinwei Tang, Yu Cao, Caiwen Ding2026-05-07下载LLM-based agents are increasingly applied to the "last mile" of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC) violations and converging Power-Performance-Area...
CARMEN: CORDIC-Accelerated Resource-Efficient Multi-Precision Inference Engine for Deep LearningSonu Kumar, Mukul Lokhande, Santosh Kumar Vishvakarma, Adam Teman2026-05-07下载This paper presents CARMEN, a runtime-adaptive, CORDIC-accelerated multi-precision vector engine for resource-efficient deep learning inference.
EULER-ADAS: Energy-Efficient & SIMD-Unified Logarithmic-Posit Engine for Precision-Reconfigurable Approximate ADAS AccelerationMukul Lokhande, Ratko Pilipovic, Omkar Kokane, Adam Teman, Santosh Kumar Vishvakarma2026-05-07下载Advanced driver-assistance systems (ADAS) require neural compute engines that deliver low-latency inference under strict power and area constraints.
Development of embedded target detection system based on FPGA and YOLOv3-TinyZihan Jiang, Fanghao Liu, Huawei Wang, Mamataziz Mattohti, Xiangquan Chen, Jingfu Guo, Xiaotian Wu, Yongjun Dong2026-05-07下载Computational complexity and storage requirements are crucial factors influencing the performance and efficiency of convolutional neural networks (CNNs) in resource-constrained environments.
On-Orbit Real-Time Wildfire Detection Under On-Board ConstraintsMatthias Rötzer, Veronika Pörtge, Martin Ickerott, Jayendra Praveen Kumar Chorapalli, Dimitri Scheftelowitsch, Max Bereczky, Dmitry Rashkovetsky, Sai Manoj Appalla, Julia Gottfriedsen2026-05-07下载We present a deployed system for on-orbit wildfire detection aboard a nine-satellite commercial thermal infrared constellation, operating under demanding joint constraints: sub-megabyte model footprin...
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNsRappy Saha, Jude Haris, Nicolas Bohm Agostini, David Kaeli, José Cano2026-05-07下载Power-of-two (PoT) quantization significantly reduces the size of deep neural networks (DNNs) and replaces multiplications with bit-shift operations for inference.
XtraMAC: An Efficient MAC Architecture for Mixed-Precision LLM Inference on FPGAFeng Yu, Hongshi Tan, Yao Chen, Weng-Fai Wong, Bingsheng He2026-05-07下载The widespread adoption of mixed-precision quantization in large language models (LLMs) has created demand for hardware that can efficiently perform multiply-accumulate (MAC) operations across mixed d...
A virtually connected probabilistic computer as a solver for higher-order, densely connected, or reconfigurable combinatorial optimisation problemsAmy J. Searle, Harry Youel, Fredrik Hasselgren, Annika Möslein, Ramy Aboushelbaya, Marko von der Leyen2026-05-07下载Recently, there has been growing interest in unconventional computing as an approach for solving NP-hard problems, by developing dedicated hardware to find solutions more efficiently than conventional...
LLM-Driven Design Space Exploration of FPGA-based AcceleratorsVinamra Sharma, Xingjian Fu, Jude Haris, José Cano2026-05-07下载Designing field-programmable gate array (FPGA)-based accelerators for modern artificial intelligence workloads requires navigating a large and complex hardware design space encompassing architectural ...
MoE-Hub: Taming Software Complexity for Seamless MoE Overlap with Hardware-Accelerated Communication on Multi-GPU SystemsZhuoshan Zhou, Chen Zhang, Shuyi Zhang, Qijun Zhang, Haibo Wang, Zhe Zhou, Zhipeng Tu, Guangyu Sun, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He, Minyi Guo2026-05-07下载The Mixture-of-Experts (MoE) architecture is crucial for scaling large language models, but its scalability is severely limited by inter-GPU communication bottlenecks in multi-GPU systems.
TokenStack: A Heterogeneous HBM-PIM Architecture and Runtime for Efficient LLM InferenceZhuoran Li, Zhuohang Bian, Zihao Huang, Guangyu Sun, Yun Liang, Youwei Zhuo2026-05-07下载Large language model (LLM) serving is now limited by the key-value (KV) cache. During decode, each new token rereads prior KV state, so attention becomes a bandwidth- and capacity-heavy memory task.
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU SystemsChen Zhang, Qijun Zhang, Zhuoshan Zhou, Yijia Diao, Haibo Wang, Zhe Zhou, Zhipeng Tu, Zhiyao Li, Guangyu Sun, Zhuoran Song, Zhigang Ji, Jingwen Leng, Minyi Guo2026-05-07下载Tensor parallelism (TP) in large-scale LLM inference and training introduces frequent collective operations that dominate inter-GPU communication.
Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUsQijun Zhang, Chen Zhang, Zhuoshan Zhou, Haibo Wang, Zhe Zhou, Zhipeng Tu, Guangyu Sun, Zhiyao Xie, Yijia Diao, Zhigang Ji, Jingwen Leng, Guanghui He, Minyi Guo2026-05-07下载Mixture-of-Experts (MoE) has been adopted by many leading large models to reduce computational requirements. However, frequent inter-GPU communication in MoE expert parallelism (EP) becomes a performa...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Pomegranate: A Lightweight Compartmentalization Architecture using Virtualization ExtensionsShriram Raja, Zhiyuan Ruan, Richard West2026-05-07下载The monolithic nature of widely used commodity operating systems means that vulnerabilities in one software component potentially compromise the entire kernel.

cs.PF - Performance ​

标题作者发布日期PDF摘要
ADELIA: Automatic Differentiation for Efficient Laplace Inference ApproximationsAfif Boudaoud, Lisa Gaedke-Merzhäuser, Alexandros Nikolaos Ziogas, Vincent Maillou, Alexandru Calotoiu, Marcin Copik, Håvard Rue, Mathieu Luisier, Torsten Hoefler2026-05-07下载Spatio-temporal Bayesian inference drives environmental and health sciences using latent Gaussian models. Integrated Nested Laplace Approximations (INLA) enable inference for these models at HPC scale...
PoTAcc: A Pipeline for End-to-End Acceleration of Power-of-Two Quantized DNNsRappy Saha, Jude Haris, Nicolas Bohm Agostini, David Kaeli, José Cano2026-05-07下载Power-of-two (PoT) quantization significantly reduces the size of deep neural networks (DNNs) and replaces multiplications with bit-shift operations for inference.
LLM-Driven Design Space Exploration of FPGA-based AcceleratorsVinamra Sharma, Xingjian Fu, Jude Haris, José Cano2026-05-07下载Designing field-programmable gate array (FPGA)-based accelerators for modern artificial intelligence workloads requires navigating a large and complex hardware design space encompassing architectural ...
When Quantization Is Free: An int4 KV Cache That Outruns fp16 on Apple SiliconMohamed Amine Bergach2026-05-07下载KV-cache quantization is framed as a quality--latency trade-off. We show it is \emph{inverted} on Apple Silicon's unified memory: a single fused Metal kernel (sign-randomized FFT ++ per-channel λ $...

基于 VitePress 构建 · 使用本地搜索查找论文