2026-05-12
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Search Your Block Floating Point Scales! | Tanmaey Gupta, Hayden Prairie, Xiaoxia Wu, Reyna Abhyankar, Qingyang Wu, Austin Silveria, Pragaash Ponnusamy, Jue Wang, Ben Athiwaratkun, Leon Song, Tri Dao, Daniel Y. Fu, Chris De Sa | 2026-05-12 | 下载 | Quantization has emerged as a standard technique for accelerating inference for generative models by enabling faster low-precision computations and reduced memory transfers. |
| Enhancing Instruction Prefetching via Cache and TLB Management | Alexandre Valentin Jamet, Georgios Vavouliotis, Marti Torrents, Dimitrios Chasapis, Marc Casas | 2026-05-12 | 下载 | Modern server workloads exhibit massive instruction footprints that heavily pressure the processor front-end, making L1 instruction (L1I) prefetching critical for sustaining performance. |
| Heterogeneous SoC Integrating an Open-Source Recurrent SNN Accelerator for Neuromorphic Edge Computing on FPGA | Michelangelo Barocci, Vittorio Fra, Enrico Macii, Gianvito Urgese | 2026-05-12 | 下载 | The growing popularity of Spiking Neural Networks (SNNs) and their applications has led to a significant fast-paced increase of neuromorphic architectures capable of mimicking the spike-based data pro... |
| Runtime Calibration as State-Trajectory Feedback Control in Quantum-Classical Workflows | Xiaolong Deng | 2026-05-12 | 下载 | In superconducting devices running variational workloads, gate and readout fidelities drift on hour timescales, while existing runtime schedulers treat backend quality as static. |
| Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications | Julien Brandoit, Arthur Fyon, Damien Ernst, Guillaume Drion | 2026-05-12 | 下载 | Sequence learning is dominated by Transformers and parallelizable recurrent neural networks (RNNs) such as state-space models, yet learning long-term dependencies remains challenging, and state-of-the... |