2026-09-10
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Vortex: Bridging Extreme Compression and Efficient LLM Inference | Haoxuan Shan, Cong Guo, Bowen Duan, Chiyue Wei, Feng Cheng, Yuzhe Fu, Yintao He, Hai "Helen" Li, Yiran Chen | 2026-09-10 | 下载 | Extreme compression techniques, including vector quantization (VQ) and input-dependent sparsity, can significantly reduce the memory footprint of large language models (LLMs). |
| Efficient Vision-Language-Action Management and Serving for Robot Factories | Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri, Christina Giannoula | 2026-09-10 | 下载 | Vision-Language-Action (VLA) models show high robotic manipulation capabilities via a two-stage design: a Vision-Language Model (VLM) stage followed by an Action Diffusion Transformer (ADiT) stage. |
| AccelForge: Comprehensive Modeling and Co-Design Framework for AI Accelerators | Tanner Andrulis, Michael Gilbert, Vivienne Sze, Joel S. Emer | 2026-09-10 | 下载 | Tensor algebra workloads, of which deep neural networks are prominent examples, are energy-intensive workloads in modern datacenter and edge deployments, making accelerators necessary to achieve energ... |
| A Time-Based Readout for Vector-Matrix Multiplication in Fully Analog Memristive SNNs | Elia Mateu-Barriendos, Álvaro Gómez-Pau, Josep Rius, Daniel Arumí, Rosa Rodríguez-Montañés, Salvador Manich | 2026-09-10 | 下载 | Artificial neural networks rely on vector-matrix multiplications (VMMs), whose implementation in von Neumann architectures is dominated by costly data movement between memory and processing units. |
| From Grid to Chip: Power Architecture, Stability, and Flexibility of AI Data Centers | Yubo Song, Rui Kong, Takuro Umihara, Pooya Davari, Frede Blaabjerg, Subham Sahoo | 2026-09-10 | 下载 | The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-... |
| PHAT: PHotonic Accelerator for TFHE | Guowei Yang, Farbin Fayza, Beren Aydoğan, Carlos A. Ríos Ocampo, Ayse K. Coskun, Ajay Joshi | 2026-09-10 | 下载 | Fully Homomorphic Encryption (FHE) enables secure computation on encrypted data, making it a promising solution for privacy-preserving applications in the cloud. |
| CHERI-D Reincarnate: efficient multicore CHERI temporal memory safety through allocation reincarnation (draft version) | Yuecheng Wang, Jonathan Woodruff, Simon W. Moore | 2026-09-10 | 下载 | We propose CHERI-D Reincarnate (Reinc), an architectural extension to CHERI for scalable and efficient temporal memory safety. Prior work CHERI-D has a finite-width generation ID stored at a fixed loc... |
| Entwine: Coordinating Tiled Computation and Fine-Grained Communication across GPUs | Kai Ma, Quanfeng Lv, Jingguo Ge, Bowei Dai, Kefan Ruan | 2026-09-10 | 下载 | Modern high-performance GPU computations partition tensors into tiles to exploit data reuse and parallelism. Individual tile computations complete earlier than the full tensor computation, creating op... |
| PATTON: Enabling Commodity PIM for Production LLM Serving | Hangyeol Kim, Sanghyun Lee, Teokkyu Suh, Joo-Young Kim | 2026-09-10 | 下载 | Processing-in-Memory (PIM) is promising for accelerating memory-bound decode attention, but attention acceleration alone is insufficient for production LLM serving, where engines dynamically allocate,... |
| Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 2 | Thomas Dalgaty, Eiji Kawasaki, Miguel de Prado, Tommaso Salvatori, Germain Haugou, Eric Flamand | 2026-09-10 | 下载 | This report extends our previous work (Part 1), which introduced an energy-based model for learning and decision-making under uncertainty. The model leverages stochastic Langevin dynamics to continuou... |
| BEACON: A Versatile Accelerator for Computational Pathology Applications | Sumanth Gudaparthi, Ananth Krishna Prasad, Lin Jia, Rajeev Balasubramonian, Srinivasan Parthasarathy | 2026-09-10 | 下载 | While accelerators for AI have seen great commercial success, it is challenging to replicate that success for other specialized domains due to a number of factors. |
| Fengshui: Demystifying Chiplet Ecosystem and Bespoke Neural Network Accelerator Codesign | Haoran Jin, Jirong Yang, Zhiheng Zhang, Justin Shin, Barry Lyu, Kangqi Zhang, Yunpeng Liu, Nathan Bleier | 2026-09-10 | 下载 | Modern ML workloads, with stringent latency and energy constraints, are increasingly hard to run efficiently on homogeneous commodity hardware. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AKTS: Sub-Microsecond Kernel Policy Switching for Language-Model Agents | Mohammadali Khodabandehlou, Mahdi Alizadeh | 2026-09-10 | 下载 | GPU-backed LLM servers often multiplex interactive requests with background batch work on the same CPUs. During a request burst, the scheduler should protect time-to-first-token; between bursts, it sh... |
| Memory Compression for High-Fanout Agent Sandboxes | Mengming Li, Ceyu XU, Qijun Zhang, Jiangnan Yu, Xiangfeng Sun, Haohui Mai, Zhiyao Xie | 2026-09-10 | 下载 | High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a sha... |