Skip to content

2026-09-10 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Vortex: Bridging Extreme Compression and Efficient LLM InferenceHaoxuan Shan, Cong Guo, Bowen Duan, Chiyue Wei, Feng Cheng, Yuzhe Fu, Yintao He, Hai "Helen" Li, Yiran Chen2026-09-10下载Extreme compression techniques, including vector quantization (VQ) and input-dependent sparsity, can significantly reduce the memory footprint of large language models (LLMs).
Efficient Vision-Language-Action Management and Serving for Robot FactoriesDionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri, Christina Giannoula2026-09-10下载Vision-Language-Action (VLA) models show high robotic manipulation capabilities via a two-stage design: a Vision-Language Model (VLM) stage followed by an Action Diffusion Transformer (ADiT) stage.
AccelForge: Comprehensive Modeling and Co-Design Framework for AI AcceleratorsTanner Andrulis, Michael Gilbert, Vivienne Sze, Joel S. Emer2026-09-10下载Tensor algebra workloads, of which deep neural networks are prominent examples, are energy-intensive workloads in modern datacenter and edge deployments, making accelerators necessary to achieve energ...
A Time-Based Readout for Vector-Matrix Multiplication in Fully Analog Memristive SNNsElia Mateu-Barriendos, Álvaro Gómez-Pau, Josep Rius, Daniel Arumí, Rosa Rodríguez-Montañés, Salvador Manich2026-09-10下载Artificial neural networks rely on vector-matrix multiplications (VMMs), whose implementation in von Neumann architectures is dominated by costly data movement between memory and processing units.
From Grid to Chip: Power Architecture, Stability, and Flexibility of AI Data CentersYubo Song, Rui Kong, Takuro Umihara, Pooya Davari, Frede Blaabjerg, Subham Sahoo2026-09-10下载The rapid growth of artificial intelligence (AI) computing is transforming data centers into large, dynamic electrical loads. Their deployment is primarily constrained by energy availability and grid-...
PHAT: PHotonic Accelerator for TFHEGuowei Yang, Farbin Fayza, Beren Aydoğan, Carlos A. Ríos Ocampo, Ayse K. Coskun, Ajay Joshi2026-09-10下载Fully Homomorphic Encryption (FHE) enables secure computation on encrypted data, making it a promising solution for privacy-preserving applications in the cloud.
CHERI-D Reincarnate: efficient multicore CHERI temporal memory safety through allocation reincarnation (draft version)Yuecheng Wang, Jonathan Woodruff, Simon W. Moore2026-09-10下载We propose CHERI-D Reincarnate (Reinc), an architectural extension to CHERI for scalable and efficient temporal memory safety. Prior work CHERI-D has a finite-width generation ID stored at a fixed loc...
Entwine: Coordinating Tiled Computation and Fine-Grained Communication across GPUsKai Ma, Quanfeng Lv, Jingguo Ge, Bowei Dai, Kefan Ruan2026-09-10下载Modern high-performance GPU computations partition tensors into tiles to exploit data reuse and parallelism. Individual tile computations complete earlier than the full tensor computation, creating op...
PATTON: Enabling Commodity PIM for Production LLM ServingHangyeol Kim, Sanghyun Lee, Teokkyu Suh, Joo-Young Kim2026-09-10下载Processing-in-Memory (PIM) is promising for accelerating memory-bound decode attention, but attention acceleration alone is insufficient for production LLM serving, where engines dynamically allocate,...
Bio-inspired Learning and Decision-Making with Probabilistic In-Memory Computing Hardware: Part 2Thomas Dalgaty, Eiji Kawasaki, Miguel de Prado, Tommaso Salvatori, Germain Haugou, Eric Flamand2026-09-10下载This report extends our previous work (Part 1), which introduced an energy-based model for learning and decision-making under uncertainty. The model leverages stochastic Langevin dynamics to continuou...
BEACON: A Versatile Accelerator for Computational Pathology ApplicationsSumanth Gudaparthi, Ananth Krishna Prasad, Lin Jia, Rajeev Balasubramonian, Srinivasan Parthasarathy2026-09-10下载While accelerators for AI have seen great commercial success, it is challenging to replicate that success for other specialized domains due to a number of factors.
Fengshui: Demystifying Chiplet Ecosystem and Bespoke Neural Network Accelerator CodesignHaoran Jin, Jirong Yang, Zhiheng Zhang, Justin Shin, Barry Lyu, Kangqi Zhang, Yunpeng Liu, Nathan Bleier2026-09-10下载Modern ML workloads, with stringent latency and energy constraints, are increasingly hard to run efficiently on homogeneous commodity hardware.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
AKTS: Sub-Microsecond Kernel Policy Switching for Language-Model AgentsMohammadali Khodabandehlou, Mahdi Alizadeh2026-09-10下载GPU-backed LLM servers often multiplex interactive requests with background batch work on the same CPUs. During a request burst, the scheduler should protect time-to-first-token; between bursts, it sh...
Memory Compression for High-Fanout Agent SandboxesMengming Li, Ceyu XU, Qijun Zhang, Jiangnan Yu, Xiangfeng Sun, Haohui Mai, Zhiyao Xie2026-09-10下载High-fanout agent workloads create a growing memory bottleneck because a single task may spawn many concurrent sandbox sessions. Yet these sandboxes are far from independent: they originate from a sha...

基于 VitePress 构建 · 使用本地搜索查找论文