Skip to content

2026-06-21 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Architecture for Health Initiative (Arch4Health): Computational Challenges in Health-Related Applications and the Role of Computer Architecture in Addressing ThemNika Mansouri Ghiasi, Konstantina Koliogeorgi, Onur Mutlu2026-06-21下载Recent biotechnological advances enable high-throughput, low-cost, and accurate biological data generation. This wealth of data enables unique opportunities for advancing healthcare.
Design and Development of a Neuromorphic Silicon Suite: PVT Sensing, Stochastic LIF Inference, On-Chip STDP Learning, and Crossbar ProgrammingPoornima Kumaresan, Santhosh Sivasubramani2026-06-21下载Edge neuromorphic systems need compact, configurable hardware that combines probabilistic inference, local learning, and an interface to emerging analogue memory.
ColumnKeeper: Efficient Solutions to the ColumnDisturb Vulnerability in DRAM-based SystemsAndreas Kosmas Kakolyris, F. Nisa Bostanci, Ataberk Olgun, Ismail Emir Yuksel, Harsh Songara, Konstantinos Marios Sgouras, Umut Baser, Konstantinos Kanellopoulos, A. Giray Yaglikci, Onut Mutlu2026-06-21下载Modern DRAM chips are vulnerable to read disturbance phenomena such as RowHammer and RowPress, which induce bitflips after accessing nearby rows a certain number of times (the read disturbance thresho...
Multi-Level Resistive Synapses for On-Chip Neural Networks: A Physics-Based Design of a Memristive Crossbar Fabric with Quasi-Continuous Conductance StatesDavid Alejandro Trejo Pizzo2026-06-21下载Building on resistive communication, this paper presents a physics-based design of an on-chip neural network with multi-level memristive synapses supporting a dense spectrum of conductance states.
Non-Uniform L2 Cache Latency Across the Streaming Multiprocessors of an NVIDIA L40Faruk Alpay, Baris Basaran2026-06-21下载The NVIDIA L40 exposes a 96 MiB L2 cache usually modeled as one uniform pool with a single hit latency. We show this is wrong at the granularity a kernel sees: L2-hit latency depends strongly and repr...
NeutronSparse: Coordinating Heterogeneous Engines for Sparse Matrix Multiplication on NPUsXin Ai, Zeyu Ling, Hao Yuan, Qiange Wang, Yanfeng Zhang, Yutao Peng, Ge Yu2026-06-21下载Sparse matrix-matrix multiplication (SpMM) is a fundamental data operation for large-scale sparse data processing. With NPUs increasingly deployed in data centers for their performance and energy effi...
DejaVu: Why You Should Write to Your DRAM Rows Twice, CarefullyHaocong Luo, İsmail Emir Yüksel, Ataberk Olgun, Nisa Bostanci, Orhun Ecemiş, Abdullah Giray Yağlıkçı, Onur Mutlu2026-06-21下载We provide the first experimental demonstration of DejaVu, a phenomenon where the data previously written to DRAM cells affects DRAM's vulnerability to read disturbance.
Apple Neural Engine: Architecture, Programming, and PerformanceSpencer H. Bryngelson2026-06-21下载The Apple Neural Engine (ANE) is the fixed-function matrix accelerator that has shipped in Apple systems-on-chip since the A11-class iPhone and iPad chips and the M1-class Mac chips, exposed to applic...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Architecture for Health Initiative (Arch4Health): Computational Challenges in Health-Related Applications and the Role of Computer Architecture in Addressing ThemNika Mansouri Ghiasi, Konstantina Koliogeorgi, Onur Mutlu2026-06-21下载Recent biotechnological advances enable high-throughput, low-cost, and accurate biological data generation. This wealth of data enables unique opportunities for advancing healthcare.
ASAP: A Disaggregated and Asynchronous Inference System for MoE PrefillWeiwei Chen, Shuang Chen, Lele Li, Qiang Hu, Han Li, Xin Ye, Ming Yan, Zhibin Yu2026-06-21下载Mixture-of-Experts (MoE) models have become the de facto standard for scaling large language models. To maintain computational efficiency, modern MoE serving systems typically employ a hybrid parallel...
Fed-CausalDiff: Decoupled Synchronization for Federated Do-Simulation and Policy EvaluationPengfei Li, Mohammad Khalil2026-06-21下载While federated learning enables collaborative modelling on decentralised data, standard methods merely fit historical observations. This purely observational approach is fundamentally insufficient fo...
Hardwired Pattern Formation by Mobile Robots with Common Unit DistanceYuta Kojima, Sébastien Tixeuil, Yukiko Yamauchi2026-06-21下载The pattern formation (PTF) problem requires mobile robots to form a specified target pattern. Existing papers investigated the PTF problem and revealed the effect of obliviousness and synchronization...
Semantic Non-Assembly: Privacy by Architectural Inertness Under Component ExposureSam Ryan2026-06-21下载Existing privacy frameworks emphasize confidentiality, access control, appropriate information flow, or statistical disclosure limitation. We introduce a complementary class of privacy guarantee (Sema...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
One-Prompt Censorship Evasion via Generative Diffusion ModelsShiyi Ling, Yuhang Gan, Chen Qian2026-06-21下载The escalating arms race between Internet censorship and evasion has driven censors to evolve from static rule-based filtering to sophisticated deep learning-based traffic analysis.
Radio Resource Management for the Uplink of Hybrid Beamforming SystemsYuan Quan, Haseen Rahman, Catherine Rosenberg2026-06-21下载This paper studies radio resource management (RRM) for the uplink of a multi-channel cellular system with hybrid beamforming based on analog beamforming using predefined codebooks and zero-forcing dig...
Making Quantum Networks Work: Routing, Calibration, and Programmable Quantum RepeatersVinay Kumar2026-06-21下载The quantum internet enables distribution of quantum states across distant nodes, supporting secure communication, distributed computing, and quantum sensing.
SHACR: A Graph-Augmented Semi-Autonomous Framework for Multi-Class Conflict Resolution in Smart Home IoT AutomationLeena Marghalani, Walid Aljoby, Suayb S. Arslan2026-06-21下载Smart home automation increasingly relies on user-defined rules across heterogeneous IoT devices. While these rules appear harmless in isolation, their concurrent execution creates hidden, cross-rule ...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Apple Neural Engine: Architecture, Programming, and PerformanceSpencer H. Bryngelson2026-06-21下载The Apple Neural Engine (ANE) is the fixed-function matrix accelerator that has shipped in Apple systems-on-chip since the A11-class iPhone and iPad chips and the M1-class Mac chips, exposed to applic...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Enabling Cloud-Level Accuracy in Edge AI through IoT Data PreprocessingAygün Varol, Katarzyna Kołodziej, Łukasz Sobczak, Michał Romaszewski, Przemysław Głomb, Naser Hossein Motlagh, Mirka Leino, Johanna Virkki2026-06-21下载Large language models (LLMs) offer a natural-language interface for interpreting Internet of Things (IoT) sensor data in smart environments; however, cloud deployment introduces latency, privacy, and ...
When Is a Columnar Scan Bandwidth-Bound? A Decode-Throughput Law and Its Cross-Hardware ValidationMadhulatha Mandarapu, Sandeep Kunkunuru2026-06-21下载A columnar scan that decompresses, filters, and aggregates should be limited only by memory bandwidth (the roofline floor T >= BytesRead/beta), yet real kernels are often compute-bound and leave band...
Apple Neural Engine: Architecture, Programming, and PerformanceSpencer H. Bryngelson2026-06-21下载The Apple Neural Engine (ANE) is the fixed-function matrix accelerator that has shipped in Apple systems-on-chip since the A11-class iPhone and iPad chips and the M1-class Mac chips, exposed to applic...

基于 VitePress 构建 · 使用本地搜索查找论文