Skip to content

2026-06-05 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
ScaleDisturb: Exploiting Temporal Asymmetry to Amplify Read Disturbance in Modern DRAM ChipsJikun Wang, Haocong Luo, Ataberk Olgun, İsmail Emir Yüksel, A. Giray Yağlıkçı, Yu Liang, F. Nisa Bostancı, Mohammad Sadrosadati, Onur Mutlu2026-06-05下载DRAM suffers from read disturbance phenomena (e.g., RowHammer and RowPress), where repeatedly accessing or continuously keeping open a DRAM row (aggressor row) induces bitflips in other physically nea...
A 65 nm Trustworthy Hypoglycemia Forecasting Engine Achieving 11.3 nJ per InferenceBoyang Cheng, Jianbo Liu, Pengyu Ren, Xueji Zhao, Steven Davis, Likai Pei, Zephan M. Enciso, Kai Ni, Ningyuan Cao2026-06-05下载Diabetes affects millions of people and requires reliable continuous glucose monitoring for early hypoglycemia warning. However, medical AI systems must be not only accurate and energy efficient, but ...
A 65-nm Privacy-Preserving Neuromorphic Encoder With 7.13-nJ Efficiency, 2.38-Mb/mm^2 Item-Memory Density, and Federated Learning SupportBoyang Cheng, Jianbo Liu, Steven Davis, Zephan M. Enciso, Likai Pei, Xueji Zhao, Muya Chang, Ningyuan Cao2026-06-05下载The increasing demand for privacy-preserving personal data analytics in smart assistants, wearable health monitors, and context-aware systems calls for hardware that is both energy-efficient and secur...
A 65 nm Multi-Modal Bayesian Inference Engine with 16.3 fJ/Sample Calibration-Free GRNG for Risk-Aware At-Home Skin Lesion ScreeningSteven Davis, Likai Pei, Jianbo Liu, Zephan M. Enciso, Boyang Cheng, Xueji Zhao, Danny Z. Chen, Ningyuan Cao2026-06-05下载We present a 65-nm risk-aware multimodal Bayesian inference engine for privacy-preserving, fully on-device skin lesion screening under uncontrolled at-home conditions.
MailoHLS: Multi-Adapter Structure-Aware Learning for Pareto-Driven HLS Pragma OptimizationElena Vouvali, Dimosthenis Masouros, Aggelos Ferikoglou, Dimitrios Soudris, Sotirios Xydis2026-06-05下载High-Level Synthesis (HLS) enables rapid development of FPGA accelerators, yet achieving high-quality results (QoR) remains challenging due to the large and irregular design space induced by compiler ...
Distributed Persistence Domain for Persistent Memory PoolingKhan Shaikhul Hadi, Andres David Delgado, Naveed Ul Mustafa, Mark Heinrich, Hao Zheng, Yan Solihin2026-06-05下载Compute Express Link (CXL) enables memory pooling over disaggregated memory, offering the potential to improve resource utilization in persistent memory systems.
Terastal: Layer-Variant-based Scheduling for Real-Time Multi-DNN Workloads on Heterogeneous AcceleratorsSing-Yao Wu, Fengshuo Song, Eli Bozorgzadeh2026-06-05下载Heterogeneous DNN accelerators improve soft real-time multi-DNN execution by mapping each layer to its preferred accelerator to reduce latency.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Cost-Aware Speculative Execution for LLM-Agent Workflows: An Integrated Five-Dimension MethodFaisal Fareed2026-06-05下载LLM-agent workflows chain model calls and tool invocations, and spend most of their wall-clock time waiting on upstream operations before downstream ones can start.
Large-Scale Regularized Matching on GPU ClustersAida Rahmattalabi, Gregory Dexter, Sanjana Garg, Qinquan Song, Shenyinying Tu, Yuan Gao, Zhipeng Wang, Rahul Mazumder2026-06-05下载Production decision systems such as ad allocation or content matching involve millions of users and thousands of items, reducing to large-scale linear programs with sparse block-diagonal structure acr...
Twelve quick tips for designing AI-driven HPC workflowsJamie J. Alnasir2026-06-05下载High-performance computing (HPC) clusters remain the backbone of large-scale scientific computation, traditionally executing deterministic, linear pipelines optimised for predictable performance.
Hierarchical Certified Semantic Commitment for Byzantine-Resilient LLM-Agent CollaborationHaoran Xu, Lei Zhang, Iadh Ounis, Xianbin Wang2026-06-05下载Byzantine collaboration among large-language-model agents requires a finality-control primitive: given delivered stochastic, structured natural-language proposals, the protocol must decide whether the...
Clairvoyant: Predictive SJF Scheduling to Mitigate Head-of-Line Blocking in Serial LLM BackendsAravind Sundaresan2026-06-05下载Serial LLM inference backends -- such as Ollama -- process requests one at a time under FCFS admission, causing Head-of-Line Blocking (HOLB) under mixed workloads at high utilisation: short factual qu...
Predictive Autoscaling in Cloud-Native and Federated Cloud-Edge Computing Environments: A Taxonomy and Future DirectionsBablu Kumar, Anshul Verma, Rajkumar Buyya2026-06-05下载Autoscaling is a key capability in cloud-native systems, where dynamic workloads, heterogeneous environments, and latency-sensitive applications require efficient and adaptive resource management.
PCCL: Process Group-Aware Scalable and Generic Collective Algorithm SynthesizerWilliam Won, Kartik Lakhotia, Madhu Kumar, Sudarshan Srinivasan, Tushar Krishna2026-06-05下载Distributed machine learning has become increasingly important due to the massive scale of large-scale generative models. Both model parameters and data are distributed across many compute devices, wh...
Mission-Level Runtime Assurance Framework for Autonomous DrivingChieh Tsai, Salim Hariri2026-06-05下载This paper studies runtime safety for autonomous driving when high-level driving commands become faulty or unreliable. Unlike conventional runtime-safety approaches that mainly focus on immediate vehi...
Communication Strategy Selection for Multi-GPU 3D FDTD with Convolutional Perfectly Matched Boundary LayersVictory C. Obieke2026-06-05下载In this paper we describe a communication-strategy study for multi-GPU three-dimensional finite-difference time-domain computation with convolutional perfectly matched layer boundary conditions using ...
Terastal: Layer-Variant-based Scheduling for Real-Time Multi-DNN Workloads on Heterogeneous AcceleratorsSing-Yao Wu, Fengshuo Song, Eli Bozorgzadeh2026-06-05下载Heterogeneous DNN accelerators improve soft real-time multi-DNN execution by mapping each layer to its preferred accelerator to reduce latency.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Rethinking IoT Intrusion Detection: Augmenting Routing Metrics with Radio FeaturesYichang Sun, Andreas Johnsson, Sourasekhar Banerjee2026-06-05下载Machine learning-based intrusion detection systems (IDS) for RPL-based IoT networks often rely solely on routing layer features, which provide only a partial view of network behaviour.
From Privacy to Workflow Integrity: Communication-Graph Metadata in Autonomous Agent InteroperabilityBijaya Dangol2026-06-05下载Agent-interoperability protocols such as A2A and MCP standardize what agents say to one another, but assume address-based transport over HTTP(S).
DIFFRACT: Neuralized Utility Maximization for Wireless Networks by Differentiable ProgrammingChee Wei Tan, Siya Chen2026-06-05下载Next-generation wireless networks, including satellite-to-Open RAN systems, demand agile and intelligent resource management capable of handling dynamic multi-user interference under stochastic qualit...
i2Slicer: Enabling Flexible and Automated Orchestration of 5G SA End-to-End Network SlicesM. Catalan-Cid, A. Fernandez, D. Camps-Mur, S. Siddiqui2026-06-05下载5G network slicing implies a step forward in customizing radio access and core networks by allowing the creation of logical networks adapted to service requirements.
Federated Foundation Models over Vehicular NetworksKasra Borazjani, Fardis Nadimi, Payam Abdisarabshali, Owen Palinski, Allan Salihovic, Dinh Nguyen, Minghui Liwang, Seyyedali Hosseinalipour2026-06-05下载This paper presents a forward-looking vision for integrating the emerging multi-modal multi-task federated foundation models (M3T FedFMs) into vehicular networks, with the goal of unifying the express...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer KernelsLenore Mullin, Gaetan Hains2026-06-05下载The attention mechanism is the dominant computational bottleneck in modern transformer-based AI. Its standard implementation incurs quadratic memory traffic in the sequence length~nn, and DRAM access...
ANNS-AMP: Accelerating Approximate Nearest Neighbor Search via Adaptive Mixed-Precision ComputingMingkai Chen, Cheng Liu, Shengwen Liang, Lei Zhang, Xiaowei Li, Huawei Li2026-06-05下载Approximate nearest neighbor search(ANNS) is a critical kernel in modern applications such as LLM and recommendation systems.However,its efficiency is fundamentally limited by the need to compute dist...
Dependencies and Dataflow in Seed-Filter-Extend PipelinesShiv Sundram2026-06-05下载Comparing genomes is critical for discovering mutations, tracking evolutionary lineages, and advancing cross-species genomics. Fundamentally, this reduces to an O(n^2) string-matching dynamic programm...

基于 VitePress 构建 · 使用本地搜索查找论文