2026-06-08
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| GRAFT: Graphlet-Triggered Backdoor Attack on GNN-Based Hardware Security Systems | Sanaz Kazemi Abharian, Sai Manoj Pudukotai Dinakarrao | 2026-06-08 | 下载 | The globalization of the integrated circuit (IC) supply chain increases the risk of security threats, such as hardware Trojans (HTs) and the theft of intellectual property (IP). |
| Fault Characterization and Hardening of Combinational Standard Cells Using 3D-TCAD Simulations for Cyber-Physical Systems | Ali Zarei, Amir M. Hajisadeghi, Hamid R. Zarandi | 2026-06-08 | 下载 | Cyber-physical systems (CPSs) are increasingly employed in applications with various levels of mission criticality, making the reliability of digital system components essential for maintaining servic... |
| A Generic Modulo-(2^n\pmδ) RNS Multiplier Based on Twit Representation | Saeid Gorgin, Amirhossein Sadr, Behzad Salami, Dara Rahmati | 2026-06-08 | 下载 | Modular multiplication is a fundamental arithmetic primitive in Residue Number Systems (RNS) and is often the dominant source of delay, area, and energy consumption in RNS datapaths used in cryptograp... |
| An 84-Format Numeric Catalog with Bit-Exact Conformance Vectors: A Vendor-Neutral Reference for FP8, BF16, MXFP4, and Microscaling Formats | Dmitrii Vasilev | 2026-06-08 | 下载 | Numeric format proliferation in machine learning hardware -- FP8 (E4M3 and E5M2), BF16, MXFP4, microscaling block formats, and dozens of research variants -- has outpaced the availability of vendor-ne... |
| SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance | Rya Sanovar, Srikant Bharadwaj, Hritvik Taneja, Moinuddin Qureshi | 2026-06-08 | 下载 | Retrieval-Augmented Generation (RAG) injects LLM queries with relevant documents to improve response quality. This injection increases prompt length and slows time to first token (TTFT). |
| Toward Intelligent Prefetching: A Survey on Complex Memory Access Prediction Techniques | Sheel Sindhu Manohar | 2026-06-08 | 下载 | Data prefetching is a critical technique for bridging the processor-memory performance gap by predicting future memory accesses and retrieving data into on-chip caches before demand. |
| OpenOpt: An Open-Source SRAM Optimizer Based on Equivalent Circuit Model | Yikai Wang, Yiheng Wu, Can Wang, Bohao Liu, Junhao Ma, Zhuohua Liu, Qinxin Mei, Shan Shen | 2026-06-08 | 下载 | This paper proposes a co-optimization framework that jointly optimizes SRAM architecture and transistor sizing using equivalent circuit models. |
| SPARX: Secure and Privacy-Aware Approximate CNN Acceleration with Edge RISC-V SoC | Sonu Kumar, Akash Sankhe, Mukul Lokhande, Santosh Kumar Vishvakarma | 2026-06-08 | 下载 | Edge-AI systems increasingly require real-time CNN inference under strict energy, performance, security, and privacy constraints. Approximate computing improves hardware efficiency by exploiting the e... |
| NeuDW-CIM: a 65-nm 0.8-pJ/Sop Reconfigurable Neuromorphic Compute-in-Memory Macro with Nonlinear Dendrites and K-Winners | Junyi Yang, Yahan Yang, Shuai Dong, Biyan Zhou, Ye Ke, Zhengnan Fu, Xin Si, An Guo, Peng Zhou, Arindam Basu | 2026-06-08 | 下载 | This work presents NeuDW-CIM, a highly efficient neuromorphic Compute-in-Memory (CIM) macro for Spiking Neural Networks (SNNs) implemented in 65 nm CMOS. |
| LongRTL: Graph-Similarity-Guided LLM-driven Long Context RTL Optimization | Yuyang Ye, Che-Kuan Shen, Xiangfei Hu, Yuchen Liu, Shuo Yin, Xufeng Yao, Bei Yu, Tsung-Yi Ho | 2026-06-08 | 下载 | Large Language Models (LLMs) show great promise in RTL code generation and optimization. However, real-world RTL designs are typically long, entangled, and poorly modularized, posing a major challenge... |
| PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM Inference | Runyang Tian, Yanru Chen, Weihong Xu, Tajana Šimunić Rosing | 2026-06-08 | 下载 | Large language models are increasingly deployed on edge devices with tight power and area budgets. While mixed-precision GEMM reduces arithmetic complexity, quantized inference is often dominated by d... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Hardware-accelerated Aggregation: Unification and Specialization | Alireza Shateri, Hongshi Tan, Michael Ng, Bingsheng He, Qizhen Zhang | 2026-06-08 | 下载 | The high efficiency of domain-specific hardware has sparked substantial interest in adopting accelerators in data analytics systems. Among many choices, GPUs and FPGAs thrived as two popular solutions... |
| AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis | Jaber Jaber, Osama Jaber | 2026-06-08 | 下载 | AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one launch, with no per-model hand-written CUDA. |
| FMplex: Model Virtualization for Serving Extensible Foundation Models | Hetvi Shastri, Pragya Sharma, Walid A. Hanafy, David Irwin, Mani Srivastava, Prashant Shenoy | 2026-06-08 | 下载 | Foundation models (FMs) are increasingly used as backbones for downstream tasks across language, vision, time-series, and multimodal applications. |
| Parent-Hash DAG: A Cost Analysis of Constant-Time Append for On-Chain Registries | Ian C. Moore, Fernando Paredes Garcia | 2026-06-08 | 下载 | Provenance trees are append-only directed acyclic graphs of artifact registrations anchored on a public blockchain, recently introduced as the data substrate of operator-gated provenance infrastructur... |
| Coupling Complementary Simulations for Combined Performance and Energy Optimization | Adel Dabah, Gregor Häfner, Sonja Happ, Simon Pickartz, Marcus Müller, Andreas Herten | 2026-06-08 | 下载 | Polymer simulations are among the most computationally demanding workloads in soft-matter research, often requiring days of execution and high energy consumption to achieve physically meaningful resul... |
| Engineering Scalable Distributed List Ranking | Peter Sanders, Matthias Schimek, Tim Niklas Uhl, Thomas Weidmann | 2026-06-08 | 下载 | The list ranking problem is one of the classical problems of parallel computing, with nontrivial algorithms and many applications as a subroutine for solving other problems. |
| Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads | Minyu Cui, Miquel Pericas | 2026-06-08 | 下载 | The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems. |
| CANS: Accelerating Multiuser Collaborative Edge Inference via Cooperative Autodidactic NeuroSurgeon | Zheshun Wu, Ziyang Zhang, Changyao Lin, Zenglin Xu, Jie Liu | 2026-06-08 | 下载 | Recently, mobile edge computing (MEC)-enabled collaborative deep neural network (DNN) inference has emerged as a promising approach for delivering intelligent services to resource-constrained mobile d... |
| AutoPilot: Learning to Steer High Speed Robust BFT | Liangrong Chen, Yue Zhang, Eric Zhou, Mohammad Javad Amiri, Ryan Marcus, Chenyuan Wu | 2026-06-08 | 下载 | Recent Byzantine Fault Tolerant (BFT) protocols achieve strong performance by combining the low-latency advantages of leader-based BFT protocols with the high-throughput benefits of DAG-based data dis... |
| Concepts in Practice: C++ MPI Bindings for the HPC Ecosystem. From a Standardizable Core to a Composable Interface | Tim Niklas Uhl, Matthias Schimek, Daniel Brommer | 2026-06-08 | 下载 | The official C++ MPI bindings were removed from the standard in 2008, leaving a gap that numerous third-party libraries have attempted to fill. |
| Chimera: Protocol-Aware Recovery for Confidential BFT Consensus | Tong Liu, Xiaoqing Wen, Ziwei Zhou, Si Liu, Jianyu Niu, Cong Wang, Yinqian Zhang | 2026-06-08 | 下载 | Trusted Execution Environments (TEEs) have enabled confidential Byzantine Fault-Tolerant (BFT) consensus systems with confidentiality and improved scalability. |
| Fairness-Aware and Latency-Controllable Scheduling for Chunked-Prefill LLM Serving | Haoxin Liu, Jiayi Wang, Yueshen Xu, Rui Li | 2026-06-08 | 下载 | As large language models (LLMs) are increasingly deployed with highly heterogeneous workloads, chunked-prefill execution has emerged as a mainstream serving architecture. |
| When More Cores Hurts: The Vector Database Scaling Paradox in HPC | Seth Ockerman, Song Young Oh, Amal Gueroudji, Rochana Chaturvedi, Philip Carns, Nicholas Chia, Matthieu Dorier, Robert Latham, Tanwi Mallick, Swan Perarnau, Robert Underwood, Kyle Chard, Ian Foster, Robert Ross, Shivaram Venkataraman | 2026-06-08 | 下载 | Vector databases have been designed and optimized for cloud environments; however, emerging scientific AI workloads (e.g., molecular search, meteorological trajectory detection, and literature-driven ... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Secrets Best Not Shared: DNS Privacy Enhancements for the Constrained IoT | Martine S. Lenders, Thomas C. Schmidt, Matthias Wählisch | 2026-06-08 | 下载 | Attackers often identify DNS traffic to disrupt or compromise Internet services. While prior work has focused on encrypting queries using DNS over TLS, HTTPS, or QUIC to counter such attacks, we consi... |
| Zero Touch Predictive Orchestration: Automating Time-Series Models for the Cloud-Edge Continuum | Abd Elghani Meliani, Arora Sagar, Adlen Ksentini, Raymond Knopp | 2026-06-08 | 下载 | The Cloud-Edge Continuum (CEC) enables latency-critical applications by distributing resources to the far edge, but its extreme volatility makes proactive Zero Touch Management via time-series forecas... |
| Strict-Priority Packet Delay in Switches with Transmit-Ring Buffering | Yash Deshpande, Quirin Vogel, Wolfgang Kellerer | 2026-06-08 | 下载 | Strict Priority (SP) scheduling is widely used at switch egress to provide low-latency service to high-priority (HP) traffic. Existing deterministic and stochastic latency models typically account for... |
| STEPS: Semantic-Contract-Guided Scheduling for LLM-Assisted Natural-Language-Driven Edge AI Services | Houyi Qi, Minghui Liwang, Xianbin Wang, Seyyedali Hosseinalipour | 2026-06-08 | 下载 | Networked AI services are increasingly delivered through edge infrastructures to support latency-sensitive applications. Edge scheduling is critical for deciding where and how AI services are executed... |
| Just-in-time Restoration with Distributed Fiber Sensing in Metropolitan Optical Networks | Sleman Mouammar, Italo B. Brasileiro, Andre C. Drummond | 2026-06-08 | 下载 | Distributed Fiber Sensing (DFS) leverages optical backscattering signals to predict failure events and enable just-in-time restoration in metropolitan optical networks, i.e. |
| Semantic and Task-Oriented V2X Communications: Pushing the Limits of V2X Networks Scalability | Luca Lusvarghi, Javier Gozalvez, Mohammad Irfan Khan, Seyhan Ucar, Miguel Sepulcre, Onur Altintas | 2026-06-08 | 下载 | Scalable Vehicle-to-Everything (V2X) networks are key to support the large-scale deployment of connected and automated mobility. However, the scalability of V2X networks is currently challenged by the... |
| Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network Operations | Arun Malik | 2026-06-08 | 下载 | Cloud network infrastructure at hyperscale presents unique operational challenges where traditional human-driven incident response cannot keep pace with the volume, velocity, and complexity of failure... |
| Block-A-Mole: The Sustainability Frontier of Moving-Target Censorship Resistance | Anindya Maiti | 2026-06-08 | 下载 | Internet censorship affects over four billion people, and deployed circumvention systems share a common weakness: their endpoints are fixed and discoverable, so a patient censor can enumerate and bloc... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| FMplex: Model Virtualization for Serving Extensible Foundation Models | Hetvi Shastri, Pragya Sharma, Walid A. Hanafy, David Irwin, Mani Srivastava, Prashant Shenoy | 2026-06-08 | 下载 | Foundation models (FMs) are increasingly used as backbones for downstream tasks across language, vision, time-series, and multimodal applications. |
| TinyContainer: Container Runtime Middleware Enabling Multi-tenant Microcontrollers with Built-in Security | Bastien Buil, Chrystel Gaber, Samuel Legouix, Emmanuel Baccelli, Samia Bouzefrane | 2026-06-08 | 下载 | Software containerization technologies for resource-limited devices enable multi-tenant microcontrollers, which allow running multiple applications with different permission levels. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| An 84-Format Numeric Catalog with Bit-Exact Conformance Vectors: A Vendor-Neutral Reference for FP8, BF16, MXFP4, and Microscaling Formats | Dmitrii Vasilev | 2026-06-08 | 下载 | Numeric format proliferation in machine learning hardware -- FP8 (E4M3 and E5M2), BF16, MXFP4, microscaling block formats, and dozens of research variants -- has outpaced the availability of vendor-ne... |
| AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis | Jaber Jaber, Osama Jaber | 2026-06-08 | 下载 | AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one launch, with no per-model hand-written CUDA. |
| Correlation Is Not Enough: Embedding Human Metadata for Individual Causal Discovery | Suraj Biswas, Saurabh Gupta, Pritam Mukherjee | 2026-06-08 | 下载 | Ask a pretrained biomedical language model whether "cortisol 28 ug/dL" and "stock-market volatility" are related, and it returns a cosine similarity of 0.83 on a scale where 1.0 means identical. |
| Fairness-Aware and Latency-Controllable Scheduling for Chunked-Prefill LLM Serving | Haoxin Liu, Jiayi Wang, Yueshen Xu, Rui Li | 2026-06-08 | 下载 | As large language models (LLMs) are increasingly deployed with highly heterogeneous workloads, chunked-prefill execution has emerged as a mainstream serving architecture. |