Skip to content

2026-06-08 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
GRAFT: Graphlet-Triggered Backdoor Attack on GNN-Based Hardware Security SystemsSanaz Kazemi Abharian, Sai Manoj Pudukotai Dinakarrao2026-06-08下载The globalization of the integrated circuit (IC) supply chain increases the risk of security threats, such as hardware Trojans (HTs) and the theft of intellectual property (IP).
Fault Characterization and Hardening of Combinational Standard Cells Using 3D-TCAD Simulations for Cyber-Physical SystemsAli Zarei, Amir M. Hajisadeghi, Hamid R. Zarandi2026-06-08下载Cyber-physical systems (CPSs) are increasingly employed in applications with various levels of mission criticality, making the reliability of digital system components essential for maintaining servic...
A Generic Modulo-(2^n\pmδ) RNS Multiplier Based on Twit RepresentationSaeid Gorgin, Amirhossein Sadr, Behzad Salami, Dara Rahmati2026-06-08下载Modular multiplication is a fundamental arithmetic primitive in Residue Number Systems (RNS) and is often the dominant source of delay, area, and energy consumption in RNS datapaths used in cryptograp...
An 84-Format Numeric Catalog with Bit-Exact Conformance Vectors: A Vendor-Neutral Reference for FP8, BF16, MXFP4, and Microscaling FormatsDmitrii Vasilev2026-06-08下载Numeric format proliferation in machine learning hardware -- FP8 (E4M3 and E5M2), BF16, MXFP4, microscaling block formats, and dozens of research variants -- has outpaced the availability of vendor-ne...
SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention InvarianceRya Sanovar, Srikant Bharadwaj, Hritvik Taneja, Moinuddin Qureshi2026-06-08下载Retrieval-Augmented Generation (RAG) injects LLM queries with relevant documents to improve response quality. This injection increases prompt length and slows time to first token (TTFT).
Toward Intelligent Prefetching: A Survey on Complex Memory Access Prediction TechniquesSheel Sindhu Manohar2026-06-08下载Data prefetching is a critical technique for bridging the processor-memory performance gap by predicting future memory accesses and retrieving data into on-chip caches before demand.
OpenOpt: An Open-Source SRAM Optimizer Based on Equivalent Circuit ModelYikai Wang, Yiheng Wu, Can Wang, Bohao Liu, Junhao Ma, Zhuohua Liu, Qinxin Mei, Shan Shen2026-06-08下载This paper proposes a co-optimization framework that jointly optimizes SRAM architecture and transistor sizing using equivalent circuit models.
SPARX: Secure and Privacy-Aware Approximate CNN Acceleration with Edge RISC-V SoCSonu Kumar, Akash Sankhe, Mukul Lokhande, Santosh Kumar Vishvakarma2026-06-08下载Edge-AI systems increasingly require real-time CNN inference under strict energy, performance, security, and privacy constraints. Approximate computing improves hardware efficiency by exploiting the e...
NeuDW-CIM: a 65-nm 0.8-pJ/Sop Reconfigurable Neuromorphic Compute-in-Memory Macro with Nonlinear Dendrites and K-WinnersJunyi Yang, Yahan Yang, Shuai Dong, Biyan Zhou, Ye Ke, Zhengnan Fu, Xin Si, An Guo, Peng Zhou, Arindam Basu2026-06-08下载This work presents NeuDW-CIM, a highly efficient neuromorphic Compute-in-Memory (CIM) macro for Spiking Neural Networks (SNNs) implemented in 65 nm CMOS.
LongRTL: Graph-Similarity-Guided LLM-driven Long Context RTL OptimizationYuyang Ye, Che-Kuan Shen, Xiangfei Hu, Yuchen Liu, Shuo Yin, Xufeng Yao, Bei Yu, Tsung-Yi Ho2026-06-08下载Large Language Models (LLMs) show great promise in RTL code generation and optimization. However, real-world RTL designs are typically long, entangled, and poorly modularized, posing a major challenge...
PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM InferenceRunyang Tian, Yanru Chen, Weihong Xu, Tajana Šimunić Rosing2026-06-08下载Large language models are increasingly deployed on edge devices with tight power and area budgets. While mixed-precision GEMM reduces arithmetic complexity, quantized inference is often dominated by d...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Hardware-accelerated Aggregation: Unification and SpecializationAlireza Shateri, Hongshi Tan, Michael Ng, Bingsheng He, Qizhen Zhang2026-06-08下载The high efficiency of domain-specific hardware has sparked substantial interest in adopting accelerators in data analytics systems. Among many choices, GPUs and FPGAs thrived as two popular solutions...
AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel SynthesisJaber Jaber, Osama Jaber2026-06-08下载AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one launch, with no per-model hand-written CUDA.
FMplex: Model Virtualization for Serving Extensible Foundation ModelsHetvi Shastri, Pragya Sharma, Walid A. Hanafy, David Irwin, Mani Srivastava, Prashant Shenoy2026-06-08下载Foundation models (FMs) are increasingly used as backbones for downstream tasks across language, vision, time-series, and multimodal applications.
Parent-Hash DAG: A Cost Analysis of Constant-Time Append for On-Chain RegistriesIan C. Moore, Fernando Paredes Garcia2026-06-08下载Provenance trees are append-only directed acyclic graphs of artifact registrations anchored on a public blockchain, recently introduced as the data substrate of operator-gated provenance infrastructur...
Coupling Complementary Simulations for Combined Performance and Energy OptimizationAdel Dabah, Gregor Häfner, Sonja Happ, Simon Pickartz, Marcus Müller, Andreas Herten2026-06-08下载Polymer simulations are among the most computationally demanding workloads in soft-matter research, often requiring days of execution and high energy consumption to achieve physically meaningful resul...
Engineering Scalable Distributed List RankingPeter Sanders, Matthias Schimek, Tim Niklas Uhl, Thomas Weidmann2026-06-08下载The list ranking problem is one of the classical problems of parallel computing, with nontrivial algorithms and many applications as a subroutine for solving other problems.
Resource-aware Computation-Communication Overlap for multi-GPU ML WorkloadsMinyu Cui, Miquel Pericas2026-06-08下载The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern ML systems.
CANS: Accelerating Multiuser Collaborative Edge Inference via Cooperative Autodidactic NeuroSurgeonZheshun Wu, Ziyang Zhang, Changyao Lin, Zenglin Xu, Jie Liu2026-06-08下载Recently, mobile edge computing (MEC)-enabled collaborative deep neural network (DNN) inference has emerged as a promising approach for delivering intelligent services to resource-constrained mobile d...
AutoPilot: Learning to Steer High Speed Robust BFTLiangrong Chen, Yue Zhang, Eric Zhou, Mohammad Javad Amiri, Ryan Marcus, Chenyuan Wu2026-06-08下载Recent Byzantine Fault Tolerant (BFT) protocols achieve strong performance by combining the low-latency advantages of leader-based BFT protocols with the high-throughput benefits of DAG-based data dis...
Concepts in Practice: C++ MPI Bindings for the HPC Ecosystem. From a Standardizable Core to a Composable InterfaceTim Niklas Uhl, Matthias Schimek, Daniel Brommer2026-06-08下载The official C++ MPI bindings were removed from the standard in 2008, leaving a gap that numerous third-party libraries have attempted to fill.
Chimera: Protocol-Aware Recovery for Confidential BFT ConsensusTong Liu, Xiaoqing Wen, Ziwei Zhou, Si Liu, Jianyu Niu, Cong Wang, Yinqian Zhang2026-06-08下载Trusted Execution Environments (TEEs) have enabled confidential Byzantine Fault-Tolerant (BFT) consensus systems with confidentiality and improved scalability.
Fairness-Aware and Latency-Controllable Scheduling for Chunked-Prefill LLM ServingHaoxin Liu, Jiayi Wang, Yueshen Xu, Rui Li2026-06-08下载As large language models (LLMs) are increasingly deployed with highly heterogeneous workloads, chunked-prefill execution has emerged as a mainstream serving architecture.
When More Cores Hurts: The Vector Database Scaling Paradox in HPCSeth Ockerman, Song Young Oh, Amal Gueroudji, Rochana Chaturvedi, Philip Carns, Nicholas Chia, Matthieu Dorier, Robert Latham, Tanwi Mallick, Swan Perarnau, Robert Underwood, Kyle Chard, Ian Foster, Robert Ross, Shivaram Venkataraman2026-06-08下载Vector databases have been designed and optimized for cloud environments; however, emerging scientific AI workloads (e.g., molecular search, meteorological trajectory detection, and literature-driven ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Secrets Best Not Shared: DNS Privacy Enhancements for the Constrained IoTMartine S. Lenders, Thomas C. Schmidt, Matthias Wählisch2026-06-08下载Attackers often identify DNS traffic to disrupt or compromise Internet services. While prior work has focused on encrypting queries using DNS over TLS, HTTPS, or QUIC to counter such attacks, we consi...
Zero Touch Predictive Orchestration: Automating Time-Series Models for the Cloud-Edge ContinuumAbd Elghani Meliani, Arora Sagar, Adlen Ksentini, Raymond Knopp2026-06-08下载The Cloud-Edge Continuum (CEC) enables latency-critical applications by distributing resources to the far edge, but its extreme volatility makes proactive Zero Touch Management via time-series forecas...
Strict-Priority Packet Delay in Switches with Transmit-Ring BufferingYash Deshpande, Quirin Vogel, Wolfgang Kellerer2026-06-08下载Strict Priority (SP) scheduling is widely used at switch egress to provide low-latency service to high-priority (HP) traffic. Existing deterministic and stochastic latency models typically account for...
STEPS: Semantic-Contract-Guided Scheduling for LLM-Assisted Natural-Language-Driven Edge AI ServicesHouyi Qi, Minghui Liwang, Xianbin Wang, Seyyedali Hosseinalipour2026-06-08下载Networked AI services are increasingly delivered through edge infrastructures to support latency-sensitive applications. Edge scheduling is critical for deciding where and how AI services are executed...
Just-in-time Restoration with Distributed Fiber Sensing in Metropolitan Optical NetworksSleman Mouammar, Italo B. Brasileiro, Andre C. Drummond2026-06-08下载Distributed Fiber Sensing (DFS) leverages optical backscattering signals to predict failure events and enable just-in-time restoration in metropolitan optical networks, i.e.
Semantic and Task-Oriented V2X Communications: Pushing the Limits of V2X Networks ScalabilityLuca Lusvarghi, Javier Gozalvez, Mohammad Irfan Khan, Seyhan Ucar, Miguel Sepulcre, Onur Altintas2026-06-08下载Scalable Vehicle-to-Everything (V2X) networks are key to support the large-scale deployment of connected and automated mobility. However, the scalability of V2X networks is currently challenged by the...
Autonomous Incident Resolution at Hyperscale: An Agentic AI Architecture for Network OperationsArun Malik2026-06-08下载Cloud network infrastructure at hyperscale presents unique operational challenges where traditional human-driven incident response cannot keep pace with the volume, velocity, and complexity of failure...
Block-A-Mole: The Sustainability Frontier of Moving-Target Censorship ResistanceAnindya Maiti2026-06-08下载Internet censorship affects over four billion people, and deployed circumvention systems share a common weakness: their endpoints are fixed and discoverable, so a patient censor can enumerate and bloc...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
FMplex: Model Virtualization for Serving Extensible Foundation ModelsHetvi Shastri, Pragya Sharma, Walid A. Hanafy, David Irwin, Mani Srivastava, Prashant Shenoy2026-06-08下载Foundation models (FMs) are increasingly used as backbones for downstream tasks across language, vision, time-series, and multimodal applications.
TinyContainer: Container Runtime Middleware Enabling Multi-tenant Microcontrollers with Built-in SecurityBastien Buil, Chrystel Gaber, Samuel Legouix, Emmanuel Baccelli, Samia Bouzefrane2026-06-08下载Software containerization technologies for resource-limited devices enable multi-tenant microcontrollers, which allow running multiple applications with different permission levels.

cs.PF - Performance ​

标题作者发布日期PDF摘要
An 84-Format Numeric Catalog with Bit-Exact Conformance Vectors: A Vendor-Neutral Reference for FP8, BF16, MXFP4, and Microscaling FormatsDmitrii Vasilev2026-06-08下载Numeric format proliferation in machine learning hardware -- FP8 (E4M3 and E5M2), BF16, MXFP4, microscaling block formats, and dozens of research variants -- has outpaced the availability of vendor-ne...
AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel SynthesisJaber Jaber, Osama Jaber2026-06-08下载AutoMegaKernel (AMK) compiles a HuggingFace Llama-family model into a single persistent cooperative CUDA kernel that runs the whole forward pass in one launch, with no per-model hand-written CUDA.
Correlation Is Not Enough: Embedding Human Metadata for Individual Causal DiscoverySuraj Biswas, Saurabh Gupta, Pritam Mukherjee2026-06-08下载Ask a pretrained biomedical language model whether "cortisol 28 ug/dL" and "stock-market volatility" are related, and it returns a cosine similarity of 0.83 on a scale where 1.0 means identical.
Fairness-Aware and Latency-Controllable Scheduling for Chunked-Prefill LLM ServingHaoxin Liu, Jiayi Wang, Yueshen Xu, Rui Li2026-06-08下载As large language models (LLMs) are increasingly deployed with highly heterogeneous workloads, chunked-prefill execution has emerged as a mainstream serving architecture.

基于 VitePress 构建 · 使用本地搜索查找论文