2026-06-24
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Nanoelectromechanical Systems (NEMS) for Hardware Security in Advanced Packaging | Himanandhan Reddy Kottur, Pavanbabu Arjunamahanthi, M. Shafkat M. Khan, Liton Kumar Biswas, Nitin Varshney, Navid Asadizanjani | 2026-06-24 | 下载 | As hardware security threats escalate across semiconductor manufacturing and advanced packaging, there is a growing need for novel physical mechanisms to counter sophisticated attacks such as tamperin... |
| Query Cost Model Calibration in Confidential Virtual Machines | Qihan Zhang, Mengyuan Li, Ibrahim Sabek | 2026-06-24 | 下载 | With the growing adoption of Confidential Computing, running databases in confidential virtual machines (CVMs) such as AMD SEV-SNP has become an attractive way to protect sensitive cloud data with min... |
| SOLAR: AI-Powered Speed-of-Light Performance Analysis | Qijing Huang, Sana Damani, Zhifan Ye, Athinagoras Skiadopoulos, Siva Kumar Sastry Hari, Jason Clemons, Sahil Modi, Jingquan Wang, Aditya Kane, Edward C Lin, Humphrey Shi, Christos Kozyrakis | 2026-06-24 | 下载 | How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are central to software, hardware, and algorithm optimizations. |
| CVA6-RT: an Open-Source Time-Predictable RV64 Processor for Mixed-Criticality Systems | Enrico Zelioli, Christopher Reinwardt, Nils Wistoff, Robert Balas, Alessandro Ottaviano, Luca Benini, Angelo Garofalo | 2026-06-24 | 下载 | This work presents CVA6-RT, a real-time micro-architectural extension of the CVA6 core to bound worst-case latency and reduce task's timing execution variability. |
| Toward Mitigating Process-Induced Performance Degradation in 3.5D Heterogeneous Packages via Pre-Silicon Firmware Co-Optimization | Chi Fei Chung, Nikolai Nedovodin | 2026-06-24 | 下载 | This paper presents a pre-silicon analysis of XRM-SSD V24/V7.0, a physics-aware predictive firmware scheduling layer for Intel's 3.5D heterogeneous integrated packages (Foveros Direct 3D + PowerVia + ... |
| Croc: Training the Next Generation Chip Designers on Domain-Specific End-to-End Open Source Silicon | Enrico Zelioli, Philippe Sauter, Thomas Benz, Hannah Pochert, Luisa Wüthrich, Beat Muheim, Frank K. Gürkaynak, Luca Benini | 2026-06-24 | 下载 | The demand for domain-specific systems-on-chip (SoCs) in artificial intelligence, robotics, and automotive systems is increasing the need for engineers with hands-on expertise on very-large-scale inte... |
| Energy-Efficient CNN Acceleration with MSDF Digit-Serial Arithmetic on FPGA | Muhammad Usman, Yousef Sadegheih, Dorit Merhof | 2026-06-24 | 下载 | This paper presents an energy-efficient hardware acceleration of the convolutional layers in the U-Net architecture for image segmentation, implemented on FPGA. |
| Agentic evolution of physically constrained foundation models | Jiangwei Zhang, Wen Sun, Chong Wang, Shiyao Li, Cheng Che, Chunjing Han, Dan Meng, Jian Yang, Yu Wang, Rui Hou | 2026-06-24 | 下载 | Artificial intelligence increasingly drives automated scientific discovery, yet contemporary generalist agents lack physical grounding, frequently hallucinating hardware-incompatible designs. |
| Cache-Resident LLM Inference in GB-Scale Last-Level Caches | Wanning Zhang, Tongzhou Gu, Marco Canini, Ceyu Xu, Jian Weng | 2026-06-24 | 下载 | Large language model (LLM) inference is increasingly dominated by data movement across the memory hierarchy. Recent 3D-stacked cache technologies have enabled GB-scale last-level caches in modern serv... |
| Programmable Probabilistic Computer with 1,000,000 p-bits | Navid Anjum Aadit, Xiuqi Zhang, Shuvro Chowdhury, Kevin Callahan-Coray, Kyle Lee, Saleh Bunaiyan, Sanjay Seshan, Clayton Thomas, Jason Twigg, Andrew Seawright, Forrest Brewer, Tathagata Srimani, Kerem Y. Camsari | 2026-06-24 | 下载 | Probabilistic computers built from p-bits have been proposed as hardware accelerators for sampling and optimizing Ising models, but existing systems have been confined to a single chip, capped by its ... |
| SafeGen: LLM-Driven Assertion Generation and Fault Criticality Evaluation for Functional Safety | Xuanyi Tan, Arjun Chaudhuri, Rubin Parekhji, Krishnendu Chakrabarty | 2026-06-24 | 下载 | With advances in autonomous driving and electric vehicle technologies, functional safety has become a critical requirement in automotive chip design. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| GPUSparse: GPU-Accelerated Learned Sparse Retrieval with Parallel Inverted Indices | Ashutosh Sharma | 2026-06-24 | 下载 | Learned sparse retrieval models such as SPLADE achieve retrieval quality competitive with dense models while preserving the interpretability and exact-match advantages of sparse representations. |
| TileMaxSim: IO-Aware GPU MaxSim Scoring with Dimension Tiling and Fused Product Quantization | Ashutosh Sharma | 2026-06-24 | 下载 | Multi-vector retrieval models such as ColBERT achieve state-of-the-art accuracy through fine-grained token-level MaxSim scoring, yet existing GPU implementations leave most hardware performance unused... |
| Priceless: An examination of Serverless Functions-as-a-Service (FaaS) pricing models | Nnamdi Ekwe-Ekwe | 2026-06-24 | 下载 | Serverless Functions-as-a-Service providers have grown in their offering since inception a decade ago, with a myriad of new functionalities offered to end-users. |
| A Distributed Quantum Approximate Optimization Algorithm Simulator for Engineering Design Optimization | Ali Rajabi, Milad Hasanzadeh, Amin Kargarian | 2026-06-24 | 下载 | This paper presents a Qiskit-compatible distributed quantum approximate optimization algorithm (DQAOA) simulator for quadratic unconstrained binary optimization (QUBO) problems arising in engineering ... |
| FinWhale: An Optimally Resilient Two-Round Terminating DAG Protocol | Razya Ladelsky, Roy Friedman | 2026-06-24 | 下载 | DAG based Byzantine Fault Tolerant protocols provide high throughput consensus under partial synchrony but existing DAG protocols still require at least three message delays to commit decisions. |
| Interference-Aware Cross-Application Placement: A Multi-Objective Optimization Approach for Microservice Clusters | Iqra Zafar, Christian Medeiros Adriano, Holger Giese | 2026-06-24 | 下载 | In modern cloud architectures, multiple applications often run within the same clustered environment, sharing underlying resources. This resource sharing can cause interference among applications, lea... |
| AI-Assisted Computational Reproducibility on the FABRIC Testbed | Komal Thareja, Paul Ruth, Berent Aldikacti, Michael Zink | 2026-06-24 | 下载 | Computational reproducibility remains difficult despite being central to scientific research. In this paper, we show how the international FABRIC testbed, combined with large language model (LLM) codi... |
| NEURON-Fabric: Architecture-Runtime Co-Design for Controlled Low-Bit Gradient Communication | Ziqiang Wang, Changcheng Huang, Chung-Horng Lung | 2026-06-24 | 下载 | Large-scale neural-network training repeatedly aggregates gradients across devices, making communication a central cost in distributed learning. |
| Endeavor: Efficient PairHMM for Detection of DNA Variants in Genome-Scale Datasets | Miguel Graça, Aleksandar Ilic | 2026-06-24 | 下载 | DNA variant calling represents a key operation in bioinformatics pipelines that aims at identifying genetic variants. Given an evidenced explosion in genomic data availability, there is an urgent need... |
| Dynamic Load Balancing for Uncertainty Quantification with Applications in Bayesian Inversion | C. M. Loi, M. Wille, A. Reinarz | 2026-06-24 | 下载 | Uncertainty Quantification (UQ) workflows present a particular scheduling challenge in high performance computing environments, as they typically generate large numbers of heterogeneous model evaluati... |
| TL++: Accuracy and Privacy Preserving Traversal Learning for Distributed Intelligent Systems | Erdenebileg Batbaatar, Young Yoon | 2026-06-24 | 下载 | Distributed intelligent systems increasingly need to train across data silos without centralizing raw data. Federated learning keeps data local but can suffer under heterogeneous partitions and requir... |
| Optimizing Semiconductor Device Simulations through Low-Precision Arithmetic | Alexander Maeder, Denghui Lu, Nicolas Vetsch, Vincent Maillou, Anders Winka, Jiang Cao, Mauro Dossena, Alexandros Nikolaos Ziogas, Mathieu Luisier | 2026-06-24 | 下载 | Architectural changes in GPUs, especially the promotion of low-precision computational units, pose significant challenges to traditional, FP64-based high-performance computing (HPC) applications, whil... |
| TwoStepDemocracy: Prototyping of self-evolving, democratic, and decentralized systems | Stan Verlaan, Johan Pouwelse | 2026-06-24 | 下载 | Decentralised systems are often built to avoid central control, but their evolution almost always depends on centralised platforms, informal maintainer authority, and a surprising amount of unpaid goo... |
| Latency-Aware Service Placement using Neural Combinatorial Optimisers for Edge--Cloud Systems | Kimia Abedpour, Mohammadsadeq Garshasbi Herabad, Zheng Li, Javid Taheri | 2026-06-24 | 下载 | The growth of Internet of Things (IoT) applications and latency-sensitive services has increased the demand for efficient service placement across compute continuum platforms, such as edge--cloud syst... |
| EmuGEMM: Fused Tensor Core Kernels for Precision Emulation in Matrix Multiplication | Denghui Lu, Alexander Maeder, Mathieu Luisier, Alexandros Nikolaos Ziogas | 2026-06-24 | 下载 | Modern GPUs devote an increasing silicon budget to low-precision matrix-multiplication units, widening the precision-throughput gap for scientific computing workloads. |
| CV-Rules: Serializability Verification of Concurrency Control Protocols via Explicit Transaction Ordering | Takashi Hoshino, Shigeo Mitsunari, Takashi Kambayashi, Ryoji Kurosawa, Sho Nakazono | 2026-06-24 | 下载 | We present CV-rules, an alternative characterization of serializability in which a transaction order constructed by a protocol satisfies two per-read conditions, C-rule (Causality) and V-rule (View Co... |
| Programmable Probabilistic Computer with 1,000,000 p-bits | Navid Anjum Aadit, Xiuqi Zhang, Shuvro Chowdhury, Kevin Callahan-Coray, Kyle Lee, Saleh Bunaiyan, Sanjay Seshan, Clayton Thomas, Jason Twigg, Andrew Seawright, Forrest Brewer, Tathagata Srimani, Kerem Y. Camsari | 2026-06-24 | 下载 | Probabilistic computers built from p-bits have been proposed as hardware accelerators for sampling and optimizing Ising models, but existing systems have been confined to a single chip, capped by its ... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| An Evaluation of ABR Switching for Time-Shifted Clients in MoQ | Abanisenioluwa Orojo, Tanvir Redoy, Samira Afzal, Andrew C. Freeman | 2026-06-24 | 下载 | Media over QUIC enables ultra low latency video streaming over QUIC, but its default quality-switching semantics risk introducing playback gaps during periods of network congestion. |
| HALO: Hierarchical Auction-assisted Learning for Offloading in SAGIN | Xuli Cai, Poonam Lohan, Sachin Ravikant Trankatwar, Burak Kantarci | 2026-06-24 | 下载 | In this paper, we investigate delay-aware task offloading and resource scheduling in a three-tier space-air-ground integrated network (SAGIN) consisting of IoT devices, UAV edge nodes, and a high-alti... |
| Lyapunov Optimization based Queue-aware Traffic Shaping for 5G-TSN in Industrial Environments | Kouros Zanbouri, Md Noor-A-Rahim, Cormac J Sreenan, Dirk Pesch | 2026-06-24 | 下载 | Manufacturing companies look increasingly at Private 5G networks to manage Automated Guided Vehicles (AGVs). While 5G promises Ultra-Reliable Low Latency Communication (URLLC), its service quality is ... |
| Can Machine Learning Break Wi-Fi Privacy? A Study on MAC Address Randomization | Marta Puig, Costas Michaelides, Lucia Pintor, Boris Bellalta, Francesc Wilhelmi | 2026-06-24 | 下载 | Medium Access Control (MAC) address randomization has been widely adopted during the IEEE 802.11 network discovery phase as a countermeasure against passive tracking. |
| A Backward-Compatible Protocol Upgrade for HotNets | Paolo Costa, Michael Schapira | 2026-06-24 | 下载 | This document outlines the changes adopted for ACM HotNets 2026, spanning its scope, review process, and program structure. Rather than isolated adjustments, these changes form a coherent effort to cl... |
| Cellular Predictions on the Move: What about Data? | Natalia Vesselinova, Pauliina Ilmonen | 2026-06-24 | 下载 | Mobile cellular load forecasting is native to network resource optimization and delivery of services with reliability, latency and quality guarantees. |
| Dependency-Aware Dominant Resource Fairness for Multi-Tenant Multi-Resource Systems | Braik Zeidan, Francesca Fossati, Sahar Hoteit, Stefano Secci | 2026-06-24 | 下载 | Multi-resource allocation in network-congested, multi-tenant systems in which demand exceeds available capacity is challenging, as there is no straightforward way to determine how much of each resourc... |
| RQ-SAFE: Coupled Request-Resource Scheduling for Online Edge SFC-DAGs | Shengdong Gu, Hongyuan Wan, Taixin Li | 2026-06-24 | 下载 | Intent-driven edge services allow multiple virtual network function (VNF) segments in a service function chain directed acyclic graph (SFC-DAG) to be locally reordered without changing service semanti... |
| Kom8ndor: An IEEE 802.11bn-Oriented Simulator for Wi-Fi 8 and Beyond | Francesc Wilhelmi, Sergio Barrachina-Muñoz, Boris Bellalta | 2026-06-24 | 下载 | The upcoming IEEE 802.11bn amendment marks a paradigm shift in Wi-Fi, which will pose ambitious performance targets under the paradigm of Ultra-High Reliability (UHR). |
| Lightweight PCGAE-Net: Parallel CrossGate Attention and Bottleneck AutoEncoder for Efficient 5G Channel Prediction | Uma Kishore Godavarti, K. Giridhar, Vanani Prince Dharmendrabhai, Anchit Panday, Madhan Raj Kanagarathinam | 2026-06-24 | 下载 | Accurate channel state information (CSI) prediction is essential for proactive beamforming and resource management in 5G massive MIMO systems, yet the deployment of high-accuracy transformer-based pre... |
| Sponsored Group Signature and its Application to Privacy-preserving Guest Access in Smart Environments | Sepideh Avizheh, Reihaneh Safavi-Naini, Shiwei Sun | 2026-06-24 | 下载 | Group signatures are privacy preserving signature schemes in which a group member can anonymously sign messages on behalf of the group, while providing accountability, by allowing the signature of a m... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| TileMaxSim: IO-Aware GPU MaxSim Scoring with Dimension Tiling and Fused Product Quantization | Ashutosh Sharma | 2026-06-24 | 下载 | Multi-vector retrieval models such as ColBERT achieve state-of-the-art accuracy through fine-grained token-level MaxSim scoring, yet existing GPU implementations leave most hardware performance unused... |
| SOLAR: AI-Powered Speed-of-Light Performance Analysis | Qijing Huang, Sana Damani, Zhifan Ye, Athinagoras Skiadopoulos, Siva Kumar Sastry Hari, Jason Clemons, Sahil Modi, Jingquan Wang, Aditya Kane, Edward C Lin, Humphrey Shi, Christos Kozyrakis | 2026-06-24 | 下载 | How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are central to software, hardware, and algorithm optimizations. |
| Axon: A Synthesizing Superoptimizer for Tensor Programs | Akash Kothari, Shaowei Zhu, Daniel Kroening, Chungha Sung | 2026-06-24 | 下载 | Writing high performance kernels for AI accelerators requires deep expertise in tiling, instruction selection, data layout, and operator fusion placing a significant burden on programmers. |
| EmuGEMM: Fused Tensor Core Kernels for Precision Emulation in Matrix Multiplication | Denghui Lu, Alexander Maeder, Mathieu Luisier, Alexandros Nikolaos Ziogas | 2026-06-24 | 下载 | Modern GPUs devote an increasing silicon budget to low-precision matrix-multiplication units, widening the precision-throughput gap for scientific computing workloads. |
| Above the Inner Loop: Exceeding Accelerate at LLM Prefill GEMM on the M1 AMX | Deyvik Bhan | 2026-06-24 | 下载 | On Apple Silicon the fp32 GEMMs dominating LLM prefill are dispatched by Accelerate to a matrix coprocessor (AMX) on the M1-M3. We ask where a hand-written kernel's throughput over Accelerate comes fr... |