Skip to content

2026-06-24 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Nanoelectromechanical Systems (NEMS) for Hardware Security in Advanced PackagingHimanandhan Reddy Kottur, Pavanbabu Arjunamahanthi, M. Shafkat M. Khan, Liton Kumar Biswas, Nitin Varshney, Navid Asadizanjani2026-06-24下载As hardware security threats escalate across semiconductor manufacturing and advanced packaging, there is a growing need for novel physical mechanisms to counter sophisticated attacks such as tamperin...
Query Cost Model Calibration in Confidential Virtual MachinesQihan Zhang, Mengyuan Li, Ibrahim Sabek2026-06-24下载With the growing adoption of Confidential Computing, running databases in confidential virtual machines (CVMs) such as AMD SEV-SNP has become an attractive way to protect sensitive cloud data with min...
SOLAR: AI-Powered Speed-of-Light Performance AnalysisQijing Huang, Sana Damani, Zhifan Ye, Athinagoras Skiadopoulos, Siva Kumar Sastry Hari, Jason Clemons, Sahil Modi, Jingquan Wang, Aditya Kane, Edward C Lin, Humphrey Shi, Christos Kozyrakis2026-06-24下载How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are central to software, hardware, and algorithm optimizations.
CVA6-RT: an Open-Source Time-Predictable RV64 Processor for Mixed-Criticality SystemsEnrico Zelioli, Christopher Reinwardt, Nils Wistoff, Robert Balas, Alessandro Ottaviano, Luca Benini, Angelo Garofalo2026-06-24下载This work presents CVA6-RT, a real-time micro-architectural extension of the CVA6 core to bound worst-case latency and reduce task's timing execution variability.
Toward Mitigating Process-Induced Performance Degradation in 3.5D Heterogeneous Packages via Pre-Silicon Firmware Co-OptimizationChi Fei Chung, Nikolai Nedovodin2026-06-24下载This paper presents a pre-silicon analysis of XRM-SSD V24/V7.0, a physics-aware predictive firmware scheduling layer for Intel's 3.5D heterogeneous integrated packages (Foveros Direct 3D + PowerVia + ...
Croc: Training the Next Generation Chip Designers on Domain-Specific End-to-End Open Source SiliconEnrico Zelioli, Philippe Sauter, Thomas Benz, Hannah Pochert, Luisa Wüthrich, Beat Muheim, Frank K. Gürkaynak, Luca Benini2026-06-24下载The demand for domain-specific systems-on-chip (SoCs) in artificial intelligence, robotics, and automotive systems is increasing the need for engineers with hands-on expertise on very-large-scale inte...
Energy-Efficient CNN Acceleration with MSDF Digit-Serial Arithmetic on FPGAMuhammad Usman, Yousef Sadegheih, Dorit Merhof2026-06-24下载This paper presents an energy-efficient hardware acceleration of the convolutional layers in the U-Net architecture for image segmentation, implemented on FPGA.
Agentic evolution of physically constrained foundation modelsJiangwei Zhang, Wen Sun, Chong Wang, Shiyao Li, Cheng Che, Chunjing Han, Dan Meng, Jian Yang, Yu Wang, Rui Hou2026-06-24下载Artificial intelligence increasingly drives automated scientific discovery, yet contemporary generalist agents lack physical grounding, frequently hallucinating hardware-incompatible designs.
Cache-Resident LLM Inference in GB-Scale Last-Level CachesWanning Zhang, Tongzhou Gu, Marco Canini, Ceyu Xu, Jian Weng2026-06-24下载Large language model (LLM) inference is increasingly dominated by data movement across the memory hierarchy. Recent 3D-stacked cache technologies have enabled GB-scale last-level caches in modern serv...
Programmable Probabilistic Computer with 1,000,000 p-bitsNavid Anjum Aadit, Xiuqi Zhang, Shuvro Chowdhury, Kevin Callahan-Coray, Kyle Lee, Saleh Bunaiyan, Sanjay Seshan, Clayton Thomas, Jason Twigg, Andrew Seawright, Forrest Brewer, Tathagata Srimani, Kerem Y. Camsari2026-06-24下载Probabilistic computers built from p-bits have been proposed as hardware accelerators for sampling and optimizing Ising models, but existing systems have been confined to a single chip, capped by its ...
SafeGen: LLM-Driven Assertion Generation and Fault Criticality Evaluation for Functional SafetyXuanyi Tan, Arjun Chaudhuri, Rubin Parekhji, Krishnendu Chakrabarty2026-06-24下载With advances in autonomous driving and electric vehicle technologies, functional safety has become a critical requirement in automotive chip design.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
GPUSparse: GPU-Accelerated Learned Sparse Retrieval with Parallel Inverted IndicesAshutosh Sharma2026-06-24下载Learned sparse retrieval models such as SPLADE achieve retrieval quality competitive with dense models while preserving the interpretability and exact-match advantages of sparse representations.
TileMaxSim: IO-Aware GPU MaxSim Scoring with Dimension Tiling and Fused Product QuantizationAshutosh Sharma2026-06-24下载Multi-vector retrieval models such as ColBERT achieve state-of-the-art accuracy through fine-grained token-level MaxSim scoring, yet existing GPU implementations leave most hardware performance unused...
Priceless: An examination of Serverless Functions-as-a-Service (FaaS) pricing modelsNnamdi Ekwe-Ekwe2026-06-24下载Serverless Functions-as-a-Service providers have grown in their offering since inception a decade ago, with a myriad of new functionalities offered to end-users.
A Distributed Quantum Approximate Optimization Algorithm Simulator for Engineering Design OptimizationAli Rajabi, Milad Hasanzadeh, Amin Kargarian2026-06-24下载This paper presents a Qiskit-compatible distributed quantum approximate optimization algorithm (DQAOA) simulator for quadratic unconstrained binary optimization (QUBO) problems arising in engineering ...
FinWhale: An Optimally Resilient Two-Round Terminating DAG ProtocolRazya Ladelsky, Roy Friedman2026-06-24下载DAG based Byzantine Fault Tolerant protocols provide high throughput consensus under partial synchrony but existing DAG protocols still require at least three message delays to commit decisions.
Interference-Aware Cross-Application Placement: A Multi-Objective Optimization Approach for Microservice ClustersIqra Zafar, Christian Medeiros Adriano, Holger Giese2026-06-24下载In modern cloud architectures, multiple applications often run within the same clustered environment, sharing underlying resources. This resource sharing can cause interference among applications, lea...
AI-Assisted Computational Reproducibility on the FABRIC TestbedKomal Thareja, Paul Ruth, Berent Aldikacti, Michael Zink2026-06-24下载Computational reproducibility remains difficult despite being central to scientific research. In this paper, we show how the international FABRIC testbed, combined with large language model (LLM) codi...
NEURON-Fabric: Architecture-Runtime Co-Design for Controlled Low-Bit Gradient CommunicationZiqiang Wang, Changcheng Huang, Chung-Horng Lung2026-06-24下载Large-scale neural-network training repeatedly aggregates gradients across devices, making communication a central cost in distributed learning.
Endeavor: Efficient PairHMM for Detection of DNA Variants in Genome-Scale DatasetsMiguel Graça, Aleksandar Ilic2026-06-24下载DNA variant calling represents a key operation in bioinformatics pipelines that aims at identifying genetic variants. Given an evidenced explosion in genomic data availability, there is an urgent need...
Dynamic Load Balancing for Uncertainty Quantification with Applications in Bayesian InversionC. M. Loi, M. Wille, A. Reinarz2026-06-24下载Uncertainty Quantification (UQ) workflows present a particular scheduling challenge in high performance computing environments, as they typically generate large numbers of heterogeneous model evaluati...
TL++: Accuracy and Privacy Preserving Traversal Learning for Distributed Intelligent SystemsErdenebileg Batbaatar, Young Yoon2026-06-24下载Distributed intelligent systems increasingly need to train across data silos without centralizing raw data. Federated learning keeps data local but can suffer under heterogeneous partitions and requir...
Optimizing Semiconductor Device Simulations through Low-Precision ArithmeticAlexander Maeder, Denghui Lu, Nicolas Vetsch, Vincent Maillou, Anders Winka, Jiang Cao, Mauro Dossena, Alexandros Nikolaos Ziogas, Mathieu Luisier2026-06-24下载Architectural changes in GPUs, especially the promotion of low-precision computational units, pose significant challenges to traditional, FP64-based high-performance computing (HPC) applications, whil...
TwoStepDemocracy: Prototyping of self-evolving, democratic, and decentralized systemsStan Verlaan, Johan Pouwelse2026-06-24下载Decentralised systems are often built to avoid central control, but their evolution almost always depends on centralised platforms, informal maintainer authority, and a surprising amount of unpaid goo...
Latency-Aware Service Placement using Neural Combinatorial Optimisers for Edge--Cloud SystemsKimia Abedpour, Mohammadsadeq Garshasbi Herabad, Zheng Li, Javid Taheri2026-06-24下载The growth of Internet of Things (IoT) applications and latency-sensitive services has increased the demand for efficient service placement across compute continuum platforms, such as edge--cloud syst...
EmuGEMM: Fused Tensor Core Kernels for Precision Emulation in Matrix MultiplicationDenghui Lu, Alexander Maeder, Mathieu Luisier, Alexandros Nikolaos Ziogas2026-06-24下载Modern GPUs devote an increasing silicon budget to low-precision matrix-multiplication units, widening the precision-throughput gap for scientific computing workloads.
CV-Rules: Serializability Verification of Concurrency Control Protocols via Explicit Transaction OrderingTakashi Hoshino, Shigeo Mitsunari, Takashi Kambayashi, Ryoji Kurosawa, Sho Nakazono2026-06-24下载We present CV-rules, an alternative characterization of serializability in which a transaction order constructed by a protocol satisfies two per-read conditions, C-rule (Causality) and V-rule (View Co...
Programmable Probabilistic Computer with 1,000,000 p-bitsNavid Anjum Aadit, Xiuqi Zhang, Shuvro Chowdhury, Kevin Callahan-Coray, Kyle Lee, Saleh Bunaiyan, Sanjay Seshan, Clayton Thomas, Jason Twigg, Andrew Seawright, Forrest Brewer, Tathagata Srimani, Kerem Y. Camsari2026-06-24下载Probabilistic computers built from p-bits have been proposed as hardware accelerators for sampling and optimizing Ising models, but existing systems have been confined to a single chip, capped by its ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
An Evaluation of ABR Switching for Time-Shifted Clients in MoQAbanisenioluwa Orojo, Tanvir Redoy, Samira Afzal, Andrew C. Freeman2026-06-24下载Media over QUIC enables ultra low latency video streaming over QUIC, but its default quality-switching semantics risk introducing playback gaps during periods of network congestion.
HALO: Hierarchical Auction-assisted Learning for Offloading in SAGINXuli Cai, Poonam Lohan, Sachin Ravikant Trankatwar, Burak Kantarci2026-06-24下载In this paper, we investigate delay-aware task offloading and resource scheduling in a three-tier space-air-ground integrated network (SAGIN) consisting of IoT devices, UAV edge nodes, and a high-alti...
Lyapunov Optimization based Queue-aware Traffic Shaping for 5G-TSN in Industrial EnvironmentsKouros Zanbouri, Md Noor-A-Rahim, Cormac J Sreenan, Dirk Pesch2026-06-24下载Manufacturing companies look increasingly at Private 5G networks to manage Automated Guided Vehicles (AGVs). While 5G promises Ultra-Reliable Low Latency Communication (URLLC), its service quality is ...
Can Machine Learning Break Wi-Fi Privacy? A Study on MAC Address RandomizationMarta Puig, Costas Michaelides, Lucia Pintor, Boris Bellalta, Francesc Wilhelmi2026-06-24下载Medium Access Control (MAC) address randomization has been widely adopted during the IEEE 802.11 network discovery phase as a countermeasure against passive tracking.
A Backward-Compatible Protocol Upgrade for HotNetsPaolo Costa, Michael Schapira2026-06-24下载This document outlines the changes adopted for ACM HotNets 2026, spanning its scope, review process, and program structure. Rather than isolated adjustments, these changes form a coherent effort to cl...
Cellular Predictions on the Move: What about Data?Natalia Vesselinova, Pauliina Ilmonen2026-06-24下载Mobile cellular load forecasting is native to network resource optimization and delivery of services with reliability, latency and quality guarantees.
Dependency-Aware Dominant Resource Fairness for Multi-Tenant Multi-Resource SystemsBraik Zeidan, Francesca Fossati, Sahar Hoteit, Stefano Secci2026-06-24下载Multi-resource allocation in network-congested, multi-tenant systems in which demand exceeds available capacity is challenging, as there is no straightforward way to determine how much of each resourc...
RQ-SAFE: Coupled Request-Resource Scheduling for Online Edge SFC-DAGsShengdong Gu, Hongyuan Wan, Taixin Li2026-06-24下载Intent-driven edge services allow multiple virtual network function (VNF) segments in a service function chain directed acyclic graph (SFC-DAG) to be locally reordered without changing service semanti...
Kom8ndor: An IEEE 802.11bn-Oriented Simulator for Wi-Fi 8 and BeyondFrancesc Wilhelmi, Sergio Barrachina-Muñoz, Boris Bellalta2026-06-24下载The upcoming IEEE 802.11bn amendment marks a paradigm shift in Wi-Fi, which will pose ambitious performance targets under the paradigm of Ultra-High Reliability (UHR).
Lightweight PCGAE-Net: Parallel CrossGate Attention and Bottleneck AutoEncoder for Efficient 5G Channel PredictionUma Kishore Godavarti, K. Giridhar, Vanani Prince Dharmendrabhai, Anchit Panday, Madhan Raj Kanagarathinam2026-06-24下载Accurate channel state information (CSI) prediction is essential for proactive beamforming and resource management in 5G massive MIMO systems, yet the deployment of high-accuracy transformer-based pre...
Sponsored Group Signature and its Application to Privacy-preserving Guest Access in Smart EnvironmentsSepideh Avizheh, Reihaneh Safavi-Naini, Shiwei Sun2026-06-24下载Group signatures are privacy preserving signature schemes in which a group member can anonymously sign messages on behalf of the group, while providing accountability, by allowing the signature of a m...

cs.PF - Performance ​

标题作者发布日期PDF摘要
TileMaxSim: IO-Aware GPU MaxSim Scoring with Dimension Tiling and Fused Product QuantizationAshutosh Sharma2026-06-24下载Multi-vector retrieval models such as ColBERT achieve state-of-the-art accuracy through fine-grained token-level MaxSim scoring, yet existing GPU implementations leave most hardware performance unused...
SOLAR: AI-Powered Speed-of-Light Performance AnalysisQijing Huang, Sana Damani, Zhifan Ye, Athinagoras Skiadopoulos, Siva Kumar Sastry Hari, Jason Clemons, Sahil Modi, Jingquan Wang, Aditya Kane, Edward C Lin, Humphrey Shi, Christos Kozyrakis2026-06-24下载How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are central to software, hardware, and algorithm optimizations.
Axon: A Synthesizing Superoptimizer for Tensor ProgramsAkash Kothari, Shaowei Zhu, Daniel Kroening, Chungha Sung2026-06-24下载Writing high performance kernels for AI accelerators requires deep expertise in tiling, instruction selection, data layout, and operator fusion placing a significant burden on programmers.
EmuGEMM: Fused Tensor Core Kernels for Precision Emulation in Matrix MultiplicationDenghui Lu, Alexander Maeder, Mathieu Luisier, Alexandros Nikolaos Ziogas2026-06-24下载Modern GPUs devote an increasing silicon budget to low-precision matrix-multiplication units, widening the precision-throughput gap for scientific computing workloads.
Above the Inner Loop: Exceeding Accelerate at LLM Prefill GEMM on the M1 AMXDeyvik Bhan2026-06-24下载On Apple Silicon the fp32 GEMMs dominating LLM prefill are dispatched by Accelerate to a matrix coprocessor (AMX) on the M1-M3. We ask where a hand-written kernel's throughput over Accelerate comes fr...

基于 VitePress 构建 · 使用本地搜索查找论文