Skip to content

2026-09-24 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
From Routing Delay Shifts to Silent Data Corruption: Neutron-Induced SEU Effects in AXI-Based Zynq UltraScale+ MPSoCsMostafa Darvishi2026-09-24下载SRAM-based FPGA system-on-chip devices are vulnerable to single-event upsets (SEUs) in configuration memory, which may perturb programmable routing resources and degrade communication fabrics.
GRACIDIT: Graph-Circuit Digital Twin for Configuration-Induced Routing Delay Prediction in Zynq UltraScale+ FPGAsMostafa Darvishi2026-09-24下载Configuration-induced perturbations in SRAM-based FPGAs may activate dormant programmable routing branches and increase path delay without immediately producing a functional error.
Enhancing Word-Level Property Directed Reachability with LLM-Driven Semantic GuidanceGuangyu Hu, Mingkai Miao, Zhiyuan Yan, Xiaofeng Zhou, Wei Zhang, Hongce Zhang2026-09-24下载Property Directed Reachability (PDR) is a prominent algorithm for hardware formal verification. However, bit-level PDR often struggles with datapath-heavy designs because bit-blasting obscures high-le...
Compiler and Hardware Co-Design for Accelerator ArchitecturesKarl Herman Krause, Emad Jacob Maroun, Martin Schoeberl2026-09-24下载Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads. However, full-stack integration remains difficult.
An Open-Source Standard-Cell Library for IHP 130nm Developed by StudentsOscar Castañeda, Lavinia Recchioni, Tobias Senti, Flurin Cahenzli, Marco Ferroni, Thorben Heekenjann, Robert Kenter, Stefan Odermatt, Daniel Richner, Gabriel Altendorfer, Davide Cannone, Miguel Correa, Ivan Herger, Lars Kröger, Nicolas Nanzer, Darja Nonaca, Christopher Reinwardt, Lukas Winklhofer, Enrico Zelioli, Yiheng Zhang, Yingxue Zhang, Domenic Keller, Seyed Hadi Mirfarshbafan, Zerun Jiang, Beat Muheim, Arianna Rubino, Jérémy Guichemerre, Frank K. Gürkaynak, Christoph Studer2026-09-24下载Standard-cell libraries form the critical interface between transistor-level circuit design and automated digital design flows, yet their design and characterization are rarely covered in depth in dig...
VQ-LIC: Shared Vector-Quantized Learned Image Compression on a Resource-Constrained FPGAMuhammad Fahd Ibrahim Bhatti, Abdullah Bin Faisal, Ahsan Usman, Naveed Anwar Bhatti, Muhammad Ali Siddiqi2026-09-24下载Learned image compression (LIC) is hard to deploy on severely resource-constrained FPGAs, since how fast it actually runs depends not just on arithmetic count, but also on memory traffic, imbalance be...
MagiCFirm: A Runtime for Magic-State Cultivation with Algorithm-Hardware Co-DesignJubo Xu, Abbas B. Ziad, Prakash Murali, Hongxiang Fan2026-09-24下载Magic-state cultivation offers a promising alternative for lowering the cost of non-Clifford operations in fault-tolerant quantum computing (FTQC).
HBF-Sim: An Extensible HBF Simulator for Large-scale GPU Memory SystemsYaqi Li, Jing Wang, Junfeng Wang, Long Yang, Han Yan, Xiaohu Chai, Liang Shi2026-09-24下载High-bandwidth flash (HBF) is introduced to address the memory wall, which can co-package a dense NAND stack with the GPU, targeting the performance gap between near-accelerator bandwidth and flash de...
Implementation and Evaluation of NTT Arithmetic for ML-KEM on a CGLATakuto Ando, Yasuhiko Nakashima2026-09-24下载FIPS 203 standardizes ML-KEM for post-quantum key establishment. Its polynomial multiplication relies on NTT butterflies with exact modular arithmetic over q = 3329.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
FRESHLATENT: Channel-Aware Latent Adaptation for Resource-Constrained Embodied VLM PerceptionRajat Bhattacharjya, Minwoo Kim, Arnab Sarkar, Tamoghno Das, Sing-Yao Wu, Eli Bozorgzadeh, Marco Levorato, Nikil Dutt2026-09-24下载Mission-critical UAVs increasingly rely on split vision-language model (VLM) perception under tight onboard-resource and wireless-communication constraints.
Encryptability As a Coordinate Choice: Depth-One Homomorphic Federated Learning of Quantum Neural NetworksMarcel Mordarski, Nathan Mani, Arshad Patel, William Knottenbelt, Roberto Bondesan2026-09-24下载Encrypted training relies on keeping server-side updates low-degree. This constraint traditionally excludes models whose weights inhabit a compact Lie group (notably variational quantum circuits, wher...
Managing Iterative Hybrid Quantum-Classical Optimization as a First-Class Scientific WorkflowGiuliana Siddi Moreau, Maria Laura Clemente, Lorenzo Pisani, Manuela Profir, Marco Pinna, Marco Moro, Lidia Leoni2026-09-24下载Today's Quantum Processing Units (QPUs) are too small and too noisy to solve large combinatorial optimization problems directly, so practical hybrid solvers split a problem into pieces and iterate a d...
Federated Targeted Maximum Likelihood EstimationDiyang Li, Fei Wang, Kyra Gan2026-09-24下载The evidence behind a scientific or operational decision is often held by hospitals, banks, or registries that cannot pool individual observations.
Unifying In-Memory Data Analytics through Sparse CompilationAnand Jayarajan, Gennady Pekhimenko2026-09-24下载As modern data analytics workloads become increasingly heterogeneous and hardware-intensive, achieving efficient multi-core performance across diverse applications remains an open challenge.
Communication-Aware Model Distributed Inference via Latent Representation CompressionPeyman Gholami, Theodoros-Thirimachos Davarakis, Teng Li, Miquel Sirera Perelló, Salil Reddy, Ayberk Yarkın Yıldız, Anish Arora, Atilla Eryilmaz, Stratis Ioannidis, Chengzhang Li, Hulya Seferoglu, Ness Shroff2026-09-24下载We study optimization of distributed model inference over resource-constrained edge resources. We propose a framework that optimizes the trade-off between model accuracy and communication costs by con...
Proceedings 19th Interaction and Concurrency ExperienceLuc Edixhoven, Simon Fowler, Rumyana Neykova, Violet Ka I Pun2026-09-24下载This volume contains the proceedings of ICE'26, the 19th Interaction and Concurrency Experience, which was held in Urbino, Italy, as a satellite event of DisCoTec'26.
Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAGGeorge Danezis, Zeno De Angeli, Philipp Jovanovic, Lefteris Kokoris-Kogias, Markus Legner, Alberto Sonnino2026-09-24下载Dual-mode consensus protocols are fast when the network is partially synchronous and remain live under asynchrony. We introduce Steelhead, a dual-mode mechanism that composes a partially synchronous a...
MQSS-Selector: RL-Guided Pass Selection for an MLIR Compilation PipelineAndre Youssefi, Ercüment Kaya, Minh Chung, Jorge Echavarria, Laura B. Schulz, Martin Schulz2026-09-24下载High Performance Computing (HPC) and Quantum Computing (QC) systems are increasingly converging towards unified High Performance Computing-Quantum Computing (HPCQC) infrastructures, driven by a growin...
KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel OptimizationAheli Poddar, Sanskar Prasad, Arindam Samanta, Subha Chakraborty, Vishal Goyal, Rohit Singh Rathaur2026-09-24下载Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, b...
KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular ExclusivityTianyu Feng, Haoxuan Yu, Tianyuan Wu, Lingyun Yang, Daocheng Ying, Yuxiao Wang, Ruibo Fan, Yinghao Yu, Guodong Yang, Liping Zhang, Wei Wang2026-09-24下载LLM agents automate GPU kernel optimization by repeatedly composing candidates and measuring their duration on real GPUs. Existing systems preserve measurement fidelity by reserving a GPU for an entir...
Sluice: Global Invariant, Local Enforcement for Pooled Payment-Channel LiquidityYueqi Wu, Huiping Sun, Peilu Guo, Yiming Zhu, Zhong Chen2026-09-24下载A routing node on the Lightning Network holds its liquidity in separate channels, so a payment can fail at a channel whose outbound balance is exhausted while the node's other channels still hold bala...
A Block Decomposed QUBO Workflow for Chromosome-Y Phylogeny ReconstructionGiuliana Siddi Moreau, Riccardo Berutti, Manuela Profir, Lorenzo Pisani, Maria Laura Clemente, Lidia Leoni2026-09-24下载This paper sets out a computational workflow that reconstructs the phylogeny of human Y-chromosome populations from a Variant Call Format (VCF) file of biallelic Single Nucleotide Polymorphisms (SNP).
Hard Stop: Kernel-Level Preemption and Containment for Rogue Agentic ExecutionJosé Luis Pino2026-09-24下载In July 2026, an unconstrained autonomous agent participating in a frontier AI cybersecurity evaluation harness breached its evaluation sandbox, established an external command-and-control foothold, a...
Reusing Spare Vehicle Computing Capacity: Is It Viable, Profitable and Sustainable?Rosario Patanè, Nadjib Achir, Andrea Araldo, Lila Boukhatem2026-09-24下载Vehicular Cloud Computing (VCC) exploits computing hardware already embedded in vehicles for purposes unrelated to offloading and puts its idle cycles to work executing end-users' offloaded tasks, avo...
Resource-Aware Model Selection for Scalable Indoor Localization on HPC PlatformsFukuharu Tanaka, Hamada Rizk, Moustafa Youssef, Hirozumi Yamaguchi2026-09-24下载Large-scale indoor localization is increasingly needed in campuses, smart buildings, factories, and digital-twin infrastructures, where wireless conditions, access-point deployments, and spatial layou...
Concurrent Split Learning Through Stable Client ClusteringMohammad Kohankhaki, Valentin Rentschler, Anke Schmeink2026-09-24下载Training with a fixed global batch limits how many distributed clients can provide examples in any one step. We examine a way to use additional server workers without increasing the batch processed by...
MagiCFirm: A Runtime for Magic-State Cultivation with Algorithm-Hardware Co-DesignJubo Xu, Abbas B. Ziad, Prakash Murali, Hongxiang Fan2026-09-24下载Magic-state cultivation offers a promising alternative for lowering the cost of non-Clifford operations in fault-tolerant quantum computing (FTQC).
TrafficFab: An Autonomic Edge-Cloud Testbed Fabric forAI-Driven Traffic ManagementMayank Arya, Pranjal Naman, Priyanshu Pansari, Roopkatha Banerjee, Daksh Mehta, Manjil Nepal, Akash Sharma, Yogesh Simmhan2026-09-24下载Traffic management in emerging megacities requires real-time analytics over thousands of CCTV video streams under latency, bandwidth, compute and energy constraints.
Security Limits of Mining Before Validation in Nakamoto ConsensusYifan Zhou, Jiang Xiao, Kaihua Qin2026-09-24下载Mining before validation allows miners to extend a newly received block before completing its validity checks, giving them a head start in the race for the next block reward.
Cross-Model Autoscaling for Shared LLM ServingXin Zhang, Xianyan Xie, Zhen He, Xijin Yin, Xingtong Lin, Bangbo Liang, Zequn Cheng, Peihao Huang, Guo Chen2026-09-24下载Multi-model LLM serving is moving toward shared MaaS clusters, where co-hosted models compete for a fixed GPU budget while each model experiences time-varying demand and must satisfy its own latency S...
MeshHeal: Two-Timescale Self-Healing for Gray Failures in Decentralized LLM Agent NetworksKeru Chen, Sen Lin, Yingbin Liang, Nathaniel D. Bastian, Shaofeng Zou2026-09-24下载Decentralized LLM-based multi-agent systems coordinate through local interactions, but an agent can remain responsive while its task-solving quality persistently degrades.
Dependency- and Layer-Aware Microservice Workflow Offloading and Service Image Caching for Edge EnvironmentsZhongxiao Wang, Yueshen Xu, Qingshan Li, Xinkui Zhao, Wei Shao, Shuiguang Deng, Rui Li2026-09-24下载The microservice architecture has been applied broadly in many mainstream computing environments. As one of the most prevalent environments, edge computing also widely employs microservices to handle ...
When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix ReuseYiyu Liu, Minlan Yu, Juncheng Yang2026-09-24下载Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood.
pytest-gpu-proof: Enabling Cloud-CPU Continuous Integration for GPU Code with Local GPU AttestationBrian Plancher2026-09-24下载GPU acceleration is now routine across robotics, but cloud-hosted GPU continuous integration (CI) runners are expensive, resulting in severe under-testing of GPU-accelerated code.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Joint Effects of Node Density, Propagation, Wi-Fi Generation, and Transport Protocol on WLAN Performance: An ns-3 StudyLeonel Olimpio Silima2026-09-24下载Wireless local area network (WLAN) performance is jointly shaped by medium contention, propagation, Wi-Fi configuration, and transport-layer behavior. This study evaluates IEEE 802.11g and IEEE 802.
SafeNom: Data-Aware Microservice PoliciesKaruna Grewal, P. Brighten Godfrey, Justin Hsu2026-09-24下载Many cloud-based applications are organized as loosely coupled microservices, where invoking a service's API triggers a cascade of APIs across many services and leads to inter-service exchange of API ...
Packet-Level In-Network Semantic Adaptation for Unstable Mobile Emergency NetworksZhiyuan Ren, Tao Zhang, Wenchi Cheng2026-09-24下载Mobile emergency networks can experience independently changing intermediate wireless links on timescales shorter than endpoint feedback can track.
Energy-Aware Two-Sided Learning for Dynamic Matching Games in Mobile CrowdsensingSumedh J. Dongare, Anja Klein, Andrea Ortiz2026-09-24下载Mobile crowdsensing (MCS) is a promising enabler of Sensing-as-a-Service (SaaS) for next generation networks (NGNs), where sensing, communication, and computing are jointly considered as on-demand ser...
SPADE-DFL: Communication-Efficient Decentralized Federated Learning via Derivative-Free Linearized ADMMMengli Wei, Mengkai Zhu, Jiawen Chen, Wenwu Yu, Duxin Che2026-09-24下载Reducing communication in derivative-free decentralized learning requires controlling the disagreement accumulated over multiple local updates.
NebulaSD: Many-for-Many Speculative DecodingJunhao He, Hongyang Du2026-09-24下载Speculative decoding accelerates Large Language Model (LLM) inference by using a lightweight draft model to propose candidate tokens for parallel verification by a target model.
FlowAtom: Atom-Based Evidence Aggregation for Multi-Label Website FingerprintingChongru Fan, Wentao Huang, Wei Wang, Zhenquan Ding, Jinqiao Shi, Wei Cai, Zhiyu Hao2026-09-24下载Identifying the set of monitored websites in mixed encrypted traffic is challenging because an individual flow often provides only partial evidence of website identity.
From WPT to Encrypted Telemetry: A Battery-Free Backscattering-based Polarimetric Wireless SensorTaki Eddine Djidjekh, Loïc Thomas, Quentin Bernyer, Gaël Loubet, Daniela Dragomirescu, Alexandru Takacs2026-09-24下载This work introduces an indoor Battery-Free Wireless Sensing Node powered through radiative Wireless Power Transfer (WPT). The proposed platform targets secure, energyefficient active sensing and over...
Secure Polarization-Shift Backscatter Identification Applied to Battery-Free BLE Sensors Powered by Wireless Power TransferTaki Eddine Djidjekh, Quentin Bernyer, Alexandru Takacs2026-09-24下载This paper presents a lightweight and protocolindependent security mechanism for battery-free Bluetooth Low Energy (BLE) sensor nodes operating in Simultaneous Wireless Information and Power Transfer ...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular ExclusivityTianyu Feng, Haoxuan Yu, Tianyuan Wu, Lingyun Yang, Daocheng Ying, Yuxiao Wang, Ruibo Fan, Yinghao Yu, Guodong Yang, Liping Zhang, Wei Wang2026-09-24下载LLM agents automate GPU kernel optimization by repeatedly composing candidates and measuring their duration on real GPUs. Existing systems preserve measurement fidelity by reserving a GPU for an entir...
Hard Stop: Kernel-Level Preemption and Containment for Rogue Agentic ExecutionJosé Luis Pino2026-09-24下载In July 2026, an unconstrained autonomous agent participating in a frontier AI cybersecurity evaluation harness breached its evaluation sandbox, established an external command-and-control foothold, a...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Energy-efficient operation of neural operators for virtual sensingJason Yoo, Samrendra Roy, Souvik Chakraborty, Syed Bahauddin Alam2026-09-24下载Virtual sensing repeatedly reconstructs physical fields from changing observations, often on a fixed geometry. We investigate how shared spatial computation reduces the energy of these updates while r...
SR-Gadgets: Make Scan-Resistant Caching PracticalYunjia Zheng, Juncheng Yang2026-09-24下载Block caches commonly serve scan-heavy I/O workloads, motivating extensive studies on scan-resistant eviction algorithms. Many of these algorithms adopt a multi-queue structure.
To Store or To Regenerate? A Cost Model for AI-Generated Content at ScaleYunjia Zheng, Zirui Wang, Haoran Ni, Tingfeng Lan, Zhaoyuan Su, Yue Cheng, Juncheng Yang2026-09-24下载AI-generated content is becoming a rapidly growing class of digital artifacts. Because these artifacts accumulate over time, their exponential growth creates a substantial storage, energy, and infrast...
DanLing NestedTensor: Composable Multi-Ragged Tensors for Deep LearningZhiyuan Chen2026-09-24下载Variable-size inputs are common in deep learning, but dense batching allocates a shared envelope and spends computation on padding. The cost multiplies across varying axes: an explicit pair state allo...
KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular ExclusivityTianyu Feng, Haoxuan Yu, Tianyuan Wu, Lingyun Yang, Daocheng Ying, Yuxiao Wang, Ruibo Fan, Yinghao Yu, Guodong Yang, Liping Zhang, Wei Wang2026-09-24下载LLM agents automate GPU kernel optimization by repeatedly composing candidates and measuring their duration on real GPUs. Existing systems preserve measurement fidelity by reserving a GPU for an entir...
Dense Matrices Are Alike; Sparse Matrices Are Sparse in Their Own Way: A Structure-Adaptive Tile Cholesky FactorizationEsmail Abdul Fattah, Hatem Ltaief, Håvard Rue, David E. Keyes2026-09-24下载Sparse direct Cholesky solvers fix one data structure for an entire matrix, but symmetric positive definite systems range from nearly dense to irregular, sometimes mixing both within one matrix.
Evaluating the Effect of the Order of Optimization Passes in Quantum Circuit OptimizationXiao-Ting Michelle To, Nils Quetschlich, Amr Elsharkawy, Martin Schulz, Robert Wille, Dieter Kranzlmüller2026-09-24下载Quantum circuit optimization is critical for mitigating the noise inherent in current quantum hardware. Quantum compilers typically sequentially apply multiple optimizations (also called ``optimizatio...
A Rapid Pipeline for Training and Deploying ML Models on WeBe BandEhsan Kourkchi, Asmita Asmita, Houman Homayoun, Mahdi Eslamimehr2026-09-24下载Developing optimized machine-learning algorithms for edge devices with limited computational and memory resources is challenging, time-consuming, and highly dependent on device-specific constraints.
TileBench: A Controlled Benchmark for Performance Evaluation and Bottleneck Diagnosis of Tile-Based Programming ModelsBowen Cui, Zhongchun Zhou, Hao Wu, Tejas Ramesh, Junyu Yin, Jialiang Gu, Keren Zhou2026-09-24下载Tile-based programming models, such as Triton and cuTile, aim to simplify high-performance kernel development, but their practical performance, tuning behavior, and usability remain difficult to compa...
Paging the Experts: A Reproducible Characterization of Flash-Backed MoE Inference on iPhoneMusa Shams2026-09-24下载Sparse activation reduces mixture-of-experts computation without eliminating the need to store all experts. We present Routide, a Swift/MLX runtime that executes the text path of a pinned public Qwen3...

基于 VitePress 构建 · 使用本地搜索查找论文