Skip to content

2026-04-17 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Co-Design of CNN Accelerators for TinyML using Approximate Matrix DecompositionJosé Juan Hernández Morales, Georgios Mentzos, Frank Hannig, Konstantinos Balaskas, Georgios Zervakis, Jörg Henkel, Jürgen Teich2026-04-17下载The paradigm shift towards local and on-device inference under stringent resource constraints is represented by the tiny machine learning (TinyML) domain.
Characterization of Real Communication Patterns and Congestion Dynamics in HPC Interconnection NetworksMiguel Sánchez de La Rosa, Gabriel Gomez-Lopez, Alejandro Baviera, Jose Duro, Francisco J. andújar, Jesus Escudero-Sahuquillo, Pedro J. Garcia, Francisco J. Alfaro, Maria E. Gomez, Julio Sahuquillo, José L. Sánchez, Francisco J. Quiles2026-04-17下载The interconnection network is a key component of Supercomputers and Data centers, and its design must cope with the increasing communication demands of current applications and services; otherwise, i...
MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUsHaoran Wu, Zeyu Cao, Yao Lai, Binglei Lou, Jiayi Nie, Can Xiao, Timi Adeniran, Przemyslaw Forys, Kauser Johar, Catriona Wright, Junyi Liu, Kai Shi, Nicholas D. Lane, Rika Antonova, Jianyi Cheng, Timothy Jones, Aaron Zhao, Robert Mullins2026-04-17下载Emerging agentic LLM workloads are driving rapidly growing demand on both memory capacity and bandwidth, with different phases of inference (e.g., prefill and decode) imposing distinct requirements.
EquivFusion: Unifying Hardware Equivalence Checking from Algorithms to Netlists via MLIRJiaying Zhu, Baoqi Zhang, Mengxia Tao, Kezhi Li, Hao Yan, Qiang Xu, Min Li2026-04-17下载Ensuring functional consistency between high-level algorithmic models and low-level hardware implementations is a critical challenge, particularly as modern design flows increasingly span heterogeneou...
CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention accelerationBas Ahn, Xingjian Tao, Manil Dev Gomony, Marc Geilen, Henk Corporaal2026-04-17下载Large Language Models (LLMs) such as LLaMA and DeepSeek, are built on transformer architectures, which have become a standard model for achieving state-of-the-art performance in natural language proce...
Secure Authentication in Wireless IoT: Hamming Code Assisted SRAM PUF as Device FingerprintFlorian Lehn, Pascal Ahr, Hans D. Schotten2026-04-17下载Static Random Access Memory (SRAM) Physically Unclonable Functions (PUFs) make use of intrinsic manufacturing variations in memory cells to derive device-unique responses.
Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware VerificationVihaan Patel, Vidya Chhabria, Aman Arora2026-04-17下载Coverage closure is the most time-consuming phase of hardware verification, and recent large language model (LLM)-based coding agents offer a promising approach to automated stimulus generation.
HYPERHEURIST: A Simulated Annealing-Based Control Framework for LLM-Driven Code Generation in Optimized Hardware DesignShiva Ahir, Prajna Bhat, Alex Doboli2026-04-17下载Large Language Models (LLMs) have shown promising progress for generating Register Transfer Level (RTL) hardware designs, largely because they can rapidly propose alternative architectural realization...
Overmind NSA: A Unified Neuro-Symbolic Computing Architecture with Approximate Nonlinear Activations and Preemptive Memory BypassWeilun Wang, Zirui Wang, Wantong Li2026-04-17下载Neuro-symbolic AI is gaining traction in domains such as large language models, scientific discovery, and autonomous systems due to its ability to combine perception with structured reasoning.
Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware DesignsSean Lowe, Elias Hilaneh, Alma Babbit, Nakul Gopalan, Vidya Chhabria, Aman Arora2026-04-17下载Hardware verification is one of the most challenging stages of the hardware design process, requiring significant time and resources to ensure a design is fully validated and production-ready.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
FliX: Flipped-Indexing for Scalable GPU Queries and UpdatesRosina Kharal, Trevor Brown, Justus Henneberg, Felix Schuhknecht2026-04-17下载GPU-based concurrent data structures (CDSs) achieve high throughput for read-only queries, but efficient support for dynamic updates on fully GPU-resident data remains challenging. Ordered CDSs (e.g.
Scalable and Adaptive Parallel Training of Graph Transformer on Large GraphsJun-Liang Lin, Kamesh Madduri, Mahmut Taylan Kandemir2026-04-17下载Graph foundation models have demonstrated remarkable adaptability across diverse downstream tasks through large-scale pretraining on graphs. However, existing implementations of the backbone model, gr...
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference ServingYichao Yuan, Mosharaf Chowdhury, Nishil Talati2026-04-17下载Power has become a central bottleneck for AI inference. This problem is becoming more urgent as agentic AI emerges as a major workload class, yet prior power-management techniques focus almost entirel...
GreenPeas: Unlocking Adaptive Quantum Error Correction with Just-in-Time Decoding HypergraphsAbbas B. Ziad, Jubo Xu, Hongxiang Fan2026-04-17下载Circuit-level decoders are essential for the realisation of low-overhead fault-tolerant quantum computing. However, they rely on complex hypergraphs that are traditionally compiled ahead-of-time.
Training Time Prediction for Mixed Precision-based Distributed TrainingMinchul Kang, Changyong Shin, Jinwoo Jeong, Hyunho Lee, Younghun Go, Gyeongmin Kim, Gyeongsik Yang, Chuck Yoo2026-04-17下载Accurate prediction of training time in distributed deep learning is crucial for resource allocation, cost estimation, and job scheduling. We observe that the floating-point precision setting is a key...
Logarithmic-Time Geodesically Convex Decomposition in Programmable MatterHenning Hillebrandt, Andreas Padalkin, Christian Scheideler, Daniel Warner, Julian Werthmann2026-04-17下载The decomposition of complex structures into simpler substructures is a powerful technique with a wide range of applications. We study the computation of decompositions in the context of programmable ...
Compositional Design, Implementation, and Verification of Swarms (Technical Report)Florian Furbach, Lucas Clorius, Roland Kuhn, Hernán Melgratti, Alceste Scalas, Emilio Tuosto2026-04-17下载Swarm protocols are a recently introduced formalism for specifying, implementing, and verifying peer-to-peer systems called swarms. A swarm consists of distributed agents called machines that communic...
Robust Synchronisation for Federated Learning in The Face of Correlated Device FailureStefan Behfar, Richard Mortier2026-04-17下载Probabilistic Synchronous Parallel (PSP) is a technique in distributed learning systems to reduce synchronization bottlenecks by sampling a subset of participating nodes per round.
T-RBFT: A Scalable and Efficient Byzantine Consensus Based on Trusted Execution Environment for Consortium BlockchainWen Gao, Xinhong Hei, Yichuan Wang2026-04-17下载With the continuous expansion of blockchain application scenarios, consortium chains have raised higher performance and security requirements for consensus mechanisms.
Evaluating SYCL as a Unified Programming Model for Heterogeneous SystemsAmi Marowka2026-04-17下载High-performance computing (HPC) applications are increasingly executed in heterogeneous environments, introducing new challenges for programming and software portability.
Continuous benchmarking: Keeping pace with an evolving ecosystem of models and technologiesJan Vogelsang, Melissa Lober, Catherine Mia Schöfmann, José Villamar, Dennis Terhorst, Johanna Senk, Hans Ekkehard Plesser, Markus Diesmann, Susanne Kunkel, Anno C. Kurth2026-04-17下载Drawing on ideas from continuous integration, we present concepts of an automated benchmarking pipeline for high performance applications. Customization and collaboration have been key design goals ow...
New Kids: An Architecture and Performance Investigation of Second-Generation Serverless PlatformsTrever Schirmer, Aris Wiegand, Lucca di Benedetto, Linus Gustafsson, Natalie Carl, Tobias Pfandzelter, David Bermbach2026-04-17下载With the ever-increasing usage of serverless computing in both industry and academia, it is essential to understand the mechanisms that power the underlying platforms.
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic PotentialsYuanchang Zhou, Hongyu Wang, Yiming Du, Yan Wang, Mingzhen Li, Siyu Hu, Xiangyu Zhang, Weijian Liu, Chen Wang, Zhuoqiang Guo, Long Wang, Jingde Bu, Yutong Lu, Guangming Tan, Weile Jia2026-04-17下载Universal Machine Learning Interatomic Potentials (uMLIPs), pre-trained on massively diverse datasets encompassing inorganic materials and organic molecules across the entire periodic table, serve as ...
CroSatFL: Energy-Efficient Federated Learning with Cross-Aggregation for Satellite Edge ComputingNan Yang, Bahman Javadi, Rodrigo Neves Calheiros, David Boland, Philip Leong2026-04-17下载Low Earth Orbit (LEO) mega-constellations extend the cloud-to-edge continuum into space, enabling satellite edge computing. However, Federated Learning (FL) in this environment is fundamentally energy...
cuNNQS-SCI: A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum StatesDaran Sun, Bowen Kan, Haoquan Long, Hairui Zhao, Haoxu Li, Yicheng Liu, Pengyu Zhou, Ankang Feng, Wenjing Huang, Yida Gu, Zhenyu Li, Honghui Shang, Yunquan Zhang, Dingwen Tao, Ninghui Sun, Guangming Tan2026-04-17下载AI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schrödinger equation for complex many-body systems.
PoSME: Proof of Sequential Memory Execution via Latency-Bound Pointer Chasing with Causal Hash BindingDavid L. Condrey2026-04-17下载We introduce PoSME (Proof of Sequential Memory Execution), a cryptographic primitive that enforces sustained sequential computation via latency-bound pointer chasing over a mutable arena.
Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM ServingTakeshi Yoshimura, Valentijn Dymphnus van de Beek, Tatsuhiro Chiba2026-04-17下载Distributed LLM serving systems optimize per-request latency and throughput. However, under long-context workloads, inference accuracy becomes more variable.
BlockRaFT: A Distributed Framework for Fault-Tolerant and Scalable Blockchain NodesManaswini Piduguralla, Souvik Sarkar, Arunmoezhi Ramachandran, Sathya Peri2026-04-17下载Blockchain technology enhances transparency by maintaining a distributed ledger among mutually untrusting parties. Despite its advantages, scalability and availability remain critical bottlenecks that...
DataCenterGym: A Physics-Grounded Simulator for Multi-Objective Data Center SchedulingNilavra Pathak, Samadrita Biswas, Nirmalya Roy2026-04-17下载Modern datacenters schedule heterogeneous workloads across geo-distributed sites with diverse compute capacities, electricity prices, and thermal conditions.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
End-to-End Performance of Video Streaming With MPEG-DASH Over Satellite 5G IAB NetworksMuhammad Adeel Zahid, Ekram Hossain, Peng Hu2026-04-17下载We present an end-to-end performance evaluation of MPEG-DASH video streaming over a Low-Earth Orbit (LEO) satellite-based 5G Integrated Access and Backhaul (IAB) network.
Deterministic Task Offloading and Resource Allocation in the IoT-Edge-Cloud ContinuumKeyvan Aghababaiyan, Baldomero Coll-Perales, Javier Gozalvez2026-04-17下载Future cellular networks will sustainably integrate computing, intelligence and services within a network of networks ecosystem that includes IoT devices and subnetworks for local communications and d...
Deterministic Task Scheduling in In-Vehicle Networks for Software-Defined VehiclesKeyvan Aghababaiyan, Baldomero Coll-Perales, Luca Lusvarghi, Javier Gozalvez2026-04-17下载Modern vehicles are embedding increasing levels of automation, connectivity, and intelligence, which require advanced in-vehicle networks and computational platforms to support the dependability and d...
Toward EU Sovereignty in Space: A Comparative Simulation Study of IRIS 2 and StarlinkAlexander Bonora, Marco Giordani, Michele Zorzi2026-04-17下载The evolution of 6th generation (6G) networks increasingly relies on satellite-based Non-Terrestrial Networks (NTNs) to extend broadband connectivity to remote and unserved regions, and to support pub...
Characterization of Real Communication Patterns and Congestion Dynamics in HPC Interconnection NetworksMiguel Sánchez de La Rosa, Gabriel Gomez-Lopez, Alejandro Baviera, Jose Duro, Francisco J. andújar, Jesus Escudero-Sahuquillo, Pedro J. Garcia, Francisco J. Alfaro, Maria E. Gomez, Julio Sahuquillo, José L. Sánchez, Francisco J. Quiles2026-04-17下载The interconnection network is a key component of Supercomputers and Data centers, and its design must cope with the increasing communication demands of current applications and services; otherwise, i...
Radio Environment Map for Energy-Efficient User-Centric Cell-Free M-MIMO NetworkMarcin Hoffmann, Paweł Kryszkiewicz2026-04-17下载This paper proposes a Radio Environment Map (REM) for energy-efficient (EE) serving cluster formulation in a user-centric cell-free network. By incorporating the location of the user and the character...
Federated Parameter-Efficient Adaptation for Interference Mitigation at the Wireless EdgeEvar Jones, Daniel J. Jakubisin, Sanmay Das2026-04-17下载Dense wireless deployments face co-channel interference from heterogeneous sources that vary across base stations (gNBs in 5G). While centralized DNN-based approaches to interference mitigation have s...
Scalable Deterministic Task Offloading and Resource Allocation in the IoT-Edge-Cloud ContinuumKeyvan Aghababaiyan, Baldomero Coll-Perales, Javier Gozalvez2026-04-17下载Future 6 G networks are envisioned as a network of networks (NoN) ecosystem, integrating communication and computing resources across multiple domains.
A Protocol-Agnostic Backscatter-Based Security Layer for Ultra-Low-Power SWIPT IoT NetworksTaki Eddine Djidjekh, Alexandru Takacs, Gaël Loubet, Lamoussa Sanogo, Daniela Dragomirescu2026-04-17下载This paper presents a lightweight, protocol-agnostic security enhancement for Simultaneous Wireless Information and Power Transfer (SWIPT) in Internet of Things (IoT) applications.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Training Time Prediction for Mixed Precision-based Distributed TrainingMinchul Kang, Changyong Shin, Jinwoo Jeong, Hyunho Lee, Younghun Go, Gyeongmin Kim, Gyeongsik Yang, Chuck Yoo2026-04-17下载Accurate prediction of training time in distributed deep learning is crucial for resource allocation, cost estimation, and job scheduling. We observe that the floating-point precision setting is a key...
CPU Optimization of a Monocular 3D Biomechanics Pipeline for Low-Resource DeploymentYan Zhang, Xiong Zhao2026-04-17下载Markerless 3D movement analysis from monocular video enables accessible biomechanical assessment in clinical and sports settings. However, most research-grade pipelines rely on GPU acceleration, limit...

基于 VitePress 构建 · 使用本地搜索查找论文