2026-04-17
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Co-Design of CNN Accelerators for TinyML using Approximate Matrix Decomposition | José Juan Hernández Morales, Georgios Mentzos, Frank Hannig, Konstantinos Balaskas, Georgios Zervakis, Jörg Henkel, Jürgen Teich | 2026-04-17 | 下载 | The paradigm shift towards local and on-device inference under stringent resource constraints is represented by the tiny machine learning (TinyML) domain. |
| Characterization of Real Communication Patterns and Congestion Dynamics in HPC Interconnection Networks | Miguel Sánchez de La Rosa, Gabriel Gomez-Lopez, Alejandro Baviera, Jose Duro, Francisco J. andújar, Jesus Escudero-Sahuquillo, Pedro J. Garcia, Francisco J. Alfaro, Maria E. Gomez, Julio Sahuquillo, José L. Sánchez, Francisco J. Quiles | 2026-04-17 | 下载 | The interconnection network is a key component of Supercomputers and Data centers, and its design must cope with the increasing communication demands of current applications and services; otherwise, i... |
| MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs | Haoran Wu, Zeyu Cao, Yao Lai, Binglei Lou, Jiayi Nie, Can Xiao, Timi Adeniran, Przemyslaw Forys, Kauser Johar, Catriona Wright, Junyi Liu, Kai Shi, Nicholas D. Lane, Rika Antonova, Jianyi Cheng, Timothy Jones, Aaron Zhao, Robert Mullins | 2026-04-17 | 下载 | Emerging agentic LLM workloads are driving rapidly growing demand on both memory capacity and bandwidth, with different phases of inference (e.g., prefill and decode) imposing distinct requirements. |
| EquivFusion: Unifying Hardware Equivalence Checking from Algorithms to Netlists via MLIR | Jiaying Zhu, Baoqi Zhang, Mengxia Tao, Kezhi Li, Hao Yan, Qiang Xu, Min Li | 2026-04-17 | 下载 | Ensuring functional consistency between high-level algorithmic models and low-level hardware implementations is a critical challenge, particularly as modern design flows increasingly span heterogeneou... |
| CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration | Bas Ahn, Xingjian Tao, Manil Dev Gomony, Marc Geilen, Henk Corporaal | 2026-04-17 | 下载 | Large Language Models (LLMs) such as LLaMA and DeepSeek, are built on transformer architectures, which have become a standard model for achieving state-of-the-art performance in natural language proce... |
| Secure Authentication in Wireless IoT: Hamming Code Assisted SRAM PUF as Device Fingerprint | Florian Lehn, Pascal Ahr, Hans D. Schotten | 2026-04-17 | 下载 | Static Random Access Memory (SRAM) Physically Unclonable Functions (PUFs) make use of intrinsic manufacturing variations in memory cells to derive device-unique responses. |
| Understanding Inference-Time Token Allocation and Coverage Limits in Agentic Hardware Verification | Vihaan Patel, Vidya Chhabria, Aman Arora | 2026-04-17 | 下载 | Coverage closure is the most time-consuming phase of hardware verification, and recent large language model (LLM)-based coding agents offer a promising approach to automated stimulus generation. |
| HYPERHEURIST: A Simulated Annealing-Based Control Framework for LLM-Driven Code Generation in Optimized Hardware Design | Shiva Ahir, Prajna Bhat, Alex Doboli | 2026-04-17 | 下载 | Large Language Models (LLMs) have shown promising progress for generating Register Transfer Level (RTL) hardware designs, largely because they can rapidly propose alternative architectural realization... |
| Overmind NSA: A Unified Neuro-Symbolic Computing Architecture with Approximate Nonlinear Activations and Preemptive Memory Bypass | Weilun Wang, Zirui Wang, Wantong Li | 2026-04-17 | 下载 | Neuro-symbolic AI is gaining traction in domains such as large language models, scientific discovery, and autonomous systems due to its ability to combine perception with structured reasoning. |
| Spec2Cov: An Agentic Framework for Code Coverage Closure of Digital Hardware Designs | Sean Lowe, Elias Hilaneh, Alma Babbit, Nakul Gopalan, Vidya Chhabria, Aman Arora | 2026-04-17 | 下载 | Hardware verification is one of the most challenging stages of the hardware design process, requiring significant time and resources to ensure a design is fully validated and production-ready. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| FliX: Flipped-Indexing for Scalable GPU Queries and Updates | Rosina Kharal, Trevor Brown, Justus Henneberg, Felix Schuhknecht | 2026-04-17 | 下载 | GPU-based concurrent data structures (CDSs) achieve high throughput for read-only queries, but efficient support for dynamic updates on fully GPU-resident data remains challenging. Ordered CDSs (e.g. |
| Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs | Jun-Liang Lin, Kamesh Madduri, Mahmut Taylan Kandemir | 2026-04-17 | 下载 | Graph foundation models have demonstrated remarkable adaptability across diverse downstream tasks through large-scale pretraining on graphs. However, existing implementations of the backbone model, gr... |
| KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving | Yichao Yuan, Mosharaf Chowdhury, Nishil Talati | 2026-04-17 | 下载 | Power has become a central bottleneck for AI inference. This problem is becoming more urgent as agentic AI emerges as a major workload class, yet prior power-management techniques focus almost entirel... |
| GreenPeas: Unlocking Adaptive Quantum Error Correction with Just-in-Time Decoding Hypergraphs | Abbas B. Ziad, Jubo Xu, Hongxiang Fan | 2026-04-17 | 下载 | Circuit-level decoders are essential for the realisation of low-overhead fault-tolerant quantum computing. However, they rely on complex hypergraphs that are traditionally compiled ahead-of-time. |
| Training Time Prediction for Mixed Precision-based Distributed Training | Minchul Kang, Changyong Shin, Jinwoo Jeong, Hyunho Lee, Younghun Go, Gyeongmin Kim, Gyeongsik Yang, Chuck Yoo | 2026-04-17 | 下载 | Accurate prediction of training time in distributed deep learning is crucial for resource allocation, cost estimation, and job scheduling. We observe that the floating-point precision setting is a key... |
| Logarithmic-Time Geodesically Convex Decomposition in Programmable Matter | Henning Hillebrandt, Andreas Padalkin, Christian Scheideler, Daniel Warner, Julian Werthmann | 2026-04-17 | 下载 | The decomposition of complex structures into simpler substructures is a powerful technique with a wide range of applications. We study the computation of decompositions in the context of programmable ... |
| Compositional Design, Implementation, and Verification of Swarms (Technical Report) | Florian Furbach, Lucas Clorius, Roland Kuhn, Hernán Melgratti, Alceste Scalas, Emilio Tuosto | 2026-04-17 | 下载 | Swarm protocols are a recently introduced formalism for specifying, implementing, and verifying peer-to-peer systems called swarms. A swarm consists of distributed agents called machines that communic... |
| Robust Synchronisation for Federated Learning in The Face of Correlated Device Failure | Stefan Behfar, Richard Mortier | 2026-04-17 | 下载 | Probabilistic Synchronous Parallel (PSP) is a technique in distributed learning systems to reduce synchronization bottlenecks by sampling a subset of participating nodes per round. |
| T-RBFT: A Scalable and Efficient Byzantine Consensus Based on Trusted Execution Environment for Consortium Blockchain | Wen Gao, Xinhong Hei, Yichuan Wang | 2026-04-17 | 下载 | With the continuous expansion of blockchain application scenarios, consortium chains have raised higher performance and security requirements for consensus mechanisms. |
| Evaluating SYCL as a Unified Programming Model for Heterogeneous Systems | Ami Marowka | 2026-04-17 | 下载 | High-performance computing (HPC) applications are increasingly executed in heterogeneous environments, introducing new challenges for programming and software portability. |
| Continuous benchmarking: Keeping pace with an evolving ecosystem of models and technologies | Jan Vogelsang, Melissa Lober, Catherine Mia Schöfmann, José Villamar, Dennis Terhorst, Johanna Senk, Hans Ekkehard Plesser, Markus Diesmann, Susanne Kunkel, Anno C. Kurth | 2026-04-17 | 下载 | Drawing on ideas from continuous integration, we present concepts of an automated benchmarking pipeline for high performance applications. Customization and collaboration have been key design goals ow... |
| New Kids: An Architecture and Performance Investigation of Second-Generation Serverless Platforms | Trever Schirmer, Aris Wiegand, Lucca di Benedetto, Linus Gustafsson, Natalie Carl, Tobias Pfandzelter, David Bermbach | 2026-04-17 | 下载 | With the ever-increasing usage of serverless computing in both industry and academia, it is essential to understand the mechanisms that power the underlying platforms. |
| Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials | Yuanchang Zhou, Hongyu Wang, Yiming Du, Yan Wang, Mingzhen Li, Siyu Hu, Xiangyu Zhang, Weijian Liu, Chen Wang, Zhuoqiang Guo, Long Wang, Jingde Bu, Yutong Lu, Guangming Tan, Weile Jia | 2026-04-17 | 下载 | Universal Machine Learning Interatomic Potentials (uMLIPs), pre-trained on massively diverse datasets encompassing inorganic materials and organic molecules across the entire periodic table, serve as ... |
| CroSatFL: Energy-Efficient Federated Learning with Cross-Aggregation for Satellite Edge Computing | Nan Yang, Bahman Javadi, Rodrigo Neves Calheiros, David Boland, Philip Leong | 2026-04-17 | 下载 | Low Earth Orbit (LEO) mega-constellations extend the cloud-to-edge continuum into space, enabling satellite edge computing. However, Federated Learning (FL) in this environment is fundamentally energy... |
| cuNNQS-SCI: A Fully GPU-Accelerated Framework for High-Performance Configuration Interaction Selection with Neural Network Quantum States | Daran Sun, Bowen Kan, Haoquan Long, Hairui Zhao, Haoxu Li, Yicheng Liu, Pengyu Zhou, Ankang Feng, Wenjing Huang, Yida Gu, Zhenyu Li, Honghui Shang, Yunquan Zhang, Dingwen Tao, Ninghui Sun, Guangming Tan | 2026-04-17 | 下载 | AI-driven methods have demonstrated considerable success in tackling the central challenge of accurately solving the Schrödinger equation for complex many-body systems. |
| PoSME: Proof of Sequential Memory Execution via Latency-Bound Pointer Chasing with Causal Hash Binding | David L. Condrey | 2026-04-17 | 下载 | We introduce PoSME (Proof of Sequential Memory Execution), a cryptographic primitive that enforces sustained sequential computation via latency-bound pointer chasing over a mutable arena. |
| Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving | Takeshi Yoshimura, Valentijn Dymphnus van de Beek, Tatsuhiro Chiba | 2026-04-17 | 下载 | Distributed LLM serving systems optimize per-request latency and throughput. However, under long-context workloads, inference accuracy becomes more variable. |
| BlockRaFT: A Distributed Framework for Fault-Tolerant and Scalable Blockchain Nodes | Manaswini Piduguralla, Souvik Sarkar, Arunmoezhi Ramachandran, Sathya Peri | 2026-04-17 | 下载 | Blockchain technology enhances transparency by maintaining a distributed ledger among mutually untrusting parties. Despite its advantages, scalability and availability remain critical bottlenecks that... |
| DataCenterGym: A Physics-Grounded Simulator for Multi-Objective Data Center Scheduling | Nilavra Pathak, Samadrita Biswas, Nirmalya Roy | 2026-04-17 | 下载 | Modern datacenters schedule heterogeneous workloads across geo-distributed sites with diverse compute capacities, electricity prices, and thermal conditions. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| End-to-End Performance of Video Streaming With MPEG-DASH Over Satellite 5G IAB Networks | Muhammad Adeel Zahid, Ekram Hossain, Peng Hu | 2026-04-17 | 下载 | We present an end-to-end performance evaluation of MPEG-DASH video streaming over a Low-Earth Orbit (LEO) satellite-based 5G Integrated Access and Backhaul (IAB) network. |
| Deterministic Task Offloading and Resource Allocation in the IoT-Edge-Cloud Continuum | Keyvan Aghababaiyan, Baldomero Coll-Perales, Javier Gozalvez | 2026-04-17 | 下载 | Future cellular networks will sustainably integrate computing, intelligence and services within a network of networks ecosystem that includes IoT devices and subnetworks for local communications and d... |
| Deterministic Task Scheduling in In-Vehicle Networks for Software-Defined Vehicles | Keyvan Aghababaiyan, Baldomero Coll-Perales, Luca Lusvarghi, Javier Gozalvez | 2026-04-17 | 下载 | Modern vehicles are embedding increasing levels of automation, connectivity, and intelligence, which require advanced in-vehicle networks and computational platforms to support the dependability and d... |
| Toward EU Sovereignty in Space: A Comparative Simulation Study of IRIS 2 and Starlink | Alexander Bonora, Marco Giordani, Michele Zorzi | 2026-04-17 | 下载 | The evolution of 6th generation (6G) networks increasingly relies on satellite-based Non-Terrestrial Networks (NTNs) to extend broadband connectivity to remote and unserved regions, and to support pub... |
| Characterization of Real Communication Patterns and Congestion Dynamics in HPC Interconnection Networks | Miguel Sánchez de La Rosa, Gabriel Gomez-Lopez, Alejandro Baviera, Jose Duro, Francisco J. andújar, Jesus Escudero-Sahuquillo, Pedro J. Garcia, Francisco J. Alfaro, Maria E. Gomez, Julio Sahuquillo, José L. Sánchez, Francisco J. Quiles | 2026-04-17 | 下载 | The interconnection network is a key component of Supercomputers and Data centers, and its design must cope with the increasing communication demands of current applications and services; otherwise, i... |
| Radio Environment Map for Energy-Efficient User-Centric Cell-Free M-MIMO Network | Marcin Hoffmann, Paweł Kryszkiewicz | 2026-04-17 | 下载 | This paper proposes a Radio Environment Map (REM) for energy-efficient (EE) serving cluster formulation in a user-centric cell-free network. By incorporating the location of the user and the character... |
| Federated Parameter-Efficient Adaptation for Interference Mitigation at the Wireless Edge | Evar Jones, Daniel J. Jakubisin, Sanmay Das | 2026-04-17 | 下载 | Dense wireless deployments face co-channel interference from heterogeneous sources that vary across base stations (gNBs in 5G). While centralized DNN-based approaches to interference mitigation have s... |
| Scalable Deterministic Task Offloading and Resource Allocation in the IoT-Edge-Cloud Continuum | Keyvan Aghababaiyan, Baldomero Coll-Perales, Javier Gozalvez | 2026-04-17 | 下载 | Future 6 G networks are envisioned as a network of networks (NoN) ecosystem, integrating communication and computing resources across multiple domains. |
| A Protocol-Agnostic Backscatter-Based Security Layer for Ultra-Low-Power SWIPT IoT Networks | Taki Eddine Djidjekh, Alexandru Takacs, Gaël Loubet, Lamoussa Sanogo, Daniela Dragomirescu | 2026-04-17 | 下载 | This paper presents a lightweight, protocol-agnostic security enhancement for Simultaneous Wireless Information and Power Transfer (SWIPT) in Internet of Things (IoT) applications. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Training Time Prediction for Mixed Precision-based Distributed Training | Minchul Kang, Changyong Shin, Jinwoo Jeong, Hyunho Lee, Younghun Go, Gyeongmin Kim, Gyeongsik Yang, Chuck Yoo | 2026-04-17 | 下载 | Accurate prediction of training time in distributed deep learning is crucial for resource allocation, cost estimation, and job scheduling. We observe that the floating-point precision setting is a key... |
| CPU Optimization of a Monocular 3D Biomechanics Pipeline for Low-Resource Deployment | Yan Zhang, Xiong Zhao | 2026-04-17 | 下载 | Markerless 3D movement analysis from monocular video enables accessible biomechanical assessment in clinical and sports settings. However, most research-grade pipelines rely on GPU acceleration, limit... |