2026-06-26
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| KernelSight-LM: A Kernel-Level LLM Inference Simulator | Xiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt | 2026-06-26 | 下载 | As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency... |
| Agentic Hardware Design as Repository-Level Code Evolution | Cunxi Yu, Chenhui Deng, Nathaniel Pinckney, Brucek Khailany | 2026-06-26 | 下载 | We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled into a project pack containing domain knowledge, an e... |
| AI-Driven Synthesis for High-Tech System Design: Automating Innovation | Luuk Oerlemans, Steven Westerhof, Theo Hofman | 2026-06-26 | 下载 | This article addresses the combinatorial complexity inherent in modern high-tech system design by presenting automation-in-design (AiD) as a transformative paradigm. |
| Self-Verifying Measurement Records: Hash-Linked Evidence Graphs for Hardware Benchmarking | Faruk Alpay, Baris Basaran | 2026-06-26 | 下载 | Performance numbers reported for hardware are accepted on trust: the reader cannot recompute them, the apparatus is gone, and the silicon itself can be silently wrong, with fleet studies reporting on ... |
| Phase Matters: Characterizing Heterogeneous Vision-Language Inference on a Mobile SoC | Aryama V Murthy, Yashas N Kotre, Prathmesh Sharma, Pragya Mishra, Sanjith Ganapathi, Priyesh Shukla | 2026-06-26 | 下载 | Recent phone-class mobile SoCs expose practical NPU execution paths for on-device vision-language model (VLM) inference, but developers still lack phase-level guidance for mapping VLM pipelines across... |
| Co-Optimization of Analog Kolmogorov-Arnold Networks for Low-Power Function Approximation in Flexible Electronics | Paula Carolina Lozano Duarte, Georgios Zervakis, Mehdi Tahoori, Sani Nassif | 2026-06-26 | 下载 | Wearable devices and Internet of Things (IoT) sensors require on-sensor processing of biosignals and environmental data, including computationally demanding operations such as nonlinear activation fun... |
| SEADA: An efficient methodology for optimizing mixed-precision DNNs on multi-precision spatial architectures | Leandro Fiorin, Marco Ronzani, Cristina Silvano | 2026-06-26 | 下载 | Mixed-precision computation has been introduced in deep neural networks (DNNs) as an effective approach to reduce latency, energy consumption, and memory footprint. |
| MultModLM: A multi-modal benchmark for Large-Language Model based hardware schematic generation | Dhruv Kulkarni, Sai Manoj Pudukotai Dinkarrao | 2026-06-26 | 下载 | Recently, Large Language models (LLMs) find application in several fields. This extends to hardware definition and synthesis. However, most works at the intersection of LLMs and hardware generation fo... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| High-Performance Resilient Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations at Scale | Jeremy J. Williams, Stefan Costea, David Tskhakaya, Leon Kos, Ales Podolnik, Jakub Hromadka, Jordy Trilaksono, Yi Ju, Kallia Chronaki, Evangelos Gkolantas, Vassilis Papaefstathiou, Allen D. Malony, Sameer Shende, Frank Jenko, Erwin Laure, Stefano Markidis | 2026-06-26 | 下载 | The increasing demand for high-performance computing in plasma physics has driven scalable and resilient simulation methods capable of efficiently exploiting modern multi-GPU architectures. |
| CLEAR-MoE: Shared-Basis Expert Extraction from Frozen Vision Transformers via Calibration-Driven Layer Selection | Md Irtiza Hossain, Humaira Ayesha, Junaid Ahmed Sifat | 2026-06-26 | 下载 | We present CLEAR-MoE, a four-phase post-training pipeline that converts a frozen pretrained Vision Transformer (ViT) into a sparse Mixture-of-Experts (MoE) model without updating backbone weights. |
| Towards Value-Constrained Credit Assignment in Fully Delegated AI Cooperatives | Young Yoon, Jimin Kim, Soyeon Park | 2026-06-26 | 下载 | We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agents that contribute data and participate in model updates under heterogeneous value c... |
| DiStash: A Disaggregated Multi-Stash Transactional Key-Value Store | Yiming Gao, Hieu Nguyen, Jun Li, Shahram Ghandeharizadeh | 2026-06-26 | 下载 | A stash is a storage medium such as Dynamic Random Access Memory (DRAM), Solid State Disk (SSD), Hard Disk Drive (HDD), or Non-Volatile Memory (NVM). |
| RAMSES: Secure high-performance computing for sensitive data | Peter Heger, Lech Nieroda, Roland Pabel, Christoph Stollwerk, Stefan Borowski, Kamil Tokmakov, Michael Commer, Martin Peifer, Stefan Wesner, Viktor Achter | 2026-06-26 | 下载 | Traditionally, the architecture of high-performance computing (HPC) systems is tailored for speed, while highly secure computer systems must sacrifice speed for security. |
| Exploring and Exploiting Synchrony Limitations of Time-Triggered Network-Agnostic Guardians | Shreya Vithal Kulhalli, Mohammad Ibrahim Alkoudsi, Gerhard Fohler | 2026-06-26 | 下载 | Time-triggered communication protocols rely on trusted components known as guardians to enforce adherence to predetermined network schedules. Network-agnostic guardians offer an efficient and scalable... |
| Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems | Adrian P. Dieguez, Victor Conchello Vendrell, Alex Batlle, Vinnam Kim, Jordi Ros-Giralt, Harris Teague | 2026-06-26 | 下载 | Knowledge Distillation (KD) enables training smaller student models under the guidance of larger teacher models, and the widely adopted TRL library implements it. |
| Lightweight Multi-Vehicle Collaborative Perception Acceleration with Fusion Position Adjustment | Wenzhao Zhang, Shujun Han, Haixiao Gao, Mengying Sun, Bizhu Wang, Xiaodong Xu | 2026-06-26 | 下载 | Multi-vehicle collaborative perception (MvCP) is considered as a key technology to facilitate automated driving (AD), where real-time MvCP under limited resources is significant for reliable AD. |
| How far does a random forest generalize from a 54-run LAMMPS+SPICA benchmark? | Dennis Alves Pedersen, Paulo Henrique Leme Ramalho, Fábio Andrijauskas | 2026-06-26 | 下载 | Selecting near-optimal hybrid MPI+OpenMP configurations for molecular dynamics workloads on modern HPC clusters has traditionally required exhaustive empirical benchmarking, consuming allocation budge... |
| P-ARC: Exploiting Subproblem Independence for Parallel Multi-Robot Motion Planning | James D. Motes, Marco Morales, Nancy M. Amato | 2026-06-26 | 下载 | This paper presents Parallel ARC (P-ARC), a parallel variant of the Adaptive Robot Coordination (ARC) approach to multi-robot motion planning (MRMP). |
| FoggyTrust: Robust Federated Learning with Hierarchical Trust Networks | Emmanuel Rassou, Tomas Gonzalez | 2026-06-26 | 下载 | Byzantine-robust federated learning seeks to protect distributed model training from malicious or corrupted clients without requiring access to their private data. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| V-TSN: A Software-Defined TSN Overlay for General-Purpose Networks | Mohammadparsa Karimi, Majid Nabi, Ahmed Khalaf, Andrew Nelson, Kees Goossens, Twan Basten | 2026-06-26 | 下载 | Time-Sensitive Networking (TSN) extends Ethernet with deterministic communication for time-critical applications such as industrial automation, in-vehicle networks, and cyber-physical systems. |
| AB-Sync: Attention-Based Slot-Level Clock Synchronization Method for UWB-TDOA Localization Networks | Tianyi Lyu, Kefei Tian, Kangqiao Qin, Qingwen Liu, Mingqing Liu | 2026-06-26 | 下载 | Ultra-wideband (UWB) time-difference-of-arrival (TDOA) localization networks provide high-update-rate indoor location services for IoT and cyber-physical applications, but their accuracy depends on na... |
| GTI-mSEMP Framework : A Proposed Framework to Simulate Malware Propagation with Inclusion of Attacker-Defender Strategy | Shadeeb Hossain, Kristopher Wilson | 2026-06-26 | 下载 | The rapid proliferation of automated, multi-vector malware threats poses a significant risk to heterogeneous, resource constrained cyber-physical networks. |
| Host-Driven Flowlet Balancing with Segment Routing over IPv6 | Ryo Nakamura, Hiroki Kano, Tomoko Okuzawa | 2026-06-26 | 下载 | This paper proposes a fully host-driven method for flowlet balancing with Segment Routing over IPv6 (SRv6). In modern data center networks, load balancing plays a pivotal role in efficiently utilizing... |
| Real-Time State Estimation in Smart Grids over 5G Networks: Experimental Validation Using Raspberry Pis and Typhoon HIL | Biswajit Kumar Dash, Luis Herrera, Filippo Malandra | 2026-06-26 | 下载 | Reliable, low-latency communication is critical for real-time monitoring and control in modern Smart Grids (SGs). The emergence of 5G networks, with enhanced reliability, significantly lower latency, ... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| KernelSight-LM: A Kernel-Level LLM Inference Simulator | Xiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt | 2026-06-26 | 下载 | As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency... |
| High-Performance Resilient Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations at Scale | Jeremy J. Williams, Stefan Costea, David Tskhakaya, Leon Kos, Ales Podolnik, Jakub Hromadka, Jordy Trilaksono, Yi Ju, Kallia Chronaki, Evangelos Gkolantas, Vassilis Papaefstathiou, Allen D. Malony, Sameer Shende, Frank Jenko, Erwin Laure, Stefano Markidis | 2026-06-26 | 下载 | The increasing demand for high-performance computing in plasma physics has driven scalable and resilient simulation methods capable of efficiently exploiting modern multi-GPU architectures. |
| DiStash: A Disaggregated Multi-Stash Transactional Key-Value Store | Yiming Gao, Hieu Nguyen, Jun Li, Shahram Ghandeharizadeh | 2026-06-26 | 下载 | A stash is a storage medium such as Dynamic Random Access Memory (DRAM), Solid State Disk (SSD), Hard Disk Drive (HDD), or Non-Volatile Memory (NVM). |
| Mixed-Precision For Energy Efficient Computations | Gülçin Gedik, Robert Schöne, Roman Iakymchuk | 2026-06-26 | 下载 | As simulations grow more realistic, the pursuit of higher accuracy results in extended computation times and substantial power consumption. This study explores mixed-precision computing as a promising... |