Skip to content

2026-06-26 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
KernelSight-LM: A Kernel-Level LLM Inference SimulatorXiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt2026-06-26下载As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency...
Agentic Hardware Design as Repository-Level Code EvolutionCunxi Yu, Chenhui Deng, Nathaniel Pinckney, Brucek Khailany2026-06-26下载We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled into a project pack containing domain knowledge, an e...
AI-Driven Synthesis for High-Tech System Design: Automating InnovationLuuk Oerlemans, Steven Westerhof, Theo Hofman2026-06-26下载This article addresses the combinatorial complexity inherent in modern high-tech system design by presenting automation-in-design (AiD) as a transformative paradigm.
Self-Verifying Measurement Records: Hash-Linked Evidence Graphs for Hardware BenchmarkingFaruk Alpay, Baris Basaran2026-06-26下载Performance numbers reported for hardware are accepted on trust: the reader cannot recompute them, the apparatus is gone, and the silicon itself can be silently wrong, with fleet studies reporting on ...
Phase Matters: Characterizing Heterogeneous Vision-Language Inference on a Mobile SoCAryama V Murthy, Yashas N Kotre, Prathmesh Sharma, Pragya Mishra, Sanjith Ganapathi, Priyesh Shukla2026-06-26下载Recent phone-class mobile SoCs expose practical NPU execution paths for on-device vision-language model (VLM) inference, but developers still lack phase-level guidance for mapping VLM pipelines across...
Co-Optimization of Analog Kolmogorov-Arnold Networks for Low-Power Function Approximation in Flexible ElectronicsPaula Carolina Lozano Duarte, Georgios Zervakis, Mehdi Tahoori, Sani Nassif2026-06-26下载Wearable devices and Internet of Things (IoT) sensors require on-sensor processing of biosignals and environmental data, including computationally demanding operations such as nonlinear activation fun...
SEADA: An efficient methodology for optimizing mixed-precision DNNs on multi-precision spatial architecturesLeandro Fiorin, Marco Ronzani, Cristina Silvano2026-06-26下载Mixed-precision computation has been introduced in deep neural networks (DNNs) as an effective approach to reduce latency, energy consumption, and memory footprint.
MultModLM: A multi-modal benchmark for Large-Language Model based hardware schematic generationDhruv Kulkarni, Sai Manoj Pudukotai Dinkarrao2026-06-26下载Recently, Large Language models (LLMs) find application in several fields. This extends to hardware definition and synthesis. However, most works at the intersection of LLMs and hardware generation fo...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
High-Performance Resilient Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations at ScaleJeremy J. Williams, Stefan Costea, David Tskhakaya, Leon Kos, Ales Podolnik, Jakub Hromadka, Jordy Trilaksono, Yi Ju, Kallia Chronaki, Evangelos Gkolantas, Vassilis Papaefstathiou, Allen D. Malony, Sameer Shende, Frank Jenko, Erwin Laure, Stefano Markidis2026-06-26下载The increasing demand for high-performance computing in plasma physics has driven scalable and resilient simulation methods capable of efficiently exploiting modern multi-GPU architectures.
CLEAR-MoE: Shared-Basis Expert Extraction from Frozen Vision Transformers via Calibration-Driven Layer SelectionMd Irtiza Hossain, Humaira Ayesha, Junaid Ahmed Sifat2026-06-26下载We present CLEAR-MoE, a four-phase post-training pipeline that converts a frozen pretrained Vision Transformer (ViT) into a sparse Mixture-of-Experts (MoE) model without updating backbone weights.
Towards Value-Constrained Credit Assignment in Fully Delegated AI CooperativesYoung Yoon, Jimin Kim, Soyeon Park2026-06-26下载We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agents that contribute data and participate in model updates under heterogeneous value c...
DiStash: A Disaggregated Multi-Stash Transactional Key-Value StoreYiming Gao, Hieu Nguyen, Jun Li, Shahram Ghandeharizadeh2026-06-26下载A stash is a storage medium such as Dynamic Random Access Memory (DRAM), Solid State Disk (SSD), Hard Disk Drive (HDD), or Non-Volatile Memory (NVM).
RAMSES: Secure high-performance computing for sensitive dataPeter Heger, Lech Nieroda, Roland Pabel, Christoph Stollwerk, Stefan Borowski, Kamil Tokmakov, Michael Commer, Martin Peifer, Stefan Wesner, Viktor Achter2026-06-26下载Traditionally, the architecture of high-performance computing (HPC) systems is tailored for speed, while highly secure computer systems must sacrifice speed for security.
Exploring and Exploiting Synchrony Limitations of Time-Triggered Network-Agnostic GuardiansShreya Vithal Kulhalli, Mohammad Ibrahim Alkoudsi, Gerhard Fohler2026-06-26下载Time-triggered communication protocols rely on trusted components known as guardians to enforce adherence to predetermined network schedules. Network-agnostic guardians offer an efficient and scalable...
Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC SystemsAdrian P. Dieguez, Victor Conchello Vendrell, Alex Batlle, Vinnam Kim, Jordi Ros-Giralt, Harris Teague2026-06-26下载Knowledge Distillation (KD) enables training smaller student models under the guidance of larger teacher models, and the widely adopted TRL library implements it.
Lightweight Multi-Vehicle Collaborative Perception Acceleration with Fusion Position AdjustmentWenzhao Zhang, Shujun Han, Haixiao Gao, Mengying Sun, Bizhu Wang, Xiaodong Xu2026-06-26下载Multi-vehicle collaborative perception (MvCP) is considered as a key technology to facilitate automated driving (AD), where real-time MvCP under limited resources is significant for reliable AD.
How far does a random forest generalize from a 54-run LAMMPS+SPICA benchmark?Dennis Alves Pedersen, Paulo Henrique Leme Ramalho, Fábio Andrijauskas2026-06-26下载Selecting near-optimal hybrid MPI+OpenMP configurations for molecular dynamics workloads on modern HPC clusters has traditionally required exhaustive empirical benchmarking, consuming allocation budge...
P-ARC: Exploiting Subproblem Independence for Parallel Multi-Robot Motion PlanningJames D. Motes, Marco Morales, Nancy M. Amato2026-06-26下载This paper presents Parallel ARC (P-ARC), a parallel variant of the Adaptive Robot Coordination (ARC) approach to multi-robot motion planning (MRMP).
FoggyTrust: Robust Federated Learning with Hierarchical Trust NetworksEmmanuel Rassou, Tomas Gonzalez2026-06-26下载Byzantine-robust federated learning seeks to protect distributed model training from malicious or corrupted clients without requiring access to their private data.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
V-TSN: A Software-Defined TSN Overlay for General-Purpose NetworksMohammadparsa Karimi, Majid Nabi, Ahmed Khalaf, Andrew Nelson, Kees Goossens, Twan Basten2026-06-26下载Time-Sensitive Networking (TSN) extends Ethernet with deterministic communication for time-critical applications such as industrial automation, in-vehicle networks, and cyber-physical systems.
AB-Sync: Attention-Based Slot-Level Clock Synchronization Method for UWB-TDOA Localization NetworksTianyi Lyu, Kefei Tian, Kangqiao Qin, Qingwen Liu, Mingqing Liu2026-06-26下载Ultra-wideband (UWB) time-difference-of-arrival (TDOA) localization networks provide high-update-rate indoor location services for IoT and cyber-physical applications, but their accuracy depends on na...
GTI-mSEMP Framework : A Proposed Framework to Simulate Malware Propagation with Inclusion of Attacker-Defender StrategyShadeeb Hossain, Kristopher Wilson2026-06-26下载The rapid proliferation of automated, multi-vector malware threats poses a significant risk to heterogeneous, resource constrained cyber-physical networks.
Host-Driven Flowlet Balancing with Segment Routing over IPv6Ryo Nakamura, Hiroki Kano, Tomoko Okuzawa2026-06-26下载This paper proposes a fully host-driven method for flowlet balancing with Segment Routing over IPv6 (SRv6). In modern data center networks, load balancing plays a pivotal role in efficiently utilizing...
Real-Time State Estimation in Smart Grids over 5G Networks: Experimental Validation Using Raspberry Pis and Typhoon HILBiswajit Kumar Dash, Luis Herrera, Filippo Malandra2026-06-26下载Reliable, low-latency communication is critical for real-time monitoring and control in modern Smart Grids (SGs). The emergence of 5G networks, with enhanced reliability, significantly lower latency, ...

cs.PF - Performance ​

标题作者发布日期PDF摘要
KernelSight-LM: A Kernel-Level LLM Inference SimulatorXiteng Yao, Taeho Kim, Hengzhi Pei, Xinle Liu, Kyle Ulrich, Leonard Lausen, Ashish Khetan, Xiang Song, George Karypis, Martin Herbordt2026-06-26下载As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hardware, models, and serving parameters to meet cost and latency...
High-Performance Resilient Multi-GPU Hybrid Particle-in-Cell Monte Carlo Simulations at ScaleJeremy J. Williams, Stefan Costea, David Tskhakaya, Leon Kos, Ales Podolnik, Jakub Hromadka, Jordy Trilaksono, Yi Ju, Kallia Chronaki, Evangelos Gkolantas, Vassilis Papaefstathiou, Allen D. Malony, Sameer Shende, Frank Jenko, Erwin Laure, Stefano Markidis2026-06-26下载The increasing demand for high-performance computing in plasma physics has driven scalable and resilient simulation methods capable of efficiently exploiting modern multi-GPU architectures.
DiStash: A Disaggregated Multi-Stash Transactional Key-Value StoreYiming Gao, Hieu Nguyen, Jun Li, Shahram Ghandeharizadeh2026-06-26下载A stash is a storage medium such as Dynamic Random Access Memory (DRAM), Solid State Disk (SSD), Hard Disk Drive (HDD), or Non-Volatile Memory (NVM).
Mixed-Precision For Energy Efficient ComputationsGülçin Gedik, Robert Schöne, Roman Iakymchuk2026-06-26下载As simulations grow more realistic, the pursuit of higher accuracy results in extended computation times and substantial power consumption. This study explores mixed-precision computing as a promising...

基于 VitePress 构建 · 使用本地搜索查找论文