2026-06-22
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling Model | Mansour Zoubeirou a Mayaki | 2026-06-22 | 下载 | Transformer-based models underpin modern natural language processing but incur rapidly growing computational and energy costs. As training scales in both model size and parallelism, accurately predict... |
| HeteroViT: A Versatile Single-Layer Vision Transformer Concept, Co-Designed for Distributed Real-Time Data Reduction on Scientific Detectors | Abhilasha Dave, Weijian Zheng, Antonino Miceli, Dionisio Doering, Ryan Herbst, Angelo Dragone | 2026-06-22 | 下载 | Next-generation X-ray detectors generate data faster than any system can affordably store or process. LCLS-II, the upgraded Linac Coherent Light Source at SLAC, produces data on the order of terabytes... |
| An Open-Source LFSR-Based Stochastic Leaky Integrate-and-Fire Neuron in SkyWater 130 nm: Design, Stochastic Characterisation, and Rate Coding | Poornima Kumaresan, Santhosh Sivasubramani | 2026-06-22 | 下载 | Stochastic spiking neurons trade exact arithmetic for controlled randomness, lowering area and tolerating input noise, which suits event-driven edge hardware. |
| VeriPilot: An LLM-Powered Verilog Debugging Framework | Yihan Wang, Cheng Liu, Jiazheng Zhang, Lei Zhang, Long Cheng, Xiaowei Li, Huawei Li | 2026-06-22 | 下载 | Verilog debugging remains one of the most time-consuming stages in digital circuit design. Recent advances in Large Language Models (LLMs) have enabled automated debugging; however, most existing appr... |
| MOCAP: Wafer-Scale-Chip-Oriented Memory-Orchestrated Chunked Pipelining Framework for Prefill-Only LLM Inference | Zichuan Wang, Huizheng Wang, Yuheng Xiao, Haonan Zuo, Taiquan Wei, Jinyi Deng, Chao Li, Yang Hu, Shouyi Yin | 2026-06-22 | 下载 | Large language models (LLMs) are increasingly used in prefill-only workloads, where end-to-end latency is dominated by the prefill phase. For long-context prefill, communication overhead grows with se... |
| Clutch: High Performance Vector-Scalar Comparison using DRAM via Chunked Temporal Coding | Daichi Tokuda, Tatsuya Kubo, Ismail Emir Yuksel, Ataberk Olgun, Haocong Luo, Tomoya Nagatani, Geraldo F. Oliveira, Abdullah Giray Yağlıkçı, Mohammad Sadrosadati, Onur Mutlu, Shinya Takamaeda-Yamazaki | 2026-06-22 | 下载 | Vector-scalar comparison is a fundamental computation primitive that compares each element in a vector against a single scalar value. It is widely used in various data-intensive workloads from databas... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing | Hang Yin, Kevin Wang | 2026-06-22 | 下载 | GPU Confidential Computing (GPU-CC) now preserves GPU-local performance: on NVIDIA B300, BF16 matmul runs at 0.998x of non-confidential performance. |
| LMS-AR: LMS Prediction-based Adaptive Regulator for Memory Bandwidth in Multicore Systems | Sudarshan Srinivasan, Deepak Gangadharan, Dip Goswami | 2026-06-22 | 下载 | Memory bandwidth contention in multi-core systems severely impacts application performance and quality-of-service (QoS) guarantees. Regulating the shared memory bandwidth mitigates the memory performa... |
| An Efficient Construction of Completely Independent Spanning Trees in Dense Gaussian Networks | Zaid Hussain, Fawaz AlAzemi, Bader AlBdaiwi | 2026-06-22 | 下载 | Fault tolerance in routing and broadcasting is a critical aspect in ensuring the reliability and robustness of communication networks, particularly in environments prone to failures. |
| Memory Layouts for GPU-Data Transfer Buffering in SPH | Mladen Ivkovic, Abouzied M. A. Nasar, Tobias Weinzierl, Matthieu Schaller, Benedict D. Rogers, Georgios Fourtakas, Scott T. Kay | 2026-06-22 | 下载 | The rise in GPU compute speed has outpaced improvements in host-to-device memory transfer speeds, despite the advent of shared-memory superchips. |
| Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse | Bole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein | 2026-06-22 | 下载 | Multimodal agents repeatedly re-examine the same video frames, UI screenshots, and rendered artifacts as their context window slides and reasoning iterates, yet every look-back re-encodes from scratch... |
| The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling Model | Mansour Zoubeirou a Mayaki | 2026-06-22 | 下载 | Transformer-based models underpin modern natural language processing but incur rapidly growing computational and energy costs. As training scales in both model size and parallelism, accurately predict... |
| Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference | Yuhang Gan, Yiwei Yang, Yuyi Li, Xiangyu Gao, Yichen Wang, Rain Jiang, Xiaoning Ding, Andi Quinn, Chen Qian | 2026-06-22 | 下载 | Long-running LLM agents keep valuable state resident on GPUs: KV caches, request schedulers, communication state, and sometimes online adapters. |
| Development and Design of FLKit: A Structured Onboarding Toolkit for Federated Learning in Health and Life Sciences | Ashkan Pirmani, Ilse Vermeulen, Goran Vinterhalter, Lotte Geys, Axel Faes, Muhammad Quamber Ali, Nishkala Sattanathan, Geert Vandeweyer, Yves Moreau, Liesbet M. Peeters | 2026-06-22 | 下载 | Federated learning lets institutions train shared models without moving their data, which makes it a natural fit for health and life sciences research under strict privacy regulation. |
| Asymmetry PRISM: A CPU/GPU Portfolio Optimization Engine for Deadline-Bounded Institutional Rebalancing | Debdoot Ghosh | 2026-06-22 | 下载 | Institutional rebalancing is a batched optimization workload with a hard operating deadline: hundreds of accounts need new weights under budget, turnover, exposure, exclusion, and tax-aware controls b... |
| When Staking Rewards Compound: Measuring the Impact of Ethereum's Pectra Upgrade | Mohammed Benseddik, Benjamin Kraner, Claudio J. Tessone | 2026-06-22 | 下载 | Ethereum's beacon chain hosts over 920,000 active validators, a number inflated by the legacy 32 ETH stake cap. The Pectra upgrade (May 2025) addresses this by introducing 0x02 compounding validators,... |
| Node-Level Performance and Energy Characterization of Flagship Science Applications on SuperMUC-NG Phase 2 | Salvatore Cielo, Elmira Birang, Alexander Pöppl, Sajad Azizi, Plamen Dobrev, Margarita Egelhofer, Ivan Pribec, Gerald Mathias | 2026-06-22 | 下载 | We present a systematic performance and energy-efficiency characterization of five flagship scientific workloads on SuperMUC-NG phase 2, the 28 PetaFLOPs system at the Leibniz Supercomputing Center (L... |
| Solving Approximate Agreement on continuous and discrete spaces | Augustin Albert, Sergio Rajsbaum | 2026-06-22 | 下载 | We consider asynchronous processes prone to crashes, communicating via shared read-write registers, and study the wait-free solvability of approximate agreement: given inputs, processes must outpu... |
| Efficient Network Inference via Hardware-Aware Architecture Search, Model Pruning & Quantization | Lucas Heublein, Mark Deutel, Axel Plinge, Felix Ott | 2026-06-22 | 下载 | Embedded global navigation satellite system (GNSS) interference monitoring requires fast and memory-efficient inference to process large volumes of raw in-phase and quadrature (IQ) samples in real tim... |
| Nautilus: A Verifiable Hierarchical Federated Learning Framework for Vehicular-Edge-Cloud Systems | Linyang Wu, Linpeng Jia, Hanwen Zhang, Tiantian Duan, Yi Sun | 2026-06-22 | 下载 | Federated Learning (FL) enables privacy-preserving collaborative learning for Internet of Vehicles (IoV) scenarios, but the extreme heterogeneity of vehicular-edge-cloud resources severely limits syst... |
| LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMs | Xiangyu Zhi, Peiqi Yin, Sheng Guan, Chenguang Zheng, James Cheng, Xiao Yan | 2026-06-22 | 下载 | Realtime omni-modal LMs support speech-centric conversations where users stream inputs, hear generated audio, and interrupt freely. Existing Omni-LM serving systems still rely on throughput-oriented L... |
| Decentralized Operations of Decarbonized Chemical Plants with Renewable-driven Transmission Systems | Richard Reed, kazi Arman Ahmed, Saba Ghasemi, Zheyu Jiang, Paritosh Ramanan | 2026-06-22 | 下载 | Electrification of ethane cracking offers a promising pathway to industrial decarbonization, provided that the electricity is sourced from renewable energy. |
| EchoFlow: A Workload-Aware Parameter Tuning Method for Blockchain Systems | Ben Lian, Linpeng Jia, Xing Chen, Xiaofeng Chen, Yi Sun | 2026-06-22 | 下载 | Blockchain systems expose a large number of tunable parameters that significantly influence system performance. However, in practice, a single parameter configuration is often applied across different... |
| Clutch: High Performance Vector-Scalar Comparison using DRAM via Chunked Temporal Coding | Daichi Tokuda, Tatsuya Kubo, Ismail Emir Yuksel, Ataberk Olgun, Haocong Luo, Tomoya Nagatani, Geraldo F. Oliveira, Abdullah Giray Yağlıkçı, Mohammad Sadrosadati, Onur Mutlu, Shinya Takamaeda-Yamazaki | 2026-06-22 | 下载 | Vector-scalar comparison is a fundamental computation primitive that compares each element in a vector against a single scalar value. It is widely used in various data-intensive workloads from databas... |
| Learning Filters with Certainty | Yuval Banoun, Daniel Sadoc Menasche, Ori Rottenstreich | 2026-06-22 | 下载 | Hash-based data structures such as Bloom filters are widely used in network systems for tasks including caching, anomaly detection, and machine learning pipelines. |
| Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCo | Chamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Hadi Mohaghegh Dolatabadi, Shamane Siriwardhana, Gil Avraham, Violetta Shevchenko, Karol Pajak, James Snewin, Alexander Long | 2026-06-22 | 下载 | To make large-scale distributed training practical outside high-bandwidth datacenters, we must reduce blocking, high-volume synchronization. While DiLoCo communicates infrequently, its outer synchroni... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Protection Switching in Hybrid Hollow-Core and Single-Mode Fiber Networks: Challenges, Analysis, and Mitigation Strategies | Md Ghulam Saber, Zhiping Jiang | 2026-06-22 | 下载 | Hollow-core fibers (HCF) are transitioning from laboratory curiosities to production-deployed infrastructure, with cloud providers operating thousands of kilometers of hollow-core links. |
| Wireless Personal Agent: Extending Wireless Intelligence from Networks to Terminals | Jiedan Tan, Fang Liu, Jingwen Tong, Shengli Zhang, Jun Zhang, Wing Shing Wong | 2026-06-22 | 下载 | Wireless networks are evolving from connectivity-oriented infrastructures into intelligent and personalized service platforms. Existing wireless intelligence remains centered on network-side optimizat... |
| LLM-Aided A* Search in Non-Geometric Network Graphs | Nouf Alabbasi, Esraa Ghourab, Omar Alhussein | 2026-06-22 | 下载 | Finding the shortest path in non-geometric network graphs, where edge weights encode arbitrary metrics such as latency or monetary cost rather than spatial distance, poses a challenge for informed sea... |
| Performance Evaluation of Selection Strategies for Inter-Satellite Paths in Walker-Delta Constellations | Marvin Felix Braun, Moritz Flüchter, Michael Menth | 2026-06-22 | 下载 | In LEO satellite constellations, traffic between a user terminal and a gateway is carried over a satellite path. As the satellite constellation rotates around Earth, a new path must be reselected repe... |
| LOLLA: Deep Reinforcement Learning for Closed-Loop Link Adaptation Towards a GPU-Accelerated AI-RAN | Rui Wang, Linchao Zhang, Qiang Liu, Kun Yang | 2026-06-22 | 下载 | Outer-loop link adaptation (OLLA) is widely deployed in 5G NR to track channel variations, yet its reliance on first-order, single-bit feedback degrades performance significantly under high-mobility a... |
| Understanding the Stealthy BGP Hijacking Risk in the ROV Era | Yihao Chen, Qi Li, Ke Xu, Zhuotao Liu, Jianping Wu | 2026-06-22 | 下载 | The partial deployment of Route Origin Validation (ROV) poses an unexpected security threat known as stealthy BGP hijacking, i.e., a particularly elusive form of BGP hijacking where malicious routes d... |
| CITADEL: CSI-Based Jamming Detection and Open-Set Classification for IIoT Networks | Aymen Bouferroum, Ildi Alla, Valeria Loscri, Abderrahim Benslimane, Vincent Lenders | 2026-06-22 | 下载 | Radio frequency jamming poses a critical threat to the availability of wireless Industrial Internet of Things (IIoT) networks. Existing detection and classification techniques are poorly suited to thi... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| LMS-AR: LMS Prediction-based Adaptive Regulator for Memory Bandwidth in Multicore Systems | Sudarshan Srinivasan, Deepak Gangadharan, Dip Goswami | 2026-06-22 | 下载 | Memory bandwidth contention in multi-core systems severely impacts application performance and quality-of-service (QoS) guarantees. Regulating the shared memory bandwidth mitigates the memory performa... |
| AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction | Shanhui Zhao, Jiacheng Liu, Guohong Liu, Jichao Yan, Jialei Ye, Yuhao Yang, Hao Wen, Shizuo Tian, Yizhen Yuan, Yuxuan Chen, Yunxin Liu, Ju Ren, Ya-Qin Zhang, Chao Huang, Yao Guo, Yuanchun Li | 2026-06-22 | 下载 | AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span applications and data sources. |
| FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation | Yinpeng Wu, Yitong Chen, Lixiang Wang, Jinyu Gu, Zhichao Hua, Yubin Xia | 2026-06-22 | 下载 | Device-side Large Language Models (LLMs) have grown explosively, offering stronger privacy and higher availability than their cloud-side counterparts. |
| EnerInfer: Energy-Aware On-Device LLM Inference | Bohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin, Debayan Roy, Yutao Liu, Yu Peng, Ning Jia, Haibo Chen | 2026-06-22 | 下载 | On-device LLM inference is increasingly attractive for privacy-preserving, reliable, and cost-effective deployment, yet its energy and thermal costs remain a critical bottleneck. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential Computing | Hang Yin, Kevin Wang | 2026-06-22 | 下载 | GPU Confidential Computing (GPU-CC) now preserves GPU-local performance: on NVIDIA B300, BF16 matmul runs at 0.998x of non-confidential performance. |
| LMS-AR: LMS Prediction-based Adaptive Regulator for Memory Bandwidth in Multicore Systems | Sudarshan Srinivasan, Deepak Gangadharan, Dip Goswami | 2026-06-22 | 下载 | Memory bandwidth contention in multi-core systems severely impacts application performance and quality-of-service (QoS) guarantees. Regulating the shared memory bandwidth mitigates the memory performa... |
| Memory Layouts for GPU-Data Transfer Buffering in SPH | Mladen Ivkovic, Abouzied M. A. Nasar, Tobias Weinzierl, Matthieu Schaller, Benedict D. Rogers, Georgios Fourtakas, Scott T. Kay | 2026-06-22 | 下载 | The rise in GPU compute speed has outpaced improvements in host-to-device memory transfer speeds, despite the advent of shared-memory superchips. |
| Node-Level Performance and Energy Characterization of Flagship Science Applications on SuperMUC-NG Phase 2 | Salvatore Cielo, Elmira Birang, Alexander Pöppl, Sajad Azizi, Plamen Dobrev, Margarita Egelhofer, Ivan Pribec, Gerald Mathias | 2026-06-22 | 下载 | We present a systematic performance and energy-efficiency characterization of five flagship scientific workloads on SuperMUC-NG phase 2, the 28 PetaFLOPs system at the Leibniz Supercomputing Center (L... |
| Learning Filters with Certainty | Yuval Banoun, Daniel Sadoc Menasche, Ori Rottenstreich | 2026-06-22 | 下载 | Hash-based data structures such as Bloom filters are widely used in network systems for tasks including caching, anomaly detection, and machine learning pipelines. |