Skip to content

2026-06-22 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling ModelMansour Zoubeirou a Mayaki2026-06-22下载Transformer-based models underpin modern natural language processing but incur rapidly growing computational and energy costs. As training scales in both model size and parallelism, accurately predict...
HeteroViT: A Versatile Single-Layer Vision Transformer Concept, Co-Designed for Distributed Real-Time Data Reduction on Scientific DetectorsAbhilasha Dave, Weijian Zheng, Antonino Miceli, Dionisio Doering, Ryan Herbst, Angelo Dragone2026-06-22下载Next-generation X-ray detectors generate data faster than any system can affordably store or process. LCLS-II, the upgraded Linac Coherent Light Source at SLAC, produces data on the order of terabytes...
An Open-Source LFSR-Based Stochastic Leaky Integrate-and-Fire Neuron in SkyWater 130 nm: Design, Stochastic Characterisation, and Rate CodingPoornima Kumaresan, Santhosh Sivasubramani2026-06-22下载Stochastic spiking neurons trade exact arithmetic for controlled randomness, lowering area and tolerating input noise, which suits event-driven edge hardware.
VeriPilot: An LLM-Powered Verilog Debugging FrameworkYihan Wang, Cheng Liu, Jiazheng Zhang, Lei Zhang, Long Cheng, Xiaowei Li, Huawei Li2026-06-22下载Verilog debugging remains one of the most time-consuming stages in digital circuit design. Recent advances in Large Language Models (LLMs) have enabled automated debugging; however, most existing appr...
MOCAP: Wafer-Scale-Chip-Oriented Memory-Orchestrated Chunked Pipelining Framework for Prefill-Only LLM InferenceZichuan Wang, Huizheng Wang, Yuheng Xiao, Haonan Zuo, Taiquan Wei, Jinyi Deng, Chao Li, Yang Hu, Shouyi Yin2026-06-22下载Large language models (LLMs) are increasingly used in prefill-only workloads, where end-to-end latency is dominated by the prefill phase. For long-context prefill, communication overhead grows with se...
Clutch: High Performance Vector-Scalar Comparison using DRAM via Chunked Temporal CodingDaichi Tokuda, Tatsuya Kubo, Ismail Emir Yuksel, Ataberk Olgun, Haocong Luo, Tomoya Nagatani, Geraldo F. Oliveira, Abdullah Giray Yağlıkçı, Mohammad Sadrosadati, Onur Mutlu, Shinya Takamaeda-Yamazaki2026-06-22下载Vector-scalar comparison is a fundamental computation primitive that compares each element in a vector against a single scalar value. It is widely used in various data-intensive workloads from databas...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential ComputingHang Yin, Kevin Wang2026-06-22下载GPU Confidential Computing (GPU-CC) now preserves GPU-local performance: on NVIDIA B300, BF16 matmul runs at 0.998x of non-confidential performance.
LMS-AR: LMS Prediction-based Adaptive Regulator for Memory Bandwidth in Multicore SystemsSudarshan Srinivasan, Deepak Gangadharan, Dip Goswami2026-06-22下载Memory bandwidth contention in multi-core systems severely impacts application performance and quality-of-service (QoS) guarantees. Regulating the shared memory bandwidth mitigates the memory performa...
An Efficient Construction of Completely Independent Spanning Trees in Dense Gaussian NetworksZaid Hussain, Fawaz AlAzemi, Bader AlBdaiwi2026-06-22下载Fault tolerance in routing and broadcasting is a critical aspect in ensuring the reliability and robustness of communication networks, particularly in environments prone to failures.
Memory Layouts for GPU-Data Transfer Buffering in SPHMladen Ivkovic, Abouzied M. A. Nasar, Tobias Weinzierl, Matthieu Schaller, Benedict D. Rogers, Georgios Fourtakas, Scott T. Kay2026-06-22下载The rise in GPU compute speed has outpaced improvements in host-to-device memory transfer speeds, despite the advent of shared-memory superchips.
Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free ReuseBole Ma, Jan Eitzinger, Harald Koestler, Gerhard Wellein2026-06-22下载Multimodal agents repeatedly re-examine the same video frames, UI screenshots, and rendered artifacts as their context window slides and reasoning iterates, yet every look-back re-encodes from scratch...
The Energy Consumption of Transformer Fine-Tuning: A Roofline-Inspired Scaling ModelMansour Zoubeirou a Mayaki2026-06-22下载Transformer-based models underpin modern natural language processing but incur rapidly growing computational and energy costs. As training scales in both model size and parallelism, accurately predict...
Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM InferenceYuhang Gan, Yiwei Yang, Yuyi Li, Xiangyu Gao, Yichen Wang, Rain Jiang, Xiaoning Ding, Andi Quinn, Chen Qian2026-06-22下载Long-running LLM agents keep valuable state resident on GPUs: KV caches, request schedulers, communication state, and sometimes online adapters.
Development and Design of FLKit: A Structured Onboarding Toolkit for Federated Learning in Health and Life SciencesAshkan Pirmani, Ilse Vermeulen, Goran Vinterhalter, Lotte Geys, Axel Faes, Muhammad Quamber Ali, Nishkala Sattanathan, Geert Vandeweyer, Yves Moreau, Liesbet M. Peeters2026-06-22下载Federated learning lets institutions train shared models without moving their data, which makes it a natural fit for health and life sciences research under strict privacy regulation.
Asymmetry PRISM: A CPU/GPU Portfolio Optimization Engine for Deadline-Bounded Institutional RebalancingDebdoot Ghosh2026-06-22下载Institutional rebalancing is a batched optimization workload with a hard operating deadline: hundreds of accounts need new weights under budget, turnover, exposure, exclusion, and tax-aware controls b...
When Staking Rewards Compound: Measuring the Impact of Ethereum's Pectra UpgradeMohammed Benseddik, Benjamin Kraner, Claudio J. Tessone2026-06-22下载Ethereum's beacon chain hosts over 920,000 active validators, a number inflated by the legacy 32 ETH stake cap. The Pectra upgrade (May 2025) addresses this by introducing 0x02 compounding validators,...
Node-Level Performance and Energy Characterization of Flagship Science Applications on SuperMUC-NG Phase 2Salvatore Cielo, Elmira Birang, Alexander Pöppl, Sajad Azizi, Plamen Dobrev, Margarita Egelhofer, Ivan Pribec, Gerald Mathias2026-06-22下载We present a systematic performance and energy-efficiency characterization of five flagship scientific workloads on SuperMUC-NG phase 2, the 28 PetaFLOPs system at the Leibniz Supercomputing Center (L...
Solving Approximate Agreement on continuous and discrete spacesAugustin Albert, Sergio Rajsbaum2026-06-22下载We consider nn asynchronous processes prone to crashes, communicating via shared read-write registers, and study the wait-free solvability of approximate agreement: given inputs, processes must outpu...
Efficient Network Inference via Hardware-Aware Architecture Search, Model Pruning & QuantizationLucas Heublein, Mark Deutel, Axel Plinge, Felix Ott2026-06-22下载Embedded global navigation satellite system (GNSS) interference monitoring requires fast and memory-efficient inference to process large volumes of raw in-phase and quadrature (IQ) samples in real tim...
Nautilus: A Verifiable Hierarchical Federated Learning Framework for Vehicular-Edge-Cloud SystemsLinyang Wu, Linpeng Jia, Hanwen Zhang, Tiantian Duan, Yi Sun2026-06-22下载Federated Learning (FL) enables privacy-preserving collaborative learning for Internet of Vehicles (IoV) scenarios, but the extreme heterogeneity of vehicular-edge-cloud resources severely limits syst...
LiveServe: Interaction-Aware Serving for Real-Time Omni-Modal LLMsXiangyu Zhi, Peiqi Yin, Sheng Guan, Chenguang Zheng, James Cheng, Xiao Yan2026-06-22下载Realtime omni-modal LMs support speech-centric conversations where users stream inputs, hear generated audio, and interrupt freely. Existing Omni-LM serving systems still rely on throughput-oriented L...
Decentralized Operations of Decarbonized Chemical Plants with Renewable-driven Transmission SystemsRichard Reed, kazi Arman Ahmed, Saba Ghasemi, Zheyu Jiang, Paritosh Ramanan2026-06-22下载Electrification of ethane cracking offers a promising pathway to industrial decarbonization, provided that the electricity is sourced from renewable energy.
EchoFlow: A Workload-Aware Parameter Tuning Method for Blockchain SystemsBen Lian, Linpeng Jia, Xing Chen, Xiaofeng Chen, Yi Sun2026-06-22下载Blockchain systems expose a large number of tunable parameters that significantly influence system performance. However, in practice, a single parameter configuration is often applied across different...
Clutch: High Performance Vector-Scalar Comparison using DRAM via Chunked Temporal CodingDaichi Tokuda, Tatsuya Kubo, Ismail Emir Yuksel, Ataberk Olgun, Haocong Luo, Tomoya Nagatani, Geraldo F. Oliveira, Abdullah Giray Yağlıkçı, Mohammad Sadrosadati, Onur Mutlu, Shinya Takamaeda-Yamazaki2026-06-22下载Vector-scalar comparison is a fundamental computation primitive that compares each element in a vector against a single scalar value. It is widely used in various data-intensive workloads from databas...
Learning Filters with CertaintyYuval Banoun, Daniel Sadoc Menasche, Ori Rottenstreich2026-06-22下载Hash-based data structures such as Bloom filters are widely used in network systems for tasks including caching, anomaly detection, and machine learning pipelines.
Factored Gossip DiLoCo: Reducing Blocking Communication in DiLoCoChamin Hewa Koneputugodage, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Hadi Mohaghegh Dolatabadi, Shamane Siriwardhana, Gil Avraham, Violetta Shevchenko, Karol Pajak, James Snewin, Alexander Long2026-06-22下载To make large-scale distributed training practical outside high-bandwidth datacenters, we must reduce blocking, high-volume synchronization. While DiLoCo communicates infrequently, its outer synchroni...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Protection Switching in Hybrid Hollow-Core and Single-Mode Fiber Networks: Challenges, Analysis, and Mitigation StrategiesMd Ghulam Saber, Zhiping Jiang2026-06-22下载Hollow-core fibers (HCF) are transitioning from laboratory curiosities to production-deployed infrastructure, with cloud providers operating thousands of kilometers of hollow-core links.
Wireless Personal Agent: Extending Wireless Intelligence from Networks to TerminalsJiedan Tan, Fang Liu, Jingwen Tong, Shengli Zhang, Jun Zhang, Wing Shing Wong2026-06-22下载Wireless networks are evolving from connectivity-oriented infrastructures into intelligent and personalized service platforms. Existing wireless intelligence remains centered on network-side optimizat...
LLM-Aided A* Search in Non-Geometric Network GraphsNouf Alabbasi, Esraa Ghourab, Omar Alhussein2026-06-22下载Finding the shortest path in non-geometric network graphs, where edge weights encode arbitrary metrics such as latency or monetary cost rather than spatial distance, poses a challenge for informed sea...
Performance Evaluation of Selection Strategies for Inter-Satellite Paths in Walker-Delta ConstellationsMarvin Felix Braun, Moritz Flüchter, Michael Menth2026-06-22下载In LEO satellite constellations, traffic between a user terminal and a gateway is carried over a satellite path. As the satellite constellation rotates around Earth, a new path must be reselected repe...
LOLLA: Deep Reinforcement Learning for Closed-Loop Link Adaptation Towards a GPU-Accelerated AI-RANRui Wang, Linchao Zhang, Qiang Liu, Kun Yang2026-06-22下载Outer-loop link adaptation (OLLA) is widely deployed in 5G NR to track channel variations, yet its reliance on first-order, single-bit feedback degrades performance significantly under high-mobility a...
Understanding the Stealthy BGP Hijacking Risk in the ROV EraYihao Chen, Qi Li, Ke Xu, Zhuotao Liu, Jianping Wu2026-06-22下载The partial deployment of Route Origin Validation (ROV) poses an unexpected security threat known as stealthy BGP hijacking, i.e., a particularly elusive form of BGP hijacking where malicious routes d...
CITADEL: CSI-Based Jamming Detection and Open-Set Classification for IIoT NetworksAymen Bouferroum, Ildi Alla, Valeria Loscri, Abderrahim Benslimane, Vincent Lenders2026-06-22下载Radio frequency jamming poses a critical threat to the availability of wireless Industrial Internet of Things (IIoT) networks. Existing detection and classification techniques are poorly suited to thi...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
LMS-AR: LMS Prediction-based Adaptive Regulator for Memory Bandwidth in Multicore SystemsSudarshan Srinivasan, Deepak Gangadharan, Dip Goswami2026-06-22下载Memory bandwidth contention in multi-core systems severely impacts application performance and quality-of-service (QoS) guarantees. Regulating the shared memory bandwidth mitigates the memory performa...
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure InteractionShanhui Zhao, Jiacheng Liu, Guohong Liu, Jichao Yan, Jialei Ye, Yuhao Yang, Hao Wen, Shizuo Tian, Yizhen Yuan, Yuxuan Chen, Yunxin Liu, Ju Ren, Ya-Qin Zhang, Chao Huang, Yao Guo, Yuanchun Li2026-06-22下载AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span applications and data sources.
FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource IsolationYinpeng Wu, Yitong Chen, Lixiang Wang, Jinyu Gu, Zhichao Hua, Yubin Xia2026-06-22下载Device-side Large Language Models (LLMs) have grown explosively, offering stronger privacy and higher availability than their cloud-side counterparts.
EnerInfer: Energy-Aware On-Device LLM InferenceBohua Zou, Nian Liu, Binqi Sun, Matteo Mascherin, Debayan Roy, Yutao Liu, Yu Peng, Ning Jia, Haibo Chen2026-06-22下载On-device LLM inference is increasingly attractive for privacy-preserving, reliable, and cost-effective deployment, yet its energy and thermal costs remain a critical bottleneck.

cs.PF - Performance ​

标题作者发布日期PDF摘要
The Serialized Bridge: Understanding and Recovering LLM Serving Performance under Blackwell GPU Confidential ComputingHang Yin, Kevin Wang2026-06-22下载GPU Confidential Computing (GPU-CC) now preserves GPU-local performance: on NVIDIA B300, BF16 matmul runs at 0.998x of non-confidential performance.
LMS-AR: LMS Prediction-based Adaptive Regulator for Memory Bandwidth in Multicore SystemsSudarshan Srinivasan, Deepak Gangadharan, Dip Goswami2026-06-22下载Memory bandwidth contention in multi-core systems severely impacts application performance and quality-of-service (QoS) guarantees. Regulating the shared memory bandwidth mitigates the memory performa...
Memory Layouts for GPU-Data Transfer Buffering in SPHMladen Ivkovic, Abouzied M. A. Nasar, Tobias Weinzierl, Matthieu Schaller, Benedict D. Rogers, Georgios Fourtakas, Scott T. Kay2026-06-22下载The rise in GPU compute speed has outpaced improvements in host-to-device memory transfer speeds, despite the advent of shared-memory superchips.
Node-Level Performance and Energy Characterization of Flagship Science Applications on SuperMUC-NG Phase 2Salvatore Cielo, Elmira Birang, Alexander Pöppl, Sajad Azizi, Plamen Dobrev, Margarita Egelhofer, Ivan Pribec, Gerald Mathias2026-06-22下载We present a systematic performance and energy-efficiency characterization of five flagship scientific workloads on SuperMUC-NG phase 2, the 28 PetaFLOPs system at the Leibniz Supercomputing Center (L...
Learning Filters with CertaintyYuval Banoun, Daniel Sadoc Menasche, Ori Rottenstreich2026-06-22下载Hash-based data structures such as Bloom filters are widely used in network systems for tasks including caching, anomaly detection, and machine learning pipelines.

基于 VitePress 构建 · 使用本地搜索查找论文