2026-04-27
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Déjà Vu Packing: Optimizing FPGA Logic Clustering Runtime via Pattern Memoization | Milo Liebster, Amin Mohaghegh, Andrew Boutros | 2026-04-27 | 下载 | Implementing a digital circuit on an FPGA fabric requires clustering technology-mapped netlist primitives into coarser-granularity blocks that can be directly mapped to the physical resources availabl... |
| Salca: A Sparsity-Aware Hardware Accelerator for Efficient Long-Context Attention Decoding | Wang Fan, Wei Cao, Xi Zha, Kedi Ma, MingQian Sun, Jialin Chen, Fengzhe Zhang, Fan Zhang | 2026-04-27 | 下载 | Long contexts improve capabilities of large language models but pose serious hardware challenges: compute and memory footprints grow linearly with sequence length. |
| Compilation and Execution of an Embeddable YOLO-NAS on the VTA | Anthony Faure-Gignoux, Kevin Delmas, Adrien Gauffriau, Claire Pagetti | 2026-04-27 | 下载 | Deploying complex Convolutional Neural Networks (CNNs) on FPGA-based accelerators is a promising way forward for safety-critical domains such as aeronautics. |
| RowHammer Vulnerability Counter (RVC): Redefining RowHammer Detection with Victim-Centric Tracking | Lavi Jain, Venkata Kalyan Tavva | 2026-04-27 | 下载 | The Rowhammer vulnerability poses an increasing challenge with newer generations of DRAM and aggressive technology scaling. Existing mitigation techniques, such as Graphene, Twice, and Hydra, primaril... |
| Opto-Atomic Spatio-Temporal Holographic Correlators for High-Speed 3D CNNs | Xi Shen, Bowen Qi, Tabassom Hamidfar, Selim M. Shahriar | 2026-04-27 | 下载 | Three-dimensional convolutional neural networks (3D CNNs) have demonstrated remarkable performance in video recognition tasks by processing both spatial and temporal features. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Spark Policy Toolkit: Semantic Contracts and Scalable Execution for Policy Learning in Spark | Zeyu Bai | 2026-04-27 | 下载 | Custom policy-learning pipelines in Spark fail for two coupled systems reasons: rowwise Python execution makes inference impractical, and driver-side candidate materialization makes split search fragi... |
| Internet of Everything in the 6G Era: Paradigms, Enablers, Potentials and Future Directions | Driss Choukri, Essaid Sabir, Elmahdi Driouh, Abdelkrim Haqiq | 2026-04-27 | 下载 | The Internet of Everything (IoE) represents an evolution of the Internet of Things (IoT) by integrating people, data, processes, and things into a unified intelligent ecosystem. |
| A Tree-Based Repository Blockchain Framework for Shared Governance in Collaborative Fork Ecosystems | Razwan Ahmed Tanvir, Greg Speegle | 2026-04-27 | 下载 | Collaborative blockchain ecosystems allow diverse groups to cooperate on tasks while providing properties such as decentralization and transaction security. |
| PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference | Ishan Patel, Ishan Joshi | 2026-04-27 | 下载 | We present PolyKV, a system in which multiple concurrent inference agents share a single, asymmetrically compressed KV cache pool. Rather than allocating a separate KV cache per agent -- the standard ... |
| Network Impact of Post-Quantum Certificate Chain sizes on Time to First Byte in TLS Deployments | Matthew Chou, Phuong Cao | 2026-04-27 | 下载 | Post-Quantum Cryptography (PQC) is a rapidly growing deployment challenge as cryptographically relevant quantum computers (CRQC) continue to advance, leaving traditional cryptographic algorithms used ... |
| SpotVista: Availability-Aware Recommendation System for Reliable and Cost-Efficient Multi-Node Spot Instances | Taeyoon Kim, Kyumin Kim, Kyunghwan Kim, Hayoung Kim, Seungwoo Jeong, Moohyun Song, Kyungyong Lee | 2026-04-27 | 下载 | Cloud vendors offer discounted spot instances to maximize surplus resource utilization, but these instances are subject to the risk of sudden interruption. |
| A Survey on Split Learning for LLM Fine-Tuning: Models, Systems, and Privacy Optimizations | Zihan Liu, Yizhen Wang, Rui Wang, Xiu Tang, Sai Wu | 2026-04-27 | 下载 | Fine-tuning unlocks large language models (LLMs) for specialized applications, but its high computational cost often puts it out of reach for resource-constrained organizations. |
| Incisor: Ex Ante Cloud Instance Selection for HPC Jobs | Michael A. Laurenzano, Shihan Cheng, David A. B. Hyde | 2026-04-27 | 下载 | We present Incisor, a cloud HPC job submission system for the ex ante instance selection problem: choosing suitable hardware in the challenging but common setting where only the executable, inputs, an... |
| Exact, Efficient, and Reliable Multi-Objective and Multi-Constrained IoT Workflow Scheduling in Edge-Hub-Cloud Cyber-Physical Systems | Andreas Kouloumpris, Georgios L. Stavrinides, Maria K. Michael, Theocharis Theocharides | 2026-04-27 | 下载 | Emerging IoT-enabled cyber-physical applications demand low-latency, energy-efficient, and reliable execution across resource-constrained edge devices with heterogeneous multicore processors and diver... |
| ITAS: A Multi-Agent Architecture for LLM-Based Intelligent Tutoring | Iizalaarab Elhaimeur, Nikos Chrisochoides | 2026-04-27 | 下载 | Large language model tutors are easy to build in a notebook and hard to run in a real course. We describe ITAS (Intelligent Teaching Assistant System), a multi-agent tutoring system that a graduate qu... |
| Latency and Cost of Multi-Agent Intelligent Tutoring at Scale | Iizalaarab Elhaimeur, Nikos Chrisochoides | 2026-04-27 | 下载 | Multi-agent LLM tutoring systems improve response quality through agent specialization, but each student query triggers several concurrent API calls whose latencies compound through a parallel-phase m... |
| Unfolding an Atomistic World: Atomistic Simulation of Reactor Pressure Vessel Steel Across Year-and-Meter Scales | Haozhi Han, Ruge Zhang, Haoquan Chen, Yifeng Chen, Haipeng Jia, Liang Yuan, Yunquan Zhang, Ting Cao, Yunxin Liu, Ya-Qin Zhang, Kun Li | 2026-04-27 | 下载 | Lifetime prediction of reactor pressure vessel (RPV) steel requires bridging atomistic degradation mechanisms with service-scale spatial and temporal regimes, from Angstroms and picoseconds to meters ... |
| TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training | Man Liu, Xingchen Liu, Xingjian Tian, Bing Lu, Shengkay Lyu, Shengquan Yin, Wenjing Huang, Zheng Wei, Hairui Zhao, Guangming Tan, Dingwen Tao | 2026-04-27 | 下载 | Handling communication overhead in large-scale tensor-parallel training remains a critical challenge due to the dense, near-zero distributions of intermediate tensors, which exacerbate errors under fr... |
| FreeScale: Distributed Training for Sequence Recommendation Models with Minimal Scaling Cost | Chenhao Feng, Haoli Zhang, Shakhzod Ali-Zade, Yanli Zhao, Liang Luo, Jennifer Cao, Lisen Deng, Siqiao Chen, Chenyu Zhao, Tristan Rice, Daniel Johnson, Min Si, Tiantu Xu, Yi Zhang, Siqi Yan, Chuanhao Zhuge, Min Ni, Bi Xue, Qunshu Zhang, Shen Li | 2026-04-27 | 下载 | Modern industrial Deep Learning Recommendation Models typically extract user preferences through the analysis of sequential interaction histories, subsequently generating predictions based on these de... |
| KubePACS: Kubernetes Cluster Using Performant, Highly Available, and Cost Efficient Spot Instances | Taeyoon Kim, Kyumin Kim, Enrique Molina-Giménez, Pedro García-López, Kyungyong Lee | 2026-04-27 | 下载 | Cloud users aim to minimize cost while maximizing performance by selecting the most suitable instance types for their workloads. To reduce expenses, spot instances have been widely adopted due to thei... |
| FlashOverlap: Minimizing Tail Latency in Communication Overlap for Distributed LLM Training | Rezaul Karim, Austin Wen, Wang Zongzuo, Weiwei Zhang, Yang Liu, Walid Ahmed | 2026-04-27 | 下载 | The rapid growth in the size of large language models has necessitated the partitioning of computational workloads across accelerators such as GPUs, TPUs, and NPUs. |
| SDSL-Solver: Scalable Distributed Sparse Linear Solvers for Large-Scale Interior Point Methods | Shaofeng Yang, Yunting Wang, Yingying Cheng, Fan Zhang, Xin He, Guangming Tan | 2026-04-27 | 下载 | The solution of sparse linear systems constitutes the dominant computational bottleneck in interior point methods (IPMs), frequently consuming over 70% of the total solution time. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Internet of Everything in the 6G Era: Paradigms, Enablers, Potentials and Future Directions | Driss Choukri, Essaid Sabir, Elmahdi Driouh, Abdelkrim Haqiq | 2026-04-27 | 下载 | The Internet of Everything (IoE) represents an evolution of the Internet of Things (IoT) by integrating people, data, processes, and things into a unified intelligent ecosystem. |
| On the Benefits of Traffic "Reprofiling" -- The Multiple Hops Case -- Part II | Jiaming Qiu, Roch Guerin | 2026-04-27 | 下载 | Delivering hard delay guarantees over packet networks is increasingly important to applications ranging from automotive systems, avionics, industrial control, etc. |
| Balancing Quantum Memories in Asymmetric Repeaters for High-Fidelity Entanglement Distribution | Karim S. Elsayed, Amr Rizk | 2026-04-27 | 下载 | At the core of the quantum Internet lie quantum repeaters that enable remote end-to-end entanglement generation. Fundamentally, the entanglement generation rate and fidelity of quantum repeaters const... |
| DECOFFEE: Decentralized Reinforcement Learning for Time-critical Workload Offloading and Energy Efficiency across the Computing Continuum | Anastasios Giannopoulos, Sotirios Spantideas, Panagiotis Trakadas | 2026-04-27 | 下载 | The rapid proliferation of latency-sensitive and battery-constrained Internet-of-Things (IoT) applications has intensified the need for intelligent workload placement mechanisms across the Edge-Cloud ... |
| TARMM: Scaling Delay-Critical Edge AI Offloading in 5G O-RAN via Temporal Graph Mobility Management | Peihao Yan, Yun Chen, Jie Lu, Qijun Wang, Huacheng Zeng | 2026-04-27 | 下载 | Emerging delay-critical edge AI applications, such as VR perception and real-time video analytics, impose stringent latency and reliability requirements on 5G networks. |
| Large-scale wireless network management via Open-RAN Tandem Apps: Cell on/off switching use case | Paweł Kryszkiewicz, Łukasz Kułacz, Marcin Pakuła, Marcin Dryjanski, Marcin Hoffmann, Piotr Skrzypczak, Heiko Lehmann, Martin Stahn | 2026-04-27 | 下载 | With growing mobile-network complexity, management and optimization have become increasingly difficult. Centralized algorithms face high control-data overhead and computational load, while distributed... |
| Beam Scheduling for Cross-Layer ISAC: A Deep Reinforcement Learning Approach | Xiyu Wang, Gilberto Berardinelli, Hei Victor Cheng, Petar Popovski, Ramoni Adeogun | 2026-04-27 | 下载 | Resource allocation in integrated sensing and communication (ISAC) systems needs to be optimized to balance the requirements of the communication and sensing modules considering complicated cross-laye... |
| Data-Driven Adaptive Resource Allocation for Reliable Low-Latency Uplink Communications in Rural Cellular 5G Multi-Connectivity | Carlos S. Alvarez-Merino, Alejandro Ramirez-Arroyo, Rasmus Suhr Mogensen, Morten V. Pedersen, Miguel Villanueva-Fernández, Emil J. Khatib, Sergio Fortes, Raquel Barco, Preben E. Mogensen | 2026-04-27 | 下载 | Reliable low-latency communication is a key requirement for mission-critical and mobile autonomous systems, including teleoperation, autonomous navigation, and real-time uplink-dominant telemetry appl... |
| Optimizing power by selective IP card shutdown using transport slicing | Alfonso Sánchez-Macián, Óscar González de Dios, José Alberto Hernández, Liesbeth Roelens, Pablo Armingol Robles, Juan Pedro Fernadez-Palacios, Ramón Casellas, Filippo Cugini | 2026-04-27 | 下载 | The increasing energy demands of upcoming sixth-generation (6G) mobile networks and networks supporting AI applications pose significant challenges for network operators in terms of operational costs ... |
| MatchRDMA: A Segmented and Rate-Matched Long-Haul RDMA Scheme for Geo-distributed LLM Training over OTN | Jun Dai, Xiaorun Wang, Xingde Li, Zheng Yang, Kexiong Fang, Zhiqun Gu, Hongxiang Wang, Yuefeng Ji, Jiawei Zhang | 2026-04-27 | 下载 | We propose MatchRDMA, a proactive, segmented, and rate-matched long-haul RDMA scheme for geo-distributed LLM training over OTN. By coordinating source and destination OTN rates, it improves inter-DC t... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Spark Policy Toolkit: Semantic Contracts and Scalable Execution for Policy Learning in Spark | Zeyu Bai | 2026-04-27 | 下载 | Custom policy-learning pipelines in Spark fail for two coupled systems reasons: rowwise Python execution makes inference impractical, and driver-side candidate materialization makes split search fragi... |
| On the Benefits of Traffic "Reprofiling" -- The Multiple Hops Case -- Part II | Jiaming Qiu, Roch Guerin | 2026-04-27 | 下载 | Delivering hard delay guarantees over packet networks is increasingly important to applications ranging from automotive systems, avionics, industrial control, etc. |