Skip to content

2026-07-29 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Investigating reservoir computing for branch predictionin pipelined processors using emerging CMOS memristor devicesHarvey Samuel George Johnson, Sendy Phang2026-07-29下载This project aimed to develop a novel reservoir compute (RC) implementation framework targeting high-speed operation and integration with CMOS digital logic.
A Low-Power Sparse Convolution Accelerator with Idle-First-Task-Assignment for Edge VisionJingyue Zhuge, Johannes Partzsch, Christian Mayr2026-07-29下载In recent years, edge-vision monitoring systems for applications such as smart animal husbandry have faced strict tripartite constraints: maintaining input resolution under extremely limited transmiss...
NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request PlacementSookyung Choi, Seungyong Lee, Kangkyu Park, Yunseo Chun, Junseok Lee, Hyeongseok Gwak, Myunghyun Rhee, Euiseok Kim, Donguk Moon, Kwangsik Shin, Guseul Heo, Youngpyo Joo, Hoshik Kim, Jongse Park2026-07-29下载Modern LLMs and their agentic applications are broadening the range of serving workloads, spanning context lengths from a few hundred tokens to hundreds of thousands.
LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM ServingMing-Yen Lee, Hanchen Yang, Faaiq Waqar, Harsono Simka, Tushar Krishna, Muhammed Ahosan Ul Karim, Shimeng Yu2026-07-29下载The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs.
CircuitProver: Agentic Lean 4 Theorem Proving with Reusable Circuit Proof Library for Hardware VerificationZiyi Yang, Wenji Fang, Chen Chen, Zhiyao Xie, Hongce Zhang2026-07-29下载Modern integrated circuits (ICs) are becoming increasingly complex, making functional verification a major bottleneck. The dominant hardware formal verification methodology, model checking, verifies e...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
InferScale: GPU-Native KV Injection for Personalized LLM ServingPeter Li, Prashant Pandey2026-07-29下载Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histories, that is shared across a user's many requests.
Hybrid Workflow Composition for Extreme-Scale Data Processing: A Case Study on the HL-LHC (Extended Version)Alan Malta Rodrigues, Douglas Thain2026-07-29下载High-Throughput Computing (HTC) environments tailored for high-concurrency resource efficiency require sophisticated orchestration to manage petabyte-scale data across heterogeneous resources.
Mind the Gap: The Disconnect Between Synthetic and Natural Edge Weights in Parallel Single-Source Shortest PathMarco D'Antonio, Thai Son Mai, Hans Vandierendonck2026-07-29下载Scientific research works often evaluate Parallel Single-Source Shortest Path (SSSP) algorithms using synthetic, uniformly distributed edge weights.
Nix to the Rescue for a Reproducible HPC-AI Software StackWenke Du, Jean-Marc Gratien, Raphael Gayno, Bruno Raffin2026-07-29下载Reproducibility in HPC remains difficult under the constraints of production supercomputers: no root access, limited internet, and software stacks that increasingly span C/C++, Fortran, Python, MPI, a...
Unified Shared Memory in OpenMP: Implementation, Programmability, and Performance on Intel AcceleratorsHarald Servat, François Dugast, Alejandro Duran, Abhinav Gaba, Rakesh Krishnaiyer2026-07-29下载OpenMP 5.0 introduced the Unified Shared Memory (USM) feature through the requires directive. The feature simplifies the adoption of the OpenMP programming model by providing a unique and common addre...
ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless PlatformXiaoxiao Jiang, Suyi Li, Sheng Yao, Tianyu Feng, Lingyun Yang, Dapeng Nie, Haoran Yang, Wei Wang2026-07-29下载Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently.
Safety-Gated Autoscaling: A Multi-Layered Defense Architecture for Kubernetes Vertical Resource OptimizationAzra Karakaya, Erva Şengül, Ahmet Kaplan2026-07-29下载Kubernetes is the standard platform for orchestrating containerized applications, yet resource management remains difficult. To stay safe, engineers over-provision CPU and memory, leaving reserved but...
DualDecoder: Accelerate Long Context LLM Inference by Predictive PrefetchZuning Liang, Zhiyi Yao, Qi Chen, Yuedong Xu, Hao Dai, Zhiqiang Ding, Tongkai Yang, Jinlong Hou, Yuan Cheng2026-07-29下载Long-context inference is becoming a fundamental capability for modern LLM serving, especially driven by emerging agentic applications. Yet it faces a severe memory wall that the KV cache scales propo...
StrataCL: Fabric-Native Communication Library for Production SupernodesTiancheng Hu, Jin Qin, Yuzheng Wang, Ke Liu, TangShengsheng Li, Sheng Wang, Zhongzhe Hu, Tianlun Hu, Wei Wang, Lijun Li, Jingbin Zhou, Xiaoming Bao, Hongwei Sun, Jieru Zhao, Huimin Cui, Tao Xie, Chenxi Wang2026-07-29下载Modern distributed AI workloads run across hundreds of accelerators, making communication a major bottleneck. Existing communication libraries remain largely buffer-centric because user and communicat...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
O-RAN: Analysis of Latency-critical Interfaces and Overview of Time Sensitive Networking SolutionsEsteban Municio, Gines Garcia-Aviles, Andres Garcia-Saavedra, Xavier Costa-Pérez2026-07-29下载5G and B5G/6G foundations heavily rely on virtualization technologies, and virtualized Radio Access Networks (vRANs) are one of their major keystones.
Assurance-Scoped Reliability for Agentic Networks: Capturing the State That MattersBilgehan Erman, Andrea Francini, Nikos Papadis2026-07-29下载Agentic networks transform accepted intents into operational services through autonomous reasoning, adaptive planning, tool use, and cross-domain coordination, but these capabilities introduce failure...
The Price of Meaning: Quantifying Semantic Communication Overheads in PracticeXinyi Lin, Peizheng Li, Adnan Aijaz2026-07-29下载Semantic communication (SemCom) promises to reduce transmitted payloads by conveying task-relevant meaning instead of raw bits. However, practical SemCom also incurs semantic metadata, control signali...
Active Movable-Element RIS Assisted Vehicular Semantic Communications: Modeling and OptimizationMaoxin Ji, Qiong Wu, Jingbo Zhang, Pingyi Fan, Kezhi Wang, Wen Chen, Guoqiang Mao, Khaled B. Letaief2026-07-29下载Severe signal blockage and fast-varying channels in vehicular environments pose critical challenges to reliable semantic communication. To address these, this paper proposes a novel Row-Movable Active...
Harnessing Large Language Models for Intelligent Resource Allocation in the Internet of EverythingHaijun Zhang, Zhuojun Duan, Zijun Wu, Xu Ma, Yuzheng Ren2026-07-29下载The rapid development of the Internet of Everything (IoE) is accelerating the adoption of intelligent applications. However, the massive number of connected devices generates diverse and heterogeneous...
Graphene-based Hemispherical Transmitarray Antenna for Wide-Angle Beam Steering and Ultrafast Moving Target TrackingSomayeh Komeylian, Christopher Paolini2026-07-29下载This work expands the application of dynamically tunable graphene to implement a hemispherical transmitarray antenna tailored for wide-angle electronic beam steering and ultrafast moving-target tracki...
Can We Trust AI in 6G? Verifiable and Auditable AI-Driven Trustworthy Wireless NetworksGenze Jiang, Yizhou Huang, Kezhi Wang2026-07-29下载Mobile network operators are increasingly exploring the use of artificial intelligence (AI) to automate complex network tasks, such as cell selection and mobility management.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Deductive Verification for Earliest Deadline First Scheduler ImplementationsDaniel Kuhse, Junjie Shi, Jan Duy Thien Pham, Kay Heider, Marcus Völker, Kuan-Hsun Chen, Jian-Jia Chen2026-07-29下载Real-Time Operating Systems (RTOSes) rely on scheduler implementations to provide predictable task execution. For safety-critical systems, it is therefore not sufficient to reason only about the abstr...

cs.PF - Performance ​

标题作者发布日期PDF摘要
A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM InferenceJing Ding, Yash Nishant, Chandrish Ambati, Jyothsna Kamati, Trung Diep2026-07-29下载LLM inference at scale faces a memory wall. The KV cache demands tens of terabytes at hundreds of gigabytes per second, yet no current memory tier delivers both at once.
Unified Shared Memory in OpenMP: Implementation, Programmability, and Performance on Intel AcceleratorsHarald Servat, François Dugast, Alejandro Duran, Abhinav Gaba, Rakesh Krishnaiyer2026-07-29下载OpenMP 5.0 introduced the Unified Shared Memory (USM) feature through the requires directive. The feature simplifies the adoption of the OpenMP programming model by providing a unique and common addre...
Global Pass Barriers Without Per-Resource RHI Tracking: A Cross-Vendor Study with BladeDzmitry Malyshau2026-07-29下载Explicit graphics APIs expose memory dependencies, per-resource accesses, and image layouts. wgpu reconstructs and validates this state; Blade keeps Vulkan images in GENERAL, tracks no per-resource st...
An Efficient Algorithm for Computing Mountain Prominence in Almost Linear TimeGeorge Alex Dumitrescu, Paul Flavian Diac2026-07-29下载Prominence is one of the most important measurements in topography and mountaineering. This paper describes an efficient, almost linear time algorithm for computing mountain prominence for all peaks o...

基于 VitePress 构建 · 使用本地搜索查找论文