2026-08-13
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement Optimization | Ruogu Chen, Jie Han | 2026-08-13 | 下载 | Macro placement significantly affects a chip's post-route performance, power, and area (PPA). Most placement methods optimize half-perimeter wirelength (HPWL) as the primary objective. |
| YAVIN: A Unified Architecture for Secure Edge Processing in Memory | Shouzhi Fang, William C. Tegge, Md Omar Faruque, Peipei Zhou, Endadul Hoque, Alex K. Jones | 2026-08-13 | 下载 | Secure, private multi-tenant execution spanning processors, memory, and accelerators remains one of the most significant challenges in modern edge computing systems. |
| ROLoad-PMP: Securing Sensitive Operations for Kernels and Bare-Metal Firmware | Wende Tan, Chenyang Li, Yangyu Chen, Yuan Li, Chao Zhang, Jianping Wu | 2026-08-13 | 下载 | A common way for attackers to compromise victim systems is hijacking sensitive operations (e.g., control-flow transfers) with attacker-controlled inputs. |
| Potential Applications of HBF in LLM Serving Systems | Yihan Yin, Yinlun Zhao, Zhixin Yun, Guanying Wu, Feng Zhu, Kai Tao, Shu Li, Fei Huang, Zhe Zhang, Shuangchen Li, Hongzhong Zheng | 2026-08-13 | 下载 | LLM serving is increasingly constrained by memory capacity as model weights, KV caches, and the number of served model variants continue to grow. |
| Why Do Prefetchers Fail? Let Agents Answer | Xiangfeng Sun, Ceyu Xu, Ningzhi Ai, Zeyu Zhu, Yiyang Yuan, Yuan Xie | 2026-08-13 | 下载 | Hardware prefetchers are crucial to processor performance, yet their design remains labor-intensive and expert-driven. Architects inspect execution and memory-access traces, identify patterns, transla... |
| Dryas: A Reprogrammable Engine for High-Speed Interconnect Tracing and Analysis | Manuel Bröchin, Tom Kuchler, Michael Giardino, David Cock, Timothy Roscoe | 2026-08-13 | 下载 | The proliferation of heterogeneous components in modern computing systems has been accompanied by new higher bandwidth and lower latency interconnects. |
| SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization | Fangzhou Liu, Peiyi Han, Jiawei Liu, Yuan Pu, Zhuolun He, Rongliang Fu, Tsung-Yi Ho, Bei Yu | 2026-08-13 | 下载 | Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization commands, making synthesis tuning both high-dimensional and expens... |
| A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family | Rishi Shah, Rishav Shrestha | 2026-08-13 | 下载 | Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it ... |
| Spec-Driven Hardware Evolution via Executable Contract Refinement and Proof-Guided RTL Update | Shibo Zhao, Yang Zhang, Mengxia Tao, Baoqi Zhang, Kezhi Li, Qiang Xu, Binwu Zhu, Hao Yan, Min Li | 2026-08-13 | 下载 | Hardware development is inherently evolutionary: major revisions typically begin by changing intended behavior and then updating a previously validated implementation, rather than regenerating RTL fro... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Balancing Workload Performance and Slurm Stress: Four Nextflow Deployment Strategies | Nil Tianchen Mu, William Dizon, Glen Otero, Torey Battelle | 2026-08-13 | 下载 | Wide Nextflow fan-outs on shared Slurm clusters can submit tens of thousands of short tasks. Deployment choices - individual jobs, arrays, or nested schedulers within allocations - affect both workflo... |
| A Barrier-Free Synchronization Algorithm for Multi-Engine AI Accelerators | Chungha Sung, Nikil V. Shyamsunder, Hanliang Zhang, Daniel Kroening, Joonwon Choi | 2026-08-13 | 下载 | Multi-engine AI accelerators such as AWS Trainium comprise specialized compute engines that execute in parallel, and the compiler must synchronize the data dependencies between them. |
| Adaptive Snapshots Require Visible Reads | Niv Sulimany, Tomer Cory, Erez Petrank | 2026-08-13 | 下载 | Snapshots are widely used to record the state of a running execution. Snapshots have been extensively studied in the literature, with the goal of improving performance and extending functionality. |
| OpScale: Operator-level Provisioning and Autoscaling for LLM Serving | Xingqi Cui, Chieh-Jan Mike Liang, Ziang Tang, Jiarong Xing, Haoran Qiu | 2026-08-13 | 下载 | Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LLMs). |
| Fast Tendermint: Speeding Up a Foundational Consensus Protocol | Preston Vander Vos, Daniel Cason | 2026-08-13 | 下载 | Tendermint is among the most widely studied and deployed Byzantine fault-tolerant (BFT) consensus protocols, owing in part to its native leader-rotation mechanism that subsumes complex view changes. |
| Triangle-Free Coloring in LOCAL via Resilient Lovász Local Lemma | Peter Davies-Peck, Xusheng Zhang | 2026-08-13 | 下载 | The Lovász Local Lemma (LLL) is a probabilistic tool that has been shown to be of central importance in the study of distributed algorithms. For example, the constructive LLL is known to be complete f... |
| vToken: Token-Level Virtualization for Reclaimable KV Caches | Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li | 2026-08-13 | 下载 | Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragment... |
| LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service | Zhengzhe Xiang, Yinlin Chen, Fuli Ying, Binbin Zhou, Hailiang Zhao, Schahram Dustdar | 2026-08-13 | 下载 | As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models without sacrificing reliability has become a central concern... |
| Validation-Centric AI-Assisted GPU Porting of a 250,000+ Line Legacy Weather Simulation Code | Tetsuya Hoshino, Masaya Kato, Kazuhisa Tsuboki, Daichi Mukunoki, Takahiro Katagiri, Toshihiro Hanawa | 2026-08-13 | 下载 | Recent advances in large language models have made CLI-based AI agents a practical tool for accelerating GPU porting of large legacy scientific applications. |
| Meshlib: In-Process Policy Enforcement for Sidecar-less Service Meshes | Habib Mostafaei, Tom van Liempd | 2026-08-13 | 下载 | Service meshes facilitate service-to-service communication and enforce security policies in microservice architectures. However, they often depend on per-pod sidecar proxies, which introduce significa... |
| TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes | Jie Li, Chenxin Jia, Jinliang Shen, Cunzhuang Liu, Ruiyi Ding, Jianwen Xian, Kang He, Chengru Song | 2026-08-13 | 下载 | In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is lin... |
| Efficient Randomized LL/SC that Preserves History Independence | Dante Bencivenga, Homa Habashi, Philipp Woelfel | 2026-08-13 | 下载 | We study the fundamental problem of implementing linearizable LL/SC objects with constant expected step complexity in a system of processes, using bounded base objects commonly available in ha... |
| InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers | Nicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos | 2026-08-13 | 下载 | The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption... |
| A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural Settings | Hei Ting, Chan, Chenwei Wu, Xueshen Liu, Zesen Zhao, Boyuan Zheng, Luis Filipe Nakayama, Michael G. Morley, Liyue Shen, Jiasi Chen, Z. Morley Mao | 2026-08-13 | 下载 | Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limite... |
| A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family | Rishi Shah, Rishav Shrestha | 2026-08-13 | 下载 | Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it ... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Weird Machines in Transport Layer Security | Michael Collins, Jada Cumberland, Brianne Dunn, Ross Gore, Samuel Jackson, Sachin Shetty, Jonathan Takeshita | 2026-08-13 | 下载 | Weird machines are latent computational capabilities that emerge from the composition of architectural components. Prior work has studied this phenomenon extensively in software systems, including x86... |
| TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies | Xiaokang Qu, Jianliang Ma, Zao Fan, Tianshu Chu, Tianlong Fan, Linyuan Lü | 2026-08-13 | 下载 | Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundary devices, inter-zone paths, and access-control policies. |
| Age of Incorrect Information for Pull-Based State Estimation of General Markov Sources | Marco Zanni, Mohamad Assaad, Touraj Soleymani | 2026-08-13 | 下载 | We study pull-based remote state estimation of an arbitrary, multi-state Markov source while accounting for both freshness and correctness attributes of information. |
| Energy-Aware Compression-Computation Co-Adaptation for Latency Minimization in Multi-User Semantic Communication | Loc X. Nguyen, Yumin Park, Avi Deb Raha, Huy Q. Le, Zhu Han, Eui-Nam Huh, Choong Seon Hong | 2026-08-13 | 下载 | Deep joint source-channel coding-enabled (DeepJSCC) semantic communication (SemCom) has excelled at delivering high perceptual quality at low channel-bandwidth ratios, which positions it as a pillar f... |
| Radio-Optical Confluence in Intelligent Edge Networks | Akshita Gupta, Devika Dass, Agastya Raj, Carlos Natalino, Marco Ruffini, Paolo Monti, Daniel Kilper | 2026-08-13 | 下载 | Challenges associated with densification of radio access networks are motivating exploration of more efficient and scalable architectures. We examine recent progress in one direction that involves mov... |
| Pareto-Aware Hierarchical Reinforcement Learning for Online Resource Allocation in RIS-assisted Large-Scale IoT Systems | Wenhan Xu, Jiashuo Jiang, Danny H. K. Tsang | 2026-08-13 | 下载 | With the rapid evolution of 5G and emerging 6G networks, reconfigurable intelligent surfaces (RIS) have become a critical technology for enhancing wireless communication scenarios. |
| InterSAGE: The Secure and Verifiable Interoperability Protocol for An Internet of Agents | Zhenhua Zou, Sheng Guo, Qiuyang Zhan, Lepeng Zhao, Shuo Li, Zhuotao Liu | 2026-08-13 | 下载 | The emerging Internet of Agents enables LLM-powered agents to discover peers, invoke tools, and delegate tasks across organizational boundaries. |
| Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and Performance | Chukwunonso Henry Nwokoye, Blessing Oluchi Iloka, Chikwue V. Umeugoji, Christopher Anene Egemba, Nnenna D. Duroha | 2026-08-13 | 下载 | The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches. |
| Digital Twin Satellite Networks: A Paradigm for Intelligent, Efficient, and Resilient Operations | Mustafa Alhassan, Peng Hu | 2026-08-13 | 下载 | Satellite mega-constellations in Low Earth Orbit (LEO) are becoming an important part of next-generation non-terrestrial networks, but their operation remains challenging because of fast network topol... |
| ASAP: Reimagining the Data Lifecycle using Application Semantic-Aware Processing | Milind Srivastava, Zeying Zhu, Yajie Zhou, Yancheng Yuan, Fenghao Dong, Peilin Xin, Zaoxing Liu, Vyas Sekar | 2026-08-13 | 下载 | Across many domains (e.g., observability, networking, security), data processing pipelines face what we refer to as the CSP problem: achieving low Cost at large Scale, while maintaining high Performan... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A Bounded Reclaim Actuator for PSI-Guided Compressed Memory: A Controlled Ablation | Abhiyan Dhakal, Sanjog Sigdel | 2026-08-13 | 下载 | When the aggregate working set of active processes exceeds physical RAM capacity, the machine experiences memory pressure. Applications may therefore slow down before the kernel kills a process. |
| vToken: Token-Level Virtualization for Reclaimable KV Caches | Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li | 2026-08-13 | 下载 | Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragment... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Performance Reporting of Mathematical Library Installations with LAAB - An Overview | Aravind Sankaran, Paolo Bientinesi | 2026-08-13 | 下载 | We present the Linear Algebra Aware Benchmarks (LAAB) framework for systematically assessing and reporting the performance of mathematical library installations on HPC systems. |
| Performance Evaluation of an Adaptive Quadrature and a Double Exponential Formula Using Arbitrary-Precision Floating-Point Arithmetic | Tomonori Kouya | 2026-08-13 | 下载 | Using arbitrary-precision arithmetic provided by the GNU Multiple Precision Floating-Point Reliable Library, we implement AQE11D---that is, Ninomiya's adaptive 9-point Newton--Cotes rule extended with... |
| ASAP: Reimagining the Data Lifecycle using Application Semantic-Aware Processing | Milind Srivastava, Zeying Zhu, Yajie Zhou, Yancheng Yuan, Fenghao Dong, Peilin Xin, Zaoxing Liu, Vyas Sekar | 2026-08-13 | 下载 | Across many domains (e.g., observability, networking, security), data processing pipelines face what we refer to as the CSP problem: achieving low Cost at large Scale, while maintaining high Performan... |