Skip to content

2026-08-13 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
PPAPlace: Differentiable Cross-Stage Objectives for Chip Placement OptimizationRuogu Chen, Jie Han2026-08-13下载Macro placement significantly affects a chip's post-route performance, power, and area (PPA). Most placement methods optimize half-perimeter wirelength (HPWL) as the primary objective.
YAVIN: A Unified Architecture for Secure Edge Processing in MemoryShouzhi Fang, William C. Tegge, Md Omar Faruque, Peipei Zhou, Endadul Hoque, Alex K. Jones2026-08-13下载Secure, private multi-tenant execution spanning processors, memory, and accelerators remains one of the most significant challenges in modern edge computing systems.
ROLoad-PMP: Securing Sensitive Operations for Kernels and Bare-Metal FirmwareWende Tan, Chenyang Li, Yangyu Chen, Yuan Li, Chao Zhang, Jianping Wu2026-08-13下载A common way for attackers to compromise victim systems is hijacking sensitive operations (e.g., control-flow transfers) with attacker-controlled inputs.
Potential Applications of HBF in LLM Serving SystemsYihan Yin, Yinlun Zhao, Zhixin Yun, Guanying Wu, Feng Zhu, Kai Tao, Shu Li, Fei Huang, Zhe Zhang, Shuangchen Li, Hongzhong Zheng2026-08-13下载LLM serving is increasingly constrained by memory capacity as model weights, KV caches, and the number of served model variants continue to grow.
Why Do Prefetchers Fail? Let Agents AnswerXiangfeng Sun, Ceyu Xu, Ningzhi Ai, Zeyu Zhu, Yiyang Yuan, Yuan Xie2026-08-13下载Hardware prefetchers are crucial to processor performance, yet their design remains labor-intensive and expert-driven. Architects inspect execution and memory-access traces, identify patterns, transla...
Dryas: A Reprogrammable Engine for High-Speed Interconnect Tracing and AnalysisManuel Bröchin, Tom Kuchler, Michael Giardino, David Cock, Timothy Roscoe2026-08-13下载The proliferation of heterogeneous components in modern computing systems has been accompanied by new higher bandwidth and lower latency interconnects.
SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis OptimizationFangzhou Liu, Peiyi Han, Jiawei Liu, Yuan Pu, Zhuolun He, Rongliang Fu, Tsung-Yi Ho, Bei Yu2026-08-13下载Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization commands, making synthesis tuning both high-dimensional and expens...
A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence FamilyRishi Shah, Rishav Shrestha2026-08-13下载Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it ...
Spec-Driven Hardware Evolution via Executable Contract Refinement and Proof-Guided RTL UpdateShibo Zhao, Yang Zhang, Mengxia Tao, Baoqi Zhang, Kezhi Li, Qiang Xu, Binwu Zhu, Hao Yan, Min Li2026-08-13下载Hardware development is inherently evolutionary: major revisions typically begin by changing intended behavior and then updating a previously validated implementation, rather than regenerating RTL fro...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Balancing Workload Performance and Slurm Stress: Four Nextflow Deployment StrategiesNil Tianchen Mu, William Dizon, Glen Otero, Torey Battelle2026-08-13下载Wide Nextflow fan-outs on shared Slurm clusters can submit tens of thousands of short tasks. Deployment choices - individual jobs, arrays, or nested schedulers within allocations - affect both workflo...
A Barrier-Free Synchronization Algorithm for Multi-Engine AI AcceleratorsChungha Sung, Nikil V. Shyamsunder, Hanliang Zhang, Daniel Kroening, Joonwon Choi2026-08-13下载Multi-engine AI accelerators such as AWS Trainium comprise specialized compute engines that execute in parallel, and the compiler must synchronize the data dependencies between them.
Adaptive Snapshots Require Visible ReadsNiv Sulimany, Tomer Cory, Erez Petrank2026-08-13下载Snapshots are widely used to record the state of a running execution. Snapshots have been extensively studied in the literature, with the goal of improving performance and extending functionality.
OpScale: Operator-level Provisioning and Autoscaling for LLM ServingXingqi Cui, Chieh-Jan Mike Liang, Ziang Tang, Jiarong Xing, Haoran Qiu2026-08-13下载Achieving cost efficiency while meeting strict user-facing SLOs (e.g., time-to-first-token) remains a fundamental challenge for cloud GPU clusters serving large language models (LLMs).
Fast Tendermint: Speeding Up a Foundational Consensus ProtocolPreston Vander Vos, Daniel Cason2026-08-13下载Tendermint is among the most widely studied and deployed Byzantine fault-tolerant (BFT) consensus protocols, owing in part to its native leader-rotation mechanism that subsumes complex view changes.
Triangle-Free Coloring in LOCAL via Resilient Lovász Local LemmaPeter Davies-Peck, Xusheng Zhang2026-08-13下载The Lovász Local Lemma (LLL) is a probabilistic tool that has been shown to be of central importance in the study of distributed algorithms. For example, the constructive LLL is known to be complete f...
vToken: Token-Level Virtualization for Reclaimable KV CachesYuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li2026-08-13下载Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragment...
LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification ServiceZhengzhe Xiang, Yinlin Chen, Fuli Ying, Binbin Zhou, Hailiang Zhao, Schahram Dustdar2026-08-13下载As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models without sacrificing reliability has become a central concern...
Validation-Centric AI-Assisted GPU Porting of a 250,000+ Line Legacy Weather Simulation CodeTetsuya Hoshino, Masaya Kato, Kazuhisa Tsuboki, Daichi Mukunoki, Takahiro Katagiri, Toshihiro Hanawa2026-08-13下载Recent advances in large language models have made CLI-based AI agents a practical tool for accelerating GPU porting of large legacy scientific applications.
Meshlib: In-Process Policy Enforcement for Sidecar-less Service MeshesHabib Mostafaei, Tom van Liempd2026-08-13下载Service meshes facilitate service-to-service communication and enforce security policies in microservice architectures. However, they often depend on per-pod sidecar proxies, which introduce significa...
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound RegimesJie Li, Chenxin Jia, Jinliang Shen, Cunzhuang Liu, Ruiyi Ding, Jianwen Xian, Kang He, Chengru Song2026-08-13下载In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is lin...
Efficient Randomized LL/SC that Preserves History IndependenceDante Bencivenga, Homa Habashi, Philipp Woelfel2026-08-13下载We study the fundamental problem of implementing mm linearizable LL/SC objects with constant expected step complexity in a system of nn processes, using bounded base objects commonly available in ha...
InFactPlanner: Planning Sustainable Geo-Distributed LLM Data CentersNicoletta Tsiopani, Moysis Symeonides, George Pallis, Marios D. Dikaiakos2026-08-13下载The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure decisions shape energy use, carbon emissions, water consumption...
A Cloud-Edge System for Multimodal Clinical Screening in Resource-Constrained Rural SettingsHei Ting, Chan, Chenwei Wu, Xueshen Liu, Zesen Zhao, Boyuan Zheng, Luis Filipe Nakayama, Michael G. Morley, Liyue Shen, Jiasi Chen, Z. Morley Mao2026-08-13下载Medical AI has demonstrated specialist-level diagnostic accuracy, yet these capabilities remain largely inaccessible in resource-constrained rural settings where bandwidth is scarce, compute is limite...
A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence FamilyRishi Shah, Rishav Shrestha2026-08-13下载Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Weird Machines in Transport Layer SecurityMichael Collins, Jada Cumberland, Brianne Dunn, Ross Gore, Samuel Jackson, Sachin Shetty, Jonathan Takeshita2026-08-13下载Weird machines are latent computational capabilities that emerge from the composition of architectural components. Prior work has studied this phenomenon extensively in software systems, including x86...
TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network TopologiesXiaokang Qu, Jianliang Ma, Zao Fan, Tianshu Chu, Tianlong Fan, Linyuan Lü2026-08-13下载Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundary devices, inter-zone paths, and access-control policies.
Age of Incorrect Information for Pull-Based State Estimation of General Markov SourcesMarco Zanni, Mohamad Assaad, Touraj Soleymani2026-08-13下载We study pull-based remote state estimation of an arbitrary, multi-state Markov source while accounting for both freshness and correctness attributes of information.
Energy-Aware Compression-Computation Co-Adaptation for Latency Minimization in Multi-User Semantic CommunicationLoc X. Nguyen, Yumin Park, Avi Deb Raha, Huy Q. Le, Zhu Han, Eui-Nam Huh, Choong Seon Hong2026-08-13下载Deep joint source-channel coding-enabled (DeepJSCC) semantic communication (SemCom) has excelled at delivering high perceptual quality at low channel-bandwidth ratios, which positions it as a pillar f...
Radio-Optical Confluence in Intelligent Edge NetworksAkshita Gupta, Devika Dass, Agastya Raj, Carlos Natalino, Marco Ruffini, Paolo Monti, Daniel Kilper2026-08-13下载Challenges associated with densification of radio access networks are motivating exploration of more efficient and scalable architectures. We examine recent progress in one direction that involves mov...
Pareto-Aware Hierarchical Reinforcement Learning for Online Resource Allocation in RIS-assisted Large-Scale IoT SystemsWenhan Xu, Jiashuo Jiang, Danny H. K. Tsang2026-08-13下载With the rapid evolution of 5G and emerging 6G networks, reconfigurable intelligent surfaces (RIS) have become a critical technology for enhancing wireless communication scenarios.
InterSAGE: The Secure and Verifiable Interoperability Protocol for An Internet of AgentsZhenhua Zou, Sheng Guo, Qiuyang Zhan, Lepeng Zhao, Shuo Li, Zhuotao Liu2026-08-13下载The emerging Internet of Agents enables LLM-powered agents to discover peers, invoke tools, and delegate tasks across organizational boundaries.
Multi-perspective Imbalance-Conscious 6G Beamforming Optimization and PerformanceChukwunonso Henry Nwokoye, Blessing Oluchi Iloka, Chikwue V. Umeugoji, Christopher Anene Egemba, Nnenna D. Duroha2026-08-13下载The study presents a systematic machine learning (ML) study of 6G-IoT beamforming optimization (6GBO) using supervised and unsupervised approaches.
Digital Twin Satellite Networks: A Paradigm for Intelligent, Efficient, and Resilient OperationsMustafa Alhassan, Peng Hu2026-08-13下载Satellite mega-constellations in Low Earth Orbit (LEO) are becoming an important part of next-generation non-terrestrial networks, but their operation remains challenging because of fast network topol...
ASAP: Reimagining the Data Lifecycle using Application Semantic-Aware ProcessingMilind Srivastava, Zeying Zhu, Yajie Zhou, Yancheng Yuan, Fenghao Dong, Peilin Xin, Zaoxing Liu, Vyas Sekar2026-08-13下载Across many domains (e.g., observability, networking, security), data processing pipelines face what we refer to as the CSP problem: achieving low Cost at large Scale, while maintaining high Performan...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
A Bounded Reclaim Actuator for PSI-Guided Compressed Memory: A Controlled AblationAbhiyan Dhakal, Sanjog Sigdel2026-08-13下载When the aggregate working set of active processes exceeds physical RAM capacity, the machine experiences memory pressure. Applications may therefore slow down before the kernel kills a process.
vToken: Token-Level Virtualization for Reclaimable KV CachesYuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li2026-08-13下载Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention uses fixed-size memory blocks to reduce allocator-level fragment...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Performance Reporting of Mathematical Library Installations with LAAB - An OverviewAravind Sankaran, Paolo Bientinesi2026-08-13下载We present the Linear Algebra Aware Benchmarks (LAAB) framework for systematically assessing and reporting the performance of mathematical library installations on HPC systems.
Performance Evaluation of an Adaptive Quadrature and a Double Exponential Formula Using Arbitrary-Precision Floating-Point ArithmeticTomonori Kouya2026-08-13下载Using arbitrary-precision arithmetic provided by the GNU Multiple Precision Floating-Point Reliable Library, we implement AQE11D---that is, Ninomiya's adaptive 9-point Newton--Cotes rule extended with...
ASAP: Reimagining the Data Lifecycle using Application Semantic-Aware ProcessingMilind Srivastava, Zeying Zhu, Yajie Zhou, Yancheng Yuan, Fenghao Dong, Peilin Xin, Zaoxing Liu, Vyas Sekar2026-08-13下载Across many domains (e.g., observability, networking, security), data processing pipelines face what we refer to as the CSP problem: achieving low Cost at large Scale, while maintaining high Performan...

基于 VitePress 构建 · 使用本地搜索查找论文