2026-09-18
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Presage: Prefetch Search via Agent-Guided Experiments | Matthew Giordano, Parthasarathy Ranganathan, Baris Kasikci, Akanksha Jain | 2026-09-18 | 下载 | Data prefetching is an established technique to mitigate cache miss latency and keep the processor saturated with data. Software exists in a unique position to issue prefetches, having algorithmic kno... |
| UniCASE: A Unified 16-bit Floating-Point Format with Criticality-Aware Selective ECC for Efficient DNN Protection | Amna Hassan, Semeen Rehman | 2026-09-18 | 下载 | Soft errors are an increasing reliability concern for Deep Neural Network execution because they can corrupt parameters, leading to accuracy degradation. |
| Scalable Packet Tracking on FPGAs for Erasure-Coded RDMA over Lossy WANs | Yicheng Qian, Konstantin Taranov, Yevgeny Yankilevich, Assaf Shacham, Mahmoud Elhaddad, Abdul Kabbani, Miriam Leeser, Nadeen Gebara | 2026-09-18 | 下载 | Modern AI workloads increasingly rely on scale across architectures that interconnect multiple datacenters to form a single "AI factory", overcoming the power and cooling constraints of individual sit... |
| Integrating Approximate Logic Synthesis into Approximate High-Level Synthesis | Jian Shi, Ruicheng Dai, Chang Meng, Yue Yang, Weikang Qian | 2026-09-18 | 下载 | Approximate high-level synthesis (HLS) and approximate logic synthesis (ALS) are two techniques for generating approximate circuits. They operate at different granularities. |
| Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case Study | Erwei Wang, Ephrem Wu, Victor J. B. Jung, Jiajie Li, Andre Rosti, Joseph Melber, Samuel Bayliss | 2026-09-18 | 下载 | Spatial NPUs such as AMD XDNA place compute tiles beside small local memories and leave data movement between them to software. Mapping a multi-stage workload onto such a device is largely a question ... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Extreme-Scale Ising Machines with Cluster Mean-Field Theory | Xiuqi Zhang, Shuvro Chowdhury, Shaila Niazi, Christian Z. Pratt, Navid Anjum Aadit, Kerem Y. Camsari | 2026-09-18 | 下载 | Scaling analog and digital Ising machines to larger problems requires overcoming finite device capacity and the cost of communication between devices. |
| Fairly Compensated Distributed Information Retrieval and Augmentation for AI Agents | Yixiang Yao, Pasha Barahimi, Srivatsan Ravi | 2026-09-18 | 下载 | The increasing reliance of autonomous AI agents on external and distributed knowledge sources introduces a fundamental challenge for decentralized information marketplaces: retrieval agents must evalu... |
| Trust-Aware Output Management for Physical Neural Network in Cloud-Continuum Systems | Maliheh Hariri, Stefan Fischer | 2026-09-18 | 下载 | Physical Neural Networks (PNNs) introduce new opportunities for cloud continuum computing, but their outputs may be affected by noise, drift, delay, and incomplete reliability information. |
| Distributed Balanced Butterfly Counting in Signed Bipartite Graphs | Kiran Mekala, Apurba Das, Suman Banerjee | 2026-09-18 | 下载 | The balanced butterfly is a fundamental primitive for analyzing signed bipartite graphs and provides a basis for studying higher-order structural properties, such as clustering coefficients and commun... |
| Verifiable Computation with Trusted Execution Environments and On-Chain Digital Rights Tokens | Bingle Stegmann Kruger, Co-Pierre Georg | 2026-09-18 | 下载 | We present an architecture that enables data owners to combine private data into data pools using Trusted Execution Environments (TEEs) and manage these pools by issuing narrowly scoped computational ... |
| PoVD: Efficient Consensus Protocol based on Verifiable Delay Function | Rui Jiang, Xintong Ling, Bin Cao, Jiaheng Wang, Xiqi Gao, Zhi Ding | 2026-09-18 | 下载 | Consensus protocols ensure the robustness and scalability of blockchains and decentralized applications built on them. However, existing consensus mechanisms often impose high computational cost or re... |
| HyperParallel-FSDP: Topology-Aware Fully Sharded Training with Layout-Driven Muon on Ascend SuperPods | Mo Sun, Yifan Yao, Yanwei Liu, Luobin Liu, Zhenzhang Yang, Kaisheng Wang, Xiangyu Meng, Chen Li, Xizheng Pang, Huilan Li, Xinglei Xu, Yushi Cui, Xinyao Lin, Kaiqi Chen, Jie Zhang, Zeke Wang, Teng Su | 2026-09-18 | 下载 | Declarative SPMD programming uses tensor sharding descriptions to drive distributed execution, separating parallelization from model code. However, the evaluated PyTorch DTensor stack dispatches every... |
| Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap | Ziyu Huang, Yangjie Zhou, Chenhao Zhu, Zihan Liu, Jinyu Liu, Shulai Zhang, Xingxun Tang, Hongzhe Yan, Xinhao Luo, Minyi Guo, Xiu Lin, Yinghao Yu, Guodong Yang, Liping Zhang, Shixuan Sun, Jingwen Leng | 2026-09-18 | 下载 | Mixture-of-Experts (MoE) inference under expert parallelism (EP) turns each MoE layer into a distributed computation with costly dispatch and combine communication. |
| TrustBOM: A Scalable Architecture for Confidentiality-Preserving SBOMs Across Organizations | Van Thang Nguyen, Frederic Rupprecht, Tom Lawrence, Lucca Di Benedetto, Sören Schubert, Amor Rezgui, Sebastian Werner, Maria C. Borges, Stefan Tai | 2026-09-18 | 下载 | Software Bills of Materials (SBOMs) have emerged as a key mechanism for software supply chain governance in enterprise architectures. However, their adoption across organizations remains limited due t... |
| TokaGLINT: A Scalable GPU-Tailored Implicit Solver for Full 3D Tokamak Electromagnetic Simulations | Zifan Yang, Haoyuan Zhang, Jialin Li, Wu Yuan, Xiazhen Liu, Jian Zhang, Jianyuan Xiao, Shan Liang | 2026-09-18 | 下载 | We introduce TokaGLINT, a GPU-accelerated implicit solver for electromagnetic field computations in full 3D tokamak simulations, aimed at efficient large-scale parallel GPU computing. |
| Brain API: An Intent-Aware Control Plane for Policy-Governed Agentic Systems | Alexander Chernov | 2026-09-18 | 下载 | Contemporary cloud and distributed systems expose control through resource-centric abstractions: services, deployments, network flows, execution graphs. |
| Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale | Hao Fu, Jichao Sun, Baiting Zhu, Qiaoling Liu, Yan Shi, Cheng Lu, Liu Liu, Yubo Wang, Xin Yao, Xiangyu Niu, Xu Dong, Wenhan Lyu, Chiyao Shen, Yinjie Huang, Minglei Chen, Shuai Ding, Li Fan, Xiao Kong | 2026-09-18 | 下载 | Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user ... |
| Adapting the Actor Model of Concurrency for High-Frequency Trading: Synchronous Message Delivery (fast_send) and a Tick-to-Book Latency Study | Vincent Maciejewski | 2026-09-18 | 下载 | The actor model - state isolation, data-race freedom, deadlock resistance, and sequential single-message reasoning - has long been dismissed as unsuitable for high-frequency trading (HFT): actors seem... |
| AI-Driven Scientific Computing Workflows: A Systems Review of Orchestration, Execution, Reproducibility and Provenance | Jamie J. Alnasir | 2026-09-18 | 下载 | Artificial intelligence (AI) is increasingly embedded within scientific computing workflows that combine simulation, data processing, optimisation, visualisation and experimental or observational comp... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| User-Level Handover Decision Making Based on Machine Learning Approaches | João Lima, Alvaro Medeiros, Eduardo Aguiar, Vicente Angelo de Sousa Junior, Tarciana Guerra | 2026-09-18 | 下载 | This letter covers a broad comparison of methods for classification and regression applications for a user-level handover decision making in scenarios with adverse propagation conditions involving bui... |
| Tick-Tock on the Open Fronthaul: Securing Synchronization in O-RAN | Yiwei Zhang, Enrico Pisanti, Imtiaz Karim, Subangkar Karmaker Shanto, Elisa Bertino | 2026-09-18 | 下载 | The Precision Time Protocol (PTP) provides the time and phase synchronization required by disaggregated Open Radio Access Networks (O-RAN). Yet, in current open fronthaul deployments, PTP traffic lack... |
| Semantics Delivery Network: Rethinking Web Retrieval Infrastructure for LLM Agents | Peichun Hua, Yunming Xiao | 2026-09-18 | 下载 | Large language models (LLMs) increasingly rely on external sources when answering questions that require proprietary information or up-to-date live web content, through both traditional single-shot re... |
| Provisional Reachability: Containing Agents by Making Every Crossing Revocable | Yoshiaki Takashita | 2026-09-18 | 下载 | A companion paper found that what a defender must block over time has units: bits per period [Takashita, 2026a]. This paper sets it. Hold every crossing in escrow for one period, audit each held item ... |
| A Multi-Cloud View of Internet Background Radiation | Nils Kempen, Ricky K. P. Mok, Bernhard Degen, Syed Mujtaba Jafri, Ralph Holz | 2026-09-18 | 下载 | As services are increasingly centralized in public clouds, understanding the nature of Internet Background Radiation (IBR) hitting these particular environments is an important part of understanding t... |
| Secure RIS-Aided Multicasting: Globally Optimal Beam Management and Discrete-Phase RIS Configuration | Luis F. Abanto-Leon, Setareh Maghsudi | 2026-09-18 | 下载 | Reconfigurable intelligent surfaces (RISs) are poised to revolutionize wireless multicasting by enabling extended coverage and reliable operation in obstructed environments. |
| Pattern-Aware Virtual Network Embedding Optimization for Cloud Data Centers | Binquan Guo, Zhou Zhang, Junfeng Zhai, Zheng Zhang, Marie Siew, Zehui Xiong | 2026-09-18 | 下载 | The network virtualization (NV) technology has enabled the sharing of multiple resources among virtual networks (VNs) in cloud data centers. One of the key challenges is to allocate resources in real-... |
| Locating and Enumerating Anycast: a Comparison of Two Approaches | Remi Hendriks, Tim Betzer, Ben Du, Raffaele Sommese, Mattijs Jonker, Roland van Rijswijk-Deij | 2026-09-18 | 下载 | Anycast allows for providing services from multiple, geographically distant Points of Presence (PoPs), using a single IP address, to, e.g., improve resilience. |
| X-SPUR: Explainable Surprisal-Based Protocol-Aware Unsupervised Reasoning for Automotive Ethernet Intrusion Detection | Jisoo Kim, Seonghoon Jeong | 2026-09-18 | 下载 | Automotive Ethernet carries heterogeneous multi-protocol traffic in modern in-vehicle networks, where labeled attack data are rarely available and the strongest prior unsupervised detector still relie... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Presage: Prefetch Search via Agent-Guided Experiments | Matthew Giordano, Parthasarathy Ranganathan, Baris Kasikci, Akanksha Jain | 2026-09-18 | 下载 | Data prefetching is an established technique to mitigate cache miss latency and keep the processor saturated with data. Software exists in a unique position to issue prefetches, having algorithmic kno... |
| The Weight Is Over - Interactive Diffusion on Consumer GPUs | Frieder Ganz, Maximilian Müller | 2026-09-18 | 下载 | On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a dec... |
| Performance Analysis of Low-Order, GPU-accelerated Finite Element Kernels using Kokkos | Fabian Böhm, Nils Kohl, Harald Köstler, Ulrich Rüde | 2026-09-18 | 下载 | We study performance portability for low-order, matrix-free finite element kernels, using the example of a vectorial, variable-coefficient PDE operator originating in geophysical models. |
| Cross-Platform vs Native Mobile Development: An Empirical Study of Software Quality Trade-offs | Alexandru Ilovan | 2026-09-18 | 下载 | Cross-platform mobile frameworks promise code reuse, shorter delivery cycles, and lower implementation effort, but their trade-offs relative to native development remain difficult to assess objectivel... |
| Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large Scale | Hao Fu, Jichao Sun, Baiting Zhu, Qiaoling Liu, Yan Shi, Cheng Lu, Liu Liu, Yubo Wang, Xin Yao, Xiangyu Niu, Xu Dong, Wenhan Lyu, Chiyao Shen, Yinjie Huang, Minglei Chen, Shuai Ding, Li Fan, Xiao Kong | 2026-09-18 | 下载 | Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user ... |