Skip to content

2026-09-18 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Presage: Prefetch Search via Agent-Guided ExperimentsMatthew Giordano, Parthasarathy Ranganathan, Baris Kasikci, Akanksha Jain2026-09-18下载Data prefetching is an established technique to mitigate cache miss latency and keep the processor saturated with data. Software exists in a unique position to issue prefetches, having algorithmic kno...
UniCASE: A Unified 16-bit Floating-Point Format with Criticality-Aware Selective ECC for Efficient DNN ProtectionAmna Hassan, Semeen Rehman2026-09-18下载Soft errors are an increasing reliability concern for Deep Neural Network execution because they can corrupt parameters, leading to accuracy degradation.
Scalable Packet Tracking on FPGAs for Erasure-Coded RDMA over Lossy WANsYicheng Qian, Konstantin Taranov, Yevgeny Yankilevich, Assaf Shacham, Mahmoud Elhaddad, Abdul Kabbani, Miriam Leeser, Nadeen Gebara2026-09-18下载Modern AI workloads increasingly rely on scale across architectures that interconnect multiple datacenters to form a single "AI factory", overcoming the power and cooling constraints of individual sit...
Integrating Approximate Logic Synthesis into Approximate High-Level SynthesisJian Shi, Ruicheng Dai, Chang Meng, Yue Yang, Weikang Qian2026-09-18下载Approximate high-level synthesis (HLS) and approximate logic synthesis (ALS) are two techniques for generating approximate circuits. They operate at different granularities.
Programming AMD XDNA NPUs with Open-source Compiler Tools: A FlashAttention Case StudyErwei Wang, Ephrem Wu, Victor J. B. Jung, Jiajie Li, Andre Rosti, Joseph Melber, Samuel Bayliss2026-09-18下载Spatial NPUs such as AMD XDNA place compute tiles beside small local memories and leave data movement between them to software. Mapping a multi-stage workload onto such a device is largely a question ...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Extreme-Scale Ising Machines with Cluster Mean-Field TheoryXiuqi Zhang, Shuvro Chowdhury, Shaila Niazi, Christian Z. Pratt, Navid Anjum Aadit, Kerem Y. Camsari2026-09-18下载Scaling analog and digital Ising machines to larger problems requires overcoming finite device capacity and the cost of communication between devices.
Fairly Compensated Distributed Information Retrieval and Augmentation for AI AgentsYixiang Yao, Pasha Barahimi, Srivatsan Ravi2026-09-18下载The increasing reliance of autonomous AI agents on external and distributed knowledge sources introduces a fundamental challenge for decentralized information marketplaces: retrieval agents must evalu...
Trust-Aware Output Management for Physical Neural Network in Cloud-Continuum SystemsMaliheh Hariri, Stefan Fischer2026-09-18下载Physical Neural Networks (PNNs) introduce new opportunities for cloud continuum computing, but their outputs may be affected by noise, drift, delay, and incomplete reliability information.
Distributed Balanced Butterfly Counting in Signed Bipartite GraphsKiran Mekala, Apurba Das, Suman Banerjee2026-09-18下载The balanced butterfly is a fundamental primitive for analyzing signed bipartite graphs and provides a basis for studying higher-order structural properties, such as clustering coefficients and commun...
Verifiable Computation with Trusted Execution Environments and On-Chain Digital Rights TokensBingle Stegmann Kruger, Co-Pierre Georg2026-09-18下载We present an architecture that enables data owners to combine private data into data pools using Trusted Execution Environments (TEEs) and manage these pools by issuing narrowly scoped computational ...
PoVD: Efficient Consensus Protocol based on Verifiable Delay FunctionRui Jiang, Xintong Ling, Bin Cao, Jiaheng Wang, Xiqi Gao, Zhi Ding2026-09-18下载Consensus protocols ensure the robustness and scalability of blockchains and decentralized applications built on them. However, existing consensus mechanisms often impose high computational cost or re...
HyperParallel-FSDP: Topology-Aware Fully Sharded Training with Layout-Driven Muon on Ascend SuperPodsMo Sun, Yifan Yao, Yanwei Liu, Luobin Liu, Zhenzhang Yang, Kaisheng Wang, Xiangyu Meng, Chen Li, Xizheng Pang, Huilan Li, Xinglei Xu, Yushi Cui, Xinyao Lin, Kaiqi Chen, Jie Zhang, Zeke Wang, Teng Su2026-09-18下载Declarative SPMD programming uses tensor sharding descriptions to drive distributed execution, separating parallelization from model code. However, the evaluated PyTorch DTensor stack dispatches every...
Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication OverlapZiyu Huang, Yangjie Zhou, Chenhao Zhu, Zihan Liu, Jinyu Liu, Shulai Zhang, Xingxun Tang, Hongzhe Yan, Xinhao Luo, Minyi Guo, Xiu Lin, Yinghao Yu, Guodong Yang, Liping Zhang, Shixuan Sun, Jingwen Leng2026-09-18下载Mixture-of-Experts (MoE) inference under expert parallelism (EP) turns each MoE layer into a distributed computation with costly dispatch and combine communication.
TrustBOM: A Scalable Architecture for Confidentiality-Preserving SBOMs Across OrganizationsVan Thang Nguyen, Frederic Rupprecht, Tom Lawrence, Lucca Di Benedetto, Sören Schubert, Amor Rezgui, Sebastian Werner, Maria C. Borges, Stefan Tai2026-09-18下载Software Bills of Materials (SBOMs) have emerged as a key mechanism for software supply chain governance in enterprise architectures. However, their adoption across organizations remains limited due t...
TokaGLINT: A Scalable GPU-Tailored Implicit Solver for Full 3D Tokamak Electromagnetic SimulationsZifan Yang, Haoyuan Zhang, Jialin Li, Wu Yuan, Xiazhen Liu, Jian Zhang, Jianyuan Xiao, Shan Liang2026-09-18下载We introduce TokaGLINT, a GPU-accelerated implicit solver for electromagnetic field computations in full 3D tokamak simulations, aimed at efficient large-scale parallel GPU computing.
Brain API: An Intent-Aware Control Plane for Policy-Governed Agentic SystemsAlexander Chernov2026-09-18下载Contemporary cloud and distributed systems expose control through resource-centric abstractions: services, deployments, network flows, execution graphs.
Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large ScaleHao Fu, Jichao Sun, Baiting Zhu, Qiaoling Liu, Yan Shi, Cheng Lu, Liu Liu, Yubo Wang, Xin Yao, Xiangyu Niu, Xu Dong, Wenhan Lyu, Chiyao Shen, Yinjie Huang, Minglei Chen, Shuai Ding, Li Fan, Xiao Kong2026-09-18下载Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user ...
Adapting the Actor Model of Concurrency for High-Frequency Trading: Synchronous Message Delivery (fast_send) and a Tick-to-Book Latency StudyVincent Maciejewski2026-09-18下载The actor model - state isolation, data-race freedom, deadlock resistance, and sequential single-message reasoning - has long been dismissed as unsuitable for high-frequency trading (HFT): actors seem...
AI-Driven Scientific Computing Workflows: A Systems Review of Orchestration, Execution, Reproducibility and ProvenanceJamie J. Alnasir2026-09-18下载Artificial intelligence (AI) is increasingly embedded within scientific computing workflows that combine simulation, data processing, optimisation, visualisation and experimental or observational comp...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
User-Level Handover Decision Making Based on Machine Learning ApproachesJoão Lima, Alvaro Medeiros, Eduardo Aguiar, Vicente Angelo de Sousa Junior, Tarciana Guerra2026-09-18下载This letter covers a broad comparison of methods for classification and regression applications for a user-level handover decision making in scenarios with adverse propagation conditions involving bui...
Tick-Tock on the Open Fronthaul: Securing Synchronization in O-RANYiwei Zhang, Enrico Pisanti, Imtiaz Karim, Subangkar Karmaker Shanto, Elisa Bertino2026-09-18下载The Precision Time Protocol (PTP) provides the time and phase synchronization required by disaggregated Open Radio Access Networks (O-RAN). Yet, in current open fronthaul deployments, PTP traffic lack...
Semantics Delivery Network: Rethinking Web Retrieval Infrastructure for LLM AgentsPeichun Hua, Yunming Xiao2026-09-18下载Large language models (LLMs) increasingly rely on external sources when answering questions that require proprietary information or up-to-date live web content, through both traditional single-shot re...
Provisional Reachability: Containing Agents by Making Every Crossing RevocableYoshiaki Takashita2026-09-18下载A companion paper found that what a defender must block over time has units: bits per period [Takashita, 2026a]. This paper sets it. Hold every crossing in escrow for one period, audit each held item ...
A Multi-Cloud View of Internet Background RadiationNils Kempen, Ricky K. P. Mok, Bernhard Degen, Syed Mujtaba Jafri, Ralph Holz2026-09-18下载As services are increasingly centralized in public clouds, understanding the nature of Internet Background Radiation (IBR) hitting these particular environments is an important part of understanding t...
Secure RIS-Aided Multicasting: Globally Optimal Beam Management and Discrete-Phase RIS ConfigurationLuis F. Abanto-Leon, Setareh Maghsudi2026-09-18下载Reconfigurable intelligent surfaces (RISs) are poised to revolutionize wireless multicasting by enabling extended coverage and reliable operation in obstructed environments.
Pattern-Aware Virtual Network Embedding Optimization for Cloud Data CentersBinquan Guo, Zhou Zhang, Junfeng Zhai, Zheng Zhang, Marie Siew, Zehui Xiong2026-09-18下载The network virtualization (NV) technology has enabled the sharing of multiple resources among virtual networks (VNs) in cloud data centers. One of the key challenges is to allocate resources in real-...
Locating and Enumerating Anycast: a Comparison of Two ApproachesRemi Hendriks, Tim Betzer, Ben Du, Raffaele Sommese, Mattijs Jonker, Roland van Rijswijk-Deij2026-09-18下载Anycast allows for providing services from multiple, geographically distant Points of Presence (PoPs), using a single IP address, to, e.g., improve resilience.
X-SPUR: Explainable Surprisal-Based Protocol-Aware Unsupervised Reasoning for Automotive Ethernet Intrusion DetectionJisoo Kim, Seonghoon Jeong2026-09-18下载Automotive Ethernet carries heterogeneous multi-protocol traffic in modern in-vehicle networks, where labeled attack data are rarely available and the strongest prior unsupervised detector still relie...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Presage: Prefetch Search via Agent-Guided ExperimentsMatthew Giordano, Parthasarathy Ranganathan, Baris Kasikci, Akanksha Jain2026-09-18下载Data prefetching is an established technique to mitigate cache miss latency and keep the processor saturated with data. Software exists in a unique position to issue prefetches, having algorithmic kno...
The Weight Is Over - Interactive Diffusion on Consumer GPUsFrieder Ganz, Maximilian Müller2026-09-18下载On-device inference is booming, but the momentum is almost all in language models. Diffusion pipelines are memory hungry, latency-sensitive, and require orchestrating an embedder, a transformer, a dec...
Performance Analysis of Low-Order, GPU-accelerated Finite Element Kernels using KokkosFabian Böhm, Nils Kohl, Harald Köstler, Ulrich Rüde2026-09-18下载We study performance portability for low-order, matrix-free finite element kernels, using the example of a vectorial, variable-coefficient PDE operator originating in geophysical models.
Cross-Platform vs Native Mobile Development: An Empirical Study of Software Quality Trade-offsAlexandru Ilovan2026-09-18下载Cross-platform mobile frameworks promise code reuse, shorter delivery cycles, and lower implementation effort, but their trade-offs relative to native development remain difficult to assess objectivel...
Hybrid GPU-CPU Retrieval for Personalized Search at Ultra-Large ScaleHao Fu, Jichao Sun, Baiting Zhu, Qiaoling Liu, Yan Shi, Cheng Lu, Liu Liu, Yubo Wang, Xin Yao, Xiangyu Niu, Xu Dong, Wenhan Lyu, Chiyao Shen, Yinjie Huang, Minglei Chen, Shuai Ding, Li Fan, Xiao Kong2026-09-18下载Embedding-based retrieval on user-generated content at the trillion-document scale exposes a sharp conflict between two production demands: deep, expressive personalization for queries with rich user ...

基于 VitePress 构建 · 使用本地搜索查找论文