2026-08-27
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Vision-centric generative AI models: A software-hardware perspective | Eleni Tselepi, Cristian Sestito, Shady Agwa, Themis Prodromakis | 2026-08-27 | 下载 | Vision generative artificial intelligence (AI) has emerged as one of the most rapidly advancing areas of deep learning. The explosion of multimodal models has made them widely associated with text-to-... |
| LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration | Matthew Youngman, Cristian Sestito, Themis Prodromakis | 2026-08-27 | 下载 | Electronic design automation (EDA) has advanced engineering productivity through successive generations of tooling that progressively automate synthesis, optimisation, and verification. |
| HOLMES: In-Context Failure-Center Localization for High-Dimensional Yield Estimation | Wei W. Xing, Xixi Zhou, Kaiqi Huang, Jiaye Pan, Hong Qiu, Xin Wang, Shan Shen | 2026-08-27 | 下载 | Importance sampling for high-sigma yield estimation requires locating the failure center from a severely imbalanced sample set. Existing surrogate-assisted methods rely on iterative gradient-based tra... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Characterizing the I/O Behavior of HPC Applications through Modeling and Simulation | Njoud O. Almaaitah, David E. Singh, Taylan Özden, Jesus Carretero, Raffaele Montella | 2026-08-27 | 下载 | Parallel applications process large amounts of data, leading to intensive parallel I/O operations. These operations can exhibit different levels of complexity, including, among others, multiple I/O ac... |
| Consensus with Stochastic Broadcast | Pierre Fraigniaud, Boaz Patt-Shamir, Sergio Rajsbaum | 2026-08-27 | 下载 | We study binary consensus in the \emph{stochastic broadcast model}, which assumes processes communicating synchronously by message broadcasts. |
| Decoupled I/O-Dominant Pipelines for Large-Scale Whole-Slide Image Embedding Extraction | Mayanka Chandrashekar, Xi Zhang, Ethan Seefried, Tirthankar Ghosal, John Gounley, Heidi Hanson | 2026-08-27 | 下载 | Whole-slide images (WSIs) are central to computational pathology but are prohibitively large, making patch-based processing the practical unit for foundation model inference. |
| Towards Reproducible Evaluation of Distributed Quantum Circuit Partitioning Algorithms | Javier Vela-Tambo, Davud Azizov, Tian Guo | 2026-08-27 | 下载 | Distributed Quantum Computing (DQC) addresses the physical scaling limitations of monolithic quantum processors by networking modular Quantum Processing Units (QPUs). |
| Sintr: Safe Interactive Transactions in the Presence of Byzantine Clients | Austin T. Li, Daniel H. Lee, Lorenzo Alvisi, Natacha Crooks, Florian Suri-Payer | 2026-08-27 | 下载 | Byzantine fault-tolerant (BFT) systems are, in principle, an appealing foundation for transactional applications involving mutually distrustful participants. |
| Performance Foundations of Parallel & Distributed Reasoning Language Models | Maciej Besta, Leonard Schmidt, Lara Nonino, Robert Gerstenberger, Pierre Pang, Patrik Okanovic, Ales Kubicek, Tiancheng Chen, Baraq Lipshitz, Torsten Hoefler | 2026-08-27 | 下载 | Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. |
| Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUs | Daniyal Khan, Amean Asad, Ansgar Grunseid | 2026-08-27 | 下载 | This paper measures the performance impact of running large language model inference and training inside a Trusted Execution Environment (TEE) on NVIDIA B200 GPUs, using Intel Trust Domain Extensions ... |
| Optimizing API Gateway Placement in Multi-Cloud Kubernetes | Vinoth Punniyamoorthy, Murali Shankar Dulam, Aswathnarayan Muthukrishnan Kirubakaran, Akshay Deshpande, Nachiappan Chockalingam, Bikesh Kumar, Naga Surya Pasupuleti, Narender Reddy Bitla | 2026-08-27 | 下载 | The use of API gateways within geographically distributed multi-cloud Kubernetes clusters poses a tradeoff between infrastructure cost, computational resources, and network latencies. |
| IBLTs Measure Before They Decode: Self-Sizing Set Reconciliation from Pre-Peeling Counts | Min Wu, Ji Qi, Chengdui Luo, Shudong Lu, Zhengsheng Ye, Zhengyang Wei | 2026-08-27 | 下载 | Set reconciliation recovers the symmetric difference with communication far below the data volume. Invertible Bloom Lookup Tables (IBLTs) are a standard tool, but their capacity must ma... |
| VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM Inference | Yan Shi, Xiaochao Wang, Jingchun Gao, Jintao Luo, Xinyi Zhou, Feng Liu, Kui Luo, Xushi Li, Xinjie Guo, Liangjun Feng | 2026-08-27 | 下载 | Chunked prefill pipeline parallelism (CPP) is a key technique for LLM inference. However, equal-size chunks exhibit imbalanced latency, as later chunks attend longer prefix KV caches and incur higher ... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Real-Time Reconstruction of Markov Sources over MPR Channels | Pansee S. Elessawy, Nikolaos Pappas | 2026-08-27 | 下载 | This paper studies the real-time reconstruction and remote actuation of two binary Markov sources over a shared wireless channel with multi-packet reception (MPR). |
| FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets | Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne | 2026-08-27 | 下载 | LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in pr... |
| Claude Code Complete User Handbook | David Soldani | 2026-08-27 | 下载 | Claude Code is an agentic work environment: a language model operating in a loop with filesystem access, shell execution, browser control, scheduled and cloud execution, external tool connections thro... |
| Extending Low Latency Service Across the Internet | Harkirat Singh, Fatih Berkay Sarpkaya, Hakan Gulec, Fraida Fund, Shivendra Panwar | 2026-08-27 | 下载 | Protocols such as L4S for low latency network services have attracted growing interest from major industry stakeholders such as Comcast, Apple, T-Mobile, and NVIDIA. |
| SFC-Aware Online Aggregated Data-Link Orchestration for SDN/NFV-Enabled SAGINs | Ziyang Guo, Bing Du | 2026-08-27 | 下载 | Civil aviation space-air-ground integrated networks (SAGINs) are expected to support heterogeneous cockpit and cabin services over dynamic air-to-air (A2A), air-to-ground (A2G), and air-to-satellite (... |
| PRO-RAN: Processor-Level Characterization of Open RAN Centralized and Distributed Units | Moojan Kamalzadeh, Larry Horner, Linqi Xiao, Abhishek Bhattacharyya, Ehsan Bahaloo Horeh, Padmapriya Patil, Venkateswarlu Gudepu, Andrea Fumagalli | 2026-08-27 | 下载 | Open Radio Access Network (O-RAN) disaggregates RAN protocol functions and enables Centralized Unit (CU) and Distributed Unit (DU) software to execute on general-purpose computing platforms. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Adaptation Fidelity of SPEC CPU2026 | Doa'a Al-Otoom, Mahesh Madhav | 2026-08-27 | 下载 | Standardized benchmarks are often criticized for not being "real workloads," but this critique is rarely backed by data. This paper provides the first systematic, quantitative analysis of the "fidelit... |
| Accelerating Data Preprocessing for Efficient Vision Model Inference on Jetson Edge Device | Tian Chen, Nawras Alnaasan, Jinghan Yao, Aamir Shafi, Hari Subramoni, Dhabaleswar K., Panda | 2026-08-27 | 下载 | Data preprocessing is a crucial part of deep learning workflows on edge devices. However, decoding data saved in JPEG format is very compute-intensive and occupies a major portion of the preprocessing... |
| Performance Foundations of Parallel & Distributed Reasoning Language Models | Maciej Besta, Leonard Schmidt, Lara Nonino, Robert Gerstenberger, Pierre Pang, Patrik Okanovic, Ales Kubicek, Tiancheng Chen, Baraq Lipshitz, Torsten Hoefler | 2026-08-27 | 下载 | Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards. |
| FoldPipe: Bounded Remote Streaming of Native Molecular Shards with Asynchronous Prefetch | Dhiren Mukesh Khatri | 2026-08-27 | 下载 | Training molecular machine-learning models on ephemeral or memory-constrained accelerator instances can require repeatedly retrieving preprocessed molecular graphs from remote storage. |
| TOPIQ: Statistical Error Propagation for Quantity-of-Interest Prediction under Lossy Compression | Youyuan Liu, Bo Jiang, Taolue Yang, Sheng Di, Robert Underwood, Sian Jin | 2026-08-27 | 下载 | Lossy compression is essential for managing massive scientific data, but per-element error bounds do not translate into bounds on downstream quantities of interest (QoIs) such as regional averages, ne... |
| Launch-Bound and Substitutable: Why Three Inference Optimizations Fail to Pay Off in Mixture-of-Experts Models | Gokulakannan Sakthivel, Jerry Wu, Amogh Rajendra, Giriprasad Radhakrishnan | 2026-08-27 | 下载 | Mixture-of-Experts (MoE) models route each token to a few of many expert networks, and that routing is data-dependent in a way standard inference optimizations do not expect. |
| PRO-RAN: Processor-Level Characterization of Open RAN Centralized and Distributed Units | Moojan Kamalzadeh, Larry Horner, Linqi Xiao, Abhishek Bhattacharyya, Ehsan Bahaloo Horeh, Padmapriya Patil, Venkateswarlu Gudepu, Andrea Fumagalli | 2026-08-27 | 下载 | Open Radio Access Network (O-RAN) disaggregates RAN protocol functions and enables Centralized Unit (CU) and Distributed Unit (DU) software to execute on general-purpose computing platforms. |