Skip to content

2026-08-27 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Vision-centric generative AI models: A software-hardware perspectiveEleni Tselepi, Cristian Sestito, Shady Agwa, Themis Prodromakis2026-08-27下载Vision generative artificial intelligence (AI) has emerged as one of the most rapidly advancing areas of deep learning. The explosion of multimodal models has made them widely associated with text-to-...
LLMs in Digital EDA: A perspective on shifting roles from Generation to OrchestrationMatthew Youngman, Cristian Sestito, Themis Prodromakis2026-08-27下载Electronic design automation (EDA) has advanced engineering productivity through successive generations of tooling that progressively automate synthesis, optimisation, and verification.
HOLMES: In-Context Failure-Center Localization for High-Dimensional Yield EstimationWei W. Xing, Xixi Zhou, Kaiqi Huang, Jiaye Pan, Hong Qiu, Xin Wang, Shan Shen2026-08-27下载Importance sampling for high-sigma yield estimation requires locating the failure center from a severely imbalanced sample set. Existing surrogate-assisted methods rely on iterative gradient-based tra...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Characterizing the I/O Behavior of HPC Applications through Modeling and SimulationNjoud O. Almaaitah, David E. Singh, Taylan Özden, Jesus Carretero, Raffaele Montella2026-08-27下载Parallel applications process large amounts of data, leading to intensive parallel I/O operations. These operations can exhibit different levels of complexity, including, among others, multiple I/O ac...
Consensus with Stochastic BroadcastPierre Fraigniaud, Boaz Patt-Shamir, Sergio Rajsbaum2026-08-27下载We study binary consensus in the \emph{stochastic broadcast model}, which assumes n≥2n\geq 2 processes communicating synchronously by message broadcasts.
Decoupled I/O-Dominant Pipelines for Large-Scale Whole-Slide Image Embedding ExtractionMayanka Chandrashekar, Xi Zhang, Ethan Seefried, Tirthankar Ghosal, John Gounley, Heidi Hanson2026-08-27下载Whole-slide images (WSIs) are central to computational pathology but are prohibitively large, making patch-based processing the practical unit for foundation model inference.
Towards Reproducible Evaluation of Distributed Quantum Circuit Partitioning AlgorithmsJavier Vela-Tambo, Davud Azizov, Tian Guo2026-08-27下载Distributed Quantum Computing (DQC) addresses the physical scaling limitations of monolithic quantum processors by networking modular Quantum Processing Units (QPUs).
Sintr: Safe Interactive Transactions in the Presence of Byzantine ClientsAustin T. Li, Daniel H. Lee, Lorenzo Alvisi, Natacha Crooks, Florian Suri-Payer2026-08-27下载Byzantine fault-tolerant (BFT) systems are, in principle, an appealing foundation for transactional applications involving mutually distrustful participants.
Performance Foundations of Parallel & Distributed Reasoning Language ModelsMaciej Besta, Leonard Schmidt, Lara Nonino, Robert Gerstenberger, Pierre Pang, Patrik Okanovic, Ales Kubicek, Tiancheng Chen, Baraq Lipshitz, Torsten Hoefler2026-08-27下载Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards.
Benchmarking Confidential Computing Performance on NVIDIA Blackwell GPUsDaniyal Khan, Amean Asad, Ansgar Grunseid2026-08-27下载This paper measures the performance impact of running large language model inference and training inside a Trusted Execution Environment (TEE) on NVIDIA B200 GPUs, using Intel Trust Domain Extensions ...
Optimizing API Gateway Placement in Multi-Cloud KubernetesVinoth Punniyamoorthy, Murali Shankar Dulam, Aswathnarayan Muthukrishnan Kirubakaran, Akshay Deshpande, Nachiappan Chockalingam, Bikesh Kumar, Naga Surya Pasupuleti, Narender Reddy Bitla2026-08-27下载The use of API gateways within geographically distributed multi-cloud Kubernetes clusters poses a tradeoff between infrastructure cost, computational resources, and network latencies.
IBLTs Measure Before They Decode: Self-Sizing Set Reconciliation from Pre-Peeling CountsMin Wu, Ji Qi, Chengdui Luo, Shudong Lu, Zhengsheng Ye, Zhengyang Wei2026-08-27下载Set reconciliation recovers the symmetric difference A△BA\triangle B with communication far below the data volume. Invertible Bloom Lookup Tables (IBLTs) are a standard tool, but their capacity must ma...
VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM InferenceYan Shi, Xiaochao Wang, Jingchun Gao, Jintao Luo, Xinyi Zhou, Feng Liu, Kui Luo, Xushi Li, Xinjie Guo, Liangjun Feng2026-08-27下载Chunked prefill pipeline parallelism (CPP) is a key technique for LLM inference. However, equal-size chunks exhibit imbalanced latency, as later chunks attend longer prefix KV caches and incur higher ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Real-Time Reconstruction of Markov Sources over MPR ChannelsPansee S. Elessawy, Nikolaos Pappas2026-08-27下载This paper studies the real-time reconstruction and remote actuation of two binary Markov sources over a shared wireless channel with multi-packet reception (MPR).
FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User TicketsKuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne2026-08-27下载LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and always assume a fault is present, conditions rarely met in pr...
Claude Code Complete User HandbookDavid Soldani2026-08-27下载Claude Code is an agentic work environment: a language model operating in a loop with filesystem access, shell execution, browser control, scheduled and cloud execution, external tool connections thro...
Extending Low Latency Service Across the InternetHarkirat Singh, Fatih Berkay Sarpkaya, Hakan Gulec, Fraida Fund, Shivendra Panwar2026-08-27下载Protocols such as L4S for low latency network services have attracted growing interest from major industry stakeholders such as Comcast, Apple, T-Mobile, and NVIDIA.
SFC-Aware Online Aggregated Data-Link Orchestration for SDN/NFV-Enabled SAGINsZiyang Guo, Bing Du2026-08-27下载Civil aviation space-air-ground integrated networks (SAGINs) are expected to support heterogeneous cockpit and cabin services over dynamic air-to-air (A2A), air-to-ground (A2G), and air-to-satellite (...
PRO-RAN: Processor-Level Characterization of Open RAN Centralized and Distributed UnitsMoojan Kamalzadeh, Larry Horner, Linqi Xiao, Abhishek Bhattacharyya, Ehsan Bahaloo Horeh, Padmapriya Patil, Venkateswarlu Gudepu, Andrea Fumagalli2026-08-27下载Open Radio Access Network (O-RAN) disaggregates RAN protocol functions and enables Centralized Unit (CU) and Distributed Unit (DU) software to execute on general-purpose computing platforms.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Adaptation Fidelity of SPEC CPU2026Doa'a Al-Otoom, Mahesh Madhav2026-08-27下载Standardized benchmarks are often criticized for not being "real workloads," but this critique is rarely backed by data. This paper provides the first systematic, quantitative analysis of the "fidelit...
Accelerating Data Preprocessing for Efficient Vision Model Inference on Jetson Edge DeviceTian Chen, Nawras Alnaasan, Jinghan Yao, Aamir Shafi, Hari Subramoni, Dhabaleswar K., Panda2026-08-27下载Data preprocessing is a crucial part of deep learning workflows on edge devices. However, decoding data saved in JPEG format is very compute-intensive and occupies a major portion of the preprocessing...
Performance Foundations of Parallel & Distributed Reasoning Language ModelsMaciej Besta, Leonard Schmidt, Lara Nonino, Robert Gerstenberger, Pierre Pang, Patrik Okanovic, Ales Kubicek, Tiancheng Chen, Baraq Lipshitz, Torsten Hoefler2026-08-27下载Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language models (LLMs) with reasoning standards.
FoldPipe: Bounded Remote Streaming of Native Molecular Shards with Asynchronous PrefetchDhiren Mukesh Khatri2026-08-27下载Training molecular machine-learning models on ephemeral or memory-constrained accelerator instances can require repeatedly retrieving preprocessed molecular graphs from remote storage.
TOPIQ: Statistical Error Propagation for Quantity-of-Interest Prediction under Lossy CompressionYouyuan Liu, Bo Jiang, Taolue Yang, Sheng Di, Robert Underwood, Sian Jin2026-08-27下载Lossy compression is essential for managing massive scientific data, but per-element error bounds do not translate into bounds on downstream quantities of interest (QoIs) such as regional averages, ne...
Launch-Bound and Substitutable: Why Three Inference Optimizations Fail to Pay Off in Mixture-of-Experts ModelsGokulakannan Sakthivel, Jerry Wu, Amogh Rajendra, Giriprasad Radhakrishnan2026-08-27下载Mixture-of-Experts (MoE) models route each token to a few of many expert networks, and that routing is data-dependent in a way standard inference optimizations do not expect.
PRO-RAN: Processor-Level Characterization of Open RAN Centralized and Distributed UnitsMoojan Kamalzadeh, Larry Horner, Linqi Xiao, Abhishek Bhattacharyya, Ehsan Bahaloo Horeh, Padmapriya Patil, Venkateswarlu Gudepu, Andrea Fumagalli2026-08-27下载Open Radio Access Network (O-RAN) disaggregates RAN protocol functions and enables Centralized Unit (CU) and Distributed Unit (DU) software to execute on general-purpose computing platforms.

基于 VitePress 构建 · 使用本地搜索查找论文