2026-08-05
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| PowerScope: ML-based Intra-Cycle Power Estimation | Jayanth Balasubramanian, Sujay Pandit, Radha Vaidya, Anand Raghunathan | 2026-08-05 | 下载 | Power estimation at sub-clock-cycle temporal resolutions is critical for tasks such as power delivery network (PDN) design, dynamic voltage droop analysis, and pre-silicon power side-channel security ... |
| EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative Decoding | Sangwoo Ha, Hyunwoo Seo, Yurim Jo, Youngjin Moon, Hoi-Jun Yoo | 2026-08-05 | 下载 | On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications. A primary bottleneck is external memory access (EMA) in feed-forward network (FFN) layers. |
| Hardware Design and Security in the Era of Chiplets and LLMs | Johann Knechtel, Ozgur Sinanoglu, Paul V. Gratz, Ramesh Karri | 2026-08-05 | 下载 | The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.5D chiplet systems and the integration of Large Language Models (LLMs) into Electronic Design Automation (E... |
| Kerckhoffs-Compliant Watermarking for Physical Design IP Protection: From Placement to Routing | Andrew B. Kahng, Yiting Liu | 2026-08-05 | 下载 | Physical design (PD) intellectual property (IP) is a valuable artifact of modern VLSI implementation. It includes optimized cell placement, clock distribution, and routing decisions produced by carefu... |
| LLM-Assisted Detection and Repair of Hardware Security Vulnerabilities in Verilog Designs | Ethen Santana, Gabriel Gyaase, Hao Zheng | 2026-08-05 | 下载 | Hardware designs, like software, are susceptible to bugs that can introduce security vulnerabilities and create opportunities for malicious exploitation. |
| A Systolic Array Architecture for Nonlinear Activation Functions and Softmax Computation using Chebyshev Polynomials | Benedikt Schaible, Anirudh Suresh Bharadwaj, Ulf Schlichtmann, Jiang Hu | 2026-08-05 | 下载 | Neural Network Accelerators have gained popularity in recent years due to their greater efficiency than CPU-based platforms. Often, these accelerators utilize different hardware units for univariate a... |
| Architectural Implications of Agentic AI Workflows | Jirong Yang, Peizhe Liu, Chaojie Zhang, Jovan Stojkovic | 2026-08-05 | 下载 | Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a pro... |
| MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing | Daijing Shi, Hongxiao Zhao, Yihan Fu, Zhan Chen, Jiayi Li, Yihang Zhu, Anjunyi Fan, Yaoyu Tao, Yuchao Yang, Bonan Yan | 2026-08-05 | 下载 | Emerging workloads, such as Multi-Agent Reinforcement Learning (MARL), large-scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel-sequential computing patte... |
| Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative Inference | Zheng Liu, Zeyu Guo, Zihan Liu, Anbang Wu, Han Zhao, Fangxin Liu, Zhezhi He, Yinhe Han, Jingwen Leng, Minyi Guo, Yiming Gan, Yu Feng | 2026-08-05 | 下载 | Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Filtered Vector Search in a Disaggregated Lakehouse: Composing Table-Format Pruning with Per-File ANN | Rakesh Jain, Thomas Griffin, Syed Zawad | 2026-08-05 | 下载 | Approximate nearest-neighbor (ANN) search increasingly runs alongside structured data - "find the 10 nearest documents where tenant='acme' AND lang='en'" - yet similarity and filtering are usually bol... |
| Application Failures and Machine Computational Efficiency | Carlo Graziani, Bethany Lusch, O. E. Bronson Messer | 2026-08-05 | 下载 | We present a framework for evaluating uptime efficiency of Exascale-class scientific computers when application failure rates are appreciable. |
| DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting | Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta | 2026-08-05 | 下载 | Federated learning repeatedly incurs local optimization and model-update transmission. We study DG-FedReuse, a simulator-level mechanism that allows selected clients to contribute age-decayed cached u... |
| Hierarchical Server Architecture for Agentic Science | Vanessa Sochat, Daniel Milroy | 2026-08-05 | 下载 | Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads require specialized hardware within and across institutions. |
| eMicro: Real-Time Multi-Hop Access Control for Microservices with eBPF | Rizky Ramadhana Putra, Osama Bajaber, Saimon Amanuel Tsegai, Teryl Taylor, Frederico Araujo, Yuede Ji, Peng Gao | 2026-08-05 | 下载 | Modern cloud applications often comprise thousands of microservices whose interactions form complex request paths. Traditional inter-service access control restricts individual service-to-service requ... |
| SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic System | Shiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding | 2026-08-05 | 下载 | Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the input sparsity pattern and execution strategy. |
| RAC: Reference-Aware Activation Compression for Communication-Efficient Split LLM Inference | Guotao Yang, Mingxi Zhao, Haopeng Li, Zhengchao Wang, Sheng Chen, Yitao Hu, Keqiu Li | 2026-08-05 | 下载 | Large language model (LLM) agents repeatedly process long, privacy-sensitive contexts, while cloud-only deployment exposes user data beyond the trusted endpoint and fully local deployment often requir... |
| AsymSpec: Efficient Cloud-Edge Speculative Decoding over Asymmetric Networks | Guotao Yang, Hao Chen, Rui Guo, Xinyu Li, Liang Zheng, Sheng Chen, Yitao Hu, Keqiu Li | 2026-08-05 | 下载 | Cloud-edge speculative decoding places a lightweight draft model at an edge gateway and a higher-quality target model in the cloud, but inserts communication into every speculative block. |
| AFD-Ledger: Deployment Provisioning for Attention--FFN Disaggregation | Chengyu Qiu, Xiao Fu, Fengcun Li, Yulei Qian, Yuchen Xie, Xunliang Cai, Yingdi Shan, Yongwei Wu, Mingxing Zhang | 2026-08-05 | 下载 | Attention--Feed-Forward Network (FFN) Disaggregation (AFD) is emerging as a promising architecture for serving Mixture-of-Experts (MoE) language models. |
| CommBench: Can LLMs Write Correct and Efficient GPU Communication Code? | Shuang Ma, Yuyi Li, Yihan Zhang, Hezhi Xie, Danyang Chen, Shuyang Ji, Ziming Mao, Cheng Ji, Ansha Prashanth, Wenting Yang, Yiran Wang, Chihan Cui, Pei Yu Lin, Ion Stoica, Yang Zhou | 2026-08-05 | 下载 | Training and serving large language models (LLMs) rely heavily on high-performance GPU communication, yet implementing efficient GPU communication primitives requires deep expertise in GPU architectur... |
| MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential Computing | Daijing Shi, Hongxiao Zhao, Yihan Fu, Zhan Chen, Jiayi Li, Yihang Zhu, Anjunyi Fan, Yaoyu Tao, Yuchao Yang, Bonan Yan | 2026-08-05 | 下载 | Emerging workloads, such as Multi-Agent Reinforcement Learning (MARL), large-scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel-sequential computing patte... |
| Zero-Instrumentation Dependency Discovery for Guided Microservice Migration Using eBPF | Eshan Trivedi, Chandrahasa Pranava | 2026-08-05 | 下载 | Migrating microservices across virtual machines (VMs) without knowledge of their runtime communication patterns risks creating cross-VM hotspots and latency spikes that are difficult to predict from s... |
| GPU-Resident CUDA Acceleration for OCUDU 5G PHY and O-RAN Fronthaul: Architecture and Preliminary Performance | Matthew Pennybacker, Wan Liu, Andriy Kharchenko, Timothy OShea | 2026-08-05 | 下载 | This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the und... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application | Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo | 2026-08-05 | 下载 | Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency requirements, such as extended reality (XR). |
| Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks | Marcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo | 2026-08-05 | 下载 | Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-sensitive applications, such as eXtended Reality (X... |
| Learning Compression Rules for Network Traffic | Quentin Lampin, Éloi Sainte-Beuve, Louis-Adrien Dufrène, Guillaume Larue, Massih-Reza Amini | 2026-08-05 | 下载 | We study the problem of learning compact rule-based compressors for structured network traffic. Each packet is a record of header fields that are highly redundant within a flow, and a compressor is a ... |
| Dart: An Automated and Reproducible Environment Toolkit for DNS Protocol Analysis | Yunyi Zhang, Xikai Xiong, Baojun Liu, Haixin Duan | 2026-08-05 | 下载 | Domain Name System (DNS) protocol analysis is fundamental to understanding and fortifying the Internet naming infrastructure. However, the lack of automated, portable, and user-friendly environment or... |
| Zero-Instrumentation Dependency Discovery for Guided Microservice Migration Using eBPF | Eshan Trivedi, Chandrahasa Pranava | 2026-08-05 | 下载 | Migrating microservices across virtual machines (VMs) without knowledge of their runtime communication patterns risks creating cross-VM hotspots and latency spikes that are difficult to predict from s... |
| GPU-Resident CUDA Acceleration for OCUDU 5G PHY and O-RAN Fronthaul: Architecture and Preliminary Performance | Matthew Pennybacker, Wan Liu, Andriy Kharchenko, Timothy OShea | 2026-08-05 | 下载 | This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the und... |
| HRRC on the Farm: Quantile Forecasting for Highly-Reliable Remote Control via LEO Networks | André Gomes, Jie Wang | 2026-08-05 | 下载 | LEO satellite networks are an attractive solution to support farm automation in Agriculture 4.0 because of their ubiquitous coverage. However, LEO networks often suffer from high latency volatility, w... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| eMicro: Real-Time Multi-Hop Access Control for Microservices with eBPF | Rizky Ramadhana Putra, Osama Bajaber, Saimon Amanuel Tsegai, Teryl Taylor, Frederico Araujo, Yuede Ji, Peng Gao | 2026-08-05 | 下载 | Modern cloud applications often comprise thousands of microservices whose interactions form complex request paths. Traditional inter-service access control restricts individual service-to-service requ... |
| Architectural Implications of Agentic AI Workflows | Jirong Yang, Peizhe Liu, Chaojie Zhang, Jovan Stojkovic | 2026-08-05 | 下载 | Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a pro... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Deployment Feasibility Analysis of Post-Quantum Digital Signatures in Safety-Critical C-V2X Communication for Urban Mobility Scenario | Akid Abrar, Sagar Dasgupta, Abdullah Al Mamun, Minhaj Uddin Ahmad, Mizanur Rahman, Mashrur Chowdhury, Ahmad Alsharif | 2026-08-05 | 下载 | The transition from the classical ECDSA to PQC creates substantially larger authentication payloads for safety-critical C-V2X sidelink communication. |