Skip to content

2026-08-05 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
PowerScope: ML-based Intra-Cycle Power EstimationJayanth Balasubramanian, Sujay Pandit, Radha Vaidya, Anand Raghunathan2026-08-05下载Power estimation at sub-clock-cycle temporal resolutions is critical for tasks such as power delivery network (PDN) design, dynamic voltage droop analysis, and pre-silicon power side-channel security ...
EdgeXpert: An Edge Device for Memory-Efficient LLM Inference with Mixture-of-Experts and Speculative DecodingSangwoo Ha, Hyunwoo Seo, Yurim Jo, Youngjin Moon, Hoi-Jun Yoo2026-08-05下载On-device deployment of Large Language Models (LLMs) has become essential for personalized edge applications. A primary bottleneck is external memory access (EMA) in feed-forward network (FFN) layers.
Hardware Design and Security in the Era of Chiplets and LLMsJohann Knechtel, Ozgur Sinanoglu, Paul V. Gratz, Ramesh Karri2026-08-05下载The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.5D chiplet systems and the integration of Large Language Models (LLMs) into Electronic Design Automation (E...
Kerckhoffs-Compliant Watermarking for Physical Design IP Protection: From Placement to RoutingAndrew B. Kahng, Yiting Liu2026-08-05下载Physical design (PD) intellectual property (IP) is a valuable artifact of modern VLSI implementation. It includes optimized cell placement, clock distribution, and routing decisions produced by carefu...
LLM-Assisted Detection and Repair of Hardware Security Vulnerabilities in Verilog DesignsEthen Santana, Gabriel Gyaase, Hao Zheng2026-08-05下载Hardware designs, like software, are susceptible to bugs that can introduce security vulnerabilities and create opportunities for malicious exploitation.
A Systolic Array Architecture for Nonlinear Activation Functions and Softmax Computation using Chebyshev PolynomialsBenedikt Schaible, Anirudh Suresh Bharadwaj, Ulf Schlichtmann, Jiang Hu2026-08-05下载Neural Network Accelerators have gained popularity in recent years due to their greater efficiency than CPU-based platforms. Often, these accelerators utilize different hardware units for univariate a...
Architectural Implications of Agentic AI WorkflowsJirong Yang, Peizhe Liu, Chaojie Zhang, Jovan Stojkovic2026-08-05下载Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a pro...
MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential ComputingDaijing Shi, Hongxiao Zhao, Yihan Fu, Zhan Chen, Jiayi Li, Yihang Zhu, Anjunyi Fan, Yaoyu Tao, Yuchao Yang, Bonan Yan2026-08-05下载Emerging workloads, such as Multi-Agent Reinforcement Learning (MARL), large-scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel-sequential computing patte...
Deltoris: Enabling Real-time VLA Inference in Embodied AI via Bit-level Sparsity and Speculative InferenceZheng Liu, Zeyu Guo, Zihan Liu, Anbang Wu, Han Zhao, Fangxin Liu, Zhezhi He, Yinhe Han, Jingwen Leng, Minyi Guo, Yiming Gan, Yu Feng2026-08-05下载Vision-language-action (VLA) models have emerged as a key component in embodied AI. Among existing approaches, diffusion-based VLA models achieve superior motion quality and generalization.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Filtered Vector Search in a Disaggregated Lakehouse: Composing Table-Format Pruning with Per-File ANNRakesh Jain, Thomas Griffin, Syed Zawad2026-08-05下载Approximate nearest-neighbor (ANN) search increasingly runs alongside structured data - "find the 10 nearest documents where tenant='acme' AND lang='en'" - yet similarity and filtering are usually bol...
Application Failures and Machine Computational EfficiencyCarlo Graziani, Bethany Lusch, O. E. Bronson Messer2026-08-05下载We present a framework for evaluating uptime efficiency of Exascale-class scientific computers when application failure rates are appreciable.
DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink AccountingRahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta2026-08-05下载Federated learning repeatedly incurs local optimization and model-update transmission. We study DG-FedReuse, a simulator-level mechanism that allows selected clients to contribute age-decayed cached u...
Hierarchical Server Architecture for Agentic ScienceVanessa Sochat, Daniel Milroy2026-08-05下载Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads require specialized hardware within and across institutions.
eMicro: Real-Time Multi-Hop Access Control for Microservices with eBPFRizky Ramadhana Putra, Osama Bajaber, Saimon Amanuel Tsegai, Teryl Taylor, Frederico Araujo, Yuede Ji, Peng Gao2026-08-05下载Modern cloud applications often comprise thousands of microservices whose interactions form complex request paths. Traditional inter-service access control restricts individual service-to-service requ...
SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic SystemShiyang Li, Guangyan Sun, Jinwei Tang, Yanzhi Wang, Mingyi Hong, Caiwen Ding2026-08-05下载Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the input sparsity pattern and execution strategy.
RAC: Reference-Aware Activation Compression for Communication-Efficient Split LLM InferenceGuotao Yang, Mingxi Zhao, Haopeng Li, Zhengchao Wang, Sheng Chen, Yitao Hu, Keqiu Li2026-08-05下载Large language model (LLM) agents repeatedly process long, privacy-sensitive contexts, while cloud-only deployment exposes user data beyond the trusted endpoint and fully local deployment often requir...
AsymSpec: Efficient Cloud-Edge Speculative Decoding over Asymmetric NetworksGuotao Yang, Hao Chen, Rui Guo, Xinyu Li, Liang Zheng, Sheng Chen, Yitao Hu, Keqiu Li2026-08-05下载Cloud-edge speculative decoding places a lightweight draft model at an edge gateway and a higher-quality target model in the cloud, but inserts communication into every speculative block.
AFD-Ledger: Deployment Provisioning for Attention--FFN DisaggregationChengyu Qiu, Xiao Fu, Fengcun Li, Yulei Qian, Yuchen Xie, Xunliang Cai, Yingdi Shan, Yongwei Wu, Mingxing Zhang2026-08-05下载Attention--Feed-Forward Network (FFN) Disaggregation (AFD) is emerging as a promising architecture for serving Mixture-of-Experts (MoE) language models.
CommBench: Can LLMs Write Correct and Efficient GPU Communication Code?Shuang Ma, Yuyi Li, Yihan Zhang, Hezhi Xie, Danyang Chen, Shuyang Ji, Ziming Mao, Cheng Ji, Ansha Prashanth, Wenting Yang, Yiran Wang, Chihan Cui, Pei Yu Lin, Ion Stoica, Yang Zhou2026-08-05下载Training and serving large language models (LLMs) rely heavily on high-performance GPU communication, yet implementing efficient GPU communication primitives requires deep expertise in GPU architectur...
MCHA: A Memory-Centric Hierarchical Architecture for Parallel-Sequential ComputingDaijing Shi, Hongxiao Zhao, Yihan Fu, Zhan Chen, Jiayi Li, Yihang Zhu, Anjunyi Fan, Yaoyu Tao, Yuchao Yang, Bonan Yan2026-08-05下载Emerging workloads, such as Multi-Agent Reinforcement Learning (MARL), large-scale neuromorphic computing, and probabilistic graphical models, intrinsically exhibit parallel-sequential computing patte...
Zero-Instrumentation Dependency Discovery for Guided Microservice Migration Using eBPFEshan Trivedi, Chandrahasa Pranava2026-08-05下载Migrating microservices across virtual machines (VMs) without knowledge of their runtime communication patterns risks creating cross-VM hotspots and latency spikes that are difficult to predict from s...
GPU-Resident CUDA Acceleration for OCUDU 5G PHY and O-RAN Fronthaul: Architecture and Preliminary PerformanceMatthew Pennybacker, Wan Liu, Andriy Kharchenko, Timothy OShea2026-08-05下载This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the und...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive ApplicationMarcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo2026-08-05下载Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency requirements, such as extended reality (XR).
Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN NetworksMarcos Carvalho, Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci, Daniel F. Macedo2026-08-05下载Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-sensitive applications, such as eXtended Reality (X...
Learning Compression Rules for Network TrafficQuentin Lampin, Éloi Sainte-Beuve, Louis-Adrien Dufrène, Guillaume Larue, Massih-Reza Amini2026-08-05下载We study the problem of learning compact rule-based compressors for structured network traffic. Each packet is a record of header fields that are highly redundant within a flow, and a compressor is a ...
Dart: An Automated and Reproducible Environment Toolkit for DNS Protocol AnalysisYunyi Zhang, Xikai Xiong, Baojun Liu, Haixin Duan2026-08-05下载Domain Name System (DNS) protocol analysis is fundamental to understanding and fortifying the Internet naming infrastructure. However, the lack of automated, portable, and user-friendly environment or...
Zero-Instrumentation Dependency Discovery for Guided Microservice Migration Using eBPFEshan Trivedi, Chandrahasa Pranava2026-08-05下载Migrating microservices across virtual machines (VMs) without knowledge of their runtime communication patterns risks creating cross-VM hotspots and latency spikes that are difficult to predict from s...
GPU-Resident CUDA Acceleration for OCUDU 5G PHY and O-RAN Fronthaul: Architecture and Preliminary PerformanceMatthew Pennybacker, Wan Liu, Andriy Kharchenko, Timothy OShea2026-08-05下载This paper describes DeepSig's CUDA-based acceleration backend for the OCUDU physical layer and O-RAN fronthaul path, integrated through acceleration interfaces that are largely independent of the und...
HRRC on the Farm: Quantile Forecasting for Highly-Reliable Remote Control via LEO NetworksAndré Gomes, Jie Wang2026-08-05下载LEO satellite networks are an attractive solution to support farm automation in Agriculture 4.0 because of their ubiquitous coverage. However, LEO networks often suffer from high latency volatility, w...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
eMicro: Real-Time Multi-Hop Access Control for Microservices with eBPFRizky Ramadhana Putra, Osama Bajaber, Saimon Amanuel Tsegai, Teryl Taylor, Frederico Araujo, Yuede Ji, Peng Gao2026-08-05下载Modern cloud applications often comprise thousands of microservices whose interactions form complex request paths. Traditional inter-service access control restricts individual service-to-service requ...
Architectural Implications of Agentic AI WorkflowsJirong Yang, Peizhe Liu, Chaojie Zhang, Jovan Stojkovic2026-08-05下载Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy and present its first architectural characterization with a pro...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Deployment Feasibility Analysis of Post-Quantum Digital Signatures in Safety-Critical C-V2X Communication for Urban Mobility ScenarioAkid Abrar, Sagar Dasgupta, Abdullah Al Mamun, Minhaj Uddin Ahmad, Mizanur Rahman, Mashrur Chowdhury, Ahmad Alsharif2026-08-05下载The transition from the classical ECDSA to PQC creates substantially larger authentication payloads for safety-critical C-V2X sidelink communication.

基于 VitePress 构建 · 使用本地搜索查找论文