Skip to content

2026-08-02 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM InferenceRuokai Yin, Priyadarshini Panda2026-08-02下载Large Language Models (LLMs) increasingly rely on sparsity to reduce inference cost, but most prior work targets a single sparsity source-either weight or activation-and optimizes for batched multi-us...
DeVIT: Low-Power Vision Transformer Acceleration Using Delta ComputationReyhaneh Hosseinzadeh, Parham Zilouchian Moghaddam, Mehdi Modarressi2026-08-02下载The emergence of transformer-based deep learning models has brought unprecedented performance across various domains, particularly in natural language processing and computer vision.
Compiler Framework for 3D Neutral-Atom Quantum ComputersChen Huang, Zhemin Zhang, Zhao Zhang, Xudong Lv, Zhiding Liang2026-08-02下载Neutral-atom quantum computers can now arrange atoms in three-dimensional tweezer arrays, yet every existing compiler assumes a flat geometry.
Forbench: Symbolic Simulation Helps Make Your Testbench More FormalZiyi Yang, Wenbin Che, Ziyue Zheng, Guangyu Hu, Hongce Zhang2026-08-02下载Simulation remains the dominant approach in pre-silicon verification due to its ease of deployment and intuitive workflow. However, as simulation only explores a limited subset of possible execution t...
On the Limits of Machine-Learned Ranking for Modern Microarchitectural PoliciesYanxin Zhang, Shayne Wadle, Yuxuan Xiong, Zheyu Fu, Trivikram Krishnamurthy, Karu Sankaralingam2026-08-02下载Machine-learning predictors estimate processor performance far faster than cycle-level simulation. For design-space exploration, however, the valuable test is not merely reproducing the usual hardware...
Beyond Static Policies: Dynamic Selection Among Modern Microarchitectural PoliciesYanxin Zhang, Ian McDougall, Junnan Li, Shayne Wadle, Vikas Singh, Karthikeyan Sankaralingam2026-08-02下载Modern processors gain performance from interacting policies: prefetchers, predictors, replacement rules, and schedulers. These policies are often evaluated one at a time, yet a policy that wins in on...
FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing?Weimin Fu, Hejia Zhang, Minghao Shao, Zeng Wang, Johann Knechtel, Ozgur Sinanoglu, Muhammad Shafique, Ramesh Karri, Xiaolong Guo2026-08-02下载Can large language models generate not just correct, but fast hardware? This paper investigates the question in financial FPGA design, where 5-10 nanoseconds of latency determines competitive advantag...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
GPU-Accelerated Multilevel Graph Clustering: A Parallel Perspective on Louvain and LeidenMichael S. Gilbert, Kamesh Madduri2026-08-02下载The sequential Louvain and Leiden algorithms are widely used techniques for modularity-optimizing clustering (or community detection) in large graphs.
NIXT: A NCCL Inspector Exporter Tool for Observability of Collective Communication in Large Model TrainingZiyang Jia, Sirshak Das, Jason Sewall, Laxmi Bhuyan, Pasha Shamis, Daniel Wong2026-08-02下载As machine learning workloads scale, it is increasingly important to gain more observability into the performance of collective communication to easily identify performance vari- ations and accelerate...
Cluster-Aware Over-the-Air Federated Learning with Energy-Harvesting Devices: From Global Training to Model PersonalizationFurkan Bagci, Busra Tegin, Mohammad Kazemi, Tolga M. Duman2026-08-02下载Federated learning (FL) enables distributed optimization and learning across decentralized edge devices while preserving data privacy, but its performance is fundamentally constrained by heterogeneous...
FedChronos: Federated Fine-Tuning of Time-Series Foundation Models for Privacy-Preserving Commodity Price ForecastingAmit Sharma, Nitin Auluck, Akramul Azim2026-08-02下载Time-series foundation models (TSFMs) such as Chronos have demonstrated strong forecasting capabilities across domains, yet adapting them to institutionally fragmented settings, where data cannot be c...
Reputation-driven Cooperation in Lattice-based Decentralized Federated Learning through Evolutionary Game TheoryPhuc Hoang Truong Huynh, Dung Tran Vinh, Khoa Duc Anh Lam, An Nghiem Nguyen Truong, Uyen Nha Tran Bui, Khang Nguyen Dinh, Bao Nguyen Le Gia, Minh Le Nguyen Nhat, Manh Hong Duong, The Anh Han, Thi Ai Thao Nguyen, and Le Hong Trang2026-08-02下载Decentralized Federated Learning (DFL) has emerged as an optimal privacy-preserving solution; however, it remains vulnerable to opportunistic behaviors due to the absence of a central coordinator.
Zellige: Moldable Sequence Placement for Mixed Image-Video DiT TrainingGuangyu Xiang, Xueze Kang, Minwei Zhao, Yuxin Wang, Shaohuai Shi, Lin Zhang, Xiaowen Chu2026-08-02下载High-quality video generation requires training Diffusion Transformers (DiTs) jointly on image and video data, posing a mixed-length sequence training problem across GPUs.
Latency-Optimal Adaptive Split Inference for Privacy-Preserving Cloud-Edge-End CollaborationYi Li, Peng Zhang, Man Ho Au2026-08-02下载Internet of Things (IoT) end devices are increasingly expected to support privacy-sensitive batch inference, yet their limited computational resources often make full local execution of convolutional ...
TIDE-MC: Two-Sided Interpolative Decomposition for Billion-Scale GPU Matrix CompletionChengying Huan, Yubo Wang, Pinhuan Wang, Lizheng Chen, Jie Zhang, Fangxin Liu, Qing Wang, Ruixuan Liu, Shaonan Ma, Mingxing Zhang, Zhibin Wang, Rong Gu, Guihai Chen, Chen Tian2026-08-02下载Matrix completion supports large-scale recommendation and scientific computing, yet existing GPU solvers commonly assume that the observed matrix or its dense factors fit in device memory.
From Cloud to Crowd: Democratizing LLM Service with Decentralized Edge Collaboration for RAGJiaxing Li, Hengzhi Wang, Feng Wang, Chi Xu, Danyang Song, Ruixiao Zhang, Edith C. H. Ngai, Jiangchuan Liu2026-08-02下载The rapid advancement of large language models (LLMs) has increased demand for scalable and cost-effective deployment, especially for mobile and edge devices.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
From Network Automation to Trustworthy Autonomous Networking in the LLM Era: A Network Control Intelligence PerspectiveTianzhu Zhang, Changgang Zheng, Shanshan Wang, Yarui Zhang, Lina Shi, Yue Jin, Xiaofei Wang, Meikang Qiu2026-08-02下载Since the inception of modern communication networks, the quest for operations automation has never ceased. Yet the evolution of network automation is difficult to characterize with a single maturity ...
An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference AgeSiddhant Ray, Nick Feamster, Junchen Jiang2026-08-02下载LLM inference has become a global-scale, heterogeneous workload spanning agents, retrieval, tool-use, code execution and multi-modal reasoning.
Clear-Weighted Bit Allocation for Satellite DownlinksAlireza Furutanpey, Qiyang Zhang, Yujie Huang, Philipp Raith, Schahram Dustdar2026-08-02下载Earth-observation satellites capture more imagery than intermittent ground contacts can transmit. Onboard systems threshold a cloud detector, discard frames or tiles, and compress the survivors with a...
Achieving Rate-Concurrency Balance for Underwater Concurrent Random AccessEnqi Zhang, Yuxuan Guo, Weining Li, Linpeng Chen, Yuetong Chen, Deqing Wang, Lizhao You, Liqun Fu2026-08-02下载Underwater acoustic networks face a fundamental rate--concurrency tradeoff: high-rate waveforms (e.g., OFDM, OTFS) are designed for point-to-point links and rely on orthogonal MAC protocols (e.g.
A Wireless Network Architecture for Monitoring of Hospitalized PatientsLuca Leonardi, Lucia Lo Bello, Gaetano Patti2026-08-02下载Monitoring of vital parameters plays a key role in better understanding the clinical condition of hospitalized patients. In this perspective, the virtual biosensor for MEDIcal WARNing precursor (MEDIW...
VeraRAN: Pre-Actuation Certification and Event-Causal Synchronization Repair for Asynchronous Multi-Interface RAN PlansYinghan Hou, Zongyou Yang2026-08-02下载Agentic RAN controllers combine mobility, energy, and resource actions across independently implemented interfaces. Even when each command is valid and the target state is safe, asynchronous actuation...
Topology-Age-Aware Cooperative Awareness in Vehicular Ad-Hoc NetworksMostafa Lotfi, Masoumeh Moradian2026-08-02下载In vehicular ad-hoc networks (VANETs), maintaining the freshness of status information among vehicles is critical for enabling timely and reliable safety-related decisions.
NoisePQC++: A Unified NIST-Compliant PQC and Hybrid-PQC Implementation of the Noise ProtocolNadeem Ahmed, Aryya Gangopadhyay, Lei Zhang2026-08-02下载The threat of quantum computers to classical public-key cryptography has created an urgent need to evolve secure communication protocols with post-quantum cryptographic (PQC) primitives.
Augmented Backpressure for Decentralized Management of Agentic NetworksZuyuan Zhang, Sizhe Tang, Tian Lan2026-08-02下载Agentic foundation-model service networks handle requests spanning retrieval, planning, generation, verification, and tool use. Unlike traditional communication networks, control performance depends o...
Learning Not to Optimize: Physics-Informed Action-Space Reshaping for Intent-Based Network ControlZuyuan Zhang, Vaneet Aggarwal, Tian Lan2026-08-02下载Modern network policy control maps intent to sequential placement-control decisions. Bellman-style policy optimization primarily asks which action to optimize, while constraints are commonly handled t...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Themis: A Filesystem Model Checker That Owns the MachineDaeyeon Son2026-08-02下载Filesystem model checkers explore an unmodified in-kernel filesystem's state space to find bugs that escape unit tests. The state of the art, Metis, runs inside the OS: it drives syscalls, and, lackin...

cs.PF - Performance ​

标题作者发布日期PDF摘要
A Unified Benchmark for Privacy-preserving Vector SearchAnne-Marie Kermarrec, Rafael Pires, Mathis Randl, Martijn de Vos2026-08-02下载Vector search powers semantic search, recommendation systems, and retrieval-augmented generation (RAG). By design, the service answering a query sees both the query embedding and, usually, the corpus ...
AgriJetsonBench: External-Power-Referenced TensorRT Benchmarking of Agricultural Vision Models on Jetson Edge PlatformsHasan Jahanifar, Hasan Mirzakhaninafchi, Wesley M. Porter, Abolfazl Najar, Glen C. Rains2026-08-02下载Agricultural vision models are often selected from validation accuracy and reported frames per second, but deployment on embedded agricultural edge-GPU systems also depends on timing boundary, numeric...

基于 VitePress 构建 · 使用本地搜索查找论文