2026-08-07
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs | Liad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson | 2026-08-07 | 下载 | Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems. |
| Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-Tuning | Vasanth Iyer | 2026-08-07 | 下载 | Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited. |
| HINT: Toward an Executable Hardware-Intent Representation Layer for LLM-Driven RTL Generation | Tairan Cheng, Yi Liu, Dongsheng Zuo, Zhengyuan Shi, Hongji Zhang, Xiangfei Hu, Maoshuo He, Hao Yan, Qiang Xu | 2026-08-07 | 下载 | Generating implementation-quality RTL with large language models (LLMs) remains difficult because direct generation must resolve microarchitecture while simultaneously producing and debugging low-leve... |
| Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLM | Shixin Zhao, Lian Liu, Tianhua Han, Mengdi Wang, Yinhe Han, Ying Wang | 2026-08-07 | 下载 | Heterogeneous architectures that combine neural processing unit (NPU) and processing-in-memory (PIM) are increasingly adopted to accelerate LLM inference. |
| QCORE: A Quantum-Control-Oriented Real-Time Execution Architecture with Extensible Closed-Loop Services and Shared AI Acceleration | Heyue Li, Yanshu Guo, Qichun Liu, Tiefu Li, Zhihua Wang, Hanjun Jiang | 2026-08-07 | 下载 | Scalable quantum processors require control, readout, feedback, calibration, and error correction to coexist under bounded latency and shared-resource constraints, whereas existing platforms typically... |
| G-Power: Architecture-level GPU Power Modeling with Aggregated Knowledge Foundations from Known GPUs | Qijun Zhang, Yao Lu, Shang Liu, Mengming Li, Chen Zhang, Dongbo Wang, Zhiyao Xie | 2026-08-07 | 下载 | Graphics Processing Units (GPUs) have been serving as critical computation resources for large-scale parallel computations. With increasing chip complexity, power efficiency has become an important de... |
| HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation | Yuebo Luo, Ahmad Sedigh Baroughi, Philip Stachura, Le Chen, Venkatram Vishwanath, Zhenman Fang, Caiwen Ding | 2026-08-07 | 下载 | Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring months of specialized e... |
| Retention-Aware RISC-V ISA Extension and Memory Controller on FPGA for MLC NVM | Mina Ibrahim, Martel Shokry, Lokesh Siddhu, Lars Bauer, Hassan Nassar, Joerg Henkel | 2026-08-07 | 下载 | Non-volatile memory (NVM) technologies, particularly Multi-Level Cell (MLC) NVMs, offer significant potential for increasing memory density. MLC NVMs provide a tradeoff between write latency and reten... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Reduce Once, Verify Many: Verifying Isolation Guarantees via Hierarchical Abstractions | Shabnam Ghasemirad, Christoph Sprenger, Si Liu, David Basin | 2026-08-07 | 下载 | We present a mathematically rigorous, systematic approach for the verification of database isolation guarantees, which (i) supports a spectrum of seven isolation levels, (ii) uncovers a fundamental di... |
| LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs | Liad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson | 2026-08-07 | 下载 | Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems. |
| Communication-efficient parallel Bruhat decomposition | Ioanna Evtushevskaya-Konovalova, Alexander Tiskin | 2026-08-07 | 下载 | The model of bulk-synchronous parallel (BSP) computation is an emerging paradigm of general-purpose parallel computing. Bruhat decomposition is an important method in numerical linear algebra, general... |
| Fast end-to-end cloud application cold-start with initscripts | Ariel Szekely, Robert Morris, M. Frans Kaashoek | 2026-08-07 | 下载 | Serverless functions are a popular way of deploying cloud applications. Because many of these functions are short- running and experience frequent cold-starts, start latencies often dominate their exe... |
| Synthesizing Voltage Ride-Through Controllers for Data Centers | Wayne Wang, Archit Bhatnagar, Tongyuan Miao, Saniya Kalamkar, Wenqi Cui, Inigo Incer, Ang Chen | 2026-08-07 | 下载 | Data centers are among the power grid's fastest-growing loads. Since data center servers are sensitive electronic components, they need to be protected against the grid's voltage disturbances during g... |
| Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated Learning | Alireza Moayedikia, Alicia Troncoso Lora | 2026-08-07 | 下载 | Sub-model federated learning lets resource-constrained clients train width-reduced versions of a global model, but existing methods allocate capacity by device resources alone. |
| Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers | Xiangyu Meng, Peng Chen, Mingzhen Li, Jianmin Wang, Sen Wang, Guangming Tan, Weile Jia, Mohamed Wahib, Tao Luo, Xun Wang | 2026-08-07 | 下载 | Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swa... |
| HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache Management | Zhiqiang Xie, Zhangheng Huang, Tingwei Huang, Ziyi Xu, Ruiyang Ma, Christos Kozyrakis | 2026-08-07 | 下载 | Top-k sparse attention makes long-context LLM decoding cheap to compute: each step reads only a few thousand selected KV entries rather than the full context. |
| FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless Networks | Majid Kundroo, Tinku Singh, Taehong Kim | 2026-08-07 | 下载 | Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. However, efficient model convergence in FL remains challenging, especially in w... |
| A Kubernetes Scheduler Plugin for Cluster-Wide Placement Optimisation | Henrik Daniel Christensen, Saverio Giallorenzo, Jacopo Mauro | 2026-08-07 | 下载 | The default scheduler of Kubernetes, the state-of-the-art container orchestrator, uses fast, local placement decisions. Unfortunately, this design leads to resource fragmentation, reduced cluster usag... |
| A Self-Adaptive Extensible CEP framework for the Cloud-Edge Continuum | Olaf Markus Link, Sukanya Bhowmik | 2026-08-07 | 下载 | Traditional complex event processing (CEP) systems focus on extracting information out of simple events in the input event streams to detect complex event patterns. |
| Stream Learning: Partition-Fair Gossip Learning Without Tokens | Fabien Mathieu, Alexandre Pham, Maria Gradinariu Potop-Butucaru, S{é}bastien Tixeuil | 2026-08-07 | 下载 | In gossip learning, a network of nodes trains a shared model collaboratively, without a central coordinator, by repeatedly exchanging parts of their local models. |
| FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition | Ghani Haider, Majid Kundroo, Boyun Eom, Dong Hwan Park, Chen Chen, Taehong Kim | 2026-08-07 | 下载 | In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR). |
| StateFlow: Sequence Pipeline Parallelism for Long-Context Modeling with Linear Recurrence | Wenxuan Zhao, Yingfa Chen, Xu Han, Wenjing Han, Tianbo Huang, Zhiyu Li, Ao Sun, Jingheng Xu, Lin Gan, Guangwen Yang | 2026-08-07 | 下载 | Long-context training is increasingly important for large language models, and linear attention and state space models have become popular for improving long-context efficiency. |
| DGEMM with Ozaki Scheme I/II on FP4 Tensor Cores: A Base-13 E2M1 Limb Representation | Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri | 2026-08-07 | 下载 | This paper proposes a method and its implementation for emulating FP64 matrix multiplication (DGEMM) by constructing, on FP4 (E2M1; 2 exponent bits and 1 mantissa bit) Tensor Cores, Ozaki schemes I an... |
| CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit Weights | Xuetian Gao | 2026-08-07 | 下载 | Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU execution. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Coordinated Spectrum Coexistence Across Heterogeneous Commercial and Federal Services | Minh Dat Nguyen, Paolo Testolina, Pedram Johari, Michele Polese, Tommaso Melodia | 2026-08-07 | 下载 | Future wireless networks are expected to support the coexistence of cellular communications, radio frequency (RF) sensing, radionavigation, and radiolocation radar-among others-over congested federal ... |
| Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes | Simone Mainardi, Kaushal Bansal, Prabhat Singh | 2026-08-07 | 下载 | We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying, and can exceed supply. |
| FedSceneX: Time-to-Target Orchestration for Same-Scene Multimodal Federated Edge Learning | Dhe Yeong Tchalla, Beining Wu, Jun Huang, Shuyang Gu, Qiang Duan | 2026-08-07 | 下载 | Federated learning at the sensing edge is typically evaluated by communication rounds, yet a round does not represent a fixed amount of work. Even on identical hardware, the methods we compare require... |
| An Analysis of Architectural and Operational Dynamics of Phishkits in the Wild | Behzad Ousat, Mohammad Ali Tofighi, Estefan Schafir, Amin Kharraz | 2026-08-07 | 下载 | Phishing attacks have always been a favored vector for adversaries to defraud users, bypass modern defense mechanisms, and penetrate critical systems. |
| LYRA: Label-Free Structural Synchronization and Resource Allocation for UAV Edge Networks | Feng He, Alireza Furutanpey, Paolo Bellavista, Yu Qiu, Jiangchuan Liu, Jiannong Cao, Schahram Dustdar | 2026-08-07 | 下载 | While deploying hierarchical vision models to process mission-critical tasks, UAV edge systems must adaptively update the models to sustain inference reliability under low-level environmental corrupti... |
| Rate-Fidelity Control for Wide-Area Quantum Links | Connor Clayton, Cory Nunn, Quinn Carmack, Wayne McKenzie, Anne Marie Richards, Xiaodi Wu, Bobby Bhattacharjee | 2026-08-07 | 下载 | Quantum network links must distribute entanglement at high rates while satisfying application-specified fidelity demands. However, wide-area deployed fiber links suffer from polarization drift which d... |
| EvoRIC: Reinforcement Learning Fine-Tuned LLM-empowered RAN Intelligent Control Toward Autonomous O-RAN | Lingyan Bao, Jemin Lee, Tony Q. S. Quek | 2026-08-07 | 下载 | Despite recent advances in applying artificial intelligence (AI) techniques to radio access network (RAN), critical challenges remain: traditional machine learning (ML) algorithms suffer from limited ... |
| A Parameter-Specific Retrieval and Knowledge-Guided Reasoning Framework for LLM-Based GPSR Optimization in FANETs | Zhipeng Lin, Bin Duo, Tong Liu, Jie Lin, Jianting Yuan, Xiaojun Yuan | 2026-08-07 | 下载 | Existing Greedy Perimeter Stateless Routing (GPSR)-based protocols for Flying Ad-Hoc Networks (FANETs) struggle to adapt routing parameters, such as hello interval, multi-path number, and greedy forwa... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Fast end-to-end cloud application cold-start with initscripts | Ariel Szekely, Robert Morris, M. Frans Kaashoek | 2026-08-07 | 下载 | Serverless functions are a popular way of deploying cloud applications. Because many of these functions are short- running and experience frequent cold-starts, start latencies often dominate their exe... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Classical Models Match or Exceed Shallow Variational Quantum Circuits on Vision Benchmarks | Christopher Fulton, Irene Tsapara, Lawrence Fulton | 2026-08-07 | 下载 | Quaternion-valued neural networks and variational quantum circuits (VQCs) both derive local transformations from geometry, yet their performance on classical supervised learning remai... |
| LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs | Liad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson | 2026-08-07 | 下载 | Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems. |
| A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy | Bhavika Jalli, Nikhil Korati Prasanna, Jayanta Choudhury | 2026-08-07 | 下载 | LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (N... |
| The Token Efficiency Index: A Peer-Benchmarked Composite Indicator for AI Token Efficiency | Caden Wong, Vikram Das, Himanshu Dhami | 2026-08-07 | 下载 | As artificial intelligence (AI) adoption accelerates across tech giants, AI-native startups, and non-technical organizations alike, a deceptively simple question remains hard to answer: is that spendi... |
| Aneto: Predicting System Performance by Exploiting Cross-Workload Regularity | Raul Taranco, Rene Mueller, Michael Giardino | 2026-08-07 | 下载 | Predicting how a workload responds to a change in memory technology requires estimating how much of each cache miss actually stalls the processor. |
| Thermodynamic Human-Computer Interaction | Uzafir Ahmad Rafaq, Muaz Hassan, Ali Muzaffar | 2026-08-07 | 下载 | Traditional human-computer interaction models rely on domain-specific techniques to model target prediction; models designed for cursor interaction prediction fail to generalize to mobile interfaces a... |