Skip to content

2026-08-07 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICsLiad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson2026-08-07下载Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems.
Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-TuningVasanth Iyer2026-08-07下载Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class accelerators remains limited.
HINT: Toward an Executable Hardware-Intent Representation Layer for LLM-Driven RTL GenerationTairan Cheng, Yi Liu, Dongsheng Zuo, Zhengyuan Shi, Hongji Zhang, Xiangfei Hu, Maoshuo He, Hao Yan, Qiang Xu2026-08-07下载Generating implementation-quality RTL with large language models (LLMs) remains difficult because direct generation must resolve microarchitecture while simultaneously producing and debugging low-leve...
Rethinking Unified Memory for NPU-PIM Systems: Dual-View Memory for Dynamic Inference of LLMShixin Zhao, Lian Liu, Tianhua Han, Mengdi Wang, Yinhe Han, Ying Wang2026-08-07下载Heterogeneous architectures that combine neural processing unit (NPU) and processing-in-memory (PIM) are increasingly adopted to accelerate LLM inference.
QCORE: A Quantum-Control-Oriented Real-Time Execution Architecture with Extensible Closed-Loop Services and Shared AI AccelerationHeyue Li, Yanshu Guo, Qichun Liu, Tiefu Li, Zhihua Wang, Hanjun Jiang2026-08-07下载Scalable quantum processors require control, readout, feedback, calibration, and error correction to coexist under bounded latency and shared-resource constraints, whereas existing platforms typically...
G-Power: Architecture-level GPU Power Modeling with Aggregated Knowledge Foundations from Known GPUsQijun Zhang, Yao Lu, Shang Liu, Mengming Li, Chen Zhang, Dongbo Wang, Zhiyao Xie2026-08-07下载Graphics Processing Units (GPUs) have been serving as critical computation resources for large-scale parallel computations. With increasing chip complexity, power efficiency has become an important de...
HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS TranslationYuebo Luo, Ahmad Sedigh Baroughi, Philip Stachura, Le Chen, Venkatram Vishwanath, Zhenman Fang, Caiwen Ding2026-08-07下载Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but developing them is costly, often requiring months of specialized e...
Retention-Aware RISC-V ISA Extension and Memory Controller on FPGA for MLC NVMMina Ibrahim, Martel Shokry, Lokesh Siddhu, Lars Bauer, Hassan Nassar, Joerg Henkel2026-08-07下载Non-volatile memory (NVM) technologies, particularly Multi-Level Cell (MLC) NVMs, offer significant potential for increasing memory density. MLC NVMs provide a tradeoff between write latency and reten...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Reduce Once, Verify Many: Verifying Isolation Guarantees via Hierarchical AbstractionsShabnam Ghasemirad, Christoph Sprenger, Si Liu, David Basin2026-08-07下载We present a mathematically rigorous, systematic approach for the verification of database isolation guarantees, which (i) supports a spectrum of seven isolation levels, (ii) uncovers a fundamental di...
LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICsLiad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson2026-08-07下载Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems.
Communication-efficient parallel Bruhat decompositionIoanna Evtushevskaya-Konovalova, Alexander Tiskin2026-08-07下载The model of bulk-synchronous parallel (BSP) computation is an emerging paradigm of general-purpose parallel computing. Bruhat decomposition is an important method in numerical linear algebra, general...
Fast end-to-end cloud application cold-start with initscriptsAriel Szekely, Robert Morris, M. Frans Kaashoek2026-08-07下载Serverless functions are a popular way of deploying cloud applications. Because many of these functions are short- running and experience frequent cold-starts, start latencies often dominate their exe...
Synthesizing Voltage Ride-Through Controllers for Data CentersWayne Wang, Archit Bhatnagar, Tongyuan Miao, Saniya Kalamkar, Wenqi Cui, Inigo Incer, Ang Chen2026-08-07下载Data centers are among the power grid's fastest-growing loads. Since data center servers are sensitive electronic components, they need to be protected against the grid's voltage disturbances during g...
Capacity Confounds and Coverage Guarantees in Adaptive Sub-model Federated LearningAlireza Moayedikia, Alicia Troncoso Lora2026-08-07下载Sub-model federated learning lets resource-constrained clients train width-reduced versions of a global model, but existing methods allocate capacity by device resources alone.
Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated SupercomputersXiangyu Meng, Peng Chen, Mingzhen Li, Jianmin Wang, Sen Wang, Guangming Tan, Weile Jia, Mohamed Wahib, Tao Luo, Xun Wang2026-08-07下载Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale. Among existing approaches, LightDock leverages Glowworm Swa...
HiSparse: Scaling Sparse-Attention Decoding with Hierarchical KV Cache ManagementZhiqiang Xie, Zhangheng Huang, Tingwei Huang, Ziyi Xu, Ruiyang Ma, Christos Kozyrakis2026-08-07下载Top-k sparse attention makes long-context LLM decoding cheap to compute: each step reads only a few thousand selected KV entries rather than the full context.
FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless NetworksMajid Kundroo, Tinku Singh, Taehong Kim2026-08-07下载Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. However, efficient model convergence in FL remains challenging, especially in w...
A Kubernetes Scheduler Plugin for Cluster-Wide Placement OptimisationHenrik Daniel Christensen, Saverio Giallorenzo, Jacopo Mauro2026-08-07下载The default scheduler of Kubernetes, the state-of-the-art container orchestrator, uses fast, local placement decisions. Unfortunately, this design leads to resource fragmentation, reduced cluster usag...
A Self-Adaptive Extensible CEP framework for the Cloud-Edge ContinuumOlaf Markus Link, Sukanya Bhowmik2026-08-07下载Traditional complex event processing (CEP) systems focus on extracting information out of simple events in the input event streams to detect complex event patterns.
Stream Learning: Partition-Fair Gossip Learning Without TokensFabien Mathieu, Alexandre Pham, Maria Gradinariu Potop-Butucaru, S{é}bastien Tixeuil2026-08-07下载In gossip learning, a network of nodes trains a shared model collaboratively, without a central coordinator, by repeatedly exchanging parts of their local models.
FedVAR: Prototype-Aligned Federated Framework for Video Anomaly RecognitionGhani Haider, Majid Kundroo, Boyun Eom, Dong Hwan Park, Chen Chen, Taehong Kim2026-08-07下载In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentralized intelligence paradigm for Video Anomaly Recognition (VAR).
StateFlow: Sequence Pipeline Parallelism for Long-Context Modeling with Linear RecurrenceWenxuan Zhao, Yingfa Chen, Xu Han, Wenjing Han, Tianbo Huang, Zhiyu Li, Ao Sun, Jingheng Xu, Lin Gan, Guangwen Yang2026-08-07下载Long-context training is increasingly important for large language models, and linear attention and state space models have become popular for improving long-context efficiency.
DGEMM with Ozaki Scheme I/II on FP4 Tensor Cores: A Base-13 E2M1 Limb RepresentationShun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri2026-08-07下载This paper proposes a method and its implementation for emulating FP64 matrix multiplication (DGEMM) by constructing, on FP4 (E2M1; 2 exponent bits and 1 mantissa bit) Tensor Cores, Ozaki schemes I an...
CubicQuant: Parametric Non-Uniform Codebooks for High-Throughput LLM Inference with 1-8-Bit WeightsXuetian Gao2026-08-07下载Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU execution.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Coordinated Spectrum Coexistence Across Heterogeneous Commercial and Federal ServicesMinh Dat Nguyen, Paolo Testolina, Pedram Johari, Michele Polese, Tommaso Melodia2026-08-07下载Future wireless networks are expected to support the coexistence of cellular communications, radio frequency (RF) sensing, radionavigation, and radiolocation radar-among others-over congested federal ...
Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service ClassesSimone Mainardi, Kaushal Bansal, Prabhat Singh2026-08-07下载We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying, and can exceed supply.
FedSceneX: Time-to-Target Orchestration for Same-Scene Multimodal Federated Edge LearningDhe Yeong Tchalla, Beining Wu, Jun Huang, Shuyang Gu, Qiang Duan2026-08-07下载Federated learning at the sensing edge is typically evaluated by communication rounds, yet a round does not represent a fixed amount of work. Even on identical hardware, the methods we compare require...
An Analysis of Architectural and Operational Dynamics of Phishkits in the WildBehzad Ousat, Mohammad Ali Tofighi, Estefan Schafir, Amin Kharraz2026-08-07下载Phishing attacks have always been a favored vector for adversaries to defraud users, bypass modern defense mechanisms, and penetrate critical systems.
LYRA: Label-Free Structural Synchronization and Resource Allocation for UAV Edge NetworksFeng He, Alireza Furutanpey, Paolo Bellavista, Yu Qiu, Jiangchuan Liu, Jiannong Cao, Schahram Dustdar2026-08-07下载While deploying hierarchical vision models to process mission-critical tasks, UAV edge systems must adaptively update the models to sustain inference reliability under low-level environmental corrupti...
Rate-Fidelity Control for Wide-Area Quantum LinksConnor Clayton, Cory Nunn, Quinn Carmack, Wayne McKenzie, Anne Marie Richards, Xiaodi Wu, Bobby Bhattacharjee2026-08-07下载Quantum network links must distribute entanglement at high rates while satisfying application-specified fidelity demands. However, wide-area deployed fiber links suffer from polarization drift which d...
EvoRIC: Reinforcement Learning Fine-Tuned LLM-empowered RAN Intelligent Control Toward Autonomous O-RANLingyan Bao, Jemin Lee, Tony Q. S. Quek2026-08-07下载Despite recent advances in applying artificial intelligence (AI) techniques to radio access network (RAN), critical challenges remain: traditional machine learning (ML) algorithms suffer from limited ...
A Parameter-Specific Retrieval and Knowledge-Guided Reasoning Framework for LLM-Based GPSR Optimization in FANETsZhipeng Lin, Bin Duo, Tong Liu, Jie Lin, Jianting Yuan, Xiaojun Yuan2026-08-07下载Existing Greedy Perimeter Stateless Routing (GPSR)-based protocols for Flying Ad-Hoc Networks (FANETs) struggle to adapt routing parameters, such as hello interval, multi-path number, and greedy forwa...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Fast end-to-end cloud application cold-start with initscriptsAriel Szekely, Robert Morris, M. Frans Kaashoek2026-08-07下载Serverless functions are a popular way of deploying cloud applications. Because many of these functions are short- running and experience frequent cold-starts, start latencies often dominate their exe...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Classical SU(2)\mathrm{SU}(2) Models Match or Exceed Shallow Variational Quantum Circuits on Vision BenchmarksChristopher Fulton, Irene Tsapara, Lawrence Fulton2026-08-07下载Quaternion-valued neural networks and variational quantum circuits (VQCs) both derive local transformations from SU(2)\mathrm{SU}(2) geometry, yet their performance on classical supervised learning remai...
LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICsLiad Gerstman, Aditya Dhakal, Dejan Milojicic, Avi Mendelson2026-08-07下载Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems.
A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving AccuracyBhavika Jalli, Nikhil Korati Prasanna, Jayanta Choudhury2026-08-07下载LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom network analytics and numerical time-series data analysis (N...
The Token Efficiency Index: A Peer-Benchmarked Composite Indicator for AI Token EfficiencyCaden Wong, Vikram Das, Himanshu Dhami2026-08-07下载As artificial intelligence (AI) adoption accelerates across tech giants, AI-native startups, and non-technical organizations alike, a deceptively simple question remains hard to answer: is that spendi...
Aneto: Predicting System Performance by Exploiting Cross-Workload RegularityRaul Taranco, Rene Mueller, Michael Giardino2026-08-07下载Predicting how a workload responds to a change in memory technology requires estimating how much of each cache miss actually stalls the processor.
Thermodynamic Human-Computer InteractionUzafir Ahmad Rafaq, Muaz Hassan, Ali Muzaffar2026-08-07下载Traditional human-computer interaction models rely on domain-specific techniques to model target prediction; models designed for cursor interaction prediction fail to generalize to mobile interfaces a...

基于 VitePress 构建 · 使用本地搜索查找论文