2026-08-12
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation Testing | Meet Bhadra | 2026-08-12 | 下载 | Benchmarks for evaluating large language models on register-transfer-level (RTL) hardware design have proliferated rapidly, yet none reports having applied mutation testing, an established hardware-ve... |
| Lonic: Algorithm-Hardware Co-Design for Energy-Efficient Fully Local Online SNN Training with INT4 Precision | Peilin Chen, Xiaoxuan Yang | 2026-08-12 | 下载 | Spiking neural networks (SNNs) have recently attracted increasing attention as an energy-efficient learning paradigm. Existing works also propose temporally and fully local online SNN training algorit... |
| FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision Trees | Zhiqiang Que, Chang Sun, Haiyang Wang, Dinesh Pamunuwa, Roshan Weerasekera, Qijia Tang, Bakhtiar Zadeh, Wayne Luk, Maria Spiropulu | 2026-08-12 | 下载 | Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or manually tuned fixed-po... |
| NITRO: High-Performance 3D NAND Flash-Based In-Storage Computing with Enhanced Activation Dataflow | Sanghun Shin, Sangyeon Kim, Gisan Ji, Sungju Ryu | 2026-08-12 | 下载 | In-storage computing (ISC) is considered a next-generation memory architecture for its ability to relieve the data bottleneck between the host and the memory. |
| Do Not Let CNOTs Overwhelm the Decoder: Scheduling Transversal Gates for Fast FTQC | Shota Ikari, Yuga Hirai, Yasunari Suzuki, Hiroshi Nakamura, Yosuke Ueno | 2026-08-12 | 下载 | Transversal CNOT (TCNOT) gates can accelerate fault-tolerant quantum computation (FTQC) in the surface code by reducing the number of syndrome extraction rounds required between logical operations fro... |
| Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra | Teng-Ruei Chen | 2026-08-12 | 下载 | NVIDIA's published specifications give the Blackwell Ultra GPU (B300) a dense-compute ratio of roughly 30:1 between FP8 and INT8 tensor-core throughput; its predecessors, H200 and B200, both provide 1... |
| APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference | Alish Kanani, Layan Badawi, Umit Y. Ogras | 2026-08-12 | 下载 | Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. |
| HBF Sucks! A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM Serving | Zhuoran Li, Zhuohang Bian, Xin Huang, Yibo Zhao, Guangyu Sun, Youwei Zhuo | 2026-08-12 | 下载 | A faster storage device should make serving faster. We find the opposite. High-Bandwidth Flash (HBF) stacks NAND behind a wide, package-local interface, promising flash-scale capacity with far lower r... |
| Uni-SFU: Algorithm-HW Co-Design for Universal SFUs via Mixed-Degree Piecewise Approximation | Miao Sun, Yucheng Huang, Mingcong Cao, Jaehyun Park, Partha Pratim Pande, Umit Y. Ogras | 2026-08-12 | 下载 | Nonlinear activation functions are essential to modern deep neural networks (DNNs), but their hardware evaluation places significant pressure on the special-function units (SFUs) of GPUs and custom ac... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Offering Microsecond-Scale Cross-VM Core Elasticity on Colocated Lightweight Virtual Machines | Yibo Yan, Seo Jin Park | 2026-08-12 | 下载 | Serverless platforms commonly colocate many diverse workloads, each in a fast-booting, memory-lean virtual machine (VM), to improve deployment density. |
| RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning | Yibo Shen, Xudong Han, Xiaowei Zhu, Gen Li, Zhenxuan Pan | 2026-08-12 | 下载 | Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, whi... |
| Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control | Josef Liyanjun Chen | 2026-08-12 | 下载 | LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. |
| Evaluating OpenMP Offloading for Intra-node Multi-GPU Programming across NVIDIA, AMD, and Intel Architectures: A 3D Heat Transfer Case Study | Ezhilmathi Krishnasamy | 2026-08-12 | 下载 | Currently, most supercomputers are equipped with GPUs from manufacturers such as NVIDIA, AMD, or Intel, which provide substantial parallelism and high throughput. |
| User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling | Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta | 2026-08-12 | 下载 | Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. |
| Enabling Differentiated QoS Degradation for Replicated Databases under Failures | Belkis Djeffal, Pierre Bourhis, Romain Rouvoy | 2026-08-12 | 下载 | Elasticity is commonly presented as the default response to capacity loss after failures, since replacement replicas can compensate for failed nodes and restore pre-incident service levels. |
| Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines | Vaishnav Raju | 2026-08-12 | 下载 | Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest... |
| Descriptive Dispatch of Computational Work | Vanessa Sochat, Daniel Milroy | 2026-08-12 | 下载 | Agents powered by AI/ML are becoming ingrained in orchestration. Dispatch of work is the task of receiving a request, transforming it for a workload manager, and successfully submitting it. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| DualPI2 Active Queue Management in ns-3: Implementation And Validation | Maria Eduarda Veras, Eduardo Freitas, Assis T. de Oliveira Filho, Djamel Sadok, Judith Kelner | 2026-08-12 | 下载 | The demand for ultra-low latency applications necessitates advanced network architectures like the Low Latency, Low Loss, and Scalable Throughput (L4S) standard. |
| Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach | Jiaao Ma, Chuan Lin, Guangjie Han, Shengchao Zhu, Qian Zhu, Ying Liu, Zhenyu Wang | 2026-08-12 | 下载 | Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic top... |
| On the Allocation of Transmit Power for Coordinated Spatial Reuse in IEEE 802.11bn Multi-Access Point Coordination | Francesc Wilhelmi, Boris Bellalta | 2026-08-12 | 下载 | IEEE 802.11bn (11bn) introduces Coordinated Spatial Reuse (Co-SR), a Multi-AP Coordination (MAPC) scheme in which two Access Points (APs) coordinate to control the transmit power for a simultaneous tr... |
| User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling | Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta | 2026-08-12 | 下载 | Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. |
| FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting | Rentao Gu, Yihang Ding, Junjie Li, Yi Ding, Weijing Sang, Xiaoli Huo, Xin Qin, Yuefeng Ji | 2026-08-12 | 下载 | Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introdu... |
| Integrated Sensing and Communication in 3GPP: Evolution from 5G-Advanced to 6G | Xingqin Lin | 2026-08-12 | 下载 | Integrated sensing and communication (ISAC) extends mobile networks from information transfer toward perception of passive objects and environments. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control | Josef Liyanjun Chen | 2026-08-12 | 下载 | LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. |
| The Ingestion Tax: Adopting File-Backed Weights in Tensor Frameworks | Yuan Si, Yufeng Lin, Daming Li, Jialu Zhang | 2026-08-12 | 下载 | Open-weight models can occupy a middle capacity regime: active weights fit in DRAM as cached file pages, but a second framework-owned representation does not fit or must be refilled as layers run, so ... |
| Who Should Own the Expert Cache? Kernel-Managed Tiering for Trillion-Parameter MoE Inference | Yuan Si, Yufeng Lin, Daming Li, Jialu Zhang | 2026-08-12 | 下载 | Mixture-of-experts models whose expert pools dwarf DRAM force every serving system to contain a cache, yet existing systems typically implement this cache in user space using expert-granular, frequenc... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling | Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta | 2026-08-12 | 下载 | Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. |
| Testing the EPYC Conjecture on Real Hardware: MoA-Guided Dense Matrix Multiplication on NCSA Delta (AMD EPYC 7763 Milan) | Lenore M Mullin | 2026-08-12 | 下载 | A companion empirical study conjectured that MoA-guided dense matrix multiplication would need per-CCD recalibration on AMD EPYC Bergamo; allocation access to that machine was declined, and this paper... |