Skip to content

2026-08-12 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
GateTruth: Auditing the Rigor of RTL Design Benchmarks via Mutation TestingMeet Bhadra2026-08-12下载Benchmarks for evaluating large language models on register-transfer-level (RTL) hardware design have proliferated rapidly, yet none reports having applied mutation testing, an established hardware-ve...
Lonic: Algorithm-Hardware Co-Design for Energy-Efficient Fully Local Online SNN Training with INT4 PrecisionPeilin Chen, Xiaoxuan Yang2026-08-12下载Spiking neural networks (SNNs) have recently attracted increasing attention as an energy-efficient learning paradigm. Existing works also propose temporally and fully local online SNN training algorit...
FQTree: Fine-grained Quantization and Hardware Generation of Boosted Decision TreesZhiqiang Que, Chang Sun, Haiyang Wang, Dinesh Pamunuwa, Roshan Weerasekera, Qijia Tang, Bakhtiar Zadeh, Wayne Luk, Maria Spiropulu2026-08-12下载Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or manually tuned fixed-po...
NITRO: High-Performance 3D NAND Flash-Based In-Storage Computing with Enhanced Activation DataflowSanghun Shin, Sangyeon Kim, Gisan Ji, Sungju Ryu2026-08-12下载In-storage computing (ISC) is considered a next-generation memory architecture for its ability to relieve the data bottleneck between the host and the memory.
Do Not Let CNOTs Overwhelm the Decoder: Scheduling Transversal Gates for Fast FTQCShota Ikari, Yuga Hirai, Yasunari Suzuki, Hiroshi Nakamura, Yosuke Ueno2026-08-12下载Transversal CNOT (TCNOT) gates can accelerate fault-tolerant quantum computation (FTQC) in the surface code by reducing the number of syndrome extraction rounds required between logical operations fro...
Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell UltraTeng-Ruei Chen2026-08-12下载NVIDIA's published specifications give the Blackwell Ultra GPU (B300) a dense-compute ratio of roughly 30:1 between FP8 and INT8 tensor-core throughput; its predecessors, H200 and B200, both provide 1...
APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE InferenceAlish Kanani, Layan Badawi, Umit Y. Ogras2026-08-12下载Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency.
HBF Sucks! A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM ServingZhuoran Li, Zhuohang Bian, Xin Huang, Yibo Zhao, Guangyu Sun, Youwei Zhuo2026-08-12下载A faster storage device should make serving faster. We find the opposite. High-Bandwidth Flash (HBF) stacks NAND behind a wide, package-local interface, promising flash-scale capacity with far lower r...
Uni-SFU: Algorithm-HW Co-Design for Universal SFUs via Mixed-Degree Piecewise ApproximationMiao Sun, Yucheng Huang, Mingcong Cao, Jaehyun Park, Partha Pratim Pande, Umit Y. Ogras2026-08-12下载Nonlinear activation functions are essential to modern deep neural networks (DNNs), but their hardware evaluation places significant pressure on the special-function units (SFUs) of GPUs and custom ac...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Offering Microsecond-Scale Cross-VM Core Elasticity on Colocated Lightweight Virtual MachinesYibo Yan, Seo Jin Park2026-08-12下载Serverless platforms commonly colocate many diverse workloads, each in a fast-booting, memory-lean virtual machine (VM), to improve deployment density.
RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement LearningYibo Shen, Xudong Han, Xiaowei Zhu, Gen Li, Zhenxuan Pan2026-08-12下载Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, whi...
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent ControlJosef Liyanjun Chen2026-08-12下载LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect.
Evaluating OpenMP Offloading for Intra-node Multi-GPU Programming across NVIDIA, AMD, and Intel Architectures: A 3D Heat Transfer Case StudyEzhilmathi Krishnasamy2026-08-12下载Currently, most supercomputers are equipped with GPUs from manufacturers such as NVIDIA, AMD, or Intel, which provide substantial parallelism and high throughput.
User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware AutoscalingAlfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta2026-08-12下载Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand.
Enabling Differentiated QoS Degradation for Replicated Databases under FailuresBelkis Djeffal, Pierre Bourhis, Romain Rouvoy2026-08-12下载Elasticity is commonly presented as the default response to capacity loss after failures, since replacement replicas can compensate for failed nodes and restore pre-incident service levels.
Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection PipelinesVaishnav Raju2026-08-12下载Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest...
Descriptive Dispatch of Computational WorkVanessa Sochat, Daniel Milroy2026-08-12下载Agents powered by AI/ML are becoming ingrained in orchestration. Dispatch of work is the task of receiving a request, transforming it for a workload manager, and successfully submitting it.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
DualPI2 Active Queue Management in ns-3: Implementation And ValidationMaria Eduarda Veras, Eduardo Freitas, Assis T. de Oliveira Filho, Djamel Sadok, Judith Kelner2026-08-12下载The demand for ultra-low latency applications necessitates advanced network architectures like the Low Latency, Low Loss, and Scalable Throughput (L4S) standard.
Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning ApproachJiaao Ma, Chuan Lin, Guangjie Han, Shengchao Zhu, Qian Zhu, Ying Liu, Zhenyu Wang2026-08-12下载Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic top...
On the Allocation of Transmit Power for Coordinated Spatial Reuse in IEEE 802.11bn Multi-Access Point CoordinationFrancesc Wilhelmi, Boris Bellalta2026-08-12下载IEEE 802.11bn (11bn) introduces Coordinated Spatial Reuse (Co-SR), a Multi-AP Coordination (MAPC) scheme in which two Access Points (APs) coordinate to control the transmit power for a simultaneous tr...
User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware AutoscalingAlfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta2026-08-12下载Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand.
FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecastingRentao Gu, Yihang Ding, Junjie Li, Yi Ding, Weijing Sang, Xiaoli Huo, Xin Qin, Yuefeng Ji2026-08-12下载Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods rely heavily on textual prompts for modality alignment-introdu...
Integrated Sensing and Communication in 3GPP: Evolution from 5G-Advanced to 6GXingqin Lin2026-08-12下载Integrated sensing and communication (ISAC) extends mobile networks from information transfer toward perception of passive objects and environments.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent ControlJosef Liyanjun Chen2026-08-12下载LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect.
The Ingestion Tax: Adopting File-Backed Weights in Tensor FrameworksYuan Si, Yufeng Lin, Daming Li, Jialu Zhang2026-08-12下载Open-weight models can occupy a middle capacity regime: active weights fit in DRAM as cached file pages, but a second framework-owned representation does not fit or must be refilled as layers run, so ...
Who Should Own the Expert Cache? Kernel-Managed Tiering for Trillion-Parameter MoE InferenceYuan Si, Yufeng Lin, Daming Li, Jialu Zhang2026-08-12下载Mixture-of-experts models whose expert pools dwarf DRAM force every serving system to contain a cache, yet existing systems typically implement this cache in user space using expert-granular, frequenc...

cs.PF - Performance ​

标题作者发布日期PDF摘要
User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware AutoscalingAlfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta2026-08-12下载Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand.
Testing the EPYC Conjecture on Real Hardware: MoA-Guided Dense Matrix Multiplication on NCSA Delta (AMD EPYC 7763 Milan)Lenore M Mullin2026-08-12下载A companion empirical study conjectured that MoA-guided dense matrix multiplication would need per-CCD recalibration on AMD EPYC Bergamo; allocation access to that machine was declined, and this paper...

基于 VitePress 构建 · 使用本地搜索查找论文