Skip to content

2026-07-08 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
ATLAS: Automated HLS for DL-Optimized FPGAsRuthwik Reddy Sunketa, Aman Arora2026-07-08下载FPGA architectures increasingly incorporate domain-specific in-fabric hardblocks to accelerate DL inference, particularly GEMM, which dominates DL computation.
Embedded Blockchain Infrastructure Management (eBIM): A RISC-V-Empowered Hardware--Software Co-Design Framework Towards Trustworthy BlockchainQinglin Yang, Yuan Liu, Yaoyao Zhang, Boya Wang, Zongjian You, Chunming Rong, Zhihong Tian2026-07-08下载Blockchain systems are undergoing a fundamental transition from decentralized ledgers for digital assets to general-purpose trust infrastructures for verifiable computation, decentralized physical res...
Vectorizing Quantum Control: A RISC-V Vector Extension Architecture for Scalable Qubit SystemsXiaorang Guo, Kun Qin, Yanbin Chen, Carsten Trinitis, Martin Schulz2026-07-08下载The Quantum Control Processor (QCP) bridges the gap between compiler toolchains and control electronics, and is responsible for translating compiled quantum circuits into executable instructions that ...
Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvencySatoshi Matsuoka2026-07-08下载We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-...
Miter-Aware LUT Mapping: Aligning Structure and Solvability for Efficient Logic Equivalence CheckingJiaying Zhu, Zhengyuan Shi, Mengxia Tao, Kezhi Li, Min Li, Qiang Xu2026-07-08下载Logic Equivalence Checking (LEC), a fundamental hardware verification task, is often bottlenecked by synthesis-induced structural perturbations and XOR-dense regions that degrade SAT solver performanc...
ThermoDSE: A Thermal-Aware and Comprehensive Design Space Exploration for Chiplet-Based DNN AcceleratorsJian Peng, Hanwei Fan, Jingbo Jiang, Lin Jiang, Wei Zhang2026-07-08下载Chiplet-based DNN accelerators provide a scalable path to balance performance and yield for modern AI workloads. However, such systems face critical challenges in area and thermal constraints.
EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAIHao Kong, Di Liu, Shuo Huai, Xiangzhong Luo, Ravi Subramaniam, Christian Makaya, Qian Lin, Weichen Liu2026-07-08下载Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resour...
Smart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded HardwareHao Kong, Di Liu, Shuo Huai, Xiangzhong Luo, Weichen Liu, Ravi Subramaniam, Christian Makaya, Qian Lin2026-07-08下载Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU SystemsTingkai Liu, Muralidhar Andoorveedu, Sanjoy Das, Sanjay Patel, Volodymyr Kindratenko2026-07-08下载The evolution of compute infrastructure has transformed multi-GPU systems into tightly integrated shared-memory structures. However, current software still mostly treats these coherent interconnects s...
Scaling WaterLily.jl with MPI and an improved geometric multigrid solverBernat Font, Marin Lauber, Tzu-Yao Huang, Gabriel D. Weymouth2026-07-08下载We present recent performance-oriented developments in WaterLily, a scale-resolving incompressible flow solver written in pure Julia that runs seamlessly on CPUs and GPUs of any vendor.
GIFT: Geometry-Informed Low-precision Gradient Communication for LLM PretrainingJieying Wang, Shuyuan Fan, Mingkai Zheng, Zhao Zhang2026-07-08下载Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can significantly reduce the ...
POO-LPSP: Parallel Osprey Optimized Least Penalty-Squared Prioritization Methods for Priority Derivation in the Analytic Hierarchy ProcessKevin Kam Fung Yuen2026-07-08下载Pairwise comparison (PC) via pairwise reciprocal matrices (PRMs) is central to the Analytic Hierarchy Process (AHP). Although the traditional eigenvector method is widely applied to derive priorities,...
Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in ProductionArun Malik2026-07-08下载AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously solved problems.
Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge IntelligenceChanwoo Cho, Wooseok Kim, Yonglak Son, Young Seo Lee, Young Geun Kim2026-07-08下载Large language models (LLMs) are widely used in intelligent services due to their remarkable capability in generative tasks. Typically, LLM-based services process the inference requests of the users i...
Multiple Double Arithmetic on NVIDIA Tensor CoresHoward Chen, Jan Verschelde2026-07-08下载A multiple double is an unevaluated sum of doubles. An NVIDIA tensor core is a specialized high performance compute core for matrix multiplication.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
PHaul: A PPO-based forwarding agent for Sub6 enhanced Integrated Access and Backhaul networksJorge Pueyo, Daniel Camps-Mur and, Miguel Catalan-Cid2026-07-08下载3GPP Integrated Access and Backhaul (IAB) allows operators to deploy outdoor mm-wave access networks in a cost-efficient manner, by reusing the same spectrum in access and backhaul.
Small Language Model-based Control for BBR over Low Earth Orbit Satellite InternetRakshitha De Silva, Shiva Raj Pokhrel, Jonathan Kua2026-07-08下载Low Earth Orbit (LEO) satellite Internet introduces rapid path variability, intermittent capacity shifts, and non-terrestrial delay dynamics that challenge transport-layer congestion control.
Unveiling TCP BBR Dominance in Starlink Internet: Experimental Insights and AnalysisRakshitha De Silva, Shiva Raj Pokhrel, Jonathan Kua2026-07-08下载This experimental study delivers a global assessment of Google's Bottleneck Bandwidth and Round-trip propagation time-version 3 (BBR-v3) Congestion Control Algorithm (CCA) over SpaceX's Starlink netwo...
How the Fusion of Onboard Sensors and V2X Data can Improve (or not) the Cooperative Perception of Connected Automated VehiclesAmir Mohammadisarab, Miguel Sepulcre, Luca Lusvarghi, Javier Gozalvez2026-07-08下载Automated vehicles rely on onboard sensors to perceive their surroundings and navigate autonomously. However, sensor performance may degrade under adverse weather conditions or when line-of-sight is o...
Bessel Beam Optimization for Near-Field THz Communications under UE Location UncertaintyAditya Jolly, Vitaly Petrov, Gábor Fodor, Emil Björnson2026-07-08下载To achieve the desired coverage and capacity levels, future terahertz (THz) wireless systems are envisioned to utilize extremely large antenna arrays.
EvoOMG: An Evolution-Oriented Multi-Agent Guidance Framework for Heterogeneous Legacy-and-MLO Wi-Fi NetworksJunjie Wu, Lingjian Zhou, Zerui Shao, Yi Zou, Tianrui Li, Yi Zhang, Ziyuan Yang2026-07-08下载The gradual deployment of Wi-Fi 7/8 multi-link operation (MLO) will lead to long-term coexistence between legacy non-MLO stations (STAs) and MLO-capable STAs in WLANs.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvencySatoshi Matsuoka2026-07-08下载We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-...

基于 VitePress 构建 · 使用本地搜索查找论文