2026-07-08
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ATLAS: Automated HLS for DL-Optimized FPGAs | Ruthwik Reddy Sunketa, Aman Arora | 2026-07-08 | 下载 | FPGA architectures increasingly incorporate domain-specific in-fabric hardblocks to accelerate DL inference, particularly GEMM, which dominates DL computation. |
| Embedded Blockchain Infrastructure Management (eBIM): A RISC-V-Empowered Hardware--Software Co-Design Framework Towards Trustworthy Blockchain | Qinglin Yang, Yuan Liu, Yaoyao Zhang, Boya Wang, Zongjian You, Chunming Rong, Zhihong Tian | 2026-07-08 | 下载 | Blockchain systems are undergoing a fundamental transition from decentralized ledgers for digital assets to general-purpose trust infrastructures for verifiable computation, decentralized physical res... |
| Vectorizing Quantum Control: A RISC-V Vector Extension Architecture for Scalable Qubit Systems | Xiaorang Guo, Kun Qin, Yanbin Chen, Carsten Trinitis, Martin Schulz | 2026-07-08 | 下载 | The Quantum Control Processor (QCP) bridges the gap between compiler toolchains and control electronics, and is responsible for translating compiled quantum circuits into executable instructions that ... |
| Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency | Satoshi Matsuoka | 2026-07-08 | 下载 | We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-... |
| Miter-Aware LUT Mapping: Aligning Structure and Solvability for Efficient Logic Equivalence Checking | Jiaying Zhu, Zhengyuan Shi, Mengxia Tao, Kezhi Li, Min Li, Qiang Xu | 2026-07-08 | 下载 | Logic Equivalence Checking (LEC), a fundamental hardware verification task, is often bottlenecked by synthesis-induced structural perturbations and XOR-dense regions that degrade SAT solver performanc... |
| ThermoDSE: A Thermal-Aware and Comprehensive Design Space Exploration for Chiplet-Based DNN Accelerators | Jian Peng, Hanwei Fan, Jingbo Jiang, Lin Jiang, Wei Zhang | 2026-07-08 | 下载 | Chiplet-based DNN accelerators provide a scalable path to balance performance and yield for modern AI workloads. However, such systems face critical challenges in area and thermal constraints. |
| EdgeCompress: Coupling Multidimensional Model Compression and Dynamic Inference for EdgeAI | Hao Kong, Di Liu, Shuo Huai, Xiangzhong Luo, Ravi Subramaniam, Christian Makaya, Qian Lin, Weichen Liu | 2026-07-08 | 下载 | Convolutional neural networks (CNNs) have demonstrated encouraging results in image classification tasks. However, the prohibitive computational cost of CNNs hinders the deployment of CNNs onto resour... |
| Smart Scissor: Coupling Spatial Redundancy Reduction and CNN Compression for Embedded Hardware | Hao Kong, Di Liu, Shuo Huai, Xiangzhong Luo, Weichen Liu, Ravi Subramaniam, Christian Makaya, Qian Lin | 2026-07-08 | 下载 | Scaling down the resolution of input images can greatly reduce the computational overhead of convolutional neural networks (CNNs), which is promising for edge AI. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems | Tingkai Liu, Muralidhar Andoorveedu, Sanjoy Das, Sanjay Patel, Volodymyr Kindratenko | 2026-07-08 | 下载 | The evolution of compute infrastructure has transformed multi-GPU systems into tightly integrated shared-memory structures. However, current software still mostly treats these coherent interconnects s... |
| Scaling WaterLily.jl with MPI and an improved geometric multigrid solver | Bernat Font, Marin Lauber, Tzu-Yao Huang, Gabriel D. Weymouth | 2026-07-08 | 下载 | We present recent performance-oriented developments in WaterLily, a scale-resolving incompressible flow solver written in pure Julia that runs seamlessly on CPUs and GPUs of any vendor. |
| GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining | Jieying Wang, Shuyuan Fan, Mingkai Zheng, Zhao Zhang | 2026-07-08 | 下载 | Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, such as FP8 and NVFP4, can significantly reduce the ... |
| POO-LPSP: Parallel Osprey Optimized Least Penalty-Squared Prioritization Methods for Priority Derivation in the Analytic Hierarchy Process | Kevin Kam Fung Yuen | 2026-07-08 | 下载 | Pairwise comparison (PC) via pairwise reciprocal matrices (PRMs) is central to the Analytic Hierarchy Process (AHP). Although the traditional eigenvector method is widely applied to derive priorities,... |
| Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production | Arun Malik | 2026-07-08 | 下载 | AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously solved problems. |
| Voltron: Enabling Elastic Multi-Device Execution of LLM Inference for Empowered Edge Intelligence | Chanwoo Cho, Wooseok Kim, Yonglak Son, Young Seo Lee, Young Geun Kim | 2026-07-08 | 下载 | Large language models (LLMs) are widely used in intelligent services due to their remarkable capability in generative tasks. Typically, LLM-based services process the inference requests of the users i... |
| Multiple Double Arithmetic on NVIDIA Tensor Cores | Howard Chen, Jan Verschelde | 2026-07-08 | 下载 | A multiple double is an unevaluated sum of doubles. An NVIDIA tensor core is a specialized high performance compute core for matrix multiplication. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| PHaul: A PPO-based forwarding agent for Sub6 enhanced Integrated Access and Backhaul networks | Jorge Pueyo, Daniel Camps-Mur and, Miguel Catalan-Cid | 2026-07-08 | 下载 | 3GPP Integrated Access and Backhaul (IAB) allows operators to deploy outdoor mm-wave access networks in a cost-efficient manner, by reusing the same spectrum in access and backhaul. |
| Small Language Model-based Control for BBR over Low Earth Orbit Satellite Internet | Rakshitha De Silva, Shiva Raj Pokhrel, Jonathan Kua | 2026-07-08 | 下载 | Low Earth Orbit (LEO) satellite Internet introduces rapid path variability, intermittent capacity shifts, and non-terrestrial delay dynamics that challenge transport-layer congestion control. |
| Unveiling TCP BBR Dominance in Starlink Internet: Experimental Insights and Analysis | Rakshitha De Silva, Shiva Raj Pokhrel, Jonathan Kua | 2026-07-08 | 下载 | This experimental study delivers a global assessment of Google's Bottleneck Bandwidth and Round-trip propagation time-version 3 (BBR-v3) Congestion Control Algorithm (CCA) over SpaceX's Starlink netwo... |
| How the Fusion of Onboard Sensors and V2X Data can Improve (or not) the Cooperative Perception of Connected Automated Vehicles | Amir Mohammadisarab, Miguel Sepulcre, Luca Lusvarghi, Javier Gozalvez | 2026-07-08 | 下载 | Automated vehicles rely on onboard sensors to perceive their surroundings and navigate autonomously. However, sensor performance may degrade under adverse weather conditions or when line-of-sight is o... |
| Bessel Beam Optimization for Near-Field THz Communications under UE Location Uncertainty | Aditya Jolly, Vitaly Petrov, Gábor Fodor, Emil Björnson | 2026-07-08 | 下载 | To achieve the desired coverage and capacity levels, future terahertz (THz) wireless systems are envisioned to utilize extremely large antenna arrays. |
| EvoOMG: An Evolution-Oriented Multi-Agent Guidance Framework for Heterogeneous Legacy-and-MLO Wi-Fi Networks | Junjie Wu, Lingjian Zhou, Zerui Shao, Yi Zou, Tianrui Li, Yi Zhang, Ziyuan Yang | 2026-07-08 | 下载 | The gradual deployment of Wi-Fi 7/8 multi-link operation (MLO) will lead to long-term coexistence between legacy non-MLO stations (STAs) and MLO-capable STAs in WLANs. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency | Satoshi Matsuoka | 2026-07-08 | 下载 | We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5.2), rapid inference-efficiency gains (near-Shannon-limit KV-... |