Skip to content

2026-07-24 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Multi-primitive in-memory computing for Monte Carlo tree searchTergel Molom-Ochir, Benjamin F. Morris, Yintao He, Archit Gajjar, Giacomo Pedretti, Hai Helen Li, Yiran Chen, Jim Ignowski, Aishwarya Natarajan2026-07-24下载Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment.
HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM DecodingChao Fang, Jun Yin, Man Shi, Marian Verhelst2026-07-24下载With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck.
Reducing Instruction-Fetch Energy in RISC-V for Embedded AI Processing via Dynamic and Static Loop CachingWiebren Wijnstra, Sameed Sohail, Berend-Jan van der Zwaag, Sabih Gerez, Amirreza Yousefzadeh2026-07-24下载Embedded RISC-V processors are increasingly deployed for on-device AI inference at the edge, where energy efficiency is a primary design constraint.
The Sparsity Tax: Weight Sparsity Trade-offs in Event-Driven SIMD and SIMT Neuromorphic CoresMattias Westerink, Sameed Sohail, Berend-Jan van der Zwaag, Sabih Gerez, Amirreza Yousefzadeh2026-07-24下载Event-driven neuromorphic inference exploits activation sparsity by updating neuron state only on spikes. However, weight sparsity introduces irregular gather-style updates that undermine lockstep Sin...
Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAsIlia Sobakinskikh, Paul Alexander Bilokon2026-07-24下载In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series.
FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple SiliconOm Mohite2026-07-24下载Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units...
Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference AcceleratorsAfzal Ahmad, Gaoyu Mao, Shoubo Hu, Hui-Ling Zhen, Mingxuan Yuan, Xinyu Chen, Wei Zhang2026-07-24下载Multi-task inference models share a single backbone across diverse tasks, yet execute identical computation regardless of which task is active - wasting energy and cycles on task-irrelevant operations...
HEMERA: A Heterogeneous Memory-Centric Accelerator with Recursive Dataflow for Edge-Constrained State-Space-Duality Models InferenceHao Ding, Ling Liang, Ruitong Qiao, Dongxue Zhao, Xiantong Qiu, Jinshan Li, Meng Li, Lei Jin, Zhiliang Xia, Zongliang Huo, Zongwei Wang, Yimao Cai2026-07-24下载Structured State Space Models (SSMs), such as Mamba, enable efficient long-sequence modeling with linear time complexity. Recent implementations realize this capability through Structured State Space ...
Unified Static-Dynamic Pruning for Efficient LLM InferenceJinhyeok Kim, Yejoon Lee, Jaeyoung Do2026-07-24下载The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dom...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
3D Gaussian Splatting for Scientific Particle Data Compression and RenderingBo Jiang, Youyuan Liu, Taolue Yang, Sheng Di, Sian Jin2026-07-24下载Large-scale particle simulations produce hundreds of millions of particles, straining storage, transfer, and interactive visualization. Existing lossy compressors such as SZ3 operate in data space and...
ggMAGNUS: Fast SpGEMM on GPUs for Irregular Matrices via Hierarchical MultisplitJordi Wolfson-Pou, Ahmed Helal, Fabrizio Petrini2026-07-24下载We present ggMAGNUS, a novel algorithm for sparse matrix-matrix multiplication (SpGEMM) of irregular matrices on GPUs. Such matrices often contain many \emph{heavy rows}, those with large intermediat...
SLA-Constrained Carbon-Aware Routing in Geo-Distributed Serverless CloudsAnmol Chaudhary, Rahul Mishra2026-07-24下载Modern cloud deployments distribute applications across multiple geographic regions, yet standard routing mechanisms prioritize latency while ignoring the fluctuating carbon intensity of local power g...
TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to ClustersZhiwen Mo, Yu Cheng, Lei Wang, Zhengju Tang, Lei Xu, Guoyu Li, Yuqi Dong, Lingxiao Ma, Yuqing Xia, Jilong Xue, Fan Yang, Luo Mai, Zhi Yang, Wayne Luk, Hongxiang Fan2026-07-24下载Recent GPU programming frameworks such as Triton, TileLang, and CUDA Tile adopt tiles as first-class primitives, making tile-centric programming the prevailing approach for high-performance GPU kernel...
NUMA balancing hampering performance of spiking network simulationsMelissa Lober, Alp Inangu, Gorka Peraza Coppola, Dennis Terhorst, Sebastian Gillessen, Jan Vogelsang, Hans Ekkehard Plesser, Brian Wylie, Benedikt Steinbusch, Guido Trensch, Susanne Kunkel, Markus Diesmann2026-07-24下载Computing centers today mostly operate conventional CPU- and GPU-based systems, where the direct way of decreasing energy consumption is a reduction in the applications' runtime.
Agentic CPU-GPU Scheduling for Heterogeneous AI WorkloadsTianxi Lu, Sherief Reda2026-07-24下载Agentic AI systems compose heterogeneous tool workloads on shared GPU/CPU infrastructure, yet existing frameworks assign all GPU-capable tools to the GPU by default.
Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAsIlia Sobakinskikh, Paul Alexander Bilokon2026-07-24下载In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series.
FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple SiliconOm Mohite2026-07-24下载Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units...
Duet: Co-Optimizing P2P Message Propagation and Rotating-Leader ConsensusYifeng Ye, Rongji Huang, Gerui Wang, Mingchao Wan, Yuxing Duan, Jingjing Zhang, Shengyun Liu2026-07-24下载In blockchain systems, peer-to-peer (P2P) overlay networks play a crucial role in providing reliable, scalable and efficient message-delivery services to upper layers.
Accountable Transaction Inclusion Lists: Enhancing Ethereum's Censorship ResistancePatrick Spiesberger, Hannes Hartenstein2026-07-24下载In Ethereum, transaction inclusion is rarely in question; what matters is the delay until inclusion. Currently, block builders could exercise censorship across consecutive blocks, threatening time-cri...
Smart Contract Tells: Aircraft Maintenance Records Are Now TrustworthyWoosuk Choi, Seungmo Kim2026-07-24下载Aircraft maintenance records are critical to airworthiness and asset valuation, yet they are often fragmented across stakeholders, creating verification bottlenecks and information asymmetry that may ...
Unified Static-Dynamic Pruning for Efficient LLM InferenceJinhyeok Kim, Yejoon Lee, Jaeyoung Do2026-07-24下载The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dom...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational ExcellenceBin Dong, Sukhada Gholba, Brooklin Gore, Shawn Kwang, David Mitchell, Samuel Oehlert, Garrett Stewart, Brendan White, Luke Baker, Ed Balas, Britt Gathright, Chin Guok, Jon-Paul Heron, John MacAuley, Scott Richmond, Chris Robb, Chris Tracy, Kesheng Wu2026-07-24下载The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and to address persistent operational pain points in t...
Let AI Agents Translate Networks, Not Reason About ThemHongyu Hè, Maria Apostolaki2026-07-24下载A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network has one, since writing a model by hand demands r...
Invariant Discovery for Networked SystemsHongyu Hè, Alexander Krentsel, Sylvia Ratnasamy, Maria Apostolaki2026-07-24下载Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generation, telemetry imputation, and input validation, yet writing t...
Twin-Fidelity-Aware Resolution of Direct xApp Conflicts in Open RANAkram Almohammedi, Mohammed Balfaqih, Sam Darshi, Rami Langar, Wael Jaafar2026-07-24下载Open Radio Access Network (O-RAN) allows independently developed xApps to control RAN functions through the Near-Real-Time RAN Intelligent Controller (Near-RT RIC).
CAPS: Fine-Tuning CCA TimingRaphael Zailer, Isaac Keslassy2026-07-24下载Data-center congestion control targets high throughput, fair bandwidth allocation, and low latency. Modern transports couple rate computation and packet scheduling into a single feedback loop, converg...
A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource AllocationFin Gentzen, Marla Grunewald, Iulisloi Zacarias, Mounir Bensalem, Admela Jukan2026-07-24下载Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from static conversational interfaces to dynamic systems capable of complex reasoning, tool execution, and de...
Predictive Lightweight MARL for Resilient Coverage in Sparse-Signaling Aerial NetworksChuan-Chi Lai, Ang-Hsun Tsai2026-07-24下载This letter proposes the Predictive Lightweight Multi-Agent Reinforcement Learning (PL-MARL) framework to ensure resilient coverage in bandwidth-constrained UAV swarms.
Neilson's Weak vs. Strong Loss Aversion: A Characterization and a Generalized CPT-Utility FunctionSymeon Vaidanis, Marios Kountouris2026-07-24下载In multi-objective and multi-criteria decision-making under risk, especially in settings involving individual behavior, risk-aware analysis based on subjective evaluation has become increasingly impor...
Location-Aware NAS Timer Optimization in NTN-TN Integrated NetworksCheng Liu, Peng Hu2026-07-24下载Efficient Non-Access Stratum (NAS) timer configuration is critical for reliable and energy-efficient Fifth Generation (5G) registration in Non-Terrestrial Network (NTN)-Terrestrial Network (TN) integr...
Cleaning the NTP Pool: Detecting and Mitigating NTP-Sourced IPv6 ScanningErik Rye, Robert Beverly2026-07-24下载The ephemeral and random nature of IPv6 client addresses presents a practical challenge to attacks that depend on Internet-wide scanning or reconnaissance -- the adversary must first \emph{find} the c...
Fewer Paths, Better Performance: Understanding the ZCube Topology through Braess's ParadoxLi Chen2026-07-24下载Datacenter networks follow a multipath doctrine: provision many paths between endpoints, hash flows across them, and let redundancy absorb both failures and load imbalance.

cs.PF - Performance ​

标题作者发布日期PDF摘要
TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to ClustersZhiwen Mo, Yu Cheng, Lei Wang, Zhengju Tang, Lei Xu, Guoyu Li, Yuqi Dong, Lingxiao Ma, Yuqing Xia, Jilong Xue, Fan Yang, Luo Mai, Zhi Yang, Wayne Luk, Hongxiang Fan2026-07-24下载Recent GPU programming frameworks such as Triton, TileLang, and CUDA Tile adopt tiles as first-class primitives, making tile-centric programming the prevailing approach for high-performance GPU kernel...
Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAsIlia Sobakinskikh, Paul Alexander Bilokon2026-07-24下载In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series.
FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple SiliconOm Mohite2026-07-24下载Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units...

基于 VitePress 构建 · 使用本地搜索查找论文