2026-07-24
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Multi-primitive in-memory computing for Monte Carlo tree search | Tergel Molom-Ochir, Benjamin F. Morris, Yintao He, Archit Gajjar, Giacomo Pedretti, Hai Helen Li, Yiran Chen, Jim Ignowski, Aishwarya Natarajan | 2026-07-24 | 下载 | Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limiting edge deployment. |
| HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding | Chao Fang, Jun Yin, Man Shi, Marian Verhelst | 2026-07-24 | 下载 | With the rapid adoption of long-context large language models (LLMs), the continuously growing KV cache during decoding has become the critical memory bottleneck. |
| Reducing Instruction-Fetch Energy in RISC-V for Embedded AI Processing via Dynamic and Static Loop Caching | Wiebren Wijnstra, Sameed Sohail, Berend-Jan van der Zwaag, Sabih Gerez, Amirreza Yousefzadeh | 2026-07-24 | 下载 | Embedded RISC-V processors are increasingly deployed for on-device AI inference at the edge, where energy efficiency is a primary design constraint. |
| The Sparsity Tax: Weight Sparsity Trade-offs in Event-Driven SIMD and SIMT Neuromorphic Cores | Mattias Westerink, Sameed Sohail, Berend-Jan van der Zwaag, Sabih Gerez, Amirreza Yousefzadeh | 2026-07-24 | 下载 | Event-driven neuromorphic inference exploits activation sparsity by updating neuron state only on spikes. However, weight sparsity introduces irregular gather-style updates that undermine lockstep Sin... |
| Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs | Ilia Sobakinskikh, Paul Alexander Bilokon | 2026-07-24 | 下载 | In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. |
| FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon | Om Mohite | 2026-07-24 | 下载 | Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units... |
| Sparse by Command: Task-Conditional Compute Skipping for Multi-Task Inference Accelerators | Afzal Ahmad, Gaoyu Mao, Shoubo Hu, Hui-Ling Zhen, Mingxuan Yuan, Xinyu Chen, Wei Zhang | 2026-07-24 | 下载 | Multi-task inference models share a single backbone across diverse tasks, yet execute identical computation regardless of which task is active - wasting energy and cycles on task-irrelevant operations... |
| HEMERA: A Heterogeneous Memory-Centric Accelerator with Recursive Dataflow for Edge-Constrained State-Space-Duality Models Inference | Hao Ding, Ling Liang, Ruitong Qiao, Dongxue Zhao, Xiantong Qiu, Jinshan Li, Meng Li, Lei Jin, Zhiliang Xia, Zongliang Huo, Zongwei Wang, Yimao Cai | 2026-07-24 | 下载 | Structured State Space Models (SSMs), such as Mamba, enable efficient long-sequence modeling with linear time complexity. Recent implementations realize this capability through Structured State Space ... |
| Unified Static-Dynamic Pruning for Efficient LLM Inference | Jinhyeok Kim, Yejoon Lee, Jaeyoung Do | 2026-07-24 | 下载 | The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dom... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| 3D Gaussian Splatting for Scientific Particle Data Compression and Rendering | Bo Jiang, Youyuan Liu, Taolue Yang, Sheng Di, Sian Jin | 2026-07-24 | 下载 | Large-scale particle simulations produce hundreds of millions of particles, straining storage, transfer, and interactive visualization. Existing lossy compressors such as SZ3 operate in data space and... |
| MAGNUS: Fast SpGEMM on GPUs for Irregular Matrices via Hierarchical Multisplit | Jordi Wolfson-Pou, Ahmed Helal, Fabrizio Petrini | 2026-07-24 | 下载 | We present MAGNUS, a novel algorithm for sparse matrix-matrix multiplication (SpGEMM) of irregular matrices on GPUs. Such matrices often contain many \emph{heavy rows}, those with large intermediat... |
| SLA-Constrained Carbon-Aware Routing in Geo-Distributed Serverless Clouds | Anmol Chaudhary, Rahul Mishra | 2026-07-24 | 下载 | Modern cloud deployments distribute applications across multiple geographic regions, yet standard routing mechanisms prioritize latency while ignoring the fluctuating carbon intensity of local power g... |
| TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters | Zhiwen Mo, Yu Cheng, Lei Wang, Zhengju Tang, Lei Xu, Guoyu Li, Yuqi Dong, Lingxiao Ma, Yuqing Xia, Jilong Xue, Fan Yang, Luo Mai, Zhi Yang, Wayne Luk, Hongxiang Fan | 2026-07-24 | 下载 | Recent GPU programming frameworks such as Triton, TileLang, and CUDA Tile adopt tiles as first-class primitives, making tile-centric programming the prevailing approach for high-performance GPU kernel... |
| NUMA balancing hampering performance of spiking network simulations | Melissa Lober, Alp Inangu, Gorka Peraza Coppola, Dennis Terhorst, Sebastian Gillessen, Jan Vogelsang, Hans Ekkehard Plesser, Brian Wylie, Benedikt Steinbusch, Guido Trensch, Susanne Kunkel, Markus Diesmann | 2026-07-24 | 下载 | Computing centers today mostly operate conventional CPU- and GPU-based systems, where the direct way of decreasing energy consumption is a reduction in the applications' runtime. |
| Agentic CPU-GPU Scheduling for Heterogeneous AI Workloads | Tianxi Lu, Sherief Reda | 2026-07-24 | 下载 | Agentic AI systems compose heterogeneous tool workloads on shared GPU/CPU infrastructure, yet existing frameworks assign all GPU-capable tools to the GPU by default. |
| Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs | Ilia Sobakinskikh, Paul Alexander Bilokon | 2026-07-24 | 下载 | In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. |
| FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon | Om Mohite | 2026-07-24 | 下载 | Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units... |
| Duet: Co-Optimizing P2P Message Propagation and Rotating-Leader Consensus | Yifeng Ye, Rongji Huang, Gerui Wang, Mingchao Wan, Yuxing Duan, Jingjing Zhang, Shengyun Liu | 2026-07-24 | 下载 | In blockchain systems, peer-to-peer (P2P) overlay networks play a crucial role in providing reliable, scalable and efficient message-delivery services to upper layers. |
| Accountable Transaction Inclusion Lists: Enhancing Ethereum's Censorship Resistance | Patrick Spiesberger, Hannes Hartenstein | 2026-07-24 | 下载 | In Ethereum, transaction inclusion is rarely in question; what matters is the delay until inclusion. Currently, block builders could exercise censorship across consecutive blocks, threatening time-cri... |
| Smart Contract Tells: Aircraft Maintenance Records Are Now Trustworthy | Woosuk Choi, Seungmo Kim | 2026-07-24 | 下载 | Aircraft maintenance records are critical to airworthiness and asset valuation, yet they are often fragmented across stakeholders, creating verification bottlenecks and information asymmetry that may ... |
| Unified Static-Dynamic Pruning for Efficient LLM Inference | Jinhyeok Kim, Yejoon Lee, Jaeyoung Do | 2026-07-24 | 下载 | The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding, where low compute intensity and bandwidth-bound kernels dom... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence | Bin Dong, Sukhada Gholba, Brooklin Gore, Shawn Kwang, David Mitchell, Samuel Oehlert, Garrett Stewart, Brendan White, Luke Baker, Ed Balas, Britt Gathright, Chin Guok, Jon-Paul Heron, John MacAuley, Scott Richmond, Chris Robb, Chris Tracy, Kesheng Wu | 2026-07-24 | 下载 | The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 initiative and to address persistent operational pain points in t... |
| Let AI Agents Translate Networks, Not Reason About Them | Hongyu Hè, Maria Apostolaki | 2026-07-24 | 下载 | A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no production network has one, since writing a model by hand demands r... |
| Invariant Discovery for Networked Systems | Hongyu Hè, Alexander Krentsel, Sylvia Ratnasamy, Maria Apostolaki | 2026-07-24 | 下载 | Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generation, telemetry imputation, and input validation, yet writing t... |
| Twin-Fidelity-Aware Resolution of Direct xApp Conflicts in Open RAN | Akram Almohammedi, Mohammed Balfaqih, Sam Darshi, Rami Langar, Wael Jaafar | 2026-07-24 | 下载 | Open Radio Access Network (O-RAN) allows independently developed xApps to control RAN functions through the Near-Real-Time RAN Intelligent Controller (Near-RT RIC). |
| CAPS: Fine-Tuning CCA Timing | Raphael Zailer, Isaac Keslassy | 2026-07-24 | 下载 | Data-center congestion control targets high throughput, fair bandwidth allocation, and low latency. Modern transports couple rate computation and packet scheduling into a single feedback loop, converg... |
| A Self-Calibrating Agentic AI Framework for Autonomous Edge Resource Allocation | Fin Gentzen, Marla Grunewald, Iulisloi Zacarias, Mounir Bensalem, Admela Jukan | 2026-07-24 | 下载 | Large Language Models (LLMs) are increasingly deployed as autonomous agents, transitioning from static conversational interfaces to dynamic systems capable of complex reasoning, tool execution, and de... |
| Predictive Lightweight MARL for Resilient Coverage in Sparse-Signaling Aerial Networks | Chuan-Chi Lai, Ang-Hsun Tsai | 2026-07-24 | 下载 | This letter proposes the Predictive Lightweight Multi-Agent Reinforcement Learning (PL-MARL) framework to ensure resilient coverage in bandwidth-constrained UAV swarms. |
| Neilson's Weak vs. Strong Loss Aversion: A Characterization and a Generalized CPT-Utility Function | Symeon Vaidanis, Marios Kountouris | 2026-07-24 | 下载 | In multi-objective and multi-criteria decision-making under risk, especially in settings involving individual behavior, risk-aware analysis based on subjective evaluation has become increasingly impor... |
| Location-Aware NAS Timer Optimization in NTN-TN Integrated Networks | Cheng Liu, Peng Hu | 2026-07-24 | 下载 | Efficient Non-Access Stratum (NAS) timer configuration is critical for reliable and energy-efficient Fifth Generation (5G) registration in Non-Terrestrial Network (NTN)-Terrestrial Network (TN) integr... |
| Cleaning the NTP Pool: Detecting and Mitigating NTP-Sourced IPv6 Scanning | Erik Rye, Robert Beverly | 2026-07-24 | 下载 | The ephemeral and random nature of IPv6 client addresses presents a practical challenge to attacks that depend on Internet-wide scanning or reconnaissance -- the adversary must first \emph{find} the c... |
| Fewer Paths, Better Performance: Understanding the ZCube Topology through Braess's Paradox | Li Chen | 2026-07-24 | 下载 | Datacenter networks follow a multipath doctrine: provision many paths between endpoints, hash flows across them, and let redundancy absorb both failures and load imbalance. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters | Zhiwen Mo, Yu Cheng, Lei Wang, Zhengju Tang, Lei Xu, Guoyu Li, Yuqi Dong, Lingxiao Ma, Yuqing Xia, Jilong Xue, Fan Yang, Luo Mai, Zhi Yang, Wayne Luk, Hongxiang Fan | 2026-07-24 | 下载 | Recent GPU programming frameworks such as Triton, TileLang, and CUDA Tile adopt tiles as first-class primitives, making tile-centric programming the prevailing approach for high-performance GPU kernel... |
| Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs | Ilia Sobakinskikh, Paul Alexander Bilokon | 2026-07-24 | 下载 | In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time anomaly detection in financial time series. |
| FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon | Om Mohite | 2026-07-24 | 下载 | Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units... |