2026-08-31
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM Inference | Chun-Ting Chen, Dongmin Han, Hangyeol Mun, Jake Hyun, Arnab Raha, Amit Agarwal, Mark Anders, Mohamed Abdelfattah, Jae-sun Seo | 2026-08-31 | 下载 | Block Quantization (BQ) is a promising approach for efficient deployment of large language models (LLMs), enabling low-precision computation with controlled accuracy degradation. |
| VARA: A Voltage-Aware ReRAM-Based Accelerator for Energy-Efficient Computing | Peng Dang, Yintao He, Huawei Li | 2026-08-31 | 下载 | ReRAM-based in-memory computing (IMC) architectures are widely regarded as a promising approach to alleviating the computational bottleneck of conventional architectures. |
| DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE Inference | Xiaoyang Lu, Belthangady Akash Vi Narayana Pai, Xian-He Sun | 2026-08-31 | 下载 | Mixture-of-Experts (MoE) models enable efficient scaling of large language model (LLM) inference but suffer from substantial data-movement overhead when deployed on neural processing unit (NPU)-based ... |
| Storage-Centric System Designs for Enabling Fast, Efficient, and Low-Cost Genomic and Metagenomic Analyses | Nika Mansouri Ghiasi | 2026-08-31 | 下载 | Genomic and metagenomic analyses play critical roles in many fields, such as precision medicine, urgent clinical settings, discovering early warnings of communicable diseases, ensuring food safety thr... |
| Clock-Gating Insertion Strategies on an Open-Source MSP430 Core: A Reproducible PPA Study and a Gate-Level Simulation Caveat | Xingran Huang, Qiming Guo, Jinwen Tang, Wenqi Jia, Dongzheng Wang | 2026-08-31 | 下载 | Clock gating, the standard technique for cutting dynamic power, is introduced either as hand-written behavioral clock gates at the register-transfer level (RTL) or as integrated clock-gating (ICG) cel... |
| Beacon: LLM Multi-Agent Driven Hardware Design Space Exploration for Heterogeneous Multi-Chiplet Deep Learning Accelerators | Boyu Li, Zongwei Zhu, Qianyue Cao, Xi Li, Xuehai Zhou | 2026-08-31 | 下载 | Heterogeneous multi-chiplet accelerators allow chiplets to be configured independently to better match different operator characteristics and improve inference efficiency. |
| LLM-based Hardware Development with Hierarchical IRs and End-to-End Multi-Agent Workflow | Chenyang Yin, Agasthi Haputhanthri, Aditya Anirudh Jonnalagadda, Zhenyu Bai, Yuanming Song, Saranyu Chattopadhyay, Mohammad Fadiheh, Tom Zelazny, Subhasish Mitra, Tulika Mitra | 2026-08-31 | 下载 | Large language models (LLMs) are increasingly used in software development, but their use in complex hardware design remains limited. This gap stems from both the scarcity of public hardware training ... |
| CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration | Yue Jiet Chong, Yimin Wang, Zhen Wu, Zixuan Wang, Wei Zhang, Xuanyao Fong | 2026-08-31 | 下载 | Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high utilization, memory ef... |
| Non-uniform Memory Partitioning For Low-Power Spiking Neural Networks | Simon Richter, Darío Fernández Khatiboun, Maryam Sadeghi, Milad Zamani, Farshad Moradi | 2026-08-31 | 下载 | Spiking Neural Networks (SNNs) naturally excel in processing temporally rich and sparse data. However, because of their time-stepped processing, memory access, specifically to synaptic weights stored ... |
| Scalable AXI4 Transaction Monitoring for Mixed-Criticality SoCs: From Phase-Level Precision to ID-Level Efficiency | Chaoqun Liang, Thomas Benz, Alessandro Ottaviano, Michael Rogenmoser, Luca Benini, Angelo Garofalo, Davide Rossi | 2026-08-31 | 下载 | Mixed-criticality Systems-on-Chip (SoCs) with on-chip interconnects based on the AXI4 open standard protocol lack a protocol-level timeout mechanism, exposing systems to deadlocks and missed real-time... |
| KORD: Breaking the Key-Generation Bottleneck in Dealerless FSS via Protocol--Hardware Co-Design | Yijing Peng, Lin Liu, Yujie Xue, Shaojing Fu, Shaoqing Li, Yaohua Wang, Rongmao Chen, Yang Guo | 2026-08-31 | 下载 | Function secret sharing (FSS) has become a core primitive in privacy-preserving computation. However, each FSS invocation requires a fresh pair of function keys generated by a trusted dealer , expands... |
| FABO: Agent-Guided Discovery of Joint Breakpoint Optimization for Timing-Driven Routing Trees | Shang Liu, Wenji Fang, Jing Wang, Hongxin Kong, Yao Lu, Zhiyao Xie | 2026-08-31 | 下载 | The topology of a routing tree determines how a multi-pin net branches and shares physical wire, directly affecting wirelength, congestion, capacitance, and delay. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| DRLM: Deep Reinforcement Learning-Based LLM Query Orchestration in Edge Environments | Reza Farahani, Zoha Azimi Ourimi, Mario Colosi, Lauri Loven, Christian Timmerer, Schahram Dustdar | 2026-08-31 | 下载 | Large language model (LLM) services increasingly process heterogeneous queries with diverse latency, accuracy, and resource requirements. While edge deployment reduces response time, the heterogeneity... |
| Client-side transparent caching for remote ROOT data analysis | Dmytro Kovalskyi, Jan Eysermans, Mariarosaria D'Alfonso, Christoph Paus | 2026-08-31 | 下载 | High-energy physics analyses often process the same data as physicists refine algorithms and test new ideas. With data increasingly read from remote storage, each iteration is subject to network laten... |
| Don't Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes | Pruthvi Davineni | 2026-08-31 | 下载 | LLM agents increasingly diagnose incidents and propose remediations. In a GitOps workflow, applying a fix means editing a version-controlled config file, and the obvious implementation, having the mod... |
| BlockMGARD: Accelerating Adaptive Scientific Data Reduction with Region-of-Interest Error Control on GPUs | Yanliang Li, Qian Gong, Qing Liu, Jaemoon Lee, Norbert Podhorszki, Scott Klasky, Xin Liang, Jieyang Chen | 2026-08-31 | 下载 | The growing scale of scientific data makes lossy compression essential for reducing data volume under controllable error. Transformation-based compressors using multilevel decomposition, such as MGARD... |
| Storage-Centric System Designs for Enabling Fast, Efficient, and Low-Cost Genomic and Metagenomic Analyses | Nika Mansouri Ghiasi | 2026-08-31 | 下载 | Genomic and metagenomic analyses play critical roles in many fields, such as precision medicine, urgent clinical settings, discovering early warnings of communicable diseases, ensuring food safety thr... |
| Beating Quadratic Time--Message Trade-off in Distributed Minimum Spanning Tree Construction | Taisuke Izumi, Naoki Kitamura, Toshimitsu Masuzawa | 2026-08-31 | 下载 | We present a new distributed algorithm for computing a minimum spanning tree (MST) in the \textsf{CONGEST-KT} model, where messages are limited to bits and each vertex initially know... |
| Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret | Xia Jiang, Lu Liu, Gang Feng | 2026-08-31 | 下载 | This paper investigates distributed online optimization for multi-agent dynamical systems with constrained inputs and time-varying cost functions. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Physiological Information Reliability: Cross-Layer Adaptive Resource Allocation for Cardiovascular Sensing | Navaneeth Krishnan Kamalakannan, Janakiraman Kamalakannan, Harinisri Velmurugan | 2026-08-31 | 下载 | Cardiovascular sensing systems must preserve clinically useful information despite signal degradation, wireless losses, energy constraints, and edge-computation latency. |
| WiSDoM: Wireless Sparse Decision Transformer with Mixture-of-Experts for Multi-Task Mobile Network Optimization | Fatih Temiz, Shavbo Salehi, Melike Erol-Kantarci | 2026-08-31 | 下载 | Emerging 6G wireless networks are expected to operate across diverse deployment scenarios, where variations in network topology, user mobility, traffic demand, and radio conditions challenge the scala... |
| Local Private Information Retrieval for Graph-Based Replicated Systems | Shreya Meel, Mohamed Nomeir, Sennur Ulukus | 2026-08-31 | 下载 | We rethink the definition of privacy in multi-server, graph-replicated private information retrieval (PIR) systems, by introducing a novel setting where the user's privacy is governed by the servers' ... |
| Semantic Freshness Optimal Sampling and Transmission for Gossiping Receivers | Irtiza Hasan, Ahmed Arafa | 2026-08-31 | 下载 | We study the optimal joint sampling and transmission policy for a transmitter communicating with two gossiping receivers that share information with each other, with the objective of tracking a source... |
| Ray Tracing-Based LoRaWAN Gateway Placement for Reliable Connectivity in Amazonian Regions | Cláudio Modesto, Lucas Mozart, Cleverson Nahum, Bruno Castro, Aldebaro Klautau | 2026-08-31 | 下载 | Network planning is an important task in wireless communications, as it helps network operators avoid unnecessary costs. In the context of the internet of things, using long-range wide-area network te... |
| Sensitivity Comparison of Microwave-Frequency and Optical Fibre Interferometry Based on State-of-the-Art Components | Georgios Aias Karydis, Marco Fasano, Paola Parolari, Pierpaolo Boffi, Charis Mesaritakis, Adonis Bogris | 2026-08-31 | 下载 | We compare the sensitivity of fibre interferometers to vibrations using microwave oscillators and state of the art lasers over distances up to 70 km. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Adaptive KV Retention for LLM Agents at Human-Approval Timescales | Minseo Choi, Ananya Joshi | 2026-08-31 | 下载 | Unlike the seconds-scale tool-call pauses targeted by prior agent-serving systems, agentic LLM requests can be suspended for minutes or hours while waiting for human approval. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Towards Stream Learning on Embedded Systems: Benchmarking the Memory Consumption of Stream Learning Methods | Sebastian Buschjäger, Nuwan Gunasekara, Heitor Murilo Gomes | 2026-08-31 | 下载 | Stream learning is commonly evaluated through predictive performance and adaptation to concept drift. However, sustained operation of a stream learner also requires predictable and bounded resource us... |
| LaMoC: Loss-Aware Modular Compression for LLMs | Mohanad Odema, Jacob Song | 2026-08-31 | 下载 | Modular compression has enabled considerable parameter reduction in LLMs while preserving strong language understanding and downstream task accuracy. |