Skip to content

2026-08-31 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
HBQ: Hierarchical Scaling Block Quantization with Hardware-Efficiency-Aware Design for Accurate LLM InferenceChun-Ting Chen, Dongmin Han, Hangyeol Mun, Jake Hyun, Arnab Raha, Amit Agarwal, Mark Anders, Mohamed Abdelfattah, Jae-sun Seo2026-08-31下载Block Quantization (BQ) is a promising approach for efficient deployment of large language models (LLMs), enabling low-precision computation with controlled accuracy degradation.
VARA: A Voltage-Aware ReRAM-Based Accelerator for Energy-Efficient ComputingPeng Dang, Yintao He, Huawei Li2026-08-31下载ReRAM-based in-memory computing (IMC) architectures are widely regarded as a promising approach to alleviating the computational bottleneck of conventional architectures.
DynaNDE: Dynamic Near-Data Expert Scheduling for Batched MoE InferenceXiaoyang Lu, Belthangady Akash Vi Narayana Pai, Xian-He Sun2026-08-31下载Mixture-of-Experts (MoE) models enable efficient scaling of large language model (LLM) inference but suffer from substantial data-movement overhead when deployed on neural processing unit (NPU)-based ...
Storage-Centric System Designs for Enabling Fast, Efficient, and Low-Cost Genomic and Metagenomic AnalysesNika Mansouri Ghiasi2026-08-31下载Genomic and metagenomic analyses play critical roles in many fields, such as precision medicine, urgent clinical settings, discovering early warnings of communicable diseases, ensuring food safety thr...
Clock-Gating Insertion Strategies on an Open-Source MSP430 Core: A Reproducible PPA Study and a Gate-Level Simulation CaveatXingran Huang, Qiming Guo, Jinwen Tang, Wenqi Jia, Dongzheng Wang2026-08-31下载Clock gating, the standard technique for cutting dynamic power, is introduced either as hand-written behavioral clock gates at the register-transfer level (RTL) or as integrated clock-gating (ICG) cel...
Beacon: LLM Multi-Agent Driven Hardware Design Space Exploration for Heterogeneous Multi-Chiplet Deep Learning AcceleratorsBoyu Li, Zongwei Zhu, Qianyue Cao, Xi Li, Xuehai Zhou2026-08-31下载Heterogeneous multi-chiplet accelerators allow chiplets to be configured independently to better match different operator characteristics and improve inference efficiency.
LLM-based Hardware Development with Hierarchical IRs and End-to-End Multi-Agent WorkflowChenyang Yin, Agasthi Haputhanthri, Aditya Anirudh Jonnalagadda, Zhenyu Bai, Yuanming Song, Saranyu Chattopadhyay, Mohammad Fadiheh, Tom Zelazny, Subhasish Mitra, Tulika Mitra2026-08-31下载Large language models (LLMs) are increasingly used in software development, but their use in complex hardware design remains limited. This gap stems from both the scarcity of public hardware training ...
CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference AccelerationYue Jiet Chong, Yimin Wang, Zhen Wu, Zixuan Wang, Wei Zhang, Xuanyao Fong2026-08-31下载Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high utilization, memory ef...
Non-uniform Memory Partitioning For Low-Power Spiking Neural NetworksSimon Richter, Darío Fernández Khatiboun, Maryam Sadeghi, Milad Zamani, Farshad Moradi2026-08-31下载Spiking Neural Networks (SNNs) naturally excel in processing temporally rich and sparse data. However, because of their time-stepped processing, memory access, specifically to synaptic weights stored ...
Scalable AXI4 Transaction Monitoring for Mixed-Criticality SoCs: From Phase-Level Precision to ID-Level EfficiencyChaoqun Liang, Thomas Benz, Alessandro Ottaviano, Michael Rogenmoser, Luca Benini, Angelo Garofalo, Davide Rossi2026-08-31下载Mixed-criticality Systems-on-Chip (SoCs) with on-chip interconnects based on the AXI4 open standard protocol lack a protocol-level timeout mechanism, exposing systems to deadlocks and missed real-time...
KORD: Breaking the Key-Generation Bottleneck in Dealerless FSS via Protocol--Hardware Co-DesignYijing Peng, Lin Liu, Yujie Xue, Shaojing Fu, Shaoqing Li, Yaohua Wang, Rongmao Chen, Yang Guo2026-08-31下载Function secret sharing (FSS) has become a core primitive in privacy-preserving computation. However, each FSS invocation requires a fresh pair of function keys generated by a trusted dealer , expands...
FABO: Agent-Guided Discovery of Joint Breakpoint Optimization for Timing-Driven Routing TreesShang Liu, Wenji Fang, Jing Wang, Hongxin Kong, Yao Lu, Zhiyao Xie2026-08-31下载The topology of a routing tree determines how a multi-pin net branches and shares physical wire, directly affecting wirelength, congestion, capacitance, and delay.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
DRLM: Deep Reinforcement Learning-Based LLM Query Orchestration in Edge EnvironmentsReza Farahani, Zoha Azimi Ourimi, Mario Colosi, Lauri Loven, Christian Timmerer, Schahram Dustdar2026-08-31下载Large language model (LLM) services increasingly process heterogeneous queries with diverse latency, accuracy, and resource requirements. While edge deployment reduces response time, the heterogeneity...
Client-side transparent caching for remote ROOT data analysisDmytro Kovalskyi, Jan Eysermans, Mariarosaria D'Alfonso, Christoph Paus2026-08-31下载High-energy physics analyses often process the same data as physicists refine algorithms and test new ideas. With data increasingly read from remote storage, each iteration is subject to network laten...
Don't Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field ChangesPruthvi Davineni2026-08-31下载LLM agents increasingly diagnose incidents and propose remediations. In a GitOps workflow, applying a fix means editing a version-controlled config file, and the obvious implementation, having the mod...
BlockMGARD: Accelerating Adaptive Scientific Data Reduction with Region-of-Interest Error Control on GPUsYanliang Li, Qian Gong, Qing Liu, Jaemoon Lee, Norbert Podhorszki, Scott Klasky, Xin Liang, Jieyang Chen2026-08-31下载The growing scale of scientific data makes lossy compression essential for reducing data volume under controllable error. Transformation-based compressors using multilevel decomposition, such as MGARD...
Storage-Centric System Designs for Enabling Fast, Efficient, and Low-Cost Genomic and Metagenomic AnalysesNika Mansouri Ghiasi2026-08-31下载Genomic and metagenomic analyses play critical roles in many fields, such as precision medicine, urgent clinical settings, discovering early warnings of communicable diseases, ensuring food safety thr...
Beating Quadratic Time--Message Trade-off in Distributed Minimum Spanning Tree ConstructionTaisuke Izumi, Naoki Kitamura, Toshimitsu Masuzawa2026-08-31下载We present a new distributed algorithm for computing a minimum spanning tree (MST) in the \textsf{CONGEST-KT1_{1}} model, where messages are limited to O(logn)O(\log n) bits and each vertex initially know...
Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic RegretXia Jiang, Lu Liu, Gang Feng2026-08-31下载This paper investigates distributed online optimization for multi-agent dynamical systems with constrained inputs and time-varying cost functions.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Physiological Information Reliability: Cross-Layer Adaptive Resource Allocation for Cardiovascular SensingNavaneeth Krishnan Kamalakannan, Janakiraman Kamalakannan, Harinisri Velmurugan2026-08-31下载Cardiovascular sensing systems must preserve clinically useful information despite signal degradation, wireless losses, energy constraints, and edge-computation latency.
WiSDoM: Wireless Sparse Decision Transformer with Mixture-of-Experts for Multi-Task Mobile Network OptimizationFatih Temiz, Shavbo Salehi, Melike Erol-Kantarci2026-08-31下载Emerging 6G wireless networks are expected to operate across diverse deployment scenarios, where variations in network topology, user mobility, traffic demand, and radio conditions challenge the scala...
Local Private Information Retrieval for Graph-Based Replicated SystemsShreya Meel, Mohamed Nomeir, Sennur Ulukus2026-08-31下载We rethink the definition of privacy in multi-server, graph-replicated private information retrieval (PIR) systems, by introducing a novel setting where the user's privacy is governed by the servers' ...
Semantic Freshness Optimal Sampling and Transmission for Gossiping ReceiversIrtiza Hasan, Ahmed Arafa2026-08-31下载We study the optimal joint sampling and transmission policy for a transmitter communicating with two gossiping receivers that share information with each other, with the objective of tracking a source...
Ray Tracing-Based LoRaWAN Gateway Placement for Reliable Connectivity in Amazonian RegionsCláudio Modesto, Lucas Mozart, Cleverson Nahum, Bruno Castro, Aldebaro Klautau2026-08-31下载Network planning is an important task in wireless communications, as it helps network operators avoid unnecessary costs. In the context of the internet of things, using long-range wide-area network te...
Sensitivity Comparison of Microwave-Frequency and Optical Fibre Interferometry Based on State-of-the-Art ComponentsGeorgios Aias Karydis, Marco Fasano, Paola Parolari, Pierpaolo Boffi, Charis Mesaritakis, Adonis Bogris2026-08-31下载We compare the sensitivity of fibre interferometers to vibrations using microwave oscillators and state of the art lasers over distances up to 70 km.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Adaptive KV Retention for LLM Agents at Human-Approval TimescalesMinseo Choi, Ananya Joshi2026-08-31下载Unlike the seconds-scale tool-call pauses targeted by prior agent-serving systems, agentic LLM requests can be suspended for minutes or hours while waiting for human approval.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Towards Stream Learning on Embedded Systems: Benchmarking the Memory Consumption of Stream Learning MethodsSebastian Buschjäger, Nuwan Gunasekara, Heitor Murilo Gomes2026-08-31下载Stream learning is commonly evaluated through predictive performance and adaptation to concept drift. However, sustained operation of a stream learner also requires predictable and bounded resource us...
LaMoC: Loss-Aware Modular Compression for LLMsMohanad Odema, Jacob Song2026-08-31下载Modular compression has enabled considerable parameter reduction in LLMs while preserving strong language understanding and downstream task accuracy.

基于 VitePress 构建 · 使用本地搜索查找论文