2026-07-09
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration | Minki Jeong, Daegun Yoon, Soohong Ahn, Seungyong Lee, Nameun Kang, Hyeonseok Ju, Ieryung Park, Joonseop Sim, Youngpyo Joo, Hoshik Kim | 2026-07-09 | 下载 | As large language models (LLMs) scale, their memory and computation demands have grown substantially, making weight-only quantization a widely adopted technique for reducing model size with minimal ac... |
| ESBMC-Arduino: Closing the Deployment Gap for Formal Verification of Open-Hardware PLCs | Pierre Dantas, Lucas Cordeiro, Waldir Junior | 2026-07-09 | 下载 | OpenPLC, Arduino OPTA, CONTROLLINO, and Industrial Shields M-Duino bring IEC 61131-3 to low-cost microcontrollers used in real automation and industrial control system (ICS) security research. |
| FPGN: Redefining Ultra-Fast Programmable Gate-based Neural Acceleration with Differentiable LUTs | Jiawei Liang, Haotong Qin, Linfeng Du, Xingyu Liu, Shangkun Li, Hui Yu, Michele Magno, Xinyu Chen, Jiang Xu, Wei Zhang | 2026-07-09 | 下载 | Achieving nanosecond-scale inference latency for deep neural networks (DNNs) has become a primary architectural concern for latency-critical applications. |
| Detecting Ladder Logic Bombs in IEC 61131-3 PLC Programs using ESBMC-PLC+: A Formal Verification Approach with Trigger Synthesis | Pierre Dantas, Lucas Cordeiro, Waldir Junior | 2026-07-09 | 下载 | A Ladder Logic Bomb (LLB) is malicious control logic in a Programmable Logic Controller (PLC) program that lies dormant until a trigger activates a payload to manipulate actuators, forge sensor readin... |
| Who Needs DRAM? We Have Fiber | Hannah Atmer, Thiemo Voigt, Yuan Yao, Stefanos Kaxiras | 2026-07-09 | 下载 | The rising pressure on DRAM availability and contract pricing reflects generative AI's massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expan... |
| CRIMP: Compact & Reliable DNN Inference on In-Memory Processing via Crossbar-Aligned Compression and Non-ideality Adaptation | Shuo Huai, Hao Kong, Xiangzhong Luo, Shiqing Li, Ravi Subramaniam, Christian Makaya, Qian Lin, Weichen Liu | 2026-07-09 | 下载 | Crossbar-based In-Memory Processing (IMP) accelerators achieve high-speed, low-power computing for deep neural networks (DNNs), but face three obstacles. |
| A Theoretical Framework for Stochastic Activity Prediction in Tensor Accelerator Wallace-Tree Multipliers | Prashanthi Metku, Chandra Gandu | 2026-07-09 | 下载 | Tensor accelerator multipliers burn dynamic power on every clock cycle, even when sparse operands require very little internal switching. No existing technique addresses this: zero-detection requires ... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM Inference | Hritvik Taneja, Anish Saxena, Abhishek Revinipati, Jae Hyung Ju, Neal C. Crago, Moinuddin Qureshi | 2026-07-09 | 下载 | The rise of reasoning models and agentic systems has made LLM token-generation latency a key bottleneck. Unlike chatbots, whose latency gains saturate at human reading speed, these systems generate in... |
| Proof-of-Continuity: A Temporal Model for Authority Propagation in Distributed Systems and AI Agents | Nicola Gallo | 2026-07-09 | 下载 | Proof-of-Possession authorization models derive authority from the possession of artifacts such as tokens, credentials, or capabilities. This paper argues that possession is insufficient for discrete ... |
| Secure Decentralized Federated Learning via Gossip and Virtual Voting | Amirhossein Taherpour, Xiaodong Wang | 2026-07-09 | 下载 | Decentralized federated learning (DFL) removes the central server by letting nodes exchange model updates through peer-to-peer gossip, but existing gossip-based methods often lack provenance finality ... |
| SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling | Jiahao Wang, Kaizhan Lin, Kaixi Zhang, Jinbo Han, Xingda Wei, Sijie Shen, Chenguang Fang, Wenyuan Yu, Rong Chen, Haibo Chen | 2026-07-09 | 下载 | LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans. |
| Coded Task Offloading for Fluid Computing: A Privacy-Aware Approach under D2D Networks | Diego Cajaraville-Aboy, Manuel Fernández-Veiga, Ana Fernández-Vilas, Rebeca P. Díaz-Redondo | 2026-07-09 | 下载 | Fluid Computing aims to support distributed applications execution across heterogeneous cloud, edge, and device resources, motivating task execution mechanisms that adapt to dynamic and privacy-sensit... |
| Who Needs DRAM? We Have Fiber | Hannah Atmer, Thiemo Voigt, Yuan Yao, Stefanos Kaxiras | 2026-07-09 | 下载 | The rising pressure on DRAM availability and contract pricing reflects generative AI's massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expan... |
| Parallel QEC Decoding Applied to Distributed Quantum Computing | Gabriele Incardona, Davide Ferrari, Michele Amoretti | 2026-07-09 | 下载 | A novel parallel approach is proposed for QEC decoding based on Belief Propagation with Ordered Statistics Decoding. The main idea is to pre-process the error vectors obtained from Belief Propagation ... |
| Computing in Anonymous Dynamic Networks with One-Bit Communications | Thibaut Blanc, Giuseppe Antonio Di Luna, Giovanni Viglietta | 2026-07-09 | 下载 | We initiate the study of deterministic computation in anonymous dynamic networks where each agent broadcasts one bit per round and receives only the number of neighbors broadcasting each bit value. |
| Adaptive Row Selection Meets Asynchrony in Randomized Kaczmarz | Evan Coleman | 2026-07-09 | 下载 | Randomized Kaczmarz is a natural fit for large sparse least-squares and tomographic reconstruction, and adaptive row selection can reduce iteration counts. |
| Empirical Analysis of GPU Frequency Behavior Under ML Workloads | Truong-Thanh Le, Hoang-Loc La, Amir Taherkordi, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan | 2026-07-09 | 下载 | This work presents ongoing research on the frequency scaling behavior of NVIDIA GPUs when executing ML/AI workloads. Our preliminary findings show that, on lower-performance GPUs, the operating freque... |
| Self-Stabilizing Algorithms in the Uniform Port Model | Liam Brinker, Yuval Emek, Oren Louidor | 2026-07-09 | 下载 | We introduce a distributed computational model referred to as the \emph{uniform port} model. An algorithm operating in this model is defined by means of local automata associated with the ports (a.k. |
| On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend | Zheng Yu | 2026-07-09 | 下载 | Non-GPU AI accelerators are increasingly adopted as alternatives to general-purpose GPUs for large-model inference, but the real engineering cost of migrating demanding workloads beyond CUDA remains p... |
| Securing Autonomous Vehicle Systems via Twin-Aware Federated Reinforcement Learning | Zifan Zhang, Minghong Fang, Dianwei Chen, Zhuqing Liu, Prashant Khanduri, Xianfeng Yang, Anupam Das, Yuchen Liu | 2026-07-09 | 下载 | Federated reinforcement learning (FRL) is crucial for enabling collaborative learning across multiple agents without sharing raw data, thereby enhancing privacy and scalability in the decision-making ... |
| Collate: Collaborative Neural Network Learning for Latency-Critical Edge Systems | Shuo Huai, Di Liu, Hao Kong, Xiangzhong Luo, Weichen Liu, Ravi Subramaniam, Christian Makaya, Qian Lin | 2026-07-09 | 下载 | Federated Learning (FL) empowers multiple clients to collaboratively learn a model, enlarging the training data of each client for high accuracy while protecting data privacy. |
| Toward a Unified GPU-Aware OpenSHMEM Specification | Naveen Ravi, Nathan Wichmann, Md. Wasi-ur- Rahman, Aurelien Bouteiller, Yıltan Hassan Temuçin, Avinash Kethineedi, Johnathan Alsop, Brandon Potter, Shubhendra Pal Singhal, Jun Shirako, Akihiro Hayashi, Vivek Sarkar, Lawrence C. Stewart, Michael Beebe, Benjamin Michalowicz, Jeongnim Kim, Thiago Teixeria, Mark F. Brown, Aaron Welch, Oscar Hernandez, Wendy Poole, Steve Poole | 2026-07-09 | 下载 | Leadership-class HPC systems are now accelerator-centric, with GPUs providing most floating-point throughput and memory bandwidth. As next-generation systems increasingly integrate accelerators throug... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Privacy-Preserving Intent Fulfilment and Assurance for 6G RAN | Joss Armstrong | 2026-07-09 | 下载 | Intent-based network management is the emerging paradigm for 6G service lifecycle automation, with the 3GPP intent management framework (TS~28. |
| Spatio-Temporal Scheduling Prediction Under Backhaul Delay for Resilient Coordinated Beamforming | Prashant Kumar Singh, Shubham Vaishnav, Ahmet Hasim Gökceoglu, Li Wang | 2026-07-09 | 下载 | Coordinated beamforming in distributed 5G networks relies on the timely exchange of inter-cell scheduling information, but backhaul latency makes this information stale. |
| ADORN: Adaptive Drift handling for Open RAN using Reinforcement Learning | Ashit Kumar Subudhi, Bhargav Chirumamilla, Shubham Vaishnav, Mduduzi C. Hlophe, Praveen Kumar Donta, Andrea Fumagalli, Venkateswarlu Gudepu, Koteswararao Kondepu | 2026-07-09 | 下载 | Dynamic traffic variations in Open Radio Access Networks (O-RAN) lead to drift, which degrades the performance of Artificial Intelligence/Machine Learning (AI/ML) models. |
| Who Needs DRAM? We Have Fiber | Hannah Atmer, Thiemo Voigt, Yuan Yao, Stefanos Kaxiras | 2026-07-09 | 下载 | The rising pressure on DRAM availability and contract pricing reflects generative AI's massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expan... |
| Securing Autonomous Vehicle Systems via Twin-Aware Federated Reinforcement Learning | Zifan Zhang, Minghong Fang, Dianwei Chen, Zhuqing Liu, Prashant Khanduri, Xianfeng Yang, Anupam Das, Yuchen Liu | 2026-07-09 | 下载 | Federated reinforcement learning (FRL) is crucial for enabling collaborative learning across multiple agents without sharing raw data, thereby enhancing privacy and scalability in the decision-making ... |
| MORES: Mobile Reasoning-as-a-Service via Distributed LLM Inference-Time Scaling | Guanchen Liu, Hongyang Du, Kaibin Huang | 2026-07-09 | 下载 | Inference-time scaling has emerged as an effective approach for enhancing the capabilities of Large Language Models (LLMs), addressing the growing demand for stronger reasoning without increasing mode... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A Quantized Native Runtime for On-Device Semantic Audio Generation | Matteo Spanio, Antonio Rodà | 2026-07-09 | 下载 | Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks. |
| Parallel QEC Decoding Applied to Distributed Quantum Computing | Gabriele Incardona, Davide Ferrari, Michele Amoretti | 2026-07-09 | 下载 | A novel parallel approach is proposed for QEC decoding based on Belief Propagation with Ordered Statistics Decoding. The main idea is to pre-process the error vectors obtained from Belief Propagation ... |