Skip to content

2026-07-09 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference AccelerationMinki Jeong, Daegun Yoon, Soohong Ahn, Seungyong Lee, Nameun Kang, Hyeonseok Ju, Ieryung Park, Joonseop Sim, Youngpyo Joo, Hoshik Kim2026-07-09下载As large language models (LLMs) scale, their memory and computation demands have grown substantially, making weight-only quantization a widely adopted technique for reducing model size with minimal ac...
ESBMC-Arduino: Closing the Deployment Gap for Formal Verification of Open-Hardware PLCsPierre Dantas, Lucas Cordeiro, Waldir Junior2026-07-09下载OpenPLC, Arduino OPTA, CONTROLLINO, and Industrial Shields M-Duino bring IEC 61131-3 to low-cost microcontrollers used in real automation and industrial control system (ICS) security research.
FPGN: Redefining Ultra-Fast Programmable Gate-based Neural Acceleration with Differentiable LUTsJiawei Liang, Haotong Qin, Linfeng Du, Xingyu Liu, Shangkun Li, Hui Yu, Michele Magno, Xinyu Chen, Jiang Xu, Wei Zhang2026-07-09下载Achieving nanosecond-scale inference latency for deep neural networks (DNNs) has become a primary architectural concern for latency-critical applications.
Detecting Ladder Logic Bombs in IEC 61131-3 PLC Programs using ESBMC-PLC+: A Formal Verification Approach with Trigger SynthesisPierre Dantas, Lucas Cordeiro, Waldir Junior2026-07-09下载A Ladder Logic Bomb (LLB) is malicious control logic in a Programmable Logic Controller (PLC) program that lies dormant until a trigger activates a payload to manipulate actuators, forge sensor readin...
Who Needs DRAM? We Have FiberHannah Atmer, Thiemo Voigt, Yuan Yao, Stefanos Kaxiras2026-07-09下载The rising pressure on DRAM availability and contract pricing reflects generative AI's massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expan...
CRIMP: Compact & Reliable DNN Inference on In-Memory Processing via Crossbar-Aligned Compression and Non-ideality AdaptationShuo Huai, Hao Kong, Xiangzhong Luo, Shiqing Li, Ravi Subramaniam, Christian Makaya, Qian Lin, Weichen Liu2026-07-09下载Crossbar-based In-Memory Processing (IMP) accelerators achieve high-speed, low-power computing for deep neural networks (DNNs), but face three obstacles.
A Theoretical Framework for Stochastic Activity Prediction in Tensor Accelerator Wallace-Tree MultipliersPrashanthi Metku, Chandra Gandu2026-07-09下载Tensor accelerator multipliers burn dynamic power on every clock cycle, even when sparse operands require very little internal switching. No existing technique addresses this: zero-detection requires ...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
SiFAR: Synchronization-Free All-Reduce for Low-Latency LLM InferenceHritvik Taneja, Anish Saxena, Abhishek Revinipati, Jae Hyung Ju, Neal C. Crago, Moinuddin Qureshi2026-07-09下载The rise of reasoning models and agentic systems has made LLM token-generation latency a key bottleneck. Unlike chatbots, whose latency gains saturate at human reading speed, these systems generate in...
Proof-of-Continuity: A Temporal Model for Authority Propagation in Distributed Systems and AI AgentsNicola Gallo2026-07-09下载Proof-of-Possession authorization models derive authority from the possession of artifacts such as tokens, credentials, or capabilities. This paper argues that possession is insufficient for discrete ...
Secure Decentralized Federated Learning via Gossip and Virtual VotingAmirhossein Taherpour, Xiaodong Wang2026-07-09下载Decentralized federated learning (DFL) removes the central server by letting nodes exchange model updates through peer-to-peer gossip, but existing gossip-based methods often lack provenance finality ...
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric SchedulingJiahao Wang, Kaizhan Lin, Kaixi Zhang, Jinbo Han, Xingda Wei, Sijie Shen, Chenguang Fang, Wenyuan Yu, Rong Chen, Haibo Chen2026-07-09下载LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by agents instead of humans.
Coded Task Offloading for Fluid Computing: A Privacy-Aware Approach under D2D NetworksDiego Cajaraville-Aboy, Manuel Fernández-Veiga, Ana Fernández-Vilas, Rebeca P. Díaz-Redondo2026-07-09下载Fluid Computing aims to support distributed applications execution across heterogeneous cloud, edge, and device resources, motivating task execution mechanisms that adapt to dynamic and privacy-sensit...
Who Needs DRAM? We Have FiberHannah Atmer, Thiemo Voigt, Yuan Yao, Stefanos Kaxiras2026-07-09下载The rising pressure on DRAM availability and contract pricing reflects generative AI's massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expan...
Parallel QEC Decoding Applied to Distributed Quantum ComputingGabriele Incardona, Davide Ferrari, Michele Amoretti2026-07-09下载A novel parallel approach is proposed for QEC decoding based on Belief Propagation with Ordered Statistics Decoding. The main idea is to pre-process the error vectors obtained from Belief Propagation ...
Computing in Anonymous Dynamic Networks with One-Bit CommunicationsThibaut Blanc, Giuseppe Antonio Di Luna, Giovanni Viglietta2026-07-09下载We initiate the study of deterministic computation in anonymous dynamic networks where each agent broadcasts one bit per round and receives only the number of neighbors broadcasting each bit value.
Adaptive Row Selection Meets Asynchrony in Randomized KaczmarzEvan Coleman2026-07-09下载Randomized Kaczmarz is a natural fit for large sparse least-squares and tomographic reconstruction, and adaptive row selection can reduce iteration counts.
Empirical Analysis of GPU Frequency Behavior Under ML WorkloadsTruong-Thanh Le, Hoang-Loc La, Amir Taherkordi, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan2026-07-09下载This work presents ongoing research on the frequency scaling behavior of NVIDIA GPUs when executing ML/AI workloads. Our preliminary findings show that, on lower-performance GPUs, the operating freque...
Self-Stabilizing Algorithms in the Uniform Port ModelLiam Brinker, Yuval Emek, Oren Louidor2026-07-09下载We introduce a distributed computational model referred to as the \emph{uniform port} model. An algorithm operating in this model is defined by means of local automata associated with the ports (a.k.
On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei AscendZheng Yu2026-07-09下载Non-GPU AI accelerators are increasingly adopted as alternatives to general-purpose GPUs for large-model inference, but the real engineering cost of migrating demanding workloads beyond CUDA remains p...
Securing Autonomous Vehicle Systems via Twin-Aware Federated Reinforcement LearningZifan Zhang, Minghong Fang, Dianwei Chen, Zhuqing Liu, Prashant Khanduri, Xianfeng Yang, Anupam Das, Yuchen Liu2026-07-09下载Federated reinforcement learning (FRL) is crucial for enabling collaborative learning across multiple agents without sharing raw data, thereby enhancing privacy and scalability in the decision-making ...
Collate: Collaborative Neural Network Learning for Latency-Critical Edge SystemsShuo Huai, Di Liu, Hao Kong, Xiangzhong Luo, Weichen Liu, Ravi Subramaniam, Christian Makaya, Qian Lin2026-07-09下载Federated Learning (FL) empowers multiple clients to collaboratively learn a model, enlarging the training data of each client for high accuracy while protecting data privacy.
Toward a Unified GPU-Aware OpenSHMEM SpecificationNaveen Ravi, Nathan Wichmann, Md. Wasi-ur- Rahman, Aurelien Bouteiller, Yıltan Hassan Temuçin, Avinash Kethineedi, Johnathan Alsop, Brandon Potter, Shubhendra Pal Singhal, Jun Shirako, Akihiro Hayashi, Vivek Sarkar, Lawrence C. Stewart, Michael Beebe, Benjamin Michalowicz, Jeongnim Kim, Thiago Teixeria, Mark F. Brown, Aaron Welch, Oscar Hernandez, Wendy Poole, Steve Poole2026-07-09下载Leadership-class HPC systems are now accelerator-centric, with GPUs providing most floating-point throughput and memory bandwidth. As next-generation systems increasingly integrate accelerators throug...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Privacy-Preserving Intent Fulfilment and Assurance for 6G RANJoss Armstrong2026-07-09下载Intent-based network management is the emerging paradigm for 6G service lifecycle automation, with the 3GPP intent management framework (TS~28.
Spatio-Temporal Scheduling Prediction Under Backhaul Delay for Resilient Coordinated BeamformingPrashant Kumar Singh, Shubham Vaishnav, Ahmet Hasim Gökceoglu, Li Wang2026-07-09下载Coordinated beamforming in distributed 5G networks relies on the timely exchange of inter-cell scheduling information, but backhaul latency makes this information stale.
ADORN: Adaptive Drift handling for Open RAN using Reinforcement LearningAshit Kumar Subudhi, Bhargav Chirumamilla, Shubham Vaishnav, Mduduzi C. Hlophe, Praveen Kumar Donta, Andrea Fumagalli, Venkateswarlu Gudepu, Koteswararao Kondepu2026-07-09下载Dynamic traffic variations in Open Radio Access Networks (O-RAN) lead to drift, which degrades the performance of Artificial Intelligence/Machine Learning (AI/ML) models.
Who Needs DRAM? We Have FiberHannah Atmer, Thiemo Voigt, Yuan Yao, Stefanos Kaxiras2026-07-09下载The rising pressure on DRAM availability and contract pricing reflects generative AI's massive high-performance memory requirements. This pressure is heavily compounded by hyperscale data center expan...
Securing Autonomous Vehicle Systems via Twin-Aware Federated Reinforcement LearningZifan Zhang, Minghong Fang, Dianwei Chen, Zhuqing Liu, Prashant Khanduri, Xianfeng Yang, Anupam Das, Yuchen Liu2026-07-09下载Federated reinforcement learning (FRL) is crucial for enabling collaborative learning across multiple agents without sharing raw data, thereby enhancing privacy and scalability in the decision-making ...
MORES: Mobile Reasoning-as-a-Service via Distributed LLM Inference-Time ScalingGuanchen Liu, Hongyang Du, Kaibin Huang2026-07-09下载Inference-time scaling has emerged as an effective approach for enhancing the capabilities of Large Language Models (LLMs), addressing the growing demand for stronger reasoning without increasing mode...

cs.PF - Performance ​

标题作者发布日期PDF摘要
A Quantized Native Runtime for On-Device Semantic Audio GenerationMatteo Spanio, Antonio Rodà2026-07-09下载Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks.
Parallel QEC Decoding Applied to Distributed Quantum ComputingGabriele Incardona, Davide Ferrari, Michele Amoretti2026-07-09下载A novel parallel approach is proposed for QEC decoding based on Belief Propagation with Ordered Statistics Decoding. The main idea is to pre-process the error vectors obtained from Belief Propagation ...

基于 VitePress 构建 · 使用本地搜索查找论文