2026-08-06
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Density-Functional Excited-State Gradients and Nonadiabatic Couplings on a Consumer GPU from a Contraction-DAG | Rubén Darío Guerrero | 2026-08-06 | 下载 | Nonadiabatic dynamics needs an excited-state gradient and an interstate nonadiabatic coupling matrix element (NACME) at every nuclear geometry, and a double-hybrid functional's accuracy has been unava... |
| Breaking Memory Bottlenecks in Quantum Control Systems for More Precise Experiments and Higher Throughput Computing | Yicheng Guang, Neel Vora, Yilun Xu, Yueqi Chen, Gang Huang | 2026-08-06 | 下载 | As quantum computing continues to demonstrate promise and attract growing attention, there is an increasing need for more precise experiments to advance the development of quantum devices, as well as ... |
| Automated Synthesis of Heterogeneous, Hierarchical, Scoped Coherence Protocols | Fletch Rydell, An Qi Zhang, Nicolai Oswald, Andres Goens, Vijay Nagarajan, Daniel Sorin | 2026-08-06 | 下载 | Processor design is converging on a new model of cache-coherent shared memory characterized by heterogeneity, hierarchy, and scopes. Protocols like CXL or AMBA CHI are used as global protocols to comb... |
| An Open-Source Power Measurement Platform for System-Level Semiconductor Testing | Linus Bantel, Sarah Rottacker, Dirk Pflüger | 2026-08-06 | 下载 | Accurate power measurement is not only essential for evaluating the energy efficiency of modern embedded and semiconductor systems, but power draw is an important proxy during stress testing. |
| A Low-Latency ASIC Architecture for Real-Time Line Segment Detection | Amir Hossein Jalilvand, Parsa Hassani Shariat Panahi, M. Hassan Najafi | 2026-08-06 | 下载 | Line segment detection is a critical preprocessing step in embedded vision applications such as autonomous navigation, visual SLAM, and industrial inspection. |
| Zero-Instruction Sensor Reads: Register-Mapped Peripherals and Hardware PWM on a Five-Stage Soft Processor | Nathanael Ren | 2026-08-06 | 下载 | We present a case study in application-driven specialization of a five-stage soft processor, evaluated on the inner control loop of a reaction-wheel self-balancing bicycle. |
| PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving | Zhongkai Yu, Ohm Rishabh Venkatachalam, Zheng Wang, Yikai Li, Yichen Lin, Zihao Yu, Yuke Wang, Liu Liu, Xulong Tang, Shuyi Pei, Yangwook Kang, Yufei Ding | 2026-08-06 | 下载 | Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| MARS: A Monte Carlo Tree Search-based Adaptive and Responsive Scheduler | Yash Kurkure, Yihe Zhang, Zhiling Lan, Michael E. Papka | 2026-08-06 | 下载 | Modern High Performance Computing systems depend on static heuristics and manual administration for job scheduling and reservation management. |
| Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving | Muhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse | 2026-08-06 | 下载 | The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests. |
| Rendezvous of Mobile Deterministic Automata in Graphs | Bibhuti Das, Andrzej Pelc | 2026-08-06 | 下载 | Two mobile agents, modeled as identical deterministic finite automata (DFA) navigating in synchronous rounds in a graph with unlabeled nodes, have to meet at some node. |
| Routing LLM Inference to the Cleanest Grid in Real Time | Aleks Bernhard, Arif Baran Yardimci | 2026-08-06 | 下载 | Large-language-model inference is a fast-growing electricity load whose marginal carbon intensity varies by more than an order of magnitude across grid regions and across the day, making request place... |
| PLB: Priority-Aware Load Balancing for Replicated Databases under Constrained Resources | Belkis Djeffal, Pierre Bourhis, Romain Rouvoy | 2026-08-06 | 下载 | Priority-differentiated services are a standard way for applications to offer different levels of performance, but database systems still often treat all sessions the same way. |
| ML-for-ML | Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat | 2026-08-06 | 下载 | AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running w... |
| TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure | Yuhan Zhou, Yuchu Luo, Hao Nie, Wangrunze Lv, Yu Zhou, Yibo Zhu, Daxin Jiang, Chenren Xu | 2026-08-06 | 下载 | Modern LLM infrastructure increasingly manages tensors not only as computation data, but also as persistent states shared across distributed components. |
| Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational Hardening | Seon Ho Kim, Ui Jeong Jeon, Su Hyeon Kim, Min Tae Hwang | 2026-08-06 | 下载 | We report operational experience full-fine-tuning a 32.76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among the first published field accounts on this accelera... |
| SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction | Guofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu, Amelie Chi Zhou | 2026-08-06 | 下载 | Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. |
| RepoOMP: Repository-Aware Hotspot OpenMP Parallelization via Dependency-Aware Context Reduction | Yongjie Qian, Ke Gao, Zhibin Zhang, Shaohui Peng, Ling Li | 2026-08-06 | 下载 | OpenMP parallelization of hotspots in mature repositories remains difficult because loop safety and optimization payoff often depend on non-local evidence. |
| Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation | Alfred M. Pastor, Maribel Castillo, Jose M. Badia | 2026-08-06 | 下载 | Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network contraction can reduce this cost by exploiting cir... |
| Wireless Linear Computation Broadcast | Shuo Tan, Syed A. Jafar | 2026-08-06 | 下载 | A linear computation broadcast (LCBC) problem comprises users (receivers) and a transmitter. The users wish to compute various (vector) linear functions of a common dataset, and possess in advance... |
| Serverless platform driven CPU loadbalancing | Abdul Rehman | 2026-08-06 | 下载 | Serverless platforms maintain a global view of function invocations and resource utilization, yet existing systems largely restrict CPU scheduling decisions to the operating system scheduler. |
| How Much Reconstruction Does Quantum Machine Learning Need? Late Fusion of Independently Trained Quantum Subcircuits | Prabhjot Singh, Adel N. Toosi, Rajkumar Buyya | 2026-08-06 | 下载 | Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling overhead exponentia... |
| Viveka: Context-Aware Sensing for Energy Efficiency in Smart Wearables | Nikhil Sreekumar, Abhishek Chandra | 2026-08-06 | 下载 | The proliferation of multi-sensor Internet of Things (IoT) systems, from Body Sensor Networks (BSNs) to industrial monitoring, is increasingly constrained by strict energy budgets and limited on-devic... |
| PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA Serving | Zhongkai Yu, Ohm Rishabh Venkatachalam, Zheng Wang, Yikai Li, Yichen Lin, Zihao Yu, Yuke Wang, Liu Liu, Xulong Tang, Shuyi Pei, Yangwook Kang, Yufei Ding | 2026-08-06 | 下载 | Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Improving the Energy Efficiency of High Throughput Computing: A Measurement-Based Case Study | Damu Ding, Xinpeng Hong, Alastair Dewhurst, James Walder, Daniel Schien, David Greenwood, Noa Zilberman | 2026-08-06 | 下载 | The significant energy consumed by data centers has become a concern both for costs and associated carbon emissions. In particular, the energy efficiency of servers is a key consideration for data cen... |
| Detecting and Characterizing Massively Shared IP Addresses | Amanda Hsu, Paul Pearce, Frank Li, Arthur Berger, Philipp Richter | 2026-08-06 | 下载 | IP addresses are commonly shared across devices and users for a variety of reasons, including NAT and proxies. These technologies operate at different scales, from residential NATs that share an IP ad... |
| From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks | Christo Kurisummoottil Thomas, Omar Hashash, Walid Saad | 2026-08-06 | 下载 | Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, still fail when embedded into physical systems, such as robots an... |
| FedTransKD-IDS: Robust Federated Transfer Learning with Knowledge Distillation for Intrusion Detection in IoT | Mohammad Hosssein Gholamrezazadeh, Ahmadreza MontazerolghaemAhmadreza Montazerolghaem | 2026-08-06 | 下载 | In modern distributed network environments, particularly in Internet of Things infrastructures and 5G networks, stringent privacy preservation and scalability requirements have created significant cha... |
| MultiMoQ: Multi-Access Media-Over-QUIC for Robust Immersive Video Streaming | Yitong Li, Xinjiao Li, Ruonan Chai, Dirk Kutscher | 2026-08-06 | 下载 | Live immersive video streaming, particularly 360-degree video, is increasingly adopted in applications such as virtual events, sports broadcasting, and remote education. |
| MARS: Multipath Adaptive Reliable Service | Yitong Li, Xinjiao Li, Dirk Kutscher | 2026-08-06 | 下载 | Multipath transport is increasingly important for Internet/WAN services that move large data volumes across heterogeneous paths, including geo-distributed analytics, content distribution, and cloud-se... |
| ML-for-ML | Yutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat | 2026-08-06 | 下载 | AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running w... |
| BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks | Guanqiao Qu, Shuo Chen, Qian Chen, Kin K. Leung, Xianhao Chen | 2026-08-06 | 下载 | Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decodi... |
| 5G ISAC-Based UAV Detection and 3-D Tracking Using Uplink Sounding Reference Signals on an End-to-End O-RAN Simulation Testbed | Arun K. Gurung, Satha K. Sathananthan, Shiva R. Pokhrel | 2026-08-06 | 下载 | Integrated Sensing and Communication (ISAC) lets cellular infrastructure serve communication users and sense on the same waveform. We present an end-to-end O-RAN simulation testbed for 5G ISAC targeti... |
| Closed-Loop Decision-Focused Learning for User-Aware Cloud Orchestration under Uncertainty | Dongbin Jiao, Xubo Zhang, Huakang Lin, Ke Shang, Shi Yan | 2026-08-06 | 下载 | Time-varying cloud workloads often cause resource under-utilization during off-peak periods and resource contention during peak periods. Existing prediction-then-optimization (PTO) frameworks suffer f... |
| DTMC-Based Analysis and Scheduling for Periodic Flows with Proactive HARQ | Haozhe Yi, Junyi Liu, Maolin Yang, Haochun Liang, Bo Liu, Feng Hong, Chaowei Liu, Hongbiao Liu | 2026-08-06 | 下载 | Ultra-Reliable Low-Latency Communication (URLLC) requires strict reliability and latency guarantees for heterogeneous periodic traffic. Proactive HARQ improves resource efficiency through early termin... |
| Enhancing Anomaly Resilience in Research Networks: A Large-Scale Forecasting Benchmark for Dynamic Security Baselining | Mohammad Arafath Uddin Shariff, Byrav Ramamurthy | 2026-08-06 | 下载 | Research and Education Networks (RENs) serve as critical infrastructure for scientific discovery, yet they face a unique security paradox: their normal traffic patterns which are characterized by mass... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Timestep-Conditioned Transformers for Global Weather Forecasting | Sam Levang, Fran Bartolic, Ty Dickinson, Chase Dwelle, Paulius Rauba, Viktor Cikojevic | 2026-08-06 | 下载 | Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. The choice of model timestep involves a fundamental trade-off: shorter timesteps (e.g. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Routing LLM Inference to the Cleanest Grid in Real Time | Aleks Bernhard, Arif Baran Yardimci | 2026-08-06 | 下载 | Large-language-model inference is a fast-growing electricity load whose marginal carbon intensity varies by more than an order of magnitude across grid regions and across the day, making request place... |
| ASGE-RR: Agentic Service Graph Embedding with Revisable Reservations for Dynamic AI-Agent Calls | Trond Vatten, Yuming Jiang | 2026-08-06 | 下载 | AI-agent workflows often involve remote calls to models, memory stores, and tools distributed across a network. As execution progresses, these dependency calls collectively form an agentic service gra... |
| Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning Inference | Jiming Su, Hantao Hua, Lujia Yin, Yiping Yao, Feng Zhu | 2026-08-06 | 下载 | In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to ... |
| An Open-Source Power Measurement Platform for System-Level Semiconductor Testing | Linus Bantel, Sarah Rottacker, Dirk Pflüger | 2026-08-06 | 下载 | Accurate power measurement is not only essential for evaluating the energy efficiency of modern embedded and semiconductor systems, but power draw is an important proxy during stress testing. |
| Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit Simulation | Alfred M. Pastor, Maribel Castillo, Jose M. Badia | 2026-08-06 | 下载 | Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network contraction can reduce this cost by exploiting cir... |