Skip to content

2026-08-06 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Density-Functional Excited-State Gradients and Nonadiabatic Couplings on a Consumer GPU from a Contraction-DAGRubén Darío Guerrero2026-08-06下载Nonadiabatic dynamics needs an excited-state gradient and an interstate nonadiabatic coupling matrix element (NACME) at every nuclear geometry, and a double-hybrid functional's accuracy has been unava...
Breaking Memory Bottlenecks in Quantum Control Systems for More Precise Experiments and Higher Throughput ComputingYicheng Guang, Neel Vora, Yilun Xu, Yueqi Chen, Gang Huang2026-08-06下载As quantum computing continues to demonstrate promise and attract growing attention, there is an increasing need for more precise experiments to advance the development of quantum devices, as well as ...
Automated Synthesis of Heterogeneous, Hierarchical, Scoped Coherence ProtocolsFletch Rydell, An Qi Zhang, Nicolai Oswald, Andres Goens, Vijay Nagarajan, Daniel Sorin2026-08-06下载Processor design is converging on a new model of cache-coherent shared memory characterized by heterogeneity, hierarchy, and scopes. Protocols like CXL or AMBA CHI are used as global protocols to comb...
An Open-Source Power Measurement Platform for System-Level Semiconductor TestingLinus Bantel, Sarah Rottacker, Dirk Pflüger2026-08-06下载Accurate power measurement is not only essential for evaluating the energy efficiency of modern embedded and semiconductor systems, but power draw is an important proxy during stress testing.
A Low-Latency ASIC Architecture for Real-Time Line Segment DetectionAmir Hossein Jalilvand, Parsa Hassani Shariat Panahi, M. Hassan Najafi2026-08-06下载Line segment detection is a critical preprocessing step in embedded vision applications such as autonomous navigation, visual SLAM, and industrial inspection.
Zero-Instruction Sensor Reads: Register-Mapped Peripherals and Hardware PWM on a Five-Stage Soft ProcessorNathanael Ren2026-08-06下载We present a case study in application-driven specialization of a five-stage soft processor, evaluated on the inner control loop of a reaction-wheel self-balancing bicycle.
PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA ServingZhongkai Yu, Ohm Rishabh Venkatachalam, Zheng Wang, Yikai Li, Yichen Lin, Zihao Yu, Yuke Wang, Liu Liu, Xulong Tang, Shuyi Pei, Yangwook Kang, Yufei Ding2026-08-06下载Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
MARS: A Monte Carlo Tree Search-based Adaptive and Responsive SchedulerYash Kurkure, Yihe Zhang, Zhiling Lan, Michael E. Papka2026-08-06下载Modern High Performance Computing systems depend on static heuristics and manual administration for job scheduling and reservation management.
Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference servingMuhammad Adnan, Rohan Mahapatra, Prashant J. Nair, Daniel Berger, Pantea Zardoshti, Rodrigo Fonseca, Esha Choukse2026-08-06下载The reasoning and agentic capabilities of large language models have expanded the range of applications they support, from short interactive exchanges to long, compute-heavy requests.
Rendezvous of Mobile Deterministic Automata in GraphsBibhuti Das, Andrzej Pelc2026-08-06下载Two mobile agents, modeled as identical deterministic finite automata (DFA) navigating in synchronous rounds in a graph with unlabeled nodes, have to meet at some node.
Routing LLM Inference to the Cleanest Grid in Real TimeAleks Bernhard, Arif Baran Yardimci2026-08-06下载Large-language-model inference is a fast-growing electricity load whose marginal carbon intensity varies by more than an order of magnitude across grid regions and across the day, making request place...
PLB: Priority-Aware Load Balancing for Replicated Databases under Constrained ResourcesBelkis Djeffal, Pierre Bourhis, Romain Rouvoy2026-08-06下载Priority-differentiated services are a standard way for applications to offer different levels of performance, but database systems still often treat all sessions the same way.
ML-for-MLYutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat2026-08-06下载AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running w...
TensorCast: The Missing Tensor Management Layer in Large Language Model InfrastructureYuhan Zhou, Yuchu Luo, Hao Nie, Wangrunze Lv, Yu Zhou, Yibo Zhu, Daxin Jiang, Chenren Xu2026-08-06下载Modern LLM infrastructure increasingly manages tensors not only as computation data, but also as persistent states shared across distributed components.
Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational HardeningSeon Ho Kim, Ui Jeong Jeon, Su Hyeon Kim, Min Tae Hwang2026-08-06下载We report operational experience full-fine-tuning a 32.76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among the first published field accounts on this accelera...
SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding PredictionGuofan Yu, Sitian Chen, Zhenheng Tang, Xiaowen Chu, Amelie Chi Zhou2026-08-06下载Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges.
RepoOMP: Repository-Aware Hotspot OpenMP Parallelization via Dependency-Aware Context ReductionYongjie Qian, Ke Gao, Zhibin Zhang, Shaohui Peng, Ling Li2026-08-06下载OpenMP parallelization of hotspots in mature repositories remains difficult because loop safety and optimization payoff often depend on non-local evidence.
Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit SimulationAlfred M. Pastor, Maribel Castillo, Jose M. Badia2026-08-06下载Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network contraction can reduce this cost by exploiting cir...
Wireless Linear Computation BroadcastShuo Tan, Syed A. Jafar2026-08-06下载A linear computation broadcast (LCBC) problem comprises KK users (receivers) and a transmitter. The users wish to compute various (vector) linear functions of a common dataset, and possess in advance...
Serverless platform driven CPU loadbalancingAbdul Rehman2026-08-06下载Serverless platforms maintain a global view of function invocations and resource utilization, yet existing systems largely restrict CPU scheduling decisions to the operating system scheduler.
How Much Reconstruction Does Quantum Machine Learning Need? Late Fusion of Independently Trained Quantum SubcircuitsPrabhjot Singh, Adel N. Toosi, Rajkumar Buyya2026-08-06下载Circuit cutting lets a large quantum neural network (QNN) run as independent subcircuits on small devices, but rebuilding its outputs by reconstruction carries a classical sampling overhead exponentia...
Viveka: Context-Aware Sensing for Energy Efficiency in Smart WearablesNikhil Sreekumar, Abhishek Chandra2026-08-06下载The proliferation of multi-sensor Internet of Things (IoT) systems, from Body Sensor Networks (BSNs) to industrial monitoring, is increasingly constrained by strict energy budgets and limited on-devic...
PLoRA: An NDP-Enhanced Pooled-Memory System for Cost-Efficient Multi-LoRA ServingZhongkai Yu, Ohm Rishabh Venkatachalam, Zheng Wang, Yikai Li, Yichen Lin, Zihao Yu, Yuke Wang, Liu Liu, Xulong Tang, Shuyi Pei, Yangwook Kang, Yufei Ding2026-08-06下载Multi-LoRA serving is how one base model becomes thousands of specialized variants, one adapter per user, task, or agent, and the deployments can hold 1000-plus adapters.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Improving the Energy Efficiency of High Throughput Computing: A Measurement-Based Case StudyDamu Ding, Xinpeng Hong, Alastair Dewhurst, James Walder, Daniel Schien, David Greenwood, Noa Zilberman2026-08-06下载The significant energy consumed by data centers has become a concern both for costs and associated carbon emissions. In particular, the energy efficiency of servers is a key consideration for data cen...
Detecting and Characterizing Massively Shared IP AddressesAmanda Hsu, Paul Pearce, Frank Li, Arthur Berger, Philipp Richter2026-08-06下载IP addresses are commonly shared across devices and users for a variety of reasons, including NAT and proxies. These technologies operate at different scales, from residential NATs that share an IP ad...
From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over NetworksChristo Kurisummoottil Thomas, Omar Hashash, Walid Saad2026-08-06下载Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, still fail when embedded into physical systems, such as robots an...
FedTransKD-IDS: Robust Federated Transfer Learning with Knowledge Distillation for Intrusion Detection in IoTMohammad Hosssein Gholamrezazadeh, Ahmadreza MontazerolghaemAhmadreza Montazerolghaem2026-08-06下载In modern distributed network environments, particularly in Internet of Things infrastructures and 5G networks, stringent privacy preservation and scalability requirements have created significant cha...
MultiMoQ: Multi-Access Media-Over-QUIC for Robust Immersive Video StreamingYitong Li, Xinjiao Li, Ruonan Chai, Dirk Kutscher2026-08-06下载Live immersive video streaming, particularly 360-degree video, is increasingly adopted in applications such as virtual events, sports broadcasting, and remote education.
MARS: Multipath Adaptive Reliable ServiceYitong Li, Xinjiao Li, Dirk Kutscher2026-08-06下载Multipath transport is increasingly important for Internet/WAN services that move large data volumes across heterogeneous paths, including geo-distributed analytics, content distribution, and cloud-se...
ML-for-MLYutong Zhao, Noga H. Rotman, Gianni Antichi, Ran Ben Basat2026-08-06下载AI training workloads are growing rapidly, making their time, energy, and infrastructure costs increasingly important. In shared cloud clusters, training and fine-tuning jobs compete with co-running w...
BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge NetworksGuanqiao Qu, Shuo Chen, Qian Chen, Kin K. Leung, Xianhao Chen2026-08-06下载Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inference mainly relies on two approaches: Autoregressive decodi...
5G ISAC-Based UAV Detection and 3-D Tracking Using Uplink Sounding Reference Signals on an End-to-End O-RAN Simulation TestbedArun K. Gurung, Satha K. Sathananthan, Shiva R. Pokhrel2026-08-06下载Integrated Sensing and Communication (ISAC) lets cellular infrastructure serve communication users and sense on the same waveform. We present an end-to-end O-RAN simulation testbed for 5G ISAC targeti...
Closed-Loop Decision-Focused Learning for User-Aware Cloud Orchestration under UncertaintyDongbin Jiao, Xubo Zhang, Huakang Lin, Ke Shang, Shi Yan2026-08-06下载Time-varying cloud workloads often cause resource under-utilization during off-peak periods and resource contention during peak periods. Existing prediction-then-optimization (PTO) frameworks suffer f...
DTMC-Based Analysis and Scheduling for Periodic Flows with Proactive HARQHaozhe Yi, Junyi Liu, Maolin Yang, Haochun Liang, Bo Liu, Feng Hong, Chaowei Liu, Hongbiao Liu2026-08-06下载Ultra-Reliable Low-Latency Communication (URLLC) requires strict reliability and latency guarantees for heterogeneous periodic traffic. Proactive HARQ improves resource efficiency through early termin...
Enhancing Anomaly Resilience in Research Networks: A Large-Scale Forecasting Benchmark for Dynamic Security BaseliningMohammad Arafath Uddin Shariff, Byrav Ramamurthy2026-08-06下载Research and Education Networks (RENs) serve as critical infrastructure for scientific discovery, yet they face a unique security paradox: their normal traffic patterns which are characterized by mass...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Timestep-Conditioned Transformers for Global Weather ForecastingSam Levang, Fran Bartolic, Ty Dickinson, Chase Dwelle, Paulius Rauba, Viktor Cikojevic2026-08-06下载Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. The choice of model timestep involves a fundamental trade-off: shorter timesteps (e.g.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Routing LLM Inference to the Cleanest Grid in Real TimeAleks Bernhard, Arif Baran Yardimci2026-08-06下载Large-language-model inference is a fast-growing electricity load whose marginal carbon intensity varies by more than an order of magnitude across grid regions and across the day, making request place...
ASGE-RR: Agentic Service Graph Embedding with Revisable Reservations for Dynamic AI-Agent CallsTrond Vatten, Yuming Jiang2026-08-06下载AI-agent workflows often involve remote calls to models, memory stores, and tools distributed across a network. As execution progresses, these dependency calls collectively form an agentic service gra...
Hybrid-Adaptive Thread Tuning to Mitigate Simulation Execution Bottlenecks in High-Performance Reinforcement Learning InferenceJiming Su, Hantao Hua, Lujia Yin, Yiping Yao, Feng Zhu2026-08-06下载In simulation-in-the-loop decision-making systems, reinforcement learning (RL) inference is often constrained by simulator-side execution overhead, where workloads are highly dynamic and sensitive to ...
An Open-Source Power Measurement Platform for System-Level Semiconductor TestingLinus Bantel, Sarah Rottacker, Dirk Pflüger2026-08-06下载Accurate power measurement is not only essential for evaluating the energy efficiency of modern embedded and semiconductor systems, but power draw is an important proxy during stress testing.
Learning to Rank Tensor Network Contraction Plans for GPU-Accelerated Quantum Circuit SimulationAlfred M. Pastor, Maribel Castillo, Jose M. Badia2026-08-06下载Classical simulation remains essential for developing and validating quantum algorithms, but its cost grows rapidly with circuit size. Tensor-network contraction can reduce this cost by exploiting cir...

基于 VitePress 构建 · 使用本地搜索查找论文