2026-09-28
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Hardware-Aware Functional Kolmogorov-Arnold Networks for Efficient Medical Image Enhancement and Segmentation | Mohammad Sadegh Sirjani | 2026-09-28 | 下载 | Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their 11.6 M parameters and 8. |
| Automated Pre-Silicon Verification of High-Speed DDR5 and LPDDR5/6 Memory Controllers: Closed-Loop Timing, Mode Register, and PHY Synchronization in UVM | Manan Patel, Anirban Majumder | 2026-09-28 | 下载 | External memory interfaces (such as LPDDR5/4 and DDR5) are essential components of modern mobile, cloud, and enterprise computing systems. While memory manufacturers focus on physical DRAM die develop... |
| GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization | Peng Xu, Nihar Koganti, Volodymyr Kindratenko, Xiaohui Chen | 2026-09-28 | 下载 | Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machine learning problems. |
| Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization | Vlad-Petru Nitu, Harsh Songara, Konstantinos Sgouras, Spiros Galanopoulos, Konstantinos Kanellopoulos, Onur Mutlu | 2026-09-28 | 下载 | Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g. |
| Analog Computing revisited: A fully analog and minimalistic Damage Detector for Ultrasonic Testing enabling Material-Integrated Structural Health Monitoring | Stefan Bosse | 2026-09-28 | 下载 | Ultrasonic Testing (UT) is commonly used to detect damage in structures, e.g., metal plates. A sensor acquires Ultrasonic waves, e.g., by using PZT transducers. |
| MEGATRON: a 28nm Analog PCM CiM/Digital System-on-Chip for Edge GenAI at 57.5 TOPS/W and 1.52 Mparam/mm | Alessandro Nadalini, Angelo Garofalo, Lorenzo Greco, Andrea Belano, Alessio Antolini, Francesco Zavalloni, Andrea Lico, Riccardo Zurla, Emanuela Calvetti, Luigi Croce, Marco Pasotti, Alessandro Cabrini, Eleonora Franchi Scarselli, Davide Rossi, Francesco Conti | 2026-09-28 | 下载 | We present MEGATRON, a heterogeneous Edge GenAI System-on-Chip in 28nm FD-SOI CMOS technology combining a non-volatile analog in-memory-computing engine based on a 4Mi-cell phase-change memory (PCM) a... |
| Efficient TCitH-Based Alternatives to SLH-DSA: Cross-Layer ASIC Design of Mirath | Hiandra Tomasi, Maximilian Schöffel, Johannes Feldmann, Norbert Wehn | 2026-09-28 | 下载 | To address the security risks posed by quantum computers, the U.S. National Institute of Standards and Technology (NIST) has standardized the post-quantum signature schemes ML-DSA, FN-DSA, and SLH-DSA... |
| BEHAVE: Functional Behavior Modeling Enables Self-Improving Agents for Hardware Design and Verification | Yuheng Wu, Berk Gokmen, Sujeeth Jinesh, Lauren McLane, Aarav Wattal, Qi Yang Huang, Zhaozhuo Xu, Thierry Tambe | 2026-09-28 | 下载 | Developing agents for hardware design and verification requires reliable correctness feedback. As a hardware specification may permit correct implementations with different latencies, matching design ... |
| Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch | Heeeon Lee, Hyunwoo Nam, Junyong Heo, Hyunmo Sung, Jay Hwan Lee, Yeonsoo Kim, Seongho Jeong, Shinhyung Yang, Bernd Burgstaller | 2026-09-28 | 下载 | Modern deep learning (DL) workloads are limited by data movement, and processing-in-memory (PIM) targets this bottleneck by placing compute units near the memory. |
| SPIMOE: Exploiting Hybrid Sparsity for Reasoning MoE Inference on Heterogeneous PIM Architectures | Rubing Yang, Cenlin Duan, Yingjie Qi, Xiaolin He, Xiao Ma, Jianlei Yang | 2026-09-28 | 下载 | Long-reasoning Mixture-of-Experts (MoE) models expose two coupled inference bottlenecks: growing KV caches shift the critical path toward attention, while sparse expert activation causes load imbalanc... |
| Coarse-to-Fine Macro Placement via Evolutionary Search and Critical Macro Tuning | Biao Liu, Zhiping Jin, Kaixuan Sun, Zengrui Lu, Qingquan Zhang, Bo Yuan | 2026-09-28 | 下载 | Macro placement is a critical stage in chip physical design that substantially affects downstream implementation quality. Recent search-based methods improve existing layouts through partial reconstru... |
| PolyCIM: Improving Data Reuse in Digital CIM Accelerators with Polyhedral-Based Compilation | Yingjie Qi, Cenlin Duan, Yiou Wang, Yikun Wang, Xiaolin He, Weisheng Zhao, Jianlei Yang | 2026-09-28 | 下载 | Digital Compute-in-Memory (CIM) presents a promising solution for accelerating deep neural networks (DNNs) through the integration of computational logic directly within memory arrays. |
| Improving Indirect Branch Prediction in Interpreters via Hardware/Software Co-Design | Linfeng Zheng, Hiroshi Sasaki | 2026-09-28 | 下载 | Interpreters have a large indirect-branch footprint, requiring large predictor capacity for accurate prediction. We propose a hardware/software co-design in which a hardware lookahead engine, running ... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Quantifying Teleportation Overhead in Distributed Unitary Coupled-Cluster Ansätze | Grier M. Jones, Hassan Tariq Shafi, Zixuan Wang, Thomas Trenty, Zachary Vernec, Hans-Arno Jacobsen | 2026-09-28 | 下载 | Distributed quantum computing (DQC) has been proposed as a way to scale quantum algorithms for practical applications beyond monolithic quantum processor architectures. |
| Encoder-Sharing Hierarchical Federated Multi-Task Learning for VANETs | M. Saeid HaghighiFard, Sinem Coleri | 2026-09-28 | 下载 | Most federated learning frameworks for vehicular ad hoc networks assume that all vehicles collaboratively train a single model for a common task. |
| Mixture-of-Kittens: MoE Megakernel for NVL72s | Stuart H. Sul, Nash Brown, Henry Wildermuth, William Lin, Federico Cassano, Christopher Ré | 2026-09-28 | 下载 | AI accelerator systems are rapidly consolidating into scale-up architectures, where tens to thousands of GPUs communicate over high-bandwidth, single-hop fabrics. |
| Making Cross-Continental Federated Learning Repeatable with FLIP: a Multi-Application Study | Rafael Garcia-Dias, Alexandre Triay Bagur, Chayanin Tangwiriyasakul, Virginia Fernandez, Parhom Esmaeili, Piyalitt Ittichaiwong, Yang Li, Lawrence Adams, Wason Buncharoen, Martin Chapman, Benjamaporn Chayanond, Sadthavud Chunrod, Tanawat Fongsri, Kass Gibson, Supat Plungprasertkul, Supawit Tangpanithandee, Kanyakorn Veerakanjana, Vicky Goh, Michela Antonelli, Joe Zhang, Kongkiat Kespechara, Sebastien Ourselin, M. Jorge Cardoso | 2026-09-28 | 下载 | Federated learning (FL) in healthcare remains challenging, as the overhead of rebuilding governance guarantees for every collaboration stops most projects at the proof-of-concept stage. |
| Dynamic Wakeup under Costly Collisions | Umesh Biswas, Maxwell Young | 2026-09-28 | 下载 | The wakeup problem captures a fundamental symmetry-breaking challenge among devices sharing a communication channel. We study the dynamic setting, where packets become active at arbitrary times on a t... |
| GPUPhysBench: Benchmarking Coding Agents for Correct and Efficient GPU Physics Simulation | Yuchen Sun, Jinjin He, Sinan Wang, Bo Zhu | 2026-09-28 | 下载 | Writing fast GPU code for physical simulation is difficult: implementations must preserve numerical accuracy while handling irregular data access, synchronization, and iterative solvers. |
| Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions | Tianyao Shi, Xipeng Shen, Yi Ding | 2026-09-28 | 下载 | Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. |
| TopoEP: Topology-Aware Load Balancing for Expert-Parallel MoE Training | Jiacheng Zhu, Xie Zhao, Gongming Zhao, Hongli Xu, Yao Fei, Jin Fang | 2026-09-28 | 下载 | Dynamic routing creates severe load imbalance in large-scale expert-parallel Mixture-of-Experts (MoE) training, turning GPUs that host hot experts into stragglers. |
| SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning | Bojian Yin, Shurong Wang, Yuqi Pan, Guoqi Li | 2026-09-28 | 下载 | Large language models are trained with backpropagation, whose global gradient coordinates all layers but forces each to hold its activations and wait for the gradient to pass back through every deeper... |
| Weaver: A System for AI-RAN Compute Sharing with Foundation Model Training | Leyang Xue, Tianxin Wang, Xin Zhe Khooi, Jiaxun Yang, Dheeraj Mahendiran, Yufeng Xia, Mun Choon Chan, Myungjin Lee, Mahesh K. Marina | 2026-09-28 | 下载 | The emergence of AI-RAN infrastructure, which equips cell sites with GPU-accelerated hardware, creates an opportunity to colocate non-RAN workloads with primary RAN processing. |
| WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse | Aaryam Sharma | 2026-09-28 | 下载 | Pipeline parallelism can improve prefill throughput by processing multiple request chunks concurrently across different stages of the model. However, keeping the pipeline fully utilized requires effic... |
| Analyzing Solana's Blocks and Transactions | Yaron Hay, Dvir David Biton, Roy Friedman | 2026-09-28 | 下载 | Solana is one of the most popular blockchains, and is arguably the most widely used blockchain for smart contracts, also known as dApps. Understanding the types of smart contracts that are being execu... |
| EdgeCraft: Automated Model Crafting for Edge IoT | Genglin Wang, Kaiwei Liu, Liekang Zeng, Wangsong Yin, Shangcheng Jin, Guoliang Xing, Zhenyu Yan | 2026-09-28 | 下载 | Machine learning (ML) increasingly powers Internet of Things (IoT) applications at the edge. Yet producing a deployable edge ML artifact for a specific scenario requires navigating a huge search space... |
| AReaL-TIK: Stateful Agentic Optimization of Unified RL Kernels through an Optimization IR | Ran Yan, Youhe Jiang, Jiayi Nie, Wenshuang Li, Yingqi Peng, Taiyi Wang, Tongkai Yang, Binhang Yuan | 2026-09-28 | 下载 | Reinforcement learning (RL) post-training often uses distinct GPU kernels for rollout and policy update. In synchronous PPO and GRPO, numerical disagreement can perturb ratios between current token pr... |
| E3J: An Efficient and Open-Source Backend for Euclidean Equivariant Operations on GPU and TPU | Olivier Peltre, Armand Picard, Adrien Pichard, Miguel Bragança, Luca Giacomoni, Valentin Heyraud, Zachary Weller-Davies, Christoph Brunken, Jules Tilly | 2026-09-28 | 下载 | We present e3j, a fast Euclid-equivariance backend for geometric deep learning applications with JAX bindings for GPU and TPU. Leveraging both optimized CUDA and Pallas kernels and algorithmic improve... |
| TempoKV: Timely Staging of LLM KV Caches for Memory-Semantic Flash | Jay H. Park, Hyungjun Kim, Dong Kim | 2026-09-28 | 下载 | Reusable prefix key-value (KV) caches can outgrow GPU memory in large language model (LLM) serving. A memory-semantic flash hierarchy offers SSD-backed capacity with a limited fast tier, but a logical... |
| Monitoring and Verification of Multitenant Kubernetes Clusters using TLA+ Trace Checking | Ioana Silaş, Adrian Crăciun | 2026-09-28 | 下载 | In distributed systems, model checking is usually used at design time for specifying an abstract model of the system and then exhaustively checking all possible behaviors. |
| Accelerator Choice Is Not Enough: AlphaFold2 Inference on Cloud TPUs | Lorenzo Pazienza, Ihab El Bani | 2026-09-28 | 下载 | AlphaFold2 is written in JAX, so the same inference code compiles and runs unchanged on CPUs, GPUs and Google Cloud TPUs. That portability makes the accelerator look like the main decision a user has ... |
| WaveAlign: Cache-Aware Query-Row Scheduling for Sparse Attention in Long-Video Generation | Zijian Dai, Sen Han, Youhui Bai, Shannon Wang, Kan Wu, Jingkai Huang, Yuhang Wang, Jing Li, Cheng Li | 2026-09-28 | 下载 | Long-video generation with diffusion transformers (DiTs) produces extremely long token sequences, making attention a dominant inference bottleneck. |
| Distributed Lower Bounds via Automatic Self-Reduction | Alkida Balliu, Francesco d'Amore, Dennis Olivetti | 2026-09-28 | 下载 | The development of round elimination into a general-purpose technique [PODC 2019] marked a turning point in our understanding of the hardness of many graph problems in the distributed setting and led ... |
| Torch-PIM: Automated Profile-Guided PIM Offloading for PyTorch | Heeeon Lee, Hyunwoo Nam, Junyong Heo, Hyunmo Sung, Jay Hwan Lee, Yeonsoo Kim, Seongho Jeong, Shinhyung Yang, Bernd Burgstaller | 2026-09-28 | 下载 | Modern deep learning (DL) workloads are limited by data movement, and processing-in-memory (PIM) targets this bottleneck by placing compute units near the memory. |
| Nereus: Adaptive Parallelism for LLM Post-Training | Songlin Jiang, Tuo Shi, Sitong Zhang, Zeke Wang, Mario Di Francesco, Bo Zhao | 2026-09-28 | 下载 | Reinforcement learning (RL) post-training for large language models (LLMs) coordinates multiple models across generation, inference, and training on GPU clusters. |
| AgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud Continuum | Michalis Kasioulis, Moysis Symeonides, George Pallis, Marios D. Dikaiakos | 2026-09-28 | 下载 | Deploying LLM-enabled agentic applications across the Edge-to-Cloud continuum remains challenging due to hardware heterogeneity, deployment complexity, limited observability, and the lack of systemati... |
| Semantics, Workflows, and Infrastructure: Understanding Agent Serving at Production Scale | Yihao Zheng, Jingzhe Jiang, Dejiang Zhu, Zhiyuan Tan, Yang Tian, Tao Wang, Minchen Yu | 2026-09-28 | 下载 | Large language model (LLM) agents execute applications through a workflow of inference requests with tool calls and user interactions. Serving these applications at production scale requires understan... |
| DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory | Xuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu, Wookeun Jung, Thanh Tuan Dao | 2026-09-28 | 下载 | Existing LLM serving systems virtualize and optimize KV-cache memory, but treat model-weight memory as fixed throughout execution. Recent work on multi-precision model representations challenges this ... |
| Before Agents Act: Assurance-Aware Semantic Scheduling for Evidence Acquisition in Distributed Systems | Jun He, Deying Yu | 2026-09-28 | 下载 | Tool-using agents can initiate consequential infrastructure changes, yet evidence required for admission may expire while other checks run or depend on a shared fault domain. |
| Spexis: Speculative Lookahead Scheduling for LLM Inference | Hyungyu Jung, Jaehyeok Yu, Hoonseo Choi, Sungkyun Kim, Jinho Lee, Jiwon Seo | 2026-09-28 | 下载 | Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. |
| The Shape of Speed: Impacts of Partition Geometry and Rank Density in Distributed Quantum Circuit Simulations | Yikai Mao, Yuan He, Shaowen Li, Masaaki Kondo | 2026-09-28 | 下载 | In distributed quantum circuit simulation, a poorly shaped partition can halve performance before computation begins. Evaluation on Fugaku across 764 validated configurations (twelve algorithms, thirt... |
| VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters | Yao Fei, Jin Fang, Size Zheng, Gongming Zhao, Hongli Xu, Jiacheng Zhu, Zhijing Xin | 2026-09-28 | 下载 | AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient infe... |
| HyDra: Demystifying and Taming Dynamic Context Parallelism at Production Scale | Zihao Fan, Yunzhuo Liu, Bo Jiang, Changgang Zheng, Lin Zheng, Ray Ying, Key Zhang | 2026-09-28 | 下载 | Long-context training runs on sequences whose lengths span orders of magnitude, and dynamic context parallelism (DCP) gives each sequence its own CP degree. |
| Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training | Junlin Chen, Daize Dong, Huanwei Di, Haolong Jia, Jiawei Wu, Haotian Xie, Mingkai Zheng, Yang Li, Leshang Chen, Huishu Wang, Eric P. Xing, Hongyi Wang | 2026-09-28 | 下载 | BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. |
| Heddle: Learning Structural Templates for Parallelism Planning on Heterogeneous GPU Clusters | Taeyoon Kim, Yonguk Song, Seoyeong Choy, Hexiao Duan, Dong Li, Seo Jin Park, Myeongjae Jeon | 2026-09-28 | 下载 | Training large machine learning models on shared GPU infrastructures faces two challenges: (1) GPU availability shifts dynamically with varying resource demands from tenants, and (2) hardware heteroge... |
| SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs | Ruijia Yang, Shiyuan Lin, Yulong Ao, Zhiyu Li, Yingli Zhao, Xianduo Li, Yonghua Lin, Zeyi Wen | 2026-09-28 | 下载 | Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. |
| Schedule Repair for DAG Workflows under Link Disruptions | Mohammadali Khodabandehlou, Jared Coleman, Bhaskar Krishnamachari, Kevin Chan | 2026-09-28 | 下载 | Schedules for directed acyclic graph (DAG) workflows in networked IoT systems are typically computed assuming a static or generally stable network. |
| Kafila: Serving Large Language Models on a Trusted Set of Heterogeneous Commodity Machines | Murtaza Rangwala, Richard O. Sinnott, Rajkumar Buyya | 2026-09-28 | 下载 | Between them, the members of a research group or a circle of friends own several consumer computers, none large enough to run a capable large language model. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| People effects on IoT indoor wireless channel characterization | Millena Michely de Medeiros Campos, Mateus de Oliveira e Mattos, Rafael da Silva Macedo, Alvaro Augusto Machado de Medeiros, Wellerson Viana de Oliveira, Vicente Angelo de Sousa Junior | 2026-09-28 | 下载 | Wireless communication under 1 GHz is suitable for Internet of Things (IoT) applications due to larger coverage capability with less power consumption. |
| Non-Ionizing Radiation Analysis in Close Proximity to Antenna Tower: A Case Study in Northeast Brazil | André B. F. Diniz, Vicente A. Souza, Marcio E. C. Rodrigues, Halysson B. Mendonça, Gutembregue S. Silva, Fred S. R. Pinheiro | 2026-09-28 | 下载 | While the amount of telecommunications services grows rapidly in the whole world, humans get potentially more exposed to Non-Ionizing Radiation (NIR) from a number of different sources. |
| Non-Ionizing Radiation Measurements for Trajectography Radars | J. Marcos Leal Barbosa Filho, Millena M. de M. Campos, Daniel L. Flor, William S. Alves, Adaildo G. D'Assunção, Marcio E. C. Rodrigues, Vicente A. de Sousa | 2026-09-28 | 下载 | This work presents a Non-Ionizing Radiation (NIR) measurement campaign and proposes a specific measurement method for trajectography radars. This kind of radar has a high gain narrow beam antenna and ... |
| Traffic Congestion Awareness and On-Demand Distribution in Vehicular Delay-Tolerant Networks in California I-210 Freeway | Xiaofei Liu, Milena Radenkovic | 2026-09-28 | 下载 | In vehicular networks under edge computing environments, vehicle-to-vehicle delay-tolerant networking (V-DTN) can disseminate congestion warnings to other vehicles via a store-carry-forward mechanism,... |
| TORQUE: Optimizing What (not) to Quantize Before and After Rotation | Ran Ben Basat, Michael Mitzenmacher, Shay Vargaftik | 2026-09-28 | 下载 | Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. |
| MINT: Modeling GenAI Impact on Network Traffic | Andrew Nguyen, Samson Kempiak, Agrim Gupta, Koushik Kar, Ish Kumar Jain | 2026-09-28 | 下载 | Generative AI (GenAI) is becoming a mainstream network workload, yet packet-level simulators lack measure-ment-driven GenAI traffic models. Currently researchers must approximate GenAI services using... |
| Energy-Driven Evaluation of Network Digital Twinning Applied to mmWave Beam Management | João Borges, Bruno Castro, Kleber Cardoso, Andrey Silva, Aldebaro Klautau | 2026-09-28 | 下载 | Network Digital Twins (NDTs) are important enablers of 6G and future networks. However, there is a lack of studies regarding practical aspects, such as the impact of simultaneously changing twinning r... |
| Resource versus Responsiveness: Benchmarking SDN Controller Runtimes for a Moving-Target-Defense Control Plane at Scale | Souhail Chakkour, Umesh Biswas, Charan Gudla | 2026-09-28 | 下载 | Network Moving Target Defense (MTD) built on Software-Defined Networking (SDN) continuously rotates host-facing addresses to invalidate an attacker's reconnaissance. |
| Hybrid QKD-PQC Network Emulation through Automated and Scalable Cloud-Native Orchestration | Iván Melijosa, Javier Pérez, Borja Nogales, Iván Vidal, Francisco Valera | 2026-09-28 | 下载 | The ongoing transition toward quantum-safe networking has motivated the development of hybrid network architectures integrating Quantum Key Distribution (QKD) and Post-Quantum Cryptography (PQC). |
| Configuration-Induced Delivery Failures in NATS JetStream: Detection and Remediation | Biplab Kumar Das | 2026-09-28 | 下载 | NATS JetStream's at-least-once delivery guarantee is conditional: five common configuration mistakes silently violate it, causing duplicate message processing, data loss, or redelivery storms with no ... |
| Weaver: A System for AI-RAN Compute Sharing with Foundation Model Training | Leyang Xue, Tianxin Wang, Xin Zhe Khooi, Jiaxun Yang, Dheeraj Mahendiran, Yufeng Xia, Mun Choon Chan, Myungjin Lee, Mahesh K. Marina | 2026-09-28 | 下载 | The emergence of AI-RAN infrastructure, which equips cell sites with GPU-accelerated hardware, creates an opportunity to colocate non-RAN workloads with primary RAN processing. |
| Implementing Data Diodes Using Commodity Hardware and Open Source Software | Peter Story, Gert-Jan den Besten | 2026-09-28 | 下载 | One-way network devices, known as data diodes, are used to defend against sophisticated cyberattacks. Partly due to their high cost, data diodes are mostly deployed in nuclear power plants and within ... |
| Task-Oriented Communications for Edge-Assisted Multi-View Localization | Zhengru Fang, Huanhuan Lou, Senkang Hu, Yihang Tao, Zongdian Li, Yiqin Deng, Jingjing Wang, Yuguang Fang | 2026-09-28 | 下载 | Unmanned aerial vehicles (UAVs) and unmanned ground vehicles (UGVs) often lose satellite positioning in urban canyons, indoor facilities, and jammed or spoofed environments, making vision-based matchi... |
| Deconstructing BLE Multi-hop: a Model-based Approach to Quantifying the Challenges | Bozheng Pang, José Alamos, Thomas C. Schmidt, Matthias Wählisch | 2026-09-28 | 下载 | Bluetooth Low Energy (BLE) was originally designed for point-to-point communication, but BLE multi-hop networks have attracted academic and industrial interest. |
| WiFi Backscatter for Green Internet of Things: Concepts, Research Trends, and Practical Challenges | Weiqi Wu, Shuo Wang, Zhaoyuan Xu | 2026-09-28 | 下载 | WiFi backscatter has emerged as a promising technology for green Internet of Things (IoT) connectivity by enabling battery-free devices to communicate through widely available WiFi signals. |
| CoRF: Cross-Scene RF Synthesis by Learning Propagation and Preserving Array Physics | Kang Yang, Duaa Nakshbandi, Wan Du, Mani Srivastava | 2026-09-28 | 下载 | Existing radio-frequency (RF) neural fields fit each scene separately, making new-scene deployment measurement- and optimization-intensive. This work studies amortized cross-scene spatial spectrum syn... |
| VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters | Yao Fei, Jin Fang, Size Zheng, Gongming Zhao, Hongli Xu, Jiacheng Zhu, Zhijing Xin | 2026-09-28 | 下载 | AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient infe... |
| HOCCL: Offloading Collective Communication from GPU Cores to Accelerate Distributed Training | Yao Fei, Gongming Zhao, Hongli Xu, Jin Fang, Jiacheng Zhu, Shuo Xu, Kun Huang, Zhuolong Yu | 2026-09-28 | 下载 | Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, thei... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Argus: Agentic, Reference-Calibrated, Tree-Guided, System-Software-Level Bottleneck Localization | Vlad-Petru Nitu, Harsh Songara, Konstantinos Sgouras, Spiros Galanopoulos, Konstantinos Kanellopoulos, Onur Mutlu | 2026-09-28 | 下载 | Operating system (OS) code can account for a substantial share of CPU execution time. First, as application logic is offloaded to heterogeneous accelerators (e.g. |
| Planarian: Managing Agent State with Statepoints | Jinnan Guo, Hao Mark Chen, Kapil Vaswani, Andrew Paverd, Peter Pietzuch | 2026-09-28 | 下载 | LLM agents solve complex tasks by iteratively changing files, invoking local tools, and interacting with remote services, which modifies state across their local environment and remote services. |
| Dynamic Flow, Static Graph: KV Cache Reuse for Efficient LLM Serving on Mobile NPUs | Zhengxiang Huang, Shengheng Chen, Chaoyue Niu, Yujie Sun, Zhaode Wang, Zeyu Zhao, Chengfei Lv, Fan Wu, Guihai Chen | 2026-09-28 | 下载 | On-device large language model (LLM) serving is a cornerstone of local-first personal intelligence, offering users data sovereignty, strong privacy guarantees, and freedom from cloud API latency and c... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| GEM-KMeans: Memory-Efficient and Accurate Clustering on Massive Scale with GPU Optimization | Peng Xu, Nihar Koganti, Volodymyr Kindratenko, Xiaohui Chen | 2026-09-28 | 下载 | Memory-efficient scaling on clustering problems without sacrificing statistical accuracy is of central interest for large-scale data analysis and machine learning problems. |
| Hardware-Aware Features for CUTLASS Kernel Selection | Shriram Chandran, Dominic Rinderer, Yakup Budanaz, Alexandru Calotoiu, Marcin Copik, Torsten Hoefler | 2026-09-28 | 下载 | GPU libraries such as CUTLASS expose tens of thousands of semantically equivalent kernels for a single operation, making exhaustive autotuning expensive and execution-free selection difficult. |
| Beyond Energy: When Sustainability Dimensions Reshape LLM Serving Decisions | Tianyao Shi, Xipeng Shen, Yi Ding | 2026-09-28 | 下载 | Large language model (LLM) serving has environmental impacts across energy consumption, carbon emission, water consumption, and biodiversity loss. |
| Beneath the Tokens: A Performance Engineering Study of Multi-Token Prediction in GPU-Accelerated LLM Inference | Suwesh Prasad Sah | 2026-09-28 | 下载 | Autoregressive large language model inference repeatedly invokes the target model to generate one token at a time, making generation sensitive to GPU memory movement and sequential execution. |
| Accelerator Choice Is Not Enough: AlphaFold2 Inference on Cloud TPUs | Lorenzo Pazienza, Ihab El Bani | 2026-09-28 | 下载 | AlphaFold2 is written in JAX, so the same inference code compiles and runs unchanged on CPUs, GPUs and Google Cloud TPUs. That portability makes the accelerator look like the main decision a user has ... |
| Tool Waiting and Re-arrival in Compile-Time-Static LLM Serving: Cost Mechanisms and Configuration Selection | Dongkyeom Jang, In-Nea Wang, Junho Jeong | 2026-09-28 | 下载 | In agentic LLM services, a session calls an external tool, waits for it, and re-arrives to continue inference. Statically compiled NPU serving can fix the batch bucket set, the maximum batch size, and... |
| DPS: Dual-Mode Precision LLM Serving with Semi-Unified Memory | Xuan Truong Nguyen, Tien Son Pham, Tuan Duc Chu, Wookeun Jung, Thanh Tuan Dao | 2026-09-28 | 下载 | Existing LLM serving systems virtualize and optimize KV-cache memory, but treat model-weight memory as fixed throughout execution. Recent work on multi-precision model representations challenges this ... |
| The Shape of Speed: Impacts of Partition Geometry and Rank Density in Distributed Quantum Circuit Simulations | Yikai Mao, Yuan He, Shaowen Li, Masaaki Kondo | 2026-09-28 | 下载 | In distributed quantum circuit simulation, a poorly shaped partition can halve performance before computation begins. Evaluation on Fugaku across 764 validated configurations (twelve algorithms, thirt... |