2026-05-11
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models | Jungwoo Kim, Rubens Lacouture, Genghan Zhang, Gina Sohn, Qizheng Zhang, Swapnil Gandhi, Christos Kozyrakis, Kunle Olukotun | 2026-05-11 | 下载 | Mixture-of-Experts (MoE) has become a dominant architecture for scaling large language models (LLMs). However, the execution characteristics of MoE inference are changing rapidly and increasingly mism... |
| TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments | Yue Guan, Hongtao Yu, Peng Chen, Daohang Shi, Karthik Manivannan, Nicholas J Riasanovsky, Manman Ren, Lei Wang, Shane Nay, Partha Kanuparthy, Zaifeng Pan, Zhengding Hu, Yufei Ding | 2026-05-11 | 下载 | Modern GPUs increasingly rely on specialized hardware units and asynchronous coordination mechanisms, so performance depends on orchestrating data movement, tensor-core computation, and synchronizatio... |
| LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges | Johann Knechtel, Ozgur Sinanoglu, Ramesh Karri | 2026-05-11 | 下载 | The integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) and hardware security is rapidly reshaping the semiconductor industry. |
| Reconfigurable Computing Challenge: Real-Time Graph Neural Networks for Online Event Selection in Big Science | Marc Neu, Frank Baptist, Thomas Lobmaier, Fabio Papagno, Torben Ferber, Jürgen Becker | 2026-05-11 | 下载 | Graph neural networks are increasingly adopted in trigger systems for collider experiments, where strict latency and throughput constraints render deployment on embedded platforms challenging. |
| ObfAx: Obfuscation and IP Piracy Detection in Approximate Circuits | Lukas Sekanina, Vojtech Mrazek | 2026-05-11 | 下载 | Approximate circuits often achieve exceptional trade-offs between computational accuracy and hardware efficiency, making them attractive for deployment as reusable Intellectual Property (IP) cores. |
| Towards an End-To-End System for Real-Time Gesture Recognition from Surface Vibrations | Florian Hettstedt, Cedric Giese, Tianheng Ling, Keiichi Yasumoto, Gregor Schiele, Andreas Erbslöh | 2026-05-11 | 下载 | Sensing surface vibrations promise unobtrusive interaction for smart home systems by enabling gesture recognition on existing everyday surfaces without disturbing living-space design. |
| Arcane: An Assertion Reduction Framework through Semantic Clustering and MCTS-Guided Rule Exploring | Hongqin Lyu, Yonghao Wang, Zhiteng Chao, Tiancheng Wang, Huawei Li | 2026-05-11 | 下载 | Assertion-based Verification (ABV) is essential for ensuring that hardware designs conform to their intended specifications. However, existing automated assertion-generation approaches, such as LLM-ba... |
| RFAmpDesigner: A Self-Evolving Multi-Agent LLM Framework for Automated Radio Frequency Amplifier Design | Hang Lu, Guochang Li, Qianyu Chen, Huiyan Gao, Shaogang Wang, Xuanyu He, Yiwei Liu, Gaopeng Chen, Nayu Li, Xiaokang Qi, Chunyi Song, Zhiwei Xu | 2026-05-11 | 下载 | Automating radio frequency (RF) amplifier design remains challenging because existing methods suffer from the curse of dimensionality, weak use of domain knowledge, and poor transferability, leading t... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer Inference | Han Meng, Danny Willow Liu, Dong Li | 2026-05-11 | 下载 | Layerwise offloading reduces the GPU memory footprint of large diffusion transformer (DiT) inference by prefetching upcoming layers from host memory, but its effectiveness hinges on hiding prefetch la... |
| MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces | Srinivas Sridharan, Andy Balogh, Bradford M. Beckmann, Brian Coutinho, Louis Feng, Sheng Fu, Sanshan Gao, Mehryar Garakani, Taekyung Heo, David Kanter, Josh Ladd, Ziwei Li, Winston Liu, Changhai Man, Dan Mihailescu, Spandan More, Joongun Park, Ashwin Ramachandran, Vinay Ramakrishnaiah, Saeed Rashidi, Vijay Janapa Reddi, Puneet Sharma, Phio Tian, William Won, Hanjiang Wu, Huan Xu, Jinsun Yoo, Tushar Krishna | 2026-05-11 | 下载 | The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload behavior in productio... |
| Byzantine Consensus in Directed Graphs with Message Authentication | Nitin H. Vaidya, Lewis Tseng | 2026-05-11 | 下载 | We consider the problem of reaching consensus in communication networks that are modeled by directed graphs. We assume the existence of a message authentication mechanism (such as digital signatures) ... |
| ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload | Ziyue Liu, Zhengyang Wang, Ruijie Zhang, Avinash Maurya, Hui Zhou, Paul Hovland, Sheng Di, Franck Cappello, Bogdan Nicolae, Zheng Zhang | 2026-05-11 | 下载 | Pre-training large language models on massive GPU clusters has made hardware faults routine rather than rare, driving the need for resilient training systems. |
| ShardTensor: Domain Parallelism for Scientific Machine Learning | Corey Adams, Peter Harrington, Akshay Subramaniam, Mohammad Shoaib Abbas, Jaideep Pathak, Mike Pritchard, Sanjay Choudhry | 2026-05-11 | 下载 | Scientific Machine Learning (SciML) faces unique challenges for extreme-resolution data, with mitigations that often fail to scale or degrade the accuracy of trained models. |
| Closer in the Gap: Towards Portable Performance on RISC-V Vector Processors | Ruimin Shi, Maya Gokhale, Pei-Hung Lin, Xavier Teruel, Ivy Peng | 2026-05-11 | 下载 | The RISC-V Vector Extension~(RVV) is a cornerstone for supporting compute throughout in scientific and machine learning workloads. Yet compiler support and performance monitoring on real RVV~1. |
| An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing Continuum | Suvi De Silva, Alfreds Lapkovskis, Alaa Saleh, Sasu Tarkoma, Praveen Kumar Donta | 2026-05-11 | 下载 | Grey failures in the computing continuum produce ambiguous overlapping symptoms that existing approaches fail to diagnose reliably, either due to a lack of causal awareness or acting under high episte... |
| Surviving Partial Rank Failures in Wide Expert-Parallel MoE Inference | Xun Sun, Shaoyuan Chen, Pingchuan Ma, Yue Chen, Ziwei Yuan, Zhanhao Cao, Han Han, Shangming Cai, Teng Ma, Xuchun Shang, Xinpeng Zhao, Ke Yang, Junlin Wei, Lianzhi Lin, Yuji Liu, Feng Ren, Haoran Hu, Cheng Wan, Yingdi Shan, Yongwei Wu, Mingxing Zhang | 2026-05-11 | 下载 | Mixture-of-Experts (MoE) serving relies on wide expert parallelism (EP) to aggregate the memory capacity and bandwidth of many GPUs within one inference instance. |
| SoK: A Systematic Bidirectional Literature Review of AI & DLT Convergence | Ali Irzam Kathia, Yimika Erinle, Abylay Satybaldy, Paolo Tasca, Nikhil Vadgama, Marco Alberto Javarone | 2026-05-11 | 下载 | The integration of Artificial Intelligence (AI) with Distributed Ledger Technology (DLT) has become a growing research area, yet contributions tend to cluster around specific application domains or ex... |
| Accelerating Compound LLM Training Workloads with Maestro | Xiulong Yuan, Hongqing Chen, Jiaxuan Peng, Fan Zhou, Zhixiang Ruan, Zekun Wang, Bo Zheng, Rui Men, Haiquan Wang, Zhipeng Zhang, Langshi Chen, Man Yuan, Jiaqi Gao, Zhengping Qian, Junyang Lin, Yong Li, Wei Lin, Junhua Wang, Jingren Zhou | 2026-05-11 | 下载 | Compound LLM training workloads-such as knowledge distillation and multimodal LLM (MLLM) training-are gaining prominence. These typically comprise heterogeneous components differing in parameter scale... |
| Privacy-preserving Chunk Scheduling in a BitTorrent Implementation of Federated Learning | Naicheng Li, Javad Dogani, Rui Wang, Kaitai Liang, Nikolaos Laoutaris | 2026-05-11 | 下载 | Traditional federated learning (FL) relies on a central aggregator server, which can create performance bottlenecks and privacy risks. Decentralized mix-and-forward designs remove the server, but repe... |
| HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge Computing | Jianyong Zhu, Hao Chen, Juan Zhang, Fangda Guo, Albert Y. Zomaya, Renyu Yang | 2026-05-11 | 下载 | Edge computing faces unprecedented resource orchestration challenges from multi-dimensional heterogeneity across device architectures, diverse task requirements in CPU-intensive, GPU-intensive, I/O-in... |
| FractalSortCPU: Bandwidth-Efficient Compressed Radix Sort on CPU | Michael Dang'ana | 2026-05-11 | 下载 | Cloud database systems, particularly their middleware and query execution layers, use sorting as a core operation in query processing, indexing and join execution. |
| Agentic Performance at the Edge: Insights from Benchmarking | Shiqiang Wang, Herbert Woisetschläger | 2026-05-11 | 下载 | Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to models around 8 billion parameters or smaller. |
| Amortized Asynchronous Byzantine Reliable Broadcast with Optimal Resilience | Michael Yiqing Hu, Alvin Hong Yao Yan, Jialin Li | 2026-05-11 | 下载 | Byzantine Reliable Broadcast (BRB) is a fundamental primitive in distributed computing and cryptographic systems. Reducing the communication complexity of BRB protocols remains an important research d... |
| Autonomous FAIR Digital Objects: From Passive Assertions to Active Knowledge | Zeyd Boukhers, Oya Beyan, Cong Yang, Christoph Lange | 2026-05-11 | 下载 | Scientific knowledge on the Web is published as passive assertions and cannot decide when to validate evidence, reconcile contradictions, or update confidence as findings accumulate. |
| Accelerating Locality-Driven Integration in Quantum Chemistry with Block-Structured Matrix Multiplication | Xinran Wei, Yan Pan, Fusong Ju, Zehao Zhou, Yihong Zhang, Lin Huang, Jianwei Zhu, Jia Zhang, Huanhuan Xia, Bin Shao, Tao Qin | 2026-05-11 | 下载 | Locality-driven integration is a pervasive computational pattern in quantum chemistry, arising whenever spatially localized basis functions interact through numerical quadrature or integral screening. |
| FusionRCG: Orchestrating Recursive Computation Graphs across GPU Memory Hierarchies | Yihong Zhang, Xinran Wei, Junshi Chen, Fusong Ju, Wei Hu, Jinlong Yang, Huanhuan Xia | 2026-05-11 | 下载 | Evaluating high-dimensional integrals via deep hierarchical recurrences is a dominant cost in quantum chemistry. While CPUs manage these efficiently, GPUs suffer a critical mismatch: limited per-threa... |
| DP-LAC: Lightweight Adaptive Clipping for Differentially Private Federated Fine-tuning of Language Models | Haaris Mehmood, Jie Xu, Karthikeyan Saravanan, Rogier Van Dalen, Mete Ozay | 2026-05-11 | 下载 | Federated learning (FL) enables the collaborative training of large-scale language models (LLMs) across edge devices while keeping user data on-device. |
| GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference | Zengzipeng Tang, Yuxuan Sun, Wei Chen, Jianwen Ding, Bo Ai | 2026-05-11 | 下载 | The recent growth of on-device Large Language Model (LLM) inference has driven significant interest in device-edge collaborative LLM inference. |
| Edge-Cloud Collaborative Pothole Detection via Onboard Event Screening and Federated Temporal Segmentation | Yingjie Wu, Kongyang Chen, Tiancai Liang | 2026-05-11 | 下载 | Road potholes threaten driving safety and increase infrastructure maintenance costs, while large-scale and timely pothole detection remains challenging in urban road networks. |
| Lakestream: A Consistent and Brokerless Data Plane for Large Foundation Model Training | Ting Sun, Junjie Zhang, Xiao Yan, Songxin Zhang, Zhuoyang Song, Jingyi Xi, Zunyao Mao, Bingyi Jing, Jiaxing Zhang, Zejian Xie | 2026-05-11 | 下载 | Modern Large Foundation Model (LFM) training has transformed the data pipeline from a static ingestion layer into a dynamic component that must co-evolve with the training process. |
| Population Protocols over Ordered Agents | Michael Blondin, Michaël Cadilhac, Benjamin Courchesne, Lucie Guillou, Corto Mascle, Isa Vialard | 2026-05-11 | 下载 | Population protocols are a distributed computation model in which a collection of anonymous, finite-state agents interact in randomly chosen pairs and update their states according to a fixed transiti... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Private Information Retrieval With Arbitrary Privacy Requirements for Graph-Based Storage | Mohamed Nomeir, Shreya Meel, Sennur Ulukus | 2026-05-11 | 下载 | We reformulate the definition of privacy in the private information retrieval (PIR) problem to accommodate flexible privacy requirements. We focus on graph-replicated PIR, with a generalized privacy r... |
| Local Private Information Retrieval: A New Privacy Perspective for Graph-Based Replicated Systems | Shreya Meel, Mohamed Nomeir, Sennur Ulukus | 2026-05-11 | 下载 | We rethink the definition of privacy in multi-server, graph-replicated private information retrieval (PIR) systems, and introduce a novel setting where the user's privacy is governed by the servers' s... |
| BEACON: A Multimodal Dataset for Learning Behavioral Fingerprints from Gameplay Data | Ishpuneet Singh, Gursmeep Kaur, Uday Pratap Singh Atwal, Guramrit Singh, Gurjot Singh, Maninder Singh | 2026-05-11 | 下载 | Continuous authentication in high-stakes digital environments requires datasets with fine-grained behavioral signals under realistic cognitive and motor demands. |
| Large Spectrum Models (LSMs): Decoder-Only Transformer-Powered Spectrum Activity Forecasting via Tokenized RF Data | Mohammad Mosiur Lunar, Mehmet C. Vuran | 2026-05-11 | 下载 | Dynamic spectrum access (DSA) has become a key pillar of next-generation wireless systems to address the spectrum scarcity due to the rapid growth of connected devices. |
| Democratizing Measurement of Critical Mobile Infrastructure: Security and Privacy in an Increasingly Centralized Communication Ecosystem | Gabriel K. Gegenhuber | 2026-05-11 | 下载 | Cellular networks serve as the backbone of global communication, providing critical access to telephony and the Internet, often in regions lacking alternatives. |
| Demystifying Deep Reinforcement Learning: A Neuro-Symbolic Framework for Interpretable Open RAN Automation | Jie Lu, Peihao Yan, Pang-Ning Tan, Y. Thomas Hou, Huacheng Zeng | 2026-05-11 | 下载 | Open Radio Access Networks (O-RAN) are increasingly adopting data-driven control through Deep Reinforcement Learning (DRL) to optimize complex tasks such as network slicing and mobility management. |
| DRIFT: Drift-Resilient Invariant-Feature Transformer for DGA Detection | Chaeyoung Lee, Chaeri Jung, Seonghoon Jeong | 2026-05-11 | 下载 | Domain Generation Algorithms (DGAs) evolve continuously to evade botnet detection, posing a persistent challenge for dependable network defense. |
| Agentic Performance at the Edge: Insights from Benchmarking | Shiqiang Wang, Herbert Woisetschläger | 2026-05-11 | 下载 | Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to models around 8 billion parameters or smaller. |
| Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An Overview | Liping Tao, Xindi Tong, Chee Wei Tan | 2026-05-11 | 下载 | Low earth orbit (LEO) satellite networks are emerging as a key infrastructure for global connectivity and space-based sensing. Many tasks in such systems can be formulated as measurement-set-to-spatia... |
| Statistical Analysis for Energy-Efficient Satellite Edge Computing with Latency Guarantees | Nicolai Dalsgaard Lyholm, Beatriz Soret, Tijana Devaja, Thomas Grundgaard Mulvad, Cedomir Stefanovic, Israel Leyva-Mayorga | 2026-05-11 | 下载 | Being able to provide latency guarantees for orbital edge computing applications through Low Earth Orbit (LEO) satellite constellations is a major milestone for their integration into 5G and 6G networ... |
| Key Encapsulation Mechanism-Based Integrated Encryption Scheme (KEM-IES) | Abel C. H. Chen | 2026-05-11 | 下载 | The Elliptic Curve Integrated Encryption Scheme (ECIES) is widely regarded as a practical method and has been adopted by multiple standards. However, the advancement of quantum computing technologies ... |
| Is DRL-based MAC Ready for Underwater Acoustic Networks? Exploring Its Practicality in Real Field Experiments | Jiani Guo, Bingwen Huangfu, Shanshan Song, Nan Sun, Miao Pan, Guangjie Han | 2026-05-11 | 下载 | Medium Access Control (MAC) protocols rely on neighbor and environment information to design collision-free access rules for Underwater Acoustic Networks (UANs). |
| GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM Inference | Zengzipeng Tang, Yuxuan Sun, Wei Chen, Jianwen Ding, Bo Ai | 2026-05-11 | 下载 | The recent growth of on-device Large Language Model (LLM) inference has driven significant interest in device-edge collaborative LLM inference. |
| Learning to Compress and Transmit: Adaptive Rate Control for Semantic Communications over LEO Satellite-to-Ground Links | Jiangtao Luo, Yongyi Ran, Guoliang Xu, Jihua Zhou | 2026-05-11 | 下载 | The bottleneck of satellite-to-ground links poses a major challenge for the timely downlink of massive on-board imagery. This paper studies adaptive image transmission over LEO satellite-to-ground lin... |
| In-Network Artificial Computing Enhanced Light Model-Switching for Emergency Communications Networks | Yuehan Li, Zhiyuan Ren, Tao Zhang, Wenchi Cheng | 2026-05-11 | 下载 | Emergency communications networks require in-network intelligence for timely traffic handling under dynamic demands and runtime constraints. In these environments, packets may need different inference... |
| GenioSim: A Novel Simulation Platform for Edge Computing over Optical Networks | Carmine Cesarano, Alessio Foggia, Roberto Natella | 2026-05-11 | 下载 | The convergence of Passive Optical Networks (PONs) and edge computing creates new opportunities: Optical Line Terminals (OLTs) and Optical Network Terminals (ONTs) can be repurposed as low-latency edg... |
| Bridging the Cognitive Gap: A Unified Memory Paradigm for 6G Agentic AI-RAN | Xijun Wang, Zhaoyang Liu, Chenyuan Feng, Xiang Chen, Howard H. Yang, Tony Q. S. Quek | 2026-05-11 | 下载 | As 6G evolves, the radio access network must transcend traditional automation to embrace agentic AI capable of perception, reasoning, and evolution. |
| CloudEmu: A Trace-Driven Cloud-Native Emulation Testbed for Vehicle Video Uplink over Cellular Networks | Takashi Torii, Soto Anno, Masaki Okada, Takuma Tsubaki, Nobuhiro Azuma, Takuya Tojo | 2026-05-11 | 下载 | We present CloudEmu, a trace-driven, cloud-native cellular-emulation testbed for vehicle video uplink communication. Reliable, low-latency video uplink over cellular networks is essential for remote m... |
| Mixed-Criticality Flow Scheduling with Low Delay and Limited Bandwidth in TSN | Wenyan Yan, Sijing Duan, Dongsheng Wei | 2026-05-11 | 下载 | Time-Sensitive Networking (TSN) is a promising Ethernet protocol with time determinism, widely used in time-critical systems such as industrial automation, automotive networks, and avionics. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces | Srinivas Sridharan, Andy Balogh, Bradford M. Beckmann, Brian Coutinho, Louis Feng, Sheng Fu, Sanshan Gao, Mehryar Garakani, Taekyung Heo, David Kanter, Josh Ladd, Ziwei Li, Winston Liu, Changhai Man, Dan Mihailescu, Spandan More, Joongun Park, Ashwin Ramachandran, Vinay Ramakrishnaiah, Saeed Rashidi, Vijay Janapa Reddi, Puneet Sharma, Phio Tian, William Won, Hanjiang Wu, Huan Xu, Jinsun Yoo, Tushar Krishna | 2026-05-11 | 下载 | The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload behavior in productio... |
| Enabling Performant and Flexible Model-Internal Observability for LLM Inference | Nengneng Yu, Sixian Xiong, Yibo Zhao, Wei Wang, Zaoxing Liu | 2026-05-11 | 下载 | Today's inference-time workloads increasingly depend on timely access to a model's internal states. We present DMI-Lib, a high-speed deep model inspector that treats internal observability as a first-... |
| An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing Continuum | Suvi De Silva, Alfreds Lapkovskis, Alaa Saleh, Sasu Tarkoma, Praveen Kumar Donta | 2026-05-11 | 下载 | Grey failures in the computing continuum produce ambiguous overlapping symptoms that existing approaches fail to diagnose reliably, either due to a lack of causal awareness or acting under high episte... |
| Geometrically Approximated Modeling for Emitter-Centric Ray-Triangle Filtering in Arbitrarily Dynamic LiDAR Simulation | Rabin Gajmer, Joonas Haapala, Zoltan Beck | 2026-05-11 | 下载 | Real-time Light Detection And Ranging (LiDAR) simulation must find, per emitted ray, the closest intersecting triangle even in dynamic scenes containing large numbers of moving and deformable objects. |
| Key Encapsulation Mechanism-Based Integrated Encryption Scheme (KEM-IES) | Abel C. H. Chen | 2026-05-11 | 下载 | The Elliptic Curve Integrated Encryption Scheme (ECIES) is widely regarded as a practical method and has been adopted by multiple standards. However, the advancement of quantum computing technologies ... |
| Muninn: Your Trajectory Diffusion Model But Faster | Gokul Puthumanaillam, Hao Jiang, Ruben Hernandez, Jose Fuentes, Paulo Padrao, Leonardo Bobadilla, Melkior Ornik | 2026-05-11 | 下载 | Diffusion-based trajectory planners can synthesize rich, multimodal robot motions, but their iterative denoising makes online planning and control prohibitively slow. |
| MambaNetBurst: Direct Byte-level Network Traffic Classification without Tokenization or Pretraining | Gayan K. Kulatilleke, Siamak Layeghy, Mahsa Baktashmotlagh, Marius Portmann | 2026-05-11 | 下载 | We present MambaNetBurst, a compact tokenizer-free byte-level sequence classifier for network burst classification based on a Mamba-2 backbone. |