Skip to content

2026-05-11 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts ModelsJungwoo Kim, Rubens Lacouture, Genghan Zhang, Gina Sohn, Qizheng Zhang, Swapnil Gandhi, Christos Kozyrakis, Kunle Olukotun2026-05-11下载Mixture-of-Experts (MoE) has become a dominant architecture for scaling large language models (LLMs). However, the execution characteristics of MoE inference are changing rapidly and increasingly mism...
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production EnvironmentsYue Guan, Hongtao Yu, Peng Chen, Daohang Shi, Karthik Manivannan, Nicholas J Riasanovsky, Manman Ren, Lei Wang, Shane Nay, Partha Kanuparthy, Zaifeng Pan, Zhengding Hu, Yufei Ding2026-05-11下载Modern GPUs increasingly rely on specialized hardware units and asynchronous coordination mechanisms, so performance depends on orchestrating data movement, tensor-core computation, and synchronizatio...
LLMs for Secure Hardware Design and Related Problems: Opportunities and ChallengesJohann Knechtel, Ozgur Sinanoglu, Ramesh Karri2026-05-11下载The integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) and hardware security is rapidly reshaping the semiconductor industry.
Reconfigurable Computing Challenge: Real-Time Graph Neural Networks for Online Event Selection in Big ScienceMarc Neu, Frank Baptist, Thomas Lobmaier, Fabio Papagno, Torben Ferber, Jürgen Becker2026-05-11下载Graph neural networks are increasingly adopted in trigger systems for collider experiments, where strict latency and throughput constraints render deployment on embedded platforms challenging.
ObfAx: Obfuscation and IP Piracy Detection in Approximate CircuitsLukas Sekanina, Vojtech Mrazek2026-05-11下载Approximate circuits often achieve exceptional trade-offs between computational accuracy and hardware efficiency, making them attractive for deployment as reusable Intellectual Property (IP) cores.
Towards an End-To-End System for Real-Time Gesture Recognition from Surface VibrationsFlorian Hettstedt, Cedric Giese, Tianheng Ling, Keiichi Yasumoto, Gregor Schiele, Andreas Erbslöh2026-05-11下载Sensing surface vibrations promise unobtrusive interaction for smart home systems by enabling gesture recognition on existing everyday surfaces without disturbing living-space design.
Arcane: An Assertion Reduction Framework through Semantic Clustering and MCTS-Guided Rule ExploringHongqin Lyu, Yonghao Wang, Zhiteng Chao, Tiancheng Wang, Huawei Li2026-05-11下载Assertion-based Verification (ABV) is essential for ensuring that hardware designs conform to their intended specifications. However, existing automated assertion-generation approaches, such as LLM-ba...
RFAmpDesigner: A Self-Evolving Multi-Agent LLM Framework for Automated Radio Frequency Amplifier DesignHang Lu, Guochang Li, Qianyu Chen, Huiyan Gao, Shaogang Wang, Xuanyu He, Yiwei Liu, Gaopeng Chen, Nayu Li, Xiaokang Qi, Chunyi Song, Zhiwei Xu2026-05-11下载Automating radio frequency (RF) amplifier design remains challenging because existing methods suffer from the curse of dimensionality, weak use of domain knowledge, and poor transferability, leading t...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
ChunkFlow: Communication-Aware Chunked Prefetching for Layerwise Offloading in Distributed Diffusion Transformer InferenceHan Meng, Danny Willow Liu, Dong Li2026-05-11下载Layerwise offloading reduces the GPU memory footprint of large diffusion transformer (DiT) inference by prefetching upcoming layers from host memory, but its effectiveness hinges on hiding prefetch la...
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution TracesSrinivas Sridharan, Andy Balogh, Bradford M. Beckmann, Brian Coutinho, Louis Feng, Sheng Fu, Sanshan Gao, Mehryar Garakani, Taekyung Heo, David Kanter, Josh Ladd, Ziwei Li, Winston Liu, Changhai Man, Dan Mihailescu, Spandan More, Joongun Park, Ashwin Ramachandran, Vinay Ramakrishnaiah, Saeed Rashidi, Vijay Janapa Reddi, Puneet Sharma, Phio Tian, William Won, Hanjiang Wu, Huan Xu, Jinsun Yoo, Tushar Krishna2026-05-11下载The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload behavior in productio...
Byzantine Consensus in Directed Graphs with Message AuthenticationNitin H. Vaidya, Lewis Tseng2026-05-11下载We consider the problem of reaching consensus in communication networks that are modeled by directed graphs. We assume the existence of a message authentication mechanism (such as digital signatures) ...
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile WorkloadZiyue Liu, Zhengyang Wang, Ruijie Zhang, Avinash Maurya, Hui Zhou, Paul Hovland, Sheng Di, Franck Cappello, Bogdan Nicolae, Zheng Zhang2026-05-11下载Pre-training large language models on massive GPU clusters has made hardware faults routine rather than rare, driving the need for resilient training systems.
ShardTensor: Domain Parallelism for Scientific Machine LearningCorey Adams, Peter Harrington, Akshay Subramaniam, Mohammad Shoaib Abbas, Jaideep Pathak, Mike Pritchard, Sanjay Choudhry2026-05-11下载Scientific Machine Learning (SciML) faces unique challenges for extreme-resolution data, with mitigations that often fail to scale or degrade the accuracy of trained models.
Closer in the Gap: Towards Portable Performance on RISC-V Vector ProcessorsRuimin Shi, Maya Gokhale, Pei-Hung Lin, Xavier Teruel, Ivy Peng2026-05-11下载The RISC-V Vector Extension~(RVV) is a cornerstone for supporting compute throughout in scientific and machine learning workloads. Yet compiler support and performance monitoring on real RVV~1.
An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing ContinuumSuvi De Silva, Alfreds Lapkovskis, Alaa Saleh, Sasu Tarkoma, Praveen Kumar Donta2026-05-11下载Grey failures in the computing continuum produce ambiguous overlapping symptoms that existing approaches fail to diagnose reliably, either due to a lack of causal awareness or acting under high episte...
Surviving Partial Rank Failures in Wide Expert-Parallel MoE InferenceXun Sun, Shaoyuan Chen, Pingchuan Ma, Yue Chen, Ziwei Yuan, Zhanhao Cao, Han Han, Shangming Cai, Teng Ma, Xuchun Shang, Xinpeng Zhao, Ke Yang, Junlin Wei, Lianzhi Lin, Yuji Liu, Feng Ren, Haoran Hu, Cheng Wan, Yingdi Shan, Yongwei Wu, Mingxing Zhang2026-05-11下载Mixture-of-Experts (MoE) serving relies on wide expert parallelism (EP) to aggregate the memory capacity and bandwidth of many GPUs within one inference instance.
SoK: A Systematic Bidirectional Literature Review of AI & DLT ConvergenceAli Irzam Kathia, Yimika Erinle, Abylay Satybaldy, Paolo Tasca, Nikhil Vadgama, Marco Alberto Javarone2026-05-11下载The integration of Artificial Intelligence (AI) with Distributed Ledger Technology (DLT) has become a growing research area, yet contributions tend to cluster around specific application domains or ex...
Accelerating Compound LLM Training Workloads with MaestroXiulong Yuan, Hongqing Chen, Jiaxuan Peng, Fan Zhou, Zhixiang Ruan, Zekun Wang, Bo Zheng, Rui Men, Haiquan Wang, Zhipeng Zhang, Langshi Chen, Man Yuan, Jiaqi Gao, Zhengping Qian, Junyang Lin, Yong Li, Wei Lin, Junhua Wang, Jingren Zhou2026-05-11下载Compound LLM training workloads-such as knowledge distillation and multimodal LLM (MLLM) training-are gaining prominence. These typically comprise heterogeneous components differing in parameter scale...
Privacy-preserving Chunk Scheduling in a BitTorrent Implementation of Federated LearningNaicheng Li, Javad Dogani, Rui Wang, Kaitai Liang, Nikolaos Laoutaris2026-05-11下载Traditional federated learning (FL) relies on a central aggregator server, which can create performance bottlenecks and privacy risks. Decentralized mix-and-forward designs remove the server, but repe...
HiRL: Hierarchical Reinforcement Learning for Coordinated Resource Management in Heterogeneous Edge ComputingJianyong Zhu, Hao Chen, Juan Zhang, Fangda Guo, Albert Y. Zomaya, Renyu Yang2026-05-11下载Edge computing faces unprecedented resource orchestration challenges from multi-dimensional heterogeneity across device architectures, diverse task requirements in CPU-intensive, GPU-intensive, I/O-in...
FractalSortCPU: Bandwidth-Efficient Compressed Radix Sort on CPUMichael Dang'ana2026-05-11下载Cloud database systems, particularly their middleware and query execution layers, use sorting as a core operation in query processing, indexing and join execution.
Agentic Performance at the Edge: Insights from BenchmarkingShiqiang Wang, Herbert Woisetschläger2026-05-11下载Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to models around 8 billion parameters or smaller.
Amortized Asynchronous Byzantine Reliable Broadcast with Optimal ResilienceMichael Yiqing Hu, Alvin Hong Yao Yan, Jialin Li2026-05-11下载Byzantine Reliable Broadcast (BRB) is a fundamental primitive in distributed computing and cryptographic systems. Reducing the communication complexity of BRB protocols remains an important research d...
Autonomous FAIR Digital Objects: From Passive Assertions to Active KnowledgeZeyd Boukhers, Oya Beyan, Cong Yang, Christoph Lange2026-05-11下载Scientific knowledge on the Web is published as passive assertions and cannot decide when to validate evidence, reconcile contradictions, or update confidence as findings accumulate.
Accelerating Locality-Driven Integration in Quantum Chemistry with Block-Structured Matrix MultiplicationXinran Wei, Yan Pan, Fusong Ju, Zehao Zhou, Yihong Zhang, Lin Huang, Jianwei Zhu, Jia Zhang, Huanhuan Xia, Bin Shao, Tao Qin2026-05-11下载Locality-driven integration is a pervasive computational pattern in quantum chemistry, arising whenever spatially localized basis functions interact through numerical quadrature or integral screening.
FusionRCG: Orchestrating Recursive Computation Graphs across GPU Memory HierarchiesYihong Zhang, Xinran Wei, Junshi Chen, Fusong Ju, Wei Hu, Jinlong Yang, Huanhuan Xia2026-05-11下载Evaluating high-dimensional integrals via deep hierarchical recurrences is a dominant cost in quantum chemistry. While CPUs manage these efficiently, GPUs suffer a critical mismatch: limited per-threa...
DP-LAC: Lightweight Adaptive Clipping for Differentially Private Federated Fine-tuning of Language ModelsHaaris Mehmood, Jie Xu, Karthikeyan Saravanan, Rogier Van Dalen, Mete Ozay2026-05-11下载Federated learning (FL) enables the collaborative training of large-scale language models (LLMs) across edge devices while keeping user data on-device.
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM InferenceZengzipeng Tang, Yuxuan Sun, Wei Chen, Jianwen Ding, Bo Ai2026-05-11下载The recent growth of on-device Large Language Model (LLM) inference has driven significant interest in device-edge collaborative LLM inference.
Edge-Cloud Collaborative Pothole Detection via Onboard Event Screening and Federated Temporal SegmentationYingjie Wu, Kongyang Chen, Tiancai Liang2026-05-11下载Road potholes threaten driving safety and increase infrastructure maintenance costs, while large-scale and timely pothole detection remains challenging in urban road networks.
Lakestream: A Consistent and Brokerless Data Plane for Large Foundation Model TrainingTing Sun, Junjie Zhang, Xiao Yan, Songxin Zhang, Zhuoyang Song, Jingyi Xi, Zunyao Mao, Bingyi Jing, Jiaxing Zhang, Zejian Xie2026-05-11下载Modern Large Foundation Model (LFM) training has transformed the data pipeline from a static ingestion layer into a dynamic component that must co-evolve with the training process.
Population Protocols over Ordered AgentsMichael Blondin, Michaël Cadilhac, Benjamin Courchesne, Lucie Guillou, Corto Mascle, Isa Vialard2026-05-11下载Population protocols are a distributed computation model in which a collection of anonymous, finite-state agents interact in randomly chosen pairs and update their states according to a fixed transiti...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Private Information Retrieval With Arbitrary Privacy Requirements for Graph-Based StorageMohamed Nomeir, Shreya Meel, Sennur Ulukus2026-05-11下载We reformulate the definition of privacy in the private information retrieval (PIR) problem to accommodate flexible privacy requirements. We focus on graph-replicated PIR, with a generalized privacy r...
Local Private Information Retrieval: A New Privacy Perspective for Graph-Based Replicated SystemsShreya Meel, Mohamed Nomeir, Sennur Ulukus2026-05-11下载We rethink the definition of privacy in multi-server, graph-replicated private information retrieval (PIR) systems, and introduce a novel setting where the user's privacy is governed by the servers' s...
BEACON: A Multimodal Dataset for Learning Behavioral Fingerprints from Gameplay DataIshpuneet Singh, Gursmeep Kaur, Uday Pratap Singh Atwal, Guramrit Singh, Gurjot Singh, Maninder Singh2026-05-11下载Continuous authentication in high-stakes digital environments requires datasets with fine-grained behavioral signals under realistic cognitive and motor demands.
Large Spectrum Models (LSMs): Decoder-Only Transformer-Powered Spectrum Activity Forecasting via Tokenized RF DataMohammad Mosiur Lunar, Mehmet C. Vuran2026-05-11下载Dynamic spectrum access (DSA) has become a key pillar of next-generation wireless systems to address the spectrum scarcity due to the rapid growth of connected devices.
Democratizing Measurement of Critical Mobile Infrastructure: Security and Privacy in an Increasingly Centralized Communication EcosystemGabriel K. Gegenhuber2026-05-11下载Cellular networks serve as the backbone of global communication, providing critical access to telephony and the Internet, often in regions lacking alternatives.
Demystifying Deep Reinforcement Learning: A Neuro-Symbolic Framework for Interpretable Open RAN AutomationJie Lu, Peihao Yan, Pang-Ning Tan, Y. Thomas Hou, Huacheng Zeng2026-05-11下载Open Radio Access Networks (O-RAN) are increasingly adopting data-driven control through Deep Reinforcement Learning (DRL) to optimize complex tasks such as network slicing and mobility management.
DRIFT: Drift-Resilient Invariant-Feature Transformer for DGA DetectionChaeyoung Lee, Chaeri Jung, Seonghoon Jeong2026-05-11下载Domain Generation Algorithms (DGAs) evolve continuously to evade botnet detection, posing a persistent challenge for dependable network defense.
Agentic Performance at the Edge: Insights from BenchmarkingShiqiang Wang, Herbert Woisetschläger2026-05-11下载Agentic artificial intelligence (AI) is a natural fit for Internet of Things (IoT) and edge systems, but edge deployments are often constrained to models around 8 billion parameters or smaller.
Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An OverviewLiping Tao, Xindi Tong, Chee Wei Tan2026-05-11下载Low earth orbit (LEO) satellite networks are emerging as a key infrastructure for global connectivity and space-based sensing. Many tasks in such systems can be formulated as measurement-set-to-spatia...
Statistical Analysis for Energy-Efficient Satellite Edge Computing with Latency GuaranteesNicolai Dalsgaard Lyholm, Beatriz Soret, Tijana Devaja, Thomas Grundgaard Mulvad, Cedomir Stefanovic, Israel Leyva-Mayorga2026-05-11下载Being able to provide latency guarantees for orbital edge computing applications through Low Earth Orbit (LEO) satellite constellations is a major milestone for their integration into 5G and 6G networ...
Key Encapsulation Mechanism-Based Integrated Encryption Scheme (KEM-IES)Abel C. H. Chen2026-05-11下载The Elliptic Curve Integrated Encryption Scheme (ECIES) is widely regarded as a practical method and has been adopted by multiple standards. However, the advancement of quantum computing technologies ...
Is DRL-based MAC Ready for Underwater Acoustic Networks? Exploring Its Practicality in Real Field ExperimentsJiani Guo, Bingwen Huangfu, Shanshan Song, Nan Sun, Miao Pan, Guangjie Han2026-05-11下载Medium Access Control (MAC) protocols rely on neighbor and environment information to design collision-free access rules for Underwater Acoustic Networks (UANs).
GELATO: Generative Entropy- and Lyapunov-based Adaptive Token Offloading for Device-Edge Speculative LLM InferenceZengzipeng Tang, Yuxuan Sun, Wei Chen, Jianwen Ding, Bo Ai2026-05-11下载The recent growth of on-device Large Language Model (LLM) inference has driven significant interest in device-edge collaborative LLM inference.
Learning to Compress and Transmit: Adaptive Rate Control for Semantic Communications over LEO Satellite-to-Ground LinksJiangtao Luo, Yongyi Ran, Guoliang Xu, Jihua Zhou2026-05-11下载The bottleneck of satellite-to-ground links poses a major challenge for the timely downlink of massive on-board imagery. This paper studies adaptive image transmission over LEO satellite-to-ground lin...
In-Network Artificial Computing Enhanced Light Model-Switching for Emergency Communications NetworksYuehan Li, Zhiyuan Ren, Tao Zhang, Wenchi Cheng2026-05-11下载Emergency communications networks require in-network intelligence for timely traffic handling under dynamic demands and runtime constraints. In these environments, packets may need different inference...
GenioSim: A Novel Simulation Platform for Edge Computing over Optical NetworksCarmine Cesarano, Alessio Foggia, Roberto Natella2026-05-11下载The convergence of Passive Optical Networks (PONs) and edge computing creates new opportunities: Optical Line Terminals (OLTs) and Optical Network Terminals (ONTs) can be repurposed as low-latency edg...
Bridging the Cognitive Gap: A Unified Memory Paradigm for 6G Agentic AI-RANXijun Wang, Zhaoyang Liu, Chenyuan Feng, Xiang Chen, Howard H. Yang, Tony Q. S. Quek2026-05-11下载As 6G evolves, the radio access network must transcend traditional automation to embrace agentic AI capable of perception, reasoning, and evolution.
CloudEmu: A Trace-Driven Cloud-Native Emulation Testbed for Vehicle Video Uplink over Cellular NetworksTakashi Torii, Soto Anno, Masaki Okada, Takuma Tsubaki, Nobuhiro Azuma, Takuya Tojo2026-05-11下载We present CloudEmu, a trace-driven, cloud-native cellular-emulation testbed for vehicle video uplink communication. Reliable, low-latency video uplink over cellular networks is essential for remote m...
Mixed-Criticality Flow Scheduling with Low Delay and Limited Bandwidth in TSNWenyan Yan, Sijing Duan, Dongsheng Wei2026-05-11下载Time-Sensitive Networking (TSN) is a promising Ethernet protocol with time determinism, widely used in time-critical systems such as industrial automation, automotive networks, and avionics.

cs.PF - Performance ​

标题作者发布日期PDF摘要
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution TracesSrinivas Sridharan, Andy Balogh, Bradford M. Beckmann, Brian Coutinho, Louis Feng, Sheng Fu, Sanshan Gao, Mehryar Garakani, Taekyung Heo, David Kanter, Josh Ladd, Ziwei Li, Winston Liu, Changhai Man, Dan Mihailescu, Spandan More, Joongun Park, Ashwin Ramachandran, Vinay Ramakrishnaiah, Saeed Rashidi, Vijay Janapa Reddi, Puneet Sharma, Phio Tian, William Won, Hanjiang Wu, Huan Xu, Jinsun Yoo, Tushar Krishna2026-05-11下载The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload behavior in productio...
Enabling Performant and Flexible Model-Internal Observability for LLM InferenceNengneng Yu, Sixian Xiong, Yibo Zhao, Wei Wang, Zaoxing Liu2026-05-11下载Today's inference-time workloads increasingly depend on timely access to a model's internal states. We present DMI-Lib, a high-speed deep model inspector that treats internal observability as a first-...
An Uncertainty-Aware Resilience Micro-Agent for Causal Observability in the Computing ContinuumSuvi De Silva, Alfreds Lapkovskis, Alaa Saleh, Sasu Tarkoma, Praveen Kumar Donta2026-05-11下载Grey failures in the computing continuum produce ambiguous overlapping symptoms that existing approaches fail to diagnose reliably, either due to a lack of causal awareness or acting under high episte...
Geometrically Approximated Modeling for Emitter-Centric Ray-Triangle Filtering in Arbitrarily Dynamic LiDAR SimulationRabin Gajmer, Joonas Haapala, Zoltan Beck2026-05-11下载Real-time Light Detection And Ranging (LiDAR) simulation must find, per emitted ray, the closest intersecting triangle even in dynamic scenes containing large numbers of moving and deformable objects.
Key Encapsulation Mechanism-Based Integrated Encryption Scheme (KEM-IES)Abel C. H. Chen2026-05-11下载The Elliptic Curve Integrated Encryption Scheme (ECIES) is widely regarded as a practical method and has been adopted by multiple standards. However, the advancement of quantum computing technologies ...
Muninn: Your Trajectory Diffusion Model But FasterGokul Puthumanaillam, Hao Jiang, Ruben Hernandez, Jose Fuentes, Paulo Padrao, Leonardo Bobadilla, Melkior Ornik2026-05-11下载Diffusion-based trajectory planners can synthesize rich, multimodal robot motions, but their iterative denoising makes online planning and control prohibitively slow.
MambaNetBurst: Direct Byte-level Network Traffic Classification without Tokenization or PretrainingGayan K. Kulatilleke, Siamak Layeghy, Mahsa Baktashmotlagh, Marius Portmann2026-05-11下载We present MambaNetBurst, a compact tokenizer-free byte-level sequence classifier for network burst classification based on a Mamba-2 backbone.

基于 VitePress 构建 · 使用本地搜索查找论文