2026-04-25
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Hybrid JIT-CUDA Graph Optimization for Low-Latency Large Language Model Inference | Divakar Kumar Yadav, Tian Zhao | 2026-04-25 | 下载 | Large Language Models (LLMs) have achieved strong performance across natural language and multimodal tasks, yet their practical deployment remains constrained by inference latency and kernel launch ov... |
| Evaluating CUDA Tile for AI Workloads on Hopper and Blackwell GPUs | Divakar Kumar Yadav, Tian Zhao, Deepak Kumar | 2026-04-25 | 下载 | NVIDIA's CUDA Tile (CuTile) introduces a Python-based, tile-centric abstraction for GPU kernel development that aims to simplify programming while retaining Tensor Core and Tensor Memory Accelerator (... |
| Tessera: Secure, Near-Line-Rate Weight Streaming for UMA Edge Accelerators | Animan Naskar | 2026-04-25 | 下载 | Deploying proprietary Deep Neural Networks (DNNs) on commodity edge devices demands hardware-backed Digital Rights Management (DRM) capable of withstanding both software-level and physical adversaries... |
| Efficient VQ-QAT and Mixed Vector/Linear quantized Neural Networks | Terry Gou, Puneet Gupta | 2026-04-25 | 下载 | In this work, we developed and tested 3 techniques for vector quantization (VQ) based model weight compression. To mitigate codebook collapse and enable end-to-end training, we adopted cosine similari... |
| Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns | Abhimanyu Bambhaniya, Geonhwa Jeong, Jason Park, Jiecao Yu, Jaewon Lee, Pengchao Wang, Changkyu Kim, Chunqiang Tang, Tushar Krishna | 2026-04-25 | 下载 | Most recent state-of-the-art (SOTA) large language models (LLMs) use Mixture-of-Experts (MoE) architectures to scale model capacity without proportional per-token compute, enabling higher-quality outp... |
| Maximizing Memory-Level Parallelism via Integrated Stochastic Logic-in-Memory Architectures | Farzad Razi, Mehran Moghadam, Sercan Aygun, M. Hassan Najafi, Marc Riedel | 2026-04-25 | 下载 | Today's high-performance architectures are increasingly constrained by data movement latency and energy overhead, as the slowdown of single-core performance scaling coincides with the rise of highly d... |
| From Language to Logic: Bridging LLMs & Formal Representations for RTL Assertion Generation | Nowfel Mashnoor, Hadi Kamali, Kimia Azar | 2026-04-25 | 下载 | SystemVerilog Assertions (SVA) are essential for formal verification of digital hardware, yet their manual creation demands significant expertise in both the design under verification and temporal log... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A Taxonomy and Resolution Strategy for Client-Level Disagreements in Federated Learning | Daan Rosendal, Ana Oprescu | 2026-04-25 | 下载 | Federated Learning (FL) typically assumes unconditional collaboration, a premise that overlooks the complexities of real-world, multi-stakeholder environments in which clients may need to exclude one ... |
| The Blockchain Execution Dilemma: Optimizing Revenue XOR Fair Ordering | Artjom Pugatsov, Can Umut Ileri, Jérémie Decouchant | 2026-04-25 | 下载 | The successive generations of consensus algorithms have progressively shifted the performance bottleneck of blockchains to the execution layer. |
| GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training | Arefin Niam, Tevfik Kosar, M. S. Q. Zulkar Nine | 2026-04-25 | 下载 | Distributed GNN training is dominated by remote feature fetching, which can be very costly. Multi-hop neighborhood sampling crosses partition boundaries and triggers fine-grained RPCs whose fixed init... |
| Usable Agent Discovery for Decentralized AI Systems | Patrizio Dazzi, Emanuele Carlini, Matteo Mordacchini, Saul Urso | 2026-04-25 | 下载 | Large-scale agentic systems run on distributed infrastructures where many software agents share physical hosts and are discovered via peer-to-peer mechanisms. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ARIstoteles -- Dissecting Apple's Baseband Interface | Tobias Kröll, Stephan Kleber, Frank Kargl, Matthias Hollick, Jiska Classen | 2026-04-25 | 下载 | Wireless chips and interfaces expose a substantial remote attack surface. As of today, most cellular baseband security research is performed on the Android ecosystem, leaving a huge gap on Apple devic... |
| ARCHES: Adaptive Real-Time Switching of AI Models for the RAN | Neagin Neasamoni Santhi, Davide Villa, Michele Polese, Salvatore D'Oro, Yunseong Lee, Koichiro Furueda, Tommaso Melodia | 2026-04-25 | 下载 | Artificial Intelligence (AI) has become a powerful tool for model-free Radio Access Network (RAN) signal processing and optimization. However, designing a single model that generalizes across all radi... |
| Sharing-oriented Resource Allocation for Multi-platoon's Groupcasting and Unicasting Communication based on the Transmission Reliability | Chung-Ming Huang, Yen-Hung Wu, Duy-Tuan Dao | 2026-04-25 | 下载 | Resource allocation in vehicular platoons is challenging due to high vehicle mobility and limited spectrum resources. To improve spectral efficiency, resource sharing is commonly adopted. |
| Advanced Anomaly Detection and Threat Intelligence in Zero Trust IoT Environments Using Machine Learning | Muhammad Umair Basharat, Jawad Hussain, Waqas Khalid, Chiew Foong Kwong | 2026-04-25 | 下载 | The growing adoption of IoT and cloud computing, combined with rapid advancements in digital technologies, has considerably increased the cyber-attack surface, resulting in increasingly complex and pe... |
| RadTwin: Generalizable Wireless Digital Twin for Dynamic Environments | Yuru Zhang, Ming Zhao, Qiang Liu, Ahmed Alkhateeb, Abhishek K. Agrawal, Qi Qu | 2026-04-25 | 下载 | Precisely modeling radio propagation in dynamic wireless environments is fundamental to the realization of wireless digital twins. Traditional ray tracing methods rely on accurate 3D models with detai... |
| An Analysis of Active Learning Algorithms using Real-World Crowd-sourced Text Annotations | Varun Totakura, Ankita Singh, Yushun Dong, Shayok Chakraborty | 2026-04-25 | 下载 | Active learning algorithms automatically identify the most informative samples from large amounts of unlabeled data and tremendously reduce human annotation effort in inducing a machine learning model... |
| An Agentic Framework for Intent Co-Creation in 6G NaaS: Architecture and Open-Source Model Evaluation | Kostis Trantzas, Besiana Agko, Christos Tranoris, Irene Denazi | 2026-04-25 | 下载 | 6G network complexity necessitates high levels of autonomy, yet current intent-based systems struggle with ambiguous or incomplete human requests. |
| Towards Agentic Test-Driven Quality Assurance for 6G Networks | Christos Tranoris, Besiana Agko, Kostis Trantzas, Irene Denazi | 2026-04-25 | 下载 | This work proposes an agentic, intent-driven end-to-end (E2E) orchestration framework that integrates intent co-creation with a Test-Driven Quality Assurance paradigm. |
| RANalyzer: Automated Continuous RAN Software Evaluation and Regression Analysis | Ravis Shirkhani, Reshma Prasad, Leonardo Bonati, Tommaso Melodia, Michele Polese | 2026-04-25 | 下载 | Software-driven O-RAN architectures enable rapid innovation through frequent, independent updates to virtualized components. However, attributing performance variations to specific software changes is... |
| Source-Code Analysis of iFogSim for Simulating Distributed IoT Architectures: Coverage, Challenges, and Enhancements | Milliam Maxime Zekeng Ndadji | 2026-04-25 | 下载 | Simulation is an indispensable tool for validating distributed IoT architectures before physical deployment, and iFogSim has emerged as one of the most widely adopted platform in the fog and edge comp... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Approximating Uniform Random Rotations by Two-Block Structured Hadamard Rotations in High Dimensions | Tomer Zilca, Gal Mendelson | 2026-04-25 | 下载 | Uniform random rotations are a useful primitive in applications such as fast Johnson-Lindenstrauss embeddings, kernel approximation, communication-efficient learning, and recent AI compression pipelin... |