2026-05-01
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline Analysis | Zhongchun Zhou, Yuhang Gu, Chengtao Lai, Ya Wang, Wei Zhang | 2026-05-01 | 下载 | To efficiently support Large Language Models (LLMs), modern GPGPU architectures have introduced new features and programming paradigms, such as warp specialization. |
| Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge | M. Grailoo, J. Núñez-Yáñez | 2026-05-01 | 下载 | Scaling laws for Large Language Models (LLMs) establish that model quality improves with computational scale, yet edge deployment imposes strict constraints on compute, memory, and power. |
| Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference | Abdurrahman Javat, Allan Kazakov | 2026-05-01 | 下载 | The operational landscape of local Large Language Model (LLM) inference has shifted from lightweight models to datacenter-class weights exceeding 70B parameters, creating profound systems challenges f... |
| VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge Devices | Zi-Wei Lin, Tian-Sheuan Chang | 2026-05-01 | 下载 | We present VitaLLM, a mixed precision accelerator that enables ternary weight large language models to run efficiently on edge devices. The design combines two compute cores, a multiplier free TINT co... |
| A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator with In-Situ Regulation for Energy-Efficient Spiking Neural Networks | Shih-Hang Kao, Yang-Chan Hung, I-Wen Wang, Bing-Han Liu, Yu-Chia Chen, Tian-Sheuan Chang, Shyh-Jye Jou, Chien-Nan Liu, Hung-Ming Chen, Wei-Zen Chen | 2026-05-01 | 下载 | This paper presents a PVT-resilient, subthreshold SRAM-based computing-in-memory (CIM) macro tailored for energy-efficient spiking neural networks (SNNs). |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A Domain-Driven Design Simulator for Business Logic-Rich Microservice Systems | Daniel da Palma Pereira, António Rito Silva | 2026-05-01 | 下载 | Developing business-logic-rich microservices requires navigating complex trade-offs between data consistency and distributed coordination. Although patterns like Sagas and Transactional Causal Consist... |
| ncsim: A Lightweight Simulator for Networked Edge Computing with Wireless Interference Modeling | Bhaskar Krishnamachari, Maya Gutierrez, Jared Coleman | 2026-05-01 | 下载 | Evaluating DAG task schedulers for wireless edge computing requires jointly modeling compute placement and wireless interference, yet existing tools treat them in isolation. |
| FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal Compression | Ben Mechels, Ryan Billmeyer, Alexander Chen, Shiyang Li, Caiwen Ding | 2026-05-01 | 下载 | Modern high-performance computing and Internet-of-Things deployments increasingly generate large volumes of signal data that must be compressed efficiently on resource-constrained acquisition devices ... |
| SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned Data | Shashank Kapadia, Deep Narayan Mishra, Sujal Reddy Alugubelli, Ajay Kumar, Swapnil Yadav, Rishi Bhatia | 2026-05-01 | 下载 | We present SURGE, a streaming GPU encoding system deployed in production to generate embeddings for over 800 million texts across 40,000 logical partitions. |
| Eliminating Hidden Serialization in Multi-Node Megakernel Communication | Byungsoo Oh, Rachee Singh | 2026-05-01 | 下载 | Recent megakernel designs for Mixture-of-Experts (MoE) inference fuse expert computation with fine-grained, GPU-initiated communication into a single persistent GPU kernel, and outperform collective-b... |
| LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven Sampling | Wei Da, Evangelia Kalyvianaki | 2026-05-01 | 下载 | Realistic evaluation of LLM serving systems requires online workloads, dynamic arrivals, queueing, and the serving engine's local scheduling for execution batching, but running such experiments on GPU... |
| AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMs | Wenxiang Lin, Juntao Huang, Luhan Zhang, Laili Li, Xiang Bao, Mengyang Zhang, Bing Wang, Shaohuai Shi | 2026-05-01 | 下载 | Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective for 4-bit activations and 8-bit gradients, which ... |
| Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge | M. Grailoo, J. Núñez-Yáñez | 2026-05-01 | 下载 | Scaling laws for Large Language Models (LLMs) establish that model quality improves with computational scale, yet edge deployment imposes strict constraints on compute, memory, and power. |
| SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters | Dongxin Guo, Jikun Wu, Siu Ming Yiu | 2026-05-01 | 下载 | AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end l... |
| Space Network of Experts: Architecture and Expert Placement | Zhanwei Wang, Huiling Yang, Min Sheng, Khaled B. Letaief, Kaibin Huang | 2026-05-01 | 下载 | Leveraging continuous solar energy harvesting at high efficiency, space data centers are envisioned as a promising platform for executing energy-intensive large language models (LLMs). |
| Adaptation of AI-accelerated CFD Simulations to the IPU platform | P. Rosciszewski, A. Krzywaniak, S. Iserte, K. Rojek, P. Gepner | 2026-05-01 | 下载 | Intelligence Processing Units (IPU) have proven useful for many AI applications. In this paper, we evaluate them within the emerging field of \emph{AI for simulation}, where traditional numerical simu... |
| Hierarchical Federated Learning for Networked AI: From Communication Saving to Architecture-Aware Design | Seyed Mohammad Azimi-Abarghouyi, Mehdi Bennis, Leandros Tassiulas | 2026-05-01 | 下载 | Federated learning (FL) is fundamentally a distributed optimization problem executed by communicating agents with local data, local computation, and partial system visibility. |
| Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference | Yuxuan Gao, Megan Wang, Yi Ling Yu | 2026-05-01 | 下载 | Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) ... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| MORPH: Multi-Environment Orchestrated Reinforcement Learning for PRB Handling in O-RAN | Alireza Ebrahimi Dorcheh, Tolunay Seyfi, Ryan Barker, Fatemeh Afghah | 2026-05-01 | 下载 | Reinforcement-learning (RL) solutions for dynamic spectrum access and radio resource management in Open Radio Access Networks (O-RAN) depend critically on the fidelity of the throughput signal used fo... |
| AIIM: Adaptive Inter-cell Interference Mitigation for Heterogeneous Multi-vendor 5G O-RAN Networks | Samuel Reinders, Alireza Ebrahimi Dorcheh, Ryan Barker, Tolunay Seyfi, Fatemeh Afghah | 2026-05-01 | 下载 | Inter-cell interference is a persistent issue in dense 5G deployments, especially in heterogeneous Open Radio Access Network (O-RAN) environments where coordination between base stations is limited. |
| ncsim: A Lightweight Simulator for Networked Edge Computing with Wireless Interference Modeling | Bhaskar Krishnamachari, Maya Gutierrez, Jared Coleman | 2026-05-01 | 下载 | Evaluating DAG task schedulers for wireless edge computing requires jointly modeling compute placement and wireless interference, yet existing tools treat them in isolation. |
| AdvNet: Revealing Performance Issues in Network Protocols by Generating Adversarial Environments | Shehab Sarar Ahmed, William Sentosa, Yinjie Zhang, Yoav Lebendiker, Michael Shnaiderman, Tomer Gilad, Nathan H. Jay, Brighten Godfrey, Michael Schapira | 2026-05-01 | 下载 | Infrastructure protocols like Congestion Control (CC) seek to provide reliable performance across a wide range of Internet environments. Currently, protocol designers assess performance through hand-d... |
| EASE: Federated Multimodal Unlearning via Entanglement-Aware Anchor Closure | Zihao Ding, Beining Wu, Jun Huang | 2026-05-01 | 下载 | Federated Multimodal Learning (FML) trains multimodal models across decentralized clients while keeping their image-text pairs private. However, joint embedding training entangles forgotten knowledge ... |
| Inductive Latent Context Persistence: Closing the Post-Handover Cold Start in 6G Radio Access Networks | Anubhab Banerjee, Daniyal Amir Awan | 2026-05-01 | 下载 | In modern radio access networks (RANs), rule-based handover (HO) decisions (e.g., A3/A5) depend on user equipment (UE) measurements only, so UEs at the same location can receive inconsistent HO outcom... |
| Beyond Per-Request QoS: Coordinating Industrial Workflows with B5G/6G Network Capabilities | Qize Guo, Bjoern Riemer, Tarik Taleb, Yan Chen, Hao Yu, Hemant Zope | 2026-05-01 | 下载 | Beyond-5G (B5G) and 6G networks are expected to enable more complex industrial services, which often operate according to multi-phase workflows with phase-specific communication requirements. |
| Space Network of Experts: Architecture and Expert Placement | Zhanwei Wang, Huiling Yang, Min Sheng, Khaled B. Letaief, Kaibin Huang | 2026-05-01 | 下载 | Leveraging continuous solar energy harvesting at high efficiency, space data centers are envisioned as a promising platform for executing energy-intensive large language models (LLMs). |
| A Policy-Driven DRL Framework for System-Level Tradeoff Control in NR-U/Wi-Fi Coexistence | Po-Heng Chou, Yi-Fang Yu, Shou-Yu Chen, Chiapin Wang | 2026-05-01 | 下载 | The coexistence of NR-U and Wi-Fi in unlicensed spectrum introduces a system-level resource coordination problem, where heterogeneous channel access mechanisms lead to a significant imbalance in spect... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU Clusters | Dongxin Guo, Jikun Wu, Siu Ming Yiu | 2026-05-01 | 下载 | AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end l... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| SoCal: A Language for Memory-Layout Factorization of Recursive Datatypes | Vidush Singhal, Mikah Kainen, Artem Pelenitsyn, Michael H. Borkowski, Mike Vollmer, Milind Kulkarni | 2026-05-01 | 下载 | Array-of-structures (AoS) to structure-of-arrays (SoA) is a classic compiler transformation that improves memory locality and enables data-parallel execution. |
| Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge | M. Grailoo, J. Núñez-Yáñez | 2026-05-01 | 下载 | Scaling laws for Large Language Models (LLMs) establish that model quality improves with computational scale, yet edge deployment imposes strict constraints on compute, memory, and power. |
| Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference | Abdurrahman Javat, Allan Kazakov | 2026-05-01 | 下载 | The operational landscape of local Large Language Model (LLM) inference has shifted from lightweight models to datacenter-class weights exceeding 70B parameters, creating profound systems challenges f... |
| How to Do Statistical Evaluations in ECE/CS Papers: A Practical Playbook for Defensible Results | Bhaskar Krishnamachari | 2026-05-01 | 下载 | Strong experimental papers in electrical and computer engineering and computer science (ECE/CS), especially in systems, networking, and applied machine learning, rest on more than a single impressive ... |
| Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI Inference | Yuxuan Gao, Megan Wang, Yi Ling Yu | 2026-05-01 | 下载 | Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) ... |