Skip to content

2026-05-01 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Sim-FA: A GPGPU Simulator Framework for Fine-Grained FlashAttention Pipeline AnalysisZhongchun Zhou, Yuhang Gu, Chengtao Lai, Ya Wang, Wei Zhang2026-05-01下载To efficiently support Large Language Models (LLMs), modern GPGPU architectures have introduced new features and programming paradigms, such as warp specialization.
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI EdgeM. Grailoo, J. Núñez-Yáñez2026-05-01下载Scaling laws for Large Language Models (LLMs) establish that model quality improves with computational scale, yet edge deployment imposes strict constraints on compute, memory, and power.
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM InferenceAbdurrahman Javat, Allan Kazakov2026-05-01下载The operational landscape of local Large Language Model (LLM) inference has shifted from lightweight models to datacenter-class weights exceeding 70B parameters, creating profound systems challenges f...
VitaLLM: A Versatile and Tiny Accelerator for Mixed-Precision LLM Inference on Edge DevicesZi-Wei Lin, Tian-Sheuan Chang2026-05-01下载We present VitaLLM, a mixed precision accelerator that enables ternary weight large language models to run efficiently on edge devices. The design combines two compute cores, a multiplier free TINT co...
A PVT-Resilient Subthreshold SRAM-Based In-Memory Computing Accelerator with In-Situ Regulation for Energy-Efficient Spiking Neural NetworksShih-Hang Kao, Yang-Chan Hung, I-Wen Wang, Bing-Han Liu, Yu-Chia Chen, Tian-Sheuan Chang, Shyh-Jye Jou, Chien-Nan Liu, Hung-Ming Chen, Wei-Zen Chen2026-05-01下载This paper presents a PVT-resilient, subthreshold SRAM-based computing-in-memory (CIM) macro tailored for energy-efficient spiking neural networks (SNNs).

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
A Domain-Driven Design Simulator for Business Logic-Rich Microservice SystemsDaniel da Palma Pereira, António Rito Silva2026-05-01下载Developing business-logic-rich microservices requires navigating complex trade-offs between data consistency and distributed coordination. Although patterns like Sagas and Transactional Causal Consist...
ncsim: A Lightweight Simulator for Networked Edge Computing with Wireless Interference ModelingBhaskar Krishnamachari, Maya Gutierrez, Jared Coleman2026-05-01下载Evaluating DAG task schedulers for wireless edge computing requires jointly modeling compute placement and wireless interference, yet existing tools treat them in isolation.
FPTC: A Fast Parallel Transform-based Codec for Efficient Asymmetric Signal CompressionBen Mechels, Ryan Billmeyer, Alexander Chen, Shiyang Li, Caiwen Ding2026-05-01下载Modern high-performance computing and Internet-of-Things deployments increasingly generate large volumes of signal data that must be compressed efficiently on resource-constrained acquisition devices ...
SURGE: SuperBatch Unified Resource-efficient GPU Encoding for Heterogeneous Partitioned DataShashank Kapadia, Deep Narayan Mishra, Sujal Reddy Alugubelli, Ajay Kumar, Swapnil Yadav, Rishi Bhatia2026-05-01下载We present SURGE, a streaming GPU encoding system deployed in production to generate embeddings for over 800 million texts across 40,000 logical partitions.
Eliminating Hidden Serialization in Multi-Node Megakernel CommunicationByungsoo Oh, Rachee Singh2026-05-01下载Recent megakernel designs for Mixture-of-Experts (MoE) inference fuse expert computation with fine-grained, GPU-initiated communication into a single persistent GPU kernel, and outperform collective-b...
LLM-Emu: Native Runtime Emulation of LLM Inference via Profile-Driven SamplingWei Da, Evangelia Kalyvianaki2026-05-01下载Realistic evaluation of LLM serving systems requires online workloads, dynamic arrivals, queueing, and the serving engine's local scheduling for execution batching, but running such experiments on GPU...
AGoQ: Activation and Gradient Quantization for Memory-Efficient Distributed Training of LLMsWenxiang Lin, Juntao Huang, Luhan Zhang, Laili Li, Xiang Bao, Mengyang Zhang, Bing Wang, Shaohuai Shi2026-05-01下载Quantization is a key method for reducing the GPU memory requirement of training large language models (LLMs). Yet, current approaches are ineffective for 4-bit activations and 8-bit gradients, which ...
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI EdgeM. Grailoo, J. Núñez-Yáñez2026-05-01下载Scaling laws for Large Language Models (LLMs) establish that model quality improves with computational scale, yet edge deployment imposes strict constraints on compute, memory, and power.
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU ClustersDongxin Guo, Jikun Wu, Siu Ming Yiu2026-05-01下载AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end l...
Space Network of Experts: Architecture and Expert PlacementZhanwei Wang, Huiling Yang, Min Sheng, Khaled B. Letaief, Kaibin Huang2026-05-01下载Leveraging continuous solar energy harvesting at high efficiency, space data centers are envisioned as a promising platform for executing energy-intensive large language models (LLMs).
Adaptation of AI-accelerated CFD Simulations to the IPU platformP. Rosciszewski, A. Krzywaniak, S. Iserte, K. Rojek, P. Gepner2026-05-01下载Intelligence Processing Units (IPU) have proven useful for many AI applications. In this paper, we evaluate them within the emerging field of \emph{AI for simulation}, where traditional numerical simu...
Hierarchical Federated Learning for Networked AI: From Communication Saving to Architecture-Aware DesignSeyed Mohammad Azimi-Abarghouyi, Mehdi Bennis, Leandros Tassiulas2026-05-01下载Federated learning (FL) is fundamentally a distributed optimization problem executed by communicating agents with local data, local computation, and partial system visibility.
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI InferenceYuxuan Gao, Megan Wang, Yi Ling Yu2026-05-01下载Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
MORPH: Multi-Environment Orchestrated Reinforcement Learning for PRB Handling in O-RANAlireza Ebrahimi Dorcheh, Tolunay Seyfi, Ryan Barker, Fatemeh Afghah2026-05-01下载Reinforcement-learning (RL) solutions for dynamic spectrum access and radio resource management in Open Radio Access Networks (O-RAN) depend critically on the fidelity of the throughput signal used fo...
AIIM: Adaptive Inter-cell Interference Mitigation for Heterogeneous Multi-vendor 5G O-RAN NetworksSamuel Reinders, Alireza Ebrahimi Dorcheh, Ryan Barker, Tolunay Seyfi, Fatemeh Afghah2026-05-01下载Inter-cell interference is a persistent issue in dense 5G deployments, especially in heterogeneous Open Radio Access Network (O-RAN) environments where coordination between base stations is limited.
ncsim: A Lightweight Simulator for Networked Edge Computing with Wireless Interference ModelingBhaskar Krishnamachari, Maya Gutierrez, Jared Coleman2026-05-01下载Evaluating DAG task schedulers for wireless edge computing requires jointly modeling compute placement and wireless interference, yet existing tools treat them in isolation.
AdvNet: Revealing Performance Issues in Network Protocols by Generating Adversarial EnvironmentsShehab Sarar Ahmed, William Sentosa, Yinjie Zhang, Yoav Lebendiker, Michael Shnaiderman, Tomer Gilad, Nathan H. Jay, Brighten Godfrey, Michael Schapira2026-05-01下载Infrastructure protocols like Congestion Control (CC) seek to provide reliable performance across a wide range of Internet environments. Currently, protocol designers assess performance through hand-d...
EASE: Federated Multimodal Unlearning via Entanglement-Aware Anchor ClosureZihao Ding, Beining Wu, Jun Huang2026-05-01下载Federated Multimodal Learning (FML) trains multimodal models across decentralized clients while keeping their image-text pairs private. However, joint embedding training entangles forgotten knowledge ...
Inductive Latent Context Persistence: Closing the Post-Handover Cold Start in 6G Radio Access NetworksAnubhab Banerjee, Daniyal Amir Awan2026-05-01下载In modern radio access networks (RANs), rule-based handover (HO) decisions (e.g., A3/A5) depend on user equipment (UE) measurements only, so UEs at the same location can receive inconsistent HO outcom...
Beyond Per-Request QoS: Coordinating Industrial Workflows with B5G/6G Network CapabilitiesQize Guo, Bjoern Riemer, Tarik Taleb, Yan Chen, Hao Yu, Hemant Zope2026-05-01下载Beyond-5G (B5G) and 6G networks are expected to enable more complex industrial services, which often operate according to multi-phase workflows with phase-specific communication requirements.
Space Network of Experts: Architecture and Expert PlacementZhanwei Wang, Huiling Yang, Min Sheng, Khaled B. Letaief, Kaibin Huang2026-05-01下载Leveraging continuous solar energy harvesting at high efficiency, space data centers are envisioned as a promising platform for executing energy-intensive large language models (LLMs).
A Policy-Driven DRL Framework for System-Level Tradeoff Control in NR-U/Wi-Fi CoexistencePo-Heng Chou, Yi-Fang Yu, Shou-Yu Chen, Chiapin Wang2026-05-01下载The coexistence of NR-U and Wi-Fi in unlicensed spectrum introduces a system-level resource coordination problem, where heterogeneous channel access mechanisms lead to a significant imbalance in spect...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
SAGA: Workflow-Atomic Scheduling for AI Agent Inference on GPU ClustersDongxin Guo, Jikun Wu, Siu Ming Yiu2026-05-01下载AI agents execute tens to hundreds of chained LLM calls per task, yet GPU schedulers treat each call as independent, discarding gigabytes of intermediate state between steps and inflating end-to-end l...

cs.PF - Performance ​

标题作者发布日期PDF摘要
SoCal: A Language for Memory-Layout Factorization of Recursive DatatypesVidush Singhal, Mikah Kainen, Artem Pelenitsyn, Michael H. Borkowski, Mike Vollmer, Milind Kulkarni2026-05-01下载Array-of-structures (AoS) to structure-of-arrays (SoA) is a classic compiler transformation that improves memory locality and enables data-parallel execution.
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI EdgeM. Grailoo, J. Núñez-Yáñez2026-05-01下载Scaling laws for Large Language Models (LLMs) establish that model quality improves with computational scale, yet edge deployment imposes strict constraints on compute, memory, and power.
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM InferenceAbdurrahman Javat, Allan Kazakov2026-05-01下载The operational landscape of local Large Language Model (LLM) inference has shifted from lightweight models to datacenter-class weights exceeding 70B parameters, creating profound systems challenges f...
How to Do Statistical Evaluations in ECE/CS Papers: A Practical Playbook for Defensible ResultsBhaskar Krishnamachari2026-05-01下载Strong experimental papers in electrical and computer engineering and computer science (ECE/CS), especially in systems, networking, and applied machine learning, rest on more than a single impressive ...
Token Arena: A Continuous Benchmark Unifying Energy and Cognition in AI InferenceYuxuan Gao, Megan Wang, Yi Ling Yu2026-05-01下载Public inference benchmarks compare AI systems at the model and provider level, but the unit at which deployment decisions are actually made is the endpoint: the (provider, model, stock-keeping-unit) ...

基于 VitePress 构建 · 使用本地搜索查找论文