Skip to content

2026-07-05 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
A Reconfigurable and Representation-Adaptive ISA-Based Architecture for Efficient DNN AccelerationVasilis Sakellariou, Vassilis Paliouras, Ioannis Kouretas, Hani Saleh, Thanos Stouraitis2026-07-05下载Domain-specific hardware accelerators provide significantly higher performance and energy efficiency for deep neural network (DNN) workloads than general-purpose processors, but often lack adaptabilit...
HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM InferenceHui Dong, Yanzhao Li, Jie Gao, Chunlu Li, Zhiyuan Zhang, Yupeng Sun, Zhenyuan Chen, Zhiqiang Zou2026-07-05下载We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on Ascend NPUs, while maintaining the online softmax...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
GORIO: GPU-Centered Remote I/O for Graph ANNS over NVMe-oFGen Zhang, Wenhao Gu2026-07-05下载Graph-based approximate nearest neighbor search (ANNS) is increasingly used in vector databases and retrieval-augmented generation services, but large vector indexes often exceed the memory capacity o...
Air-Plan: Query-Optimized Topology Selection for Over-the-Air Decentralized Federated LearningKaushal Attaluri, Rebeca P. Diaz-Redondo, Manuel Fernandez Veiga2026-07-05下载Over-the-air (OTA) aggregation exploits the superposition property of wireless multiple-access channels to aggregate model updates from multiple devices within a single transmission slot, significantl...
Sangam: Efficiently Serving Diffusion LLMs with the AR StackNitin Kedia, Saurabh Agarwal, Myungjin Lee, Aditya Akella2026-07-05下载Diffusion language models (dLLMs) generate text by iteratively denoising a masked response and can commit multiple output positions per model invocation.
CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM ServingJingfeng Wu, Yiyuan He, Minxian Xu, Xitong Gao, Chong Ma, Le Chen, Min Shen, Lin Qu, Kejiang Ye, CHengzhong Xu2026-07-05下载Online large language model (LLM) serving has become the backbone of modern AI applications, powering diverse downstream services through shared hardware clusters.
BrownoutMoE: Structure-Aware Expert Grouping for Efficient and Accurate LLM Web-based ServicesYi Ding, Minxian Xu, Zhengxin Fang, Kejiang Ye, Chengzhong Xu2026-07-05下载Mixture-of-Experts (MoE) large language models (LLMs) are increasingly deployed in Web-facing services, where inference must be both accurate and responsive under bursty demand.
FedSPM: Routing-Enabled Federated Learning under Dual Heterogeneity via Semiparametric MixtureZijian Wang, Pengfei Li, Guangyu Yang, Qiong Zhang2026-07-05下载Routing-prediction federated learning has emerged as a new paradigm that reframes inter-client heterogeneity as a resource for system-level intelligence: at inference time, the server routes each exte...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Wireless Gas Leak Detection and LocalizationFabien Chraim, Yusuf Bugra Erol, Kris Pister2026-07-05下载Thousands of industrial gas leaks occur every year, with many leading to injuries, deaths, equipment damage, and a disastrous environmental effect.
On the Physical Plausibility and Distribution Alignment for Sim-to-Real RF PositioningArarat Saribekyan, Armen Manukyan, Hrant Khachatrian, Theofanis P. Raptis2026-07-05下载Reliable radio frequency (RF) positioning from cellular measurements is limited by the high cost and limited coverage of real drive-test data, especially when models must work on streets not seen duri...
Agentic-V2X: Small Language Model Agents for Deadline-Aware V2X Scheduling in 5G/6G NetworksGerasimos Papanikolaou-Ntais, Alexandros Kaloxylos, Athanasios Kanavos2026-07-05下载Large Language Models (LLMs) are proposed as control interfaces for next-generation networks, but their latency, hallucinations, and lack of control guarantees make them unsuitable for near-real-time ...
Agentic IoT: Architectures, Applications, and Challenges Toward the Internet of AgentsRümeysa Hilal Sevinç, Bahaeddin Türkoğlu, İbrahim Kök2026-07-05下载The integration of AI into Internet of Things (AIoT) systems has gradually transformed them from passive data collection infrastructures into intelligent systems capable of anomaly detection, predicti...

cs.PF - Performance ​

标题作者发布日期PDF摘要
HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM InferenceHui Dong, Yanzhao Li, Jie Gao, Chunlu Li, Zhiyuan Zhang, Yupeng Sun, Zhenyuan Chen, Zhiqiang Zou2026-07-05下载We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM inference on Ascend NPUs, while maintaining the online softmax...

基于 VitePress 构建 · 使用本地搜索查找论文