Skip to content

2026-06-29 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and PerformanceYang Zou, Zijian Ding, Yizhou Sun, Jason Cong2026-06-29下载High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code remains challenging due to restrictive language support and the...
SpikON: A Dual-Parallel and Efficient Accelerator for Online Spiking Neural Networks LearningPeilin Chen, Xiaoxuan Yang2026-06-29下载Spiking neural networks (SNNs) have emerged as a promising paradigm for energy-efficient brain-inspired computing. However, existing online unsupervised SNN learning suffers from low training accuracy...
CryoZip: An Efficient Cryogenic Compressor for Quantum Error Correction SyndromesGuanchen Tao, Alexander Knapen, Jacob Mack, Gokul Subramanian Ravi, Qirui Zhang, Mehdi Saligane, Dennis Sylvester2026-06-29下载Scaling fault tolerant quantum computing is increasingly constrained by the limited bandwidth and power budget across the 4 K to room temperature (RT) interface.
COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile DevicesYilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao, Jian Liu, Haibing Guan, Li Jiang2026-06-29下载The development of on-device large language models (LLMs) is driven by the need for privacy and fast response times. Energy-intensive data transfer on mobile devices makes Processing-in-Memory (PIM) a...
Model Predictive Current Control with Harmonic Correction for Single-Phase AC-DC EV ChargingChanghong Li, Bharathkumar Hegde, Biswajit Basu, Shreejith Shanker2026-06-29下载The increasing integration of Electric Vehicles (EVs) has imposed a growing harmonic challenge on the power grid. For AC/DC Power Factor Correction (PFC) in single-phase On-Board Chargers (OBCs), Mode...
RQP: Resource-Oriented Quantiser Pruning for Neural Networks on FPGAsChanghong Li, Biswajit Basu, Shreejith Shanker2026-06-29下载High granularity quantisation (HGQ) exploits weight-level quantisation and pruning to design resource-efficient neural network accelerators, achieving an attractive trade-off between accuracy and hard...
Mega: A 22 nm Convolutional Spiking Neural Network Accelerator Achieving 0.375 pJ/SOP for Efficient Edge VisionRick Luiken, Manil Dev Gomony, Sander Stuijk2026-06-29下载Convolutional Spiking Neural Networks (SNN) offer the potential for highly energy-efficient vision processing by exploiting sparse, event-driven computation.
HBM Is Not All You Need: Efficient Disaggregated LLM Serving across Memory-heterogeneous AcceleratorsZhixiang Wei, Yun Wang, James Yen, Mingyuan Xia, Zhengwei Qi2026-06-29下载LLM inference comprises a compute-bound prefill phase and a memory-bound decode phase, and recent systems disaggregate them onto separate hardware.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Towards Transparent Checkpointing with AI-driven Code GenerationHai Duc Nguyen, Tekin Bicer, Kyle Chard, Ian Foster, Bogdan Nicolae2026-06-29下载Adding reliable checkpoint/restart support to an MPI scientific application is a time-consuming expert effort that requires deep knowledge of both the application and resilience.
Budget-Adaptive Routing: Skipping the Weak When the Strong Answers AnywayWei Geng, Nitinder Mohan, Jörg Ott2026-06-29下载Edge-cloud inference collaborations are often designed with a routing estimator that decides whether to offload each frame from weak models at the edge to stronger models in the cloud.
StreamGuard: Low-Overhead Resilience for Real-time HPC Data StreamsHai Duc Nguyen, Bogdan Nicolae, Tekin Bicer, Amal Gueroudji, Matthieu Dorier, Kyle Chard, Ian Foster2026-06-29下载Real-time scientific workflows operate on continuous data streams and must produce timely, high-quality results despite executing on complex, failure-prone infrastructure.
Protecting Futures against Silent Data Corruption -- Efficient Task Replication for Dynamic Data DependenciesRüdiger Nather, Claudia Fohry, Mia Reitz2026-06-29下载As the size of computational problems grows, so does the likelihood of Silent Data Corruptions (SDCs). A common defense is replication, where the computation is repeated and correct results are determ...
Data Replication Meets Function Scheduling in the Edge-Cloud ContinuumMatteo Cenzato, Dario d'Abate, Arianna Dragoni, Matteo Briscini, Alessandro Margara2026-06-29下载Serverless computing is an appealing model for the edge-cloud continuum, but its stateless assumption breaks down once functions need persistent data: fetching state from a distant cloud store erases ...
SubEdge: A Subscriber-Centric Edge Computing Subsystem in 6G Networks for AIAbdirazak Ali Asir Rage, Riccardo Pozza, Rahim Tafazolli2026-06-29下载Beyond traditional connectivity, 6G is envisioned to transform mobile networks into a distributed fabric that provides native integrated communication, computing, and intelligence services.
COSM: A Cooperative Scheduling Framework for Concurrent PIM and CPU Execution on Mobile DevicesYilong Zhao, Fangxin Liu, Onur Mutlu, Mingyu Gao, Jian Liu, Haibing Guan, Li Jiang2026-06-29下载The development of on-device large language models (LLMs) is driven by the need for privacy and fast response times. Energy-intensive data transfer on mobile devices makes Processing-in-Memory (PIM) a...
Spandana: Reconciling Strict SLOs with Low Cost under Fine-Grained Load FluctuationsDilina Dehigama, Shyam Jesalpura, Zeyu Xu, Marton Nemeth, Shengda Zhu, Marios Kogias, Boris Grot2026-06-29下载Cloud-based online services face significant sub-second load fluctuations while needing to meet strict Service Level Objectives (SLOs). Cluster operators often over-provision resources to protect SLOs...
GPU Parallelization Strategies for Forward and Backward Propagation in Shallow Neural Networks: A CUDA-Based Comparative StudyRania Zitouni, Nadine Bousdjira, Sarah Hasnaoui, Amel Sadoun, Fatma Salhi2026-06-29下载We present a comparative study of CUDA optimization strategies applied to forward and backward propagation in a shallow neural network. Three stacked optimizations are evaluated: (1) tiled shared memo...
HSAP: A Hierarchical Sequence-aware Parallelism for Hybrid-Context Generative ModelsSongxin Zhang, Zejian Xie, Zhuoyang Song, Cong lin, Junyu Lu, Jiaxing Zhang, Bingyi Jing2026-06-29下载In this paper, we aim to combine the advantages of existing sequence parallelism paradigms and overcomes their drawbacks, the most serious of which is the incapability to correctly compute causal atte...
Analyzing Linearizability in Relativistic Distributed SystemsKahbod Aeini, Wojciech Golab2026-06-29下载Einstein's theory of relativity correctly predicted that time is relative, and subject to both kinematic and gravitational dilation. Therefore, executions of distributed systems cannot always be model...
Energy-Aware Scheduling for Serverless LLM Serving on Shared GPUsTianyu Wang, Gourav Rattihalli, Aditya Dhakal, Longfei Shangguan, Dejan Milojicic2026-06-29下载As LLM inference becomes a major cloud workload, its growing energy footprint makes cluster-wide energy optimization increasingly important. Serverless LLM serving helps platforms absorb traffic volat...
FBench: A Flexible Benchmark for CFG-Based What-If Exploration of HPC I/O PatternsZhaobin Zhu, Chen Wang, Kathryn Mohror, Sarah Neuwirth2026-06-29下载The I/O performance of large-scale HPC applications depends on a complex interplay of access patterns, middleware optimizations, and file system configurations.
HBM Is Not All You Need: Efficient Disaggregated LLM Serving across Memory-heterogeneous AcceleratorsZhixiang Wei, Yun Wang, James Yen, Mingyuan Xia, Zhengwei Qi2026-06-29下载LLM inference comprises a compute-bound prefill phase and a memory-bound decode phase, and recent systems disaggregate them onto separate hardware.
Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE InferenceHui Zang, Pengfei Xia, Hong Liu, Jiajia Chu, Tuo Hao, Minghao Chen, Rui Zhang, Ziyang Zhang2026-06-29下载Mixture-of-Experts (MoE) architectures enable language models to achieve unprecedented scale via sparse activation. However, their inference performance is often limited by data movement bottlenecks.
Rethinking Collaborative Trust for Verifiably Decentralized Blockchain SystemsYunqi Zhang, Shaileshh Bojja Venkatakrishnan2026-06-29下载Despite the promise of decentralization, measurement studies have identified a conspicuous lack of decentralization in blockchains. Centralization has been observed in almost all layers of the blockch...
SMART-MIG: A Learning Framework for Scalable and Energy-Efficient GPU SchedulingWenqing Yu, Neel Karia, Tanvi Hisaria, Clifford Stein, Olivier Tardieu, Asser Tantawi2026-06-29下载The emergence of Multi-Instance GPU (MIG) technology enables us to run smaller machine learning models on partitions of a GPU rather than the entire device, thus improving utilization and reducing ene...
Demystifying the Design Space and Best Practices for Heterogeneous LLM Inference and ServingZhixin Wang, Zhengbo Wang, Fangcheng Fu, Yinhui Lu, Jinlong Hou, Yijie Chen, Xiaowei Shen, He Liu, Xiangbin Li, Jun Chen, Ruya Gu, Dian Wang, Zhou Tan, Yuan Cheng, Hongzhou Zhang, Xiangjun Huang, Ping Zhang, Xiaohe Hu2026-06-29下载Heterogeneous prefill-decode (PD) inference is now in production: prefill on cost-efficient or supply-available accelerators, decode on bandwidth-strong ones, and KV state crossing mixed interconnects...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Budget-Adaptive Routing: Skipping the Weak When the Strong Answers AnywayWei Geng, Nitinder Mohan, Jörg Ott2026-06-29下载Edge-cloud inference collaborations are often designed with a routing estimator that decides whether to offload each frame from weak models at the edge to stronger models in the cloud.
A Practical Implementation of Day-3 Cooperative Intersection with Automated Connected Mini-CarsLorenzo Farina, Vittorio Todisco, Federico Gavioli, Salvatore Iandolo, Francesco Moretti, Giuseppe Perrone, Matteo Piccoli, Francesco Raviglione, Marco Rapelli, Antonio Solida, Claudio Casetti, Paolo Burgio, Carlo Augusto Grazia, Alessandro Bazzi2026-06-29下载Cooperative driving enabled by connected and automated vehicles is expected to improve traffic efficiency and safety, particularly at intersections where traditional control mechanisms such as traffic...
CALO: Constraint-Aware Learning Optimization for Joint Resource Allocation in Double-Active RIS-Assisted Wireless NetworksAlaa S. Arabiyat, Mohammad J. Abdel-Rahman2026-06-29下载Double-active reconfigurable intelligent surface (RIS)-assisted wireless systems can improve coverage and achievable rate in blockage-dominated environments.
When and Which Sensor to Observe? Timely Tracking of a Joint Markov SourceIsmail Cosandal, Sennur Ulukus, Nail Akar2026-06-29下载We investigate the problem of remote estimation (at a monitor) of a discrete-time joint Markov process with individual components which can be observed with dedicated sensors.
Wireless Backdoor Attack and Defense for Semantic Communications over Multiple Access ChannelYalin E. Sagduyu, Tugba Erpek, Aylin Yener, Sennur Ulukus2026-06-29下载Semantic communication (SemCom) aims to preserve semantic meaning and task-oriented information beyond conventional message recovery over wireless channels.
SubEdge: A Subscriber-Centric Edge Computing Subsystem in 6G Networks for AIAbdirazak Ali Asir Rage, Riccardo Pozza, Rahim Tafazolli2026-06-29下载Beyond traditional connectivity, 6G is envisioned to transform mobile networks into a distributed fabric that provides native integrated communication, computing, and intelligence services.
COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated TopologiesChen Frydman, Aviram Zilberman, Rubin Krief, Abed Showgan, Andres Murillo, Sekiya Motoyoshi, Asaf Shabtai, Yuval Elovici, Rami Puzis2026-06-29下载Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to that adversary, validates it without breaking production, an...
LLMs and Optical Networks: A Symbiotic RelationshipMëmëdhe Ibrahimi, Qiaolun Zhang, Giovanni S. Sticca, Jiaheng Xiong, Francesco Musumeci, Massimo Tornatore2026-06-29下载This paper explores the emerging symbiosis between LLMs and optical networks. Massive LLMs require geo-distributed training, which demands advanced optical transport capabilities that require new key ...
Selective Deployment of Bidirectional Hollow-Core Fibers in Hybrid SMF/HCF Optical NetworksMëmëdhe Ibrahimi, Giovanni S. Sticca, Angelo Ferrara, Massimo Tornatore2026-06-29下载We investigate selectively deploying bidirectional transmission in hybrid Hollow-Core Fiber (HCF) networks. Upgrading 50% of links to bidirectional HCF yields at least a 40% throughput increase compar...
Scalable Intention Sharing for ETSI VAMsFelipe E. Valle, Oscar Amador, Johan Thunberg, Elena Haller, Alexey Vinel2026-06-29下载Efficient maneuver coordination in dense V2X environments requires accurate short-term prediction while maintaining low communication and computational overhead.
LEOSTP: A Spatio-Temporal Traffic Prediction Framework for LEO Satellite NetworksShaoyou Ao, Yong Niu, Zhu Han, Cheng Li, Bo Ai2026-06-29下载With the evolution of next-generation mobile communication networks and the commercial boom of Low Earth Orbit (LEO) satellites, globally covered satellite networks are gradually becoming a crucial in...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI AgentsYogeswar Reddy Thota2026-06-29下载Current operating systems expose interfaces optimized for human users but not for AI agents. Humans benefit from pixels, icons, windows, visual grouping, mouse movement, and keyboard shortcuts; AI age...

cs.PF - Performance ​

标题作者发布日期PDF摘要
The Fourth-Root Complexity of Data MovementChen Ding2026-06-29下载Time complexity typically assumes O(1)O(1) cost per data access. This paper presents an analysis based on an abstract memory hierarchy. For a common class of applications, it shows that the data-access ...
TraceLab: Characterizing Coding Agent Workloads for LLM ServingKan Zhu, Mathew Jacob, Chenxi Ma, Yi Pan, Stephanie Wang, Arvind Krishnamurthy, Baris Kasikci2026-06-29下载Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this challenge requires understanding real workload patterns, yet ...
FBench: A Flexible Benchmark for CFG-Based What-If Exploration of HPC I/O PatternsZhaobin Zhu, Chen Wang, Kathryn Mohror, Sarah Neuwirth2026-06-29下载The I/O performance of large-scale HPC applications depends on a complex interplay of access patterns, middleware optimizations, and file system configurations.

基于 VitePress 构建 · 使用本地搜索查找论文