Skip to content

2026-07-14 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
A Reality Check on Quantum Optimisation: Evidence from an Industrial Case StudyHila Safi, Karen Wintersperger, Oliver von Sicard, Christoph Niedermeier, Wolfgang Mauerer2026-07-14下载Quantum Processing Units promise speed-ups for selected computational problems, including combinatorial optimisation, but their industrial utility remains an open challenge.
Microflow: Microarchitectural Causal Observability for Deep Cross-Layer Analysis and OptimizationSaber Ganjisaffar, Chengyu Song, Nael Abu-Ghazaleh2026-07-14下载Existing architectural simulators expose aggregate metrics or raw traces, but fail to reveal complex interactions among microarchitectural events and their relationship to program execution.
A 32-channel event-based bio-signal analog front-end with adaptive delta and pulse frequency encodingNarayanan Shyam, Saptarshi Ghosh, Giacomo Indiveri2026-07-14下载Low-power event-based Analog Front-Ends (AFEs) are essential for building efficient, end-to-end neuromorphic signal processing systems. In this paper, we present an event-based AFE Application-Specifi...
HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM InferenceGregory Hyegang Jun, Wesley Pang, Eddie Richter, Mehdi Saeedi, Aporva Amarnath, Pallavi Ferrao, Deming Chen2026-07-14下载Modern edge system-on-chips (SoCs) combine CPUs, integrated GPUs (iGPUs), and neural processing units (NPUs), yet existing LLM runtimes typically make coarse device-level decisions or optimize operato...
CLIP-3D: Closed-Loop Evaluation of Performance and Physical Constraints for 3D ICsShuo Ren, Libo Shen, Yaohui Han, Leilei Jin, Chenghan Wang, Zhen Zhuang, Rongliang Fu, Bei Yu, Tsung-Yi Ho2026-07-14下载3D integration packs more power into a smaller footprint, so a candidate design's actual throughput depends on its layout: which macro sits on which tier, where the hot spot lands, and how cache geome...
No Attention, No Problem: DPU-Aware Attention Approximation in Modern YOLO on FPGASuraj Karki, Qazi Arbab Ahmed, Thorsten Jungeblut2026-07-14下载Edge-based Artificial Intelligence (AI) acceleration has recently improved progress in real-time object detection. Object detection on edge devices requires a balance between accuracy, speed, and powe...
Realizable N:M Sparse Transformer Inference via Search-Kernel Co-DesignYiming Liu, Wenqi Lou, Zhiguang Wang, Zhiwei Ke, Fengrui Zuo, Chao Wang, Xuehai Zhou2026-07-14下载Vision Transformers (ViTs) achieve strong accuracy but incur high inference latency. Semi-structured N:M sparsity can reduce arithmetic cost, yet its theoretical savings often fail to translate into p...
ArchSim: Computer Architecture Simulation as a ServiceSabila Al Jannat, Wenhan Lyu, Le Khanh Trinh Mai, Huizhi Zhao, Zhuoyan Zheng, Katherine E. Isaacs, Yifan Sun2026-07-14下载Conducting a complete computer architecture simulation study is challenging because configuration, execution, and analysis are often encoded implicitly in scripts or directory conventions rather than ...
Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the LimitXiaomi MiMo Team, Anqi Liu, Aoxin Ma, Bo Chen, Bo Yang, Chen Wang, Chen Zhang, Chengda Tang, Chengwei Wang, Chiheng Lou, Depeng Yan, Fuli Luo, Gang Wang, Hailin Zhang, Jiale Sun, Kang Zhou, Rui Huang, Shaohui Liu, Shen Huang, Shijie Cao, Shuaishuai Fan, Tianling Zhou, Xiangwei Deng, Xueyang Xie, Xuli Wang, Yingchun Lai, Yu Yang, Yuan Zhang, Zhen Tang, Zhonghua Deng, Zihan Jiang2026-07-14下载We present a full-pipeline inference optimization for the MiMo-V2.5 model family, which combines Hybrid Sliding Window Attention (Hybrid SWA), sparse Mixture-of-Experts (MoE), and multimodal encoders.
Emulated Integrity Replica: Enabling Self-Healing on FPGA SoCs via Hierarchical TwinsArsalan Ali Malik, Ali Suvizi, Guru Venkataramani, Aydin Aysu2026-07-14下载Convolutional neural networks (CNNs) are increasingly being deployed on system-on-chip (SoC) platforms, where hardware-accelerated inference enables low-latency edge computing.
ORRAM: An OpenROAD-Integrated RAM Generator Using Standard CellsBrayden Louie, Thinh P. Nguyen, Matt Liberty, Austin Rovinski2026-07-14下载Memory inference remains a significant challenge in turnkey ASIC design flows. Inferring flip-flops from RTL can create thousands of densely interconnected instances which dramatically slow down desig...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Agora: Collective and Permissionless Internet-Scale Pretraining of Large Language ModelsGil Avraham, Violetta Shevchenko, Hadi Mohaghegh Dolatabadi, Karol Pajak, James Snewin, Harry Xi, Rodney O'Donnell, Thalaiyasingam Ajanthan, Sameera Ramasinghe, Chamin Hewa Koneputugodage, Shamane Siriwardhana, Alexander Long2026-07-14下载Training large language models at the multi-billion to trillion parameter scale is confined to datacenters, where data-parallel (DP) and model-parallel (MP) techniques presume homogeneous accelerators...
Hybrid multi-objective evolutionary algorithms for service placement in the computing continuum: a comparative study with genetic traceabilitySergi Vivo, Carlos Guerrero, Isaac Lera2026-07-14下载This paper addresses multi-objective service placement in computing continuum environments through a collaborative hybrid island-model MOEA. The key innovation is not the design of a new general hybri...
Privacy Attacks on Stable MarriageStephan A. Fahrenkrog-Petersen, Aleksander Figiel, Darya Melnyk, Tijana Milentijević, Stefan Schmid2026-07-14下载The stable marriage problem appears in many privacy-sensitive domains, for example in the National Resident Matching Program in the US. In such applications, preserving the privacy of users' preferenc...
Mixed-Timescale Differential Coding for Downlink Model Broadcast in Wireless Federated LearningChung-Hsuan Hu, Zheng Chen, Erik G. Larsson2026-07-14下载In standard federated learning systems, the parameter server broadcasts the global model to the participating devices in every iteration. Motivated by the temporal correlation between consecutive glob...
Proceedings of HLPP 2026: 19th International Symposium on High-Level Parallel Programming and ApplicationsChong Li, Corinne Ancourt, Gaétan Hains2026-07-14下载This volume contains the ten peer-reviewed papers presented at HLPP 2026, the 19th International Symposium on High-Level Parallel Programming and Applications, held on 9-10 July 2026 at the Institut H...
HeteroMosaic: Exposing and Exploiting Heterogeneous Execution Opportunities for Energy-Efficient Edge LLM InferenceGregory Hyegang Jun, Wesley Pang, Eddie Richter, Mehdi Saeedi, Aporva Amarnath, Pallavi Ferrao, Deming Chen2026-07-14下载Modern edge system-on-chips (SoCs) combine CPUs, integrated GPUs (iGPUs), and neural processing units (NPUs), yet existing LLM runtimes typically make coarse device-level decisions or optimize operato...
Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-ExpertsJincheng Xie, Runheng Liu, Heyan Huang, Yawen Ling, Hanbin Dai, Yu Zheng, Wen Hu2026-07-14下载Sparse Mixture-of-Experts (MoE) models have become an important approach for scaling Large Language Models (LLMs), but their inference efficiency depends strongly on expert activation patterns.
Scaling Synthetic-Image Pre-Training for Federated Fine-Tuning of Large Vision ModelsQianpiao Ma, Xiaozhu Song, Junlong Zhou, Yue Zeng, Jianchun Liu, Huaqing Tu2026-07-14下载Federated fine-tuning (FedFT) enables adapting pre-trained large vision models (LVMs) on distributed, privacy-sensitive devices, while its practical deployment is hindered by three critical challenges...
Parallel Sampling from the Ising pp-Spin ModelNima Anari, Aniket Das, Alireza Haqi2026-07-14下载We study the parallel complexity of sampling from the high-temperature Ising mixed pp-spin Gibbs measure, a canonical instance of a mean-field spin glass on the hypercube {±1}n\{\pm 1\}^n.
Parallelizing Legacy Mesh Generation Software: Lessons Learned from a Pseudo-Constrained Parallel Data Refinement Approach for Advancing Front Local ReconnectionKevin Garner, David Marcum, Nikos Chrisochoides2026-07-14下载This paper presents lessons learned from parallelizing the legacy software known as Advancing Front Local Reconnection (AFLR) as a black box. The parallel procedure utilizes (i) a data decomposition s...
Profiling and Scheduling Complex O-RAN Applications Across the 5G Edge and CloudYoonjae Hwang, Bhaskar Krishnamachari2026-07-14下载The O-RAN paradigm decomposes intelligent RAN control into pipelines of interdependent AI/ML functions, including traffic prediction, signal quality estimation, and slice scheduling, that must execute...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
A Measurement Plane for Quantum NetworkingAbderrahim Amlou, Amar Abane, Anouar Rahmouni, Mheni Merzouki, Abdella Battou, Ahmed Lbath, Ya-Shian Li-Baboud, Oliver Slattery, Thomas Gerrits2026-07-14下载Quantum networking testbeds lack a distinct plane for coordinating distributed measurements and collecting experimental data across heterogeneous devices.
Designing a GDPR-Compliant Security Architecture for Remote Elderly Care Systems: A Privacy-by-Design ApproachMd. Rahid Parvez, Mikael Soini2026-07-14下载IoMT-based remote elderly care systems generate continuous streams of sensitive health data, yet existing security architectures have not simultaneously addressed three interdependent challenges: GDPR...
High-Precision Hybrid FA-PSO Based Inversion of Building Material Parameters for Fundamental Wireless Performance EvaluationZhuowei Li, Yalei Zhu, Hanqing Zhang, Sui Li, Meng Chen, Tong Zhang, Zi-Yang Wu, Dan Yang, Songjiang Yang, Jiliang Zhang2026-07-14下载In this paper, we propose an inversion method based on the firefly particle swarm optimization (FA-PSO) algorithm to estimate the permittivity, conductivity, and thickness of building materials using ...
Q2NSViz: An Open-source Standalone Visualizer for Quantum Network SimulationsFrancesco Mazza, Marcello Caleffi, Angela Sara Cacciapuoti2026-07-14下载The unique and non-classical features of quantum networks make their simulation and intuitive understanding inherently difficult. In this work, we present Q2NSViz, an open-source Python-based visualiz...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Cross-Core Inference Offload as an Operating-System Service on Dual-Core MicrocontrollersDimitrios Kafetzis2026-07-14下载Dual-core MCUs are asymmetric: on NXP's MCXN947, the second Cortex-M33 has no FPU, DSP extension, TrustZone, or MPU. We treat the asymmetry as a design input in the Phase 3 dual-core architecture of S...
Inference Pipelines as Operating-System Objects: Priority Scheduling and Constant-Footprint Streaming for Microcontroller Neural InferenceDimitrios Kafetzis2026-07-14下载Microcontroller runtimes treat the inference pipeline -- pre-processing, accelerator invocation, post-processing -- as application code: every project re-implements stage sequencing, buffer sizing, an...
SynapticOS: An Inference-First Runtime Architecture for Neural Processing Units on Resource-Constrained MicrocontrollersDimitrios Kafetzis2026-07-14下载Microcontrollers with on-die neural processing units (NPUs) have become mainstream, but the system software hosting them has not: production combinations of Zephyr or FreeRTOS with TensorFlow Lite Mic...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Microflow: Microarchitectural Causal Observability for Deep Cross-Layer Analysis and OptimizationSaber Ganjisaffar, Chengyu Song, Nael Abu-Ghazaleh2026-07-14下载Existing architectural simulators expose aggregate metrics or raw traces, but fail to reveal complex interactions among microarchitectural events and their relationship to program execution.
EMO: Energy Efficiency Modeling and Optimization for AI WorkloadsJiyu Luo, Shaoyu Chen, Jingwei Sun, Shengcai Liu, Ke Tang, Guangzhong Sun2026-07-14下载The massive energy consumption of GPU-accelerated AI workloads challenges sustainable computing. We observe that execution asynchrony (e.g., CPU-GPU, concurrent streams, multi-GPU) creates slack, allo...

基于 VitePress 构建 · 使用本地搜索查找论文