Skip to content

2026-06-10 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable DevicesFarough Shayeste Roodi, Parham Zilouchian Moghaddam, Mahdi Mohammadi-nasab, Mehdi Modarressi, Mostafa Ersali Salehi Nasab, Masoud Daneshtalab2026-06-10下载Wearable healthcare devices are the fastest-growing Internet of Things (IoT) sector. Many automated healthcare services rely on two crucial biological signals, namely ECG and EEG, which reflect the ac...
Eidola: Modeling Multi-GPU Network Communication Traffic in Distributed AI WorkloadsRanganath R. Selagamsetty, Matthew Poremba, Bradford M. Beckmann, Joshua San Miguel, Mikko H. Lipasti2026-06-10下载As distributed AI workloads grow in scale, multi-GPU systems have become essential for training large models. Although techniques like kernel fusion and overlapping communication with computation help...
Partitioned Tags, Shared Data: Reconciling Strict Cache Isolation with Write-Shared CoherenceKartik Ramkrishnan, Stephen McCamant, Antonia Zhai, Pen Chung Yew2026-06-10下载Cache partitioning is among the strongest structural defenses against eviction-based cache side channels, yet a decade-old design issue has blocked its widespread deployment in secure shared-OS settin...
BenDi: An Energy-Efficient Quasi-Stochastic Systolic Architecture for Edge BioelectronicsBochen Ye, Yihan Pan, Shady Agwa, Themis Prodromakis2026-06-10下载Continuous long-term monitoring and diagnosis of biomedical signals, such as electrocardiograms (ECGs), can help mitigate an increasing threat to public health.
Making Locality-aware GEMM Compatible with Page-Granularity Placement on Chiplet GPUsEuijun Chung, Jae Hyung Ju, Hyesoon Kim2026-06-10下载Multi-chiplet GPUs scale compute throughput and high-bandwidth memory (HBM) capacity, but their non-uniform memory system makes locality between chiplets and their data critical to the GPU's performan...
A Fast Locality Simulator for GEMM Design-Space Exploration on Multi-Chiplet GPUsEuijun Chung, Hyesoon Kim2026-06-10下载Multi-chiplet GPUs split memory into local and remote HBM regions across a silicon interposer, and reducing the remote HBM traffic is crucial for the performance and energy efficiency of multi-chiplet...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
On the Limits of Performance Portability in Directive-Based GPU ProgrammingAlessandro Romeo, Nitin Shukla, Stefano Truzzi, Alessio Suriano, Andrea Mignone2026-06-10下载The transition of scientific applications to GPU-accelerated exascale systems is constrained by trade-offs between performance, portability, and productivity.
M*: A Modular, Extensible, Serving System for Multimodal ModelsAtindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang2026-06-10下载We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and...
A Communication Complexity Lower Bound for Nonuniformly Convex Consensus OptimizationDemyan Yarmoshik, Maxim Klimenko2026-06-10下载We study the communication complexity of convex decentralized optimization over time-varying networks, where nn nodes hold private functions and must agree on the global minimizer using only synchron...
Eidola: Modeling Multi-GPU Network Communication Traffic in Distributed AI WorkloadsRanganath R. Selagamsetty, Matthew Poremba, Bradford M. Beckmann, Joshua San Miguel, Mikko H. Lipasti2026-06-10下载As distributed AI workloads grow in scale, multi-GPU systems have become essential for training large models. Although techniques like kernel fusion and overlapping communication with computation help...
ITME: Inference Tiered Memory Expansion with Disaggregated CXL-Hybrid MemoriesHakbeom Jang, Younghoon Min, Sunwoong Kim, Taeyoung Ahn, Hanyee Kim, Youngpyo Joo, Hoshik Kim, Jongryool Kim2026-06-10下载The rapid shift toward agentic and long-context workloads in Large Language Models (LLMs) is pushing the industry beyond the capacity of individual servers toward disaggregated shared storage to handl...
Fair Comparison of Scheduling Algorithms on Heterogeneous Edge Clusters: A Continuous Adaptive BenchmarkZihang Wang, Boris Sedlak, Juan Luis Herrera, Schahram Dustdar2026-06-10下载Modern Artificial Intelligence (AI) workloads deployed across the heterogeneous tiers of an edge--cloud continuum must satisfy multi-dimensional Service Level Objectives (SLOs) over latency, throughpu...
Efficient and Robust Online Learning to Rank in Decentralized SystemsMarcel Gregoriadis, Martijn de Vos, Sayan Biswas, Anne-Marie Kermarrec, Johan Pouwelse2026-06-10下载In Online Learning to Rank (OLTR), ranking models are trained directly from live user interactions, but existing systems rely on a trusted central server to collect and process these interactions.
The PM-EdgeMap: Towards Real-Time Process Mining on the Edge-Cloud ContinuumHendrik Reiter, Christian Imenkamp, Olaf Landsiedel, Andrea Maldonado, Patrick Rathje, Wilhelm Hasselbring2026-06-10下载Smart factories are evolving into Cyber-Physical Systems (CPS), demanding increased autonomy. This necessitates real-time decision making, facilitated by insights derived from sensor data.
Near-Optimal Distributed 2-Ruling Sets on Graphs with Low ArboricityMalte Baumecker, Rustam Latypov, Yannic Maus, Jara Uitto2026-06-10下载Given a graph G=(V,E)G=(V,E), a β-ruling set is a subset of nodes S⊆VS\subseteq V that is independent, and each node in VV is at distance at most β from some node in SS.
From Fork-Join to Asynchronous Tasks: Parallelizing Tiled Cholesky Decomposition with OpenMP and HPXAlexander Strack, Alexander Van Craen, Dirk Pflüger2026-06-10下载Fork-join parallelism, popularized by OpenMP, remains the dominant model for shared-memory parallel programming, but its implicit synchronization barriers can penalize algorithms with inhomogeneous wo...
Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-trainingYuming Zhou, Haoyang Li, Sheng Lin, Yanfeng Zhao, Tong Zhao, Xupeng Miao, Jie Jiang, Fangcheng Fu, Bin Cui2026-06-10下载Mixture-of-Experts (MoE) and reinforcement learning (RL) post-training now dominate large language model (LLM) development, yet expert load imbalance remains a critical challenge.
Optimizing Cloud Deployment: Blending of IaaS and FaaS for Microservice ArchitectureNikhil Kapoor, Sougata Mukherjea2026-06-10下载The rapid evolution of cloud computing has resulted in the adoption of hybrid deployments that blend Infrastructure-as-a-Service (IaaS) and Function-as-a-Service (FaaS) service models to optimize reso...
Consensus Time in 3-Majority and 2-Choices Is Determined by the Maximum Initial Opinion DensityNiccolò D Archivio2026-06-10下载We establish the correct parameter governing the convergence time of the 3-Majority and 2-Choices dynamics on the complete graph in the synchronous model.
MHOT: Height-Optimized Authenticated Data Structure for Blockchain State CommitmentSipeng Xie, Qianhong Wu, Minghang Li, Qiyuan Gao, Bo Qin, Qin Wang2026-06-10下载State root computation dominates (78%) blockchain block processing time. Ethereum's canonical authenticated data structure, i.e., Merkle Patricia Trie (MPT), suffers from severe tree-height growth and...
Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost EstimationChitral Patil2026-06-10下载Every public LLM cost calculator we surveyed treats GPU utilization as a fixed input -- entered by the user, baked in as a preset, or silently assumed at 100% -- never measured against the operator's ...
Sovereign Assurance Boundary: Certificate-Bound Admission for Agentic InfrastructureJun He, Deying Yu2026-06-10下载Agentic infrastructure introduces a critical control-plane authorization problem: non-deterministic reasoning systems can propose high-stakes mutations to production resources, yet existing security m...
Tensor-Network-Based Distributed Quantum Dynamics on Independent Quantum ComputersAnurag Dwivedi, Melissa C. Revelle, Daniel S. Lobser, Brian K. McFarland, Edward C. Tortorici, Christopher G. Yale, Susan M. Clark, Philip Richerme, Srinivasan S. Iyengar2026-06-10下载We present an approach based on tensor networks for distributed quantum computing simulation of chemical wavepacket dynamics in a continuous variable representation.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Free-Placement Optimization of Ground Station Locations for Low-Earth Orbit SatellitesGrace Ra Kim, Duncan Eddy, Vedant Srinivas, Mykel J. Kochenderfer2026-06-10下载Rapidly expanding low Earth orbit satellite constellations are placing increasing demands on terrestrial ground networks, motivating the development of more efficient ground station network designs.
Greenness-Driven Scheduling in Far Edge Kubernetes: A CODECO EvaluationKaikang Huang, Dalal Ali, Rute C. Sofia2026-06-10下载Energy consumption is an increasing concern in IoT-Edge-Cloud infrastructures, where containerized application orchestration must balance performance with sustainability.
Exploratory Analysis of Wi-Fi 6 Dynamic Resource Unit Sharing in Small-Scale Network ScenariosSai Mada, Anna Baron, Luigi Martino, Rute C. Sofia2026-06-10下载This paper investigates dynamic Resource Unit (RU) allocation strategies for Wi-Fi~6 (IEEE 802.11ax) networks integrated with Time-Sensitive Networking (TSN), targeting the limitations of static RU sc...
LLM-Enabled NWDAF: A Step Toward AI-Native 6G Network IntelligenceHenok Daniel, Omar Alhussein, Cheng Li, Jie Liang, Ernesto Damiani2026-06-10下载The Network Data Analytics Function (NWDAF) is central to enabling zero-touch network management in fifth-generation (5G) networks by supporting real-time analytics and closed-loop automation.
SwarmSense-DNN: A Trustworthy and Decentralized Neural Framework for Proactive Anomaly Defense in Consumer IoTJing Yang, Vijay Govindarajan, Saad Arif, Xu Xu, Mohamed Kallel, Zaffar Ahmed Shaikh, Zhe Liu, Chunhong Yuan, Lip Yee Por2026-06-10下载The rapid growth of consumer IoT devices has introduced unprecedented challenges in trustworthy anomaly detection against AI-enabled cyber threats, requiring real-time, privacy-preserving, and scalabl...
A VPN-as-a-Service Tailored Enabler for Computing-constrained EnvironmentsCarolina Fernández-Martínez, César Cajas Parra, Shuaib Siddiqui2026-06-10下载Industry has embraced Zero Trust (ZT) architectural tenets and implementations for cloud-native environments, following stricter security requirements to both internal and external tenants.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Real-Time Language Model Jamming: A Case Study for Live Music Accompaniment GenerationBowen Zheng, Andrew H. Yang, Jiaqi Ruan, Jia He, Xinyue Li, Yuan-Hsin Chen, Ziyu Wang, Xiaosong Ma2026-06-10下载Language models (LMs) have become one of the most prominent paradigms in modern generative modeling. While making them faster has been the main focus of real-time deployment, speed alone is not enough...

cs.PF - Performance ​

标题作者发布日期PDF摘要
nomp: A Framework for Building Domain Specific CompilersThilina Ratnayaka, Kaushik Kulkarni, Nipuna Fernando, Pubudu Hewavitharana, Hirumal Priyashan, Poorna Gunathilaka, Nagitha Abeywickrema, Ravindu Hirimuthugoda, Tarun Prabhu, Kirshanthan Sundararajah, Sanath Jayasena2026-06-10下载The low-level GPU programming models (CUDA, HIP, OpenCL, etc.) provide detailed control of the data flow and execution plan of a program in order to extract close-to-metal performance.
The Brain That Goes Quiet: Serving a Large Model's Knowledge at 131 Tokens per Second on an 8 GB Laptop by Removing the Large Model from the Runtime PathMyeong Jun Jo2026-06-10下载In earlier work I showed that a 35B-class Mixture-of-Experts model can be loaded and executed on a consumer laptop with 8 GB of GPU memory. That result solved a placement problem and immediately expos...
From Fork-Join to Asynchronous Tasks: Parallelizing Tiled Cholesky Decomposition with OpenMP and HPXAlexander Strack, Alexander Van Craen, Dirk Pflüger2026-06-10下载Fork-join parallelism, popularized by OpenMP, remains the dominant model for shared-memory parallel programming, but its implicit synchronization barriers can penalize algorithms with inhomogeneous wo...
Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost EstimationChitral Patil2026-06-10下载Every public LLM cost calculator we surveyed treats GPU utilization as a fixed input -- entered by the user, baked in as a preset, or silently assumed at 100% -- never measured against the operator's ...
XPR: An Extensible Cross-Platform Point-Based Differentiable RendererSteve Rhyner, Sankeerth Durvasula, Aleksandr Kovalev, Hansel Jia, Adrian Zhao, Mrutunjayya Mrutunjayya, Nilesh Ahuja, Selvakumar Panneer, Christina Giannoula, Nandita Vijaykumar2026-06-10下载Point-based differentiable rendering underpins modern 3D reconstruction, novel-view synthesis, and learning-based graphics pipelines, but developing new rendering methods often requires extensive low-...

基于 VitePress 构建 · 使用本地搜索查找论文