Skip to content

2026-05-18 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Predictive Software Scheduling as an Early-Warning Hint Layer for Optical Engine Thermal Drift in Heterogeneous SoIC PackagingChi Fei Chung2026-05-18下载As semiconductor scaling reaches the A16 / 2 nm node, the integration of co-packaged optics (CPO) via TSMC's Co-Packaged Optics Ultra Engine (COUPE) architecture introduces critical thermal-optical co...
Building Reliable Arithmetic Multipliers Under NBTI Aging and Process VariationsMasoud Heidary, Biresh Kumar Joardar2026-05-18下载Hardware aging poses a significant challenge for integrated circuits (ICs), leading to performance degradation and eventual failure. In this work, we focus on the aging of arithmetic multipliers, whic...
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and ResilienceSiddique Abubakr Muntaka, Edward Danso Ansong, Benjamin Yankson, Oliver Kornyo, Faiza Hussein, Mohammed Nadhir Muntaka, Joshua Dagadu, Prince Clement Addo, Maxwell Dorgbefu Jnr., Franco Osei-Wusu, Foster Yeboah, Michael Asante2026-05-18下载Uninterrupted system availability is a critical requirement for enterprise operations, yet traditional high-availability clusters suffer from limitations such as single points of failure and inefficie...
Enabling Agile Ambient IoT Networking via a Parameterized Hybrid RadioJiazhen Lei, Fengyuan Zhu, Tianze Cao, Yuxin Sha, Linling Zhong, Wenhui Li, Bingbing Wang, Zeming Yang, Jinyang Sun, Yibin Deng, Xiaohua Tian2026-05-18下载The emergence of Ambient IoT signals a paradigm shift toward massive batteryless networking. However, the absence of an agile physical layer substrate remains a fundamental barrier to research and sta...
CPPL: A Circuit Prompt Programming LanguageShuo Yin, Yihe Wang, Lancheng Zou, Xufeng Yao, Tinghuan Chen, Chen Bai, Zhengrong Wang, Tsung-Yi Ho, Bei Yu2026-05-18下载Large language models (LLMs) have shown promise in register-transfer level (RTL) design automation, but direct RTL generation remains difficult to validate, optimize, and integrate with compiler-based...
ROA-Based Subharmonic Injection Locking for Oscillator-Based Ising MachinesNicholas Sica, Baris Taskin2026-05-18下载This paper introduces on-chip integrated rotary traveling wave oscillators (RTWOs) organized into rotary oscillator array (ROA) bricks as an external perturbation to induce subharmonic injection locki...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI TrainingIoannis Papavasileiou, Sairam Prabhakar, Indu Kant Deo, Sergejs Makovejs2026-05-18下载We use discrete-event simulation to quantify the impact of fiber latency on the efficacy of geo-distributed AI model training with data parallelism.
Meta-Theorems for Cuttable Distributed ProblemsMarthe Bonamy, Avinandan Das, Cyril Gavoille, Timothé Picavet, Jukka Suomela, Alexandra Wesolek2026-05-18下载We prove that given any α-approximation LOCAL algorithm for Minimum Dominating Set (MDS) on planar graphs, we can construct an f(g)f(g)-round (3α+1)-approximation LOCAL algorithm for MDS on graphs e...
Unleashing the Power of Tree-of-Thoughts for Edge-Enabled AIGC Service ProvisioningZhang Liu, Shanhao Zhan, Shaowei Shen, Lianfen Huang, Qiao Xiang, Ying-Jun Angela Zhang, Dusit Niyato2026-05-18下载Delivering AI-generated content (AIGC) services fundamentally relies on the reasoning capabilities of generative AI (GenAI) models. Chain-of-Thought (CoT) enhances such reasoning by guiding models thr...
Near-Resolution of the Tradeoff Conjecture in Distributed Proof Labeling SchemesArnold Filtser, Orr Fischer2026-05-18下载In the tt-Proof Labeling Scheme model (tt-PLS model), our goal is to certify that a network of nodes satisfies a given property PP. A prover assigns a label to each node, and each node decides to a...
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime VariabilityRuitao Liu, Xinyang Tian, Shuo Chen, Tingrui Zhang, Guang Yang, Alan Zhao, Wei Xu2026-05-18下载Pipeline parallelism is a key technique for scaling large-model training, but modern workloads exhibit runtime variability in computation and communication.
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video GenerationYukang Chen, Luozhou Wang, Wei Huang, Shuai Yang, Bohan Zhang, Yicheng Xiao, Ruihang Chu, Weian Mao, Qixin Hu, Shaoteng Liu, Yuyang Zhao, Huizi Mao, Ying-Cong Chen, Enze Xie, Xiaojuan Qi, Song Han2026-05-18下载We present LongLive-2.0, an NVFP4-based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing speed and memory bottlenecks.
Mosaic: Towards Efficient Training of Multimodal Models with Spatial Resource MultiplexingYanbo Wang, Yuxuan Wang, Chen Chen, Chunyu Xue, Yu Feng, Anbang Wu, Quan Chen, Yin Chen, Qizhen Weng2026-05-18下载With the wide adoption of Multimodal Models (MMs) in real-world scenarios, it is significant to efficiently train emerging MMs that exhibit increasingly complex module architectures.
Ranking Opinions with Few States in Population ProtocolsTom-Lukas Breitkopf, Julien Dallot, Antoine El-Hayek, Stefan Schmid2026-05-18下载Population protocols are a model of distributed computing where nn agents, each a simple finite-state machine, interact in pairs to solve a common task against a (adversarial) interaction scheduler.
PopPy: Opportunistically Exploiting Parallelism in Python Compound AI ApplicationsStephen Mell, David Mell, Konstantinos Kallas, Steve Zdancewic, Osbert Bastani2026-05-18下载Compound AI applications, which compose calls to ML models using a general-purpose programming language like Python, are widely used for a variety of user-facing tasks, from software engineering to en...
EPIC: Abstraction and Polymorphism of In-Network Collectives on EthernetYitao Yuan, Jianglong Nie, Tianyu Bai, Ruizhe Zhou, Siyuan Cao, Xujie Fan, Yuchen Xu, Junkai Chen, Chenqi Zhao, Nengyuan Zhang, Shaoke Fang, Jiangyuan Chen, Yuanfeng Chen, Jiaqi Sun, Zhan Wang, Xiaohua Xu, Yuchao Zhang, Yang Liu, Xiangrui Yang, Jing Lin, Xiaohe Hu, Yang Li, Chao Jiang, Limin Xiao, Weifeng Zhang, Junjie Wang, Wei Cheng, Yazhu Lan, Jianbo Dong, Binzhang Fu, Wenfei Wu2026-05-18下载In-Network Collective (INC) acceleration holds immense potential for optimizing AI training and inference; however, its cross-layer nature has historically hindered investment and adoption within the ...
Efficient Gradient Methods for Distributed Saddle ProblemsRuichen Luo, Anton Rodomanov, Sebastian U. Stich2026-05-18下载The distributed setting for Saddle Problems (SPs) has recently emerged as a framework for various modern applications in machine learning and multiagent systems.
CB-SpMV:A Data Aggregating and Balance Algorithm for Cache-Friendly Block-Based SpMV on GPUsXing Cong, Fukai Sun, Yifan Chen, Chenhao Xie*, Yi Liu, Depei Qian2026-05-18下载Sparse matrix-vector multiplication (SpMV) is crucial in computational science, engineering, and machine learning. Despite substantial efforts to improve SpMV performance on GPUs through various techn...
Heterogeneous Tasks Offloading in Vehicular Edge Computing: A Federated Meta Deep Reinforcement Learning ApproachYaorong Huang, Jingtao Luo, Xuechao Wang2026-05-18下载Vehicular edge computing (VEC) enables latency-sensitive vehicular applications by offloading computation-intensive tasks to nearby edge servers.
Residue Number System Comparison revisited, a software perspectiveLaurent-Stéphane Didier, Léa Glandus, Nadia El Mrabet, Jean-Marc Robert2026-05-18下载This paper presents a novel method to compare two numbers in Residue Number System (RNS) using an additional modulus, which is often already available because it is required in modular computations an...
JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic PotentialsHongyu Wang, Weijian Liu, Hongtao Xu, Yan Wang, Mingzhen Li, Weile Jia, Guangming Tan2026-05-18下载Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy MD simulations, and ...
Duet instrumentation: An Agentic Approach to Improving Sensitivity in Cloud Service BenchmarkingSebastian Koch, Nils Japke, David Bermbach2026-05-18下载Continuous cloud service performance benchmarking is essential for detecting performance bugs early before deploying them to production. However, detecting performance regressions using application be...
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and ResilienceSiddique Abubakr Muntaka, Edward Danso Ansong, Benjamin Yankson, Oliver Kornyo, Faiza Hussein, Mohammed Nadhir Muntaka, Joshua Dagadu, Prince Clement Addo, Maxwell Dorgbefu Jnr., Franco Osei-Wusu, Foster Yeboah, Michael Asante2026-05-18下载Uninterrupted system availability is a critical requirement for enterprise operations, yet traditional high-availability clusters suffer from limitations such as single points of failure and inefficie...
ASSESSING THE STOCHASTIC PROPERTIES OF MODERN PSEUDO-RANDOM GENERATORS FOR PARALLEL COMPUTINGThéau Wartel, David R. C. Hill2026-05-18下载Pseudo-random number generators (PRNGs) are widely used in modern computing and are expected to exhibit excellent statistical performance and repeatability.
Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum MethodAbdurakhmon Sadiev, Artavazd Maranjyan, Ivan Ilin, Peter Richtárik2026-05-18下载Muon has recently emerged as a strong alternative to AdamW for training neural networks, with encouraging large-scale pretraining results and growing evidence that matrix-structured updates can be fas...
Early-Stabilizing CountingChristoph Lenzen, Julian Loss2026-05-18下载Synchronous Counting is the task of reaching agreement on a common round counter in a synchronous system of nn nodes with up to tt Byzantine faults in a self-stabilizing manner.
Distributed Renaming with Subquadratic Bits via Scalable Committee ElectionSirui Bai, Xinyu Fu, Yuyi Wang, Chaodong Zheng2026-05-18下载In distributed computing, the renaming problem requires nn nodes with unique identities from a large namespace [N][N] to acquire new, distinct identities from a smaller target namespace [M][M].
TIDAL: Recovering Temporal Phase for Cloud Block Storage Placement from LLM-Derived SemanticsDifan Tan, Changlin Wan, Jiawen Liu, Hua Wang, Ke Zhou2026-05-18下载Cloud Virtual Disk (CVD) placement in Cloud Block Storage (CBS) is critical for resource efficiency and performance isolation. Existing schemes prioritize spatial load balancing by dispersing disks ac...
The Task Completion Problem and its Application to Crash-Resilient ComputationOrr Fischer, Ran Gelles2026-05-18下载We study the Task Completion problem, in which MM abstract tasks must be completed by a network of nn crash-prone nodes, where up to αn nodes may crash for some constant α\<1.
A System Aware Resource Allocation for Distributed Workflows in Quantum Computing EnvironmentsAbhishek Sawaika, Udaya Parampalli, Rajkumar Buyya2026-05-18下载Rapid advancements in cloud based platforms providing access to quantum computing capabilities have opened up several challenges for efficient usage of these highly delicate and costly devices.
AdaptiveLoad: Towards Efficient Video Diffusion Transformer TrainingYucheng Guo, Yongjian Guo, Zhong Guan, Haoran Sun, Wen Huang, Wanting Xu, Jing Long, Shuai Di, Junwu Xiong2026-05-18下载In video generation models, particularly world models, training large-scale video diffusion Transformers (such as DiT and MMDiT) poses significant computational challenges due to the extreme variance ...
Guard: Scalable Straggler Detection and Node Health Management for Large-Scale TrainingGuanliang Liu, Abhinandan Patni, Congzhu Lin, Zoe Zeng, Jack Wittmayer, Josh Wu, Ashvin Nihalani, Binxuan Huang, Yinghong Liu, Rory Na, Anthony Ko, Alexander Zhipa, Cong Cheng, Mi Sun, Vijay Rajakumar, Rejith George Joseph, Parthasarathy Govindarajen2026-05-18下载Training frontier-scale foundation models involves coordinating tens of thousands of GPUs over multi-month runs, where even minor performance degradations can accumulate into substantial efficiency lo...
TierCheck: Tiered Checkpointing for Fault Tolerance in Large Language Model TrainingShujie Han, Feng Jiang, Patrick P. C. Lee, Xiao Zhang, Zhijie Huang, Nannan Zhao, Xiaonan Zhao, Lichen Pan2026-05-18下载Large Language Model (LLM) training is frequently interrupted by a heterogeneous spectrum of failures, from common GPU crashes to catastrophic cluster-wide outages.
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache QuantizationZhongzhu Zhou, Donglin Zhuang, Jisen Li, Ziyan Chen, Shuaiwen Leon Song, Ben Athiwaratkun, Xiaoxia Wu2026-05-18下载INT2 KV-cache quantization is attractive for long-context LLM serving, but it remains difficult to make both accurate and deployable. Simple rotations such as Hadamard transforms reduce outliers, but ...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
A Geometric Algebra-Informed 3D Gaussian Splatting Framework for Wireless Scene RepresentationJingzhou Shen, Tianya Zhao, Xuyu Wang2026-05-18下载In this paper, we introduce Geometric Algebra-Informed 3D Gaussian Splatting (GAI-GS), a framework for wireless modeling that couples 3D Gaussian splatting with a geometric algebra-based attention mec...
Collaborative Air-Ground Sensing, Communication, Computing, Storage, and Intelligence for Low-Altitude EconomyYiqin Deng, Junhui Gao, Zihan Fang, Yanan Ma, Xianhao Chen, Yuguang Fang2026-05-18下载Low-altitude economy (LAE) is transforming low-altitude airspace into a new cyber-physical infrastructure. Although air-ground communications have been widely studied, LAE is fundamentally different i...
Enabling Agile Ambient IoT Networking via a Parameterized Hybrid RadioJiazhen Lei, Fengyuan Zhu, Tianze Cao, Yuxin Sha, Linling Zhong, Wenhui Li, Bingbing Wang, Zeming Yang, Jinyang Sun, Yibin Deng, Xiaohua Tian2026-05-18下载The emergence of Ambient IoT signals a paradigm shift toward massive batteryless networking. However, the absence of an agile physical layer substrate remains a fundamental barrier to research and sta...
ASTRA: Asynchronous Age-Aware Satellite Random Access via Mean-Field ControlSayam Chakraborty, Aimin Li, Yigit Ince, Sajjad Baghaee, Elif Uysal2026-05-18下载Satellite Internet-of-Things (IoT) enables massive status-update services beyond terrestrial coverage, but grant-free uplink access creates a coupled freshness-control problem: increasing repetition a...
CA3D: Computing Accessibility-Aware Cooperative 3D Deployment of Multiple UAVsYiqin Deng, Zihan Fang, Yijie Wang, Qingxiao Huang, Junhui Gao, Qianyao Ren, Yuguang Fang2026-05-18下载This letter investigates computing-accessibility-aware cooperative 3D deployment of multiple UAVs for task completion enhancement, termed CA3D.
Enhancing Network Resilience via Graph-Based Anomaly Detection in Sovereign FunctionsXin Hao, Wei Ni, Chenhan Zhang, Massimo Piccardi, Raymond Owen2026-05-18下载Sovereign network functions, e.g., routing protocols, are becoming increasingly complex and susceptible to failures arising from protocol configuration anomalies and anomalous configurations.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
TIDAL: Recovering Temporal Phase for Cloud Block Storage Placement from LLM-Derived SemanticsDifan Tan, Changlin Wan, Jiawen Liu, Hua Wang, Ke Zhou2026-05-18下载Cloud Virtual Disk (CVD) placement in Cloud Block Storage (CBS) is critical for resource efficiency and performance isolation. Existing schemes prioritize spatial load balancing by dispersing disks ac...
PipeANN-Filter: An Efficient Filtered Vector Search System on SSDHao Guo, Jiwu Shu, Youyou Lu2026-05-18下载We propose PipeANN-Filter, an efficient filtered vector search system on SSD. Unlike existing systems that explore only valid vectors (i.e., those satisfying the attribute constraints) during search, ...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI TrainingIoannis Papavasileiou, Sairam Prabhakar, Indu Kant Deo, Sergejs Makovejs2026-05-18下载We use discrete-event simulation to quantify the impact of fiber latency on the efficacy of geo-distributed AI model training with data parallelism.
Reducing Waiting Time for Medical Tourists Through Hybrid Agent-Based and Discrete-Event Simulation: A Hospital Case StudyMelika Baghi, Hadi Mosadegh2026-05-18下载Medical tourists face a scheduling problem that differs from that of local patients. Treatment delays extend not just care delivery time, but also accommodation and travel costs.
On Generalized Performance Evaluation and Generalized Controller SynthesisZining Cao2026-05-18下载In this paper, we propose the frameworks of generalized performance evaluation and generalized controller synthesis. To this end, we give a true concurrent process calculus as the model of systems, an...
Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV EvictionGabriel Garcia2026-05-18下载We study KV cache eviction under a shared globally capped decode-time harness. Seven policies (LRU, H2O, SnapKV, StreamingLLM, Ada-KV, QUEST, Random) share a prompt-boundary vulnerability: without str...
OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache QuantizationZhongzhu Zhou, Donglin Zhuang, Jisen Li, Ziyan Chen, Shuaiwen Leon Song, Ben Athiwaratkun, Xiaoxia Wu2026-05-18下载INT2 KV-cache quantization is attractive for long-context LLM serving, but it remains difficult to make both accurate and deployable. Simple rotations such as Hadamard transforms reduce outliers, but ...

基于 VitePress 构建 · 使用本地搜索查找论文