Skip to content

2026-06-02 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
StepPRM-RTL: Stepwise Process-Reward Guided LLM Fine-Tuning for Enhanced RTL SynthesisPrashanth Vijayaraghavan, Apoorva Nitsure, Luyao Shi, Ehsan Degan, Vandana Mukherjee2026-06-02下载Automatic generation of RTL code for digital hardware designs remains challenging due to long-horizon reasoning, multi-step dependencies, and strict correctness constraints in Verilog and VHDL.
Feasibility of Time-Domain DNN-Based Speech Enhancement on Embedded FPGA for Hearing AidFeyisayo Olalere, Umut Altin, Kiki van der Heijden, Marcel van Gerven2026-06-02下载Hearing aids impose strict latency and power constraints that current DNN-based speech enhancement systems struggle to meet on embedded hardware.
HighTide: An Agent-Curated Open-Source VLSI Benchmark SuiteBenjamin Goldblatt, Paolo Pedroso, Farhad Modaresi, Ethan Sifferman, Matthew R. Guthaus2026-06-02下载We introduce HighTide, an evolving AI-assisted benchmark suite. Specifically, the contributions are: (i) a diverse open-source suite spanning multiple design languages and technology nodes, (ii) Bazel...
ACRONYM: Accelerated Approximate Nearest Neighbor Search in Memory for Dynamic Vector DatabasesMd Mizanur Rahaman Nayan, Tianqi Zhang, Flavio Ponzina, Tajana Rosing, Azad J Naeemi2026-06-02下载Vector database search with frequent updates is increasingly critical in applications such as retrieval augmented generation, recommendation systems, and large-scale embedding retrieval.
ZK-Flex: A Flexible and Scalable Framework for Accelerating Zero-Knowledge ProofsAdiwena Putra, Cuong Manh Duong, Anh Quang Pham, Joo-Young Kim2026-06-02下载Zero-knowledge proofs (ZKP) allows a prover to convince a verifier of computational correctness without revealing private data, ensuring both privacy and verifiability.
MOSAIC: Efficient Mixture-of-Agent Scheduling via Adaptive Aggregation and Inference ConcurrencySaptarshi Mitra, Yifan Zhang, Rachid Karami, Phyo Pyae Moe Aung, Nazmul Takbir, Sreetama Sarkar, Souvik Kundu, Sitao Huang2026-06-02下载Mixture-of-Agents (MoA) systems improve reasoning accuracy by routing each query to multiple expert LLMs and aggregating their outputs. Efficiently executing this workload on limited GPU resources has...
Glass Box at Orbit: A Constitutional AI Verification Framework for Trustworthy Autonomous CubeSat IntelligenceKarthik Barma, Anil Sanneboyina, V C Premchand Yadav2026-06-02下载The space industry is quietly building toward something nobody has fully reckoned with: orbital data centers running thousands of autonomous AI workloads with no human in the loop, 550 km above the Ea...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
ACEAPEX: Parallel LZ77 Decoding via Encode-Time Absolute Offset ResolutionYakiv Shavidze2026-06-02下载LZ77-based codecs exhibit a fundamental sequential bottleneck in decoding: each back-reference depends on previously decompressed data, preventing multi-core scaling.
Notarized Agents: Receiver-Attested Confidential Receipts for AI Agent ActionsJuan Figuera2026-06-02下载Current AI agent observability is structurally compromised: the entity producing the activity log is the same entity whose activity is being logged.
EvalStop: Using World Feedback to Detect and Correct Reward Overoptimization in Multi-Tenant RLHF PlatformsGuilin Zhang, Chuanyi Sun, Shahryar Sarkani, John M. Fossaceca2026-06-02下载Cloud LLM fine-tuning platforms increasingly serve RLHF workloads, where a learned reward model is optimized as a proxy for human quality. As Gao et al.
UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load BalancingXinming Wei, Chao Jin, Tuo Dai, Yinmin Zhong, Shan Yu, Chengxu Yang, Bingyang Wu, Zili Zhang, Jing Mai, Qianchao Zhu, Zhouyang Li, Yuliang Liu, Guojie Luo2026-06-02下载Large-scale expert parallelism (EP) is becoming pivotal for training and serving frontier MoE models, but it also amplifies device-level expert load imbalance into compute stragglers, token all-to-all...
NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM InferenceMubarak Adetunji Ojewale2026-06-02下载Disaggregated LLM inference forces the KV cache to traverse the datacenter network before decoding begins, so transfer time enters directly into the Time to First Token (TTFT) budget.
E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog EnvironmentsTruong-Thanh Le, Amir Taherkordi, Hoang-Loc La, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan2026-06-02下载Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment must address cost effic...
CADET: A Modular Platform for Evaluating Distributed Cooperative Autonomy in Connected Autonomous VehiclesPragya Sharma, Brian Wang, Mani Srivastava2026-06-02下载Deep learning models are increasingly central to autonomous vehicle (AV) pipelines, yet their integration has traditionally followed a monolithic design where perception, planning, and control execute...
Fast TetraBFT: Optimizing Latency Where It MattersAntonio J. Fernández-Pinto, Manuel Bravo, Gregory Chockler, Alexey Gotsman2026-06-02下载Unauthenticated Byzantine consensus protocols achieve optimal failure resilience while relying only on authenticated point-to-point channels, not authenticated messages.
Deterministic Distance Approximation in MPC via Improved Hitting SetsKyungjin Cho, Michal Dory, Yannic Maus, Tijn de Vos2026-06-02下载In this paper, we provide the first deterministic algorithms with sublogarithmic round complexity for spanners and approximate shortest paths in various MPC models.
SIGMA: A Versatile Streaming Graph Partitioner for Vertex- and Edge-Balanced Distributed GNN TrainingBarbara Hoffmann, Shai Dorian Peretz, Adil Chhabra, Ahmet Kadir Yalcinkaya, Ruben Mayer, Christian Schulz2026-06-02下载Distributed Graph Neural Network (GNN) training depends critically on how the underlying graph is partitioned across compute resources. Existing graph partitioners focus either on vertex partitioning ...
Demystifying Pipeline Parallelism: First Theory for PipeDreamIvan Ilin, Peter Richtárik2026-06-02下载Training modern machine learning models increasingly requires computation to be distributed across many accelerators. Data parallelism remains the default choice and is often paired with tensor-parall...
Predicting Lakehouse Performance in Clouds: An Empirical Exploration of Query Runtime VarianceJames Nurdin, Wei Liu, Richard Mccreadie, Lauritz Thamsen2026-06-02下载Data analytics increasingly runs on distributed lakehouse systems, where platform operators must optimise monetary, resource, and environmental costs.
BlobShuffle: Cost-Effective Repartitioning in Stream Processing Systems via Object Storage Exemplified with Kafka StreamsSören Henning, Otmar Ertl, Adriano Vogel2026-06-02下载Shuffling or repartitioning data streams is an essential operation of state-of-the-art stream processing frameworks to support stateful workloads in a large-scale, distributed setting.
OpenAgenet/OAN: Technical Architecture for Trust-Governed Agent Identity and DiscoveryJinliang Xu2026-06-02下载This paper describes the technical architecture of OpenAgenet / OAN. OAN is a protocol-neutral trust layer for open Agent interconnection. It specifies the role architecture, identity objects, registr...
Libra: Efficient Resource Management for Agentic RL Post-TrainingKaiwen Chen, Xin Tan, Jingzong Li, Hong Xu2026-06-02下载Reinforcement learning (RL) has become a standard post-training paradigm for large language models (LLMs), extending beyond preference alignment to complex reasoning and multi-turn agentic behaviors.
Brief Announcement: Generative Markov Model for Distributed Computing SystemsAlfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta2026-06-02下载Emerging distributed computing paradigms, such as the computing continuum, are inherently heterogeneous, stochastic, and complex. Efficiently and effectively utilizing all available resources across t...
FOLD: Fuzzy Online Deduplication for Very Large Evolving Datasets via Approximate Nearest Neighbor SearchNelson Bore, Pritish Mishra, Constantin Adam, Eyal de Lara, Oana Balmau2026-06-02下载Fuzzy deduplication is key to constructing large language model training corpora. However, classic Locality-Sensitive Hashing pipelines scale poorly as corpora grow and are ill-suited to continuous in...
DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU InferenceKathiravan Palaniappan2026-06-02下载The rapid growth of large language model (LLM) inference services has increased the demand for efficient multi-tenant GPU scheduling. While modern inference runtimes such as vLLM improve throughput th...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Anycast Performance in ContextEric Liang2026-06-02下载IP anycast lets a service advertise one address from many physical sites, leaving BGP to map each client to a site. It is central to the DNS root server system, public resolvers, and some content deli...
NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM InferenceMubarak Adetunji Ojewale2026-06-02下载Disaggregated LLM inference forces the KV cache to traverse the datacenter network before decoding begins, so transfer time enters directly into the Time to First Token (TTFT) budget.
AUGUSTE: Online-Learning dApp for Predictive URLLC SchedulingMaxime Elkael, Michele Polese, Yunseong Lee, Koichiro Furueda, Tommaso Melodia2026-06-02下载Ultra Reliable and Low Latency Communications (URLLC) was one of the main motivations behind 5G, with 3GPP advertising 1-10 ms latency targets for applications such as industrial automation, Vehicle-T...
Towards Intrusion Detection Systems for RPL-based IoT Networks using Foundation ModelsElias Lunderbye, Sourasekhar Banerjee, Christian Rohner, Andreas Johnsson2026-06-02下载AI-based intrusion detection systems (IDS) have shown promise in detecting attacks on IoT systems. In this work, we explore the use of foundation models to detect and identify attacks, with a specific...
Throughput Optimization for Multi-AP IEEE P802.11bq Networks Based on Combinatorial Multi-Armed BanditsAnshan Yuan, Mingqi Han, Xinghua Sun2026-06-02下载This paper addresses distributed throughput optimization for dense multi-AP IEEE P802.11bq networks. We develop a packet-level model that jointly captures cross-link carrier-sense multiple access with...
When BBR Meets Live StreamingXu Yan, Tong Li, Bo Wu, Cheng Luo, Jiuxiang Zhu, Laizhong Cui2026-06-02下载Recently, industrial pioneers like Amazon, Tencent, ByteDance, and Huawei have been adopting BBR as their congestion control algorithm for live-streaming applications, including TikTok Live.
Rain: RDMA-assisted In-Network Scheduling for Microsecond-scale WorkloadsZhihuang Ma, Xingming Cui, Xiaoliang Chen, Zuqing Zhu2026-06-02下载Modern data center applications increasingly require microsecond-scale service time with strict tail latency requirements, which can hardly be realized with existing in-network task schedulers due to ...
Brief Announcement: Generative Markov Model for Distributed Computing SystemsAlfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta2026-06-02下载Emerging distributed computing paradigms, such as the computing continuum, are inherently heterogeneous, stochastic, and complex. Efficiently and effectively utilizing all available resources across t...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM AgentsYingqi Zhang2026-06-02下载Large language model (LLM) agents are evolving from request-response assistants into long-running software actors: they maintain state across model calls, fork subtasks, wait for external events, requ...

基于 VitePress 构建 · 使用本地搜索查找论文