2026-06-02
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| StepPRM-RTL: Stepwise Process-Reward Guided LLM Fine-Tuning for Enhanced RTL Synthesis | Prashanth Vijayaraghavan, Apoorva Nitsure, Luyao Shi, Ehsan Degan, Vandana Mukherjee | 2026-06-02 | 下载 | Automatic generation of RTL code for digital hardware designs remains challenging due to long-horizon reasoning, multi-step dependencies, and strict correctness constraints in Verilog and VHDL. |
| Feasibility of Time-Domain DNN-Based Speech Enhancement on Embedded FPGA for Hearing Aid | Feyisayo Olalere, Umut Altin, Kiki van der Heijden, Marcel van Gerven | 2026-06-02 | 下载 | Hearing aids impose strict latency and power constraints that current DNN-based speech enhancement systems struggle to meet on embedded hardware. |
| HighTide: An Agent-Curated Open-Source VLSI Benchmark Suite | Benjamin Goldblatt, Paolo Pedroso, Farhad Modaresi, Ethan Sifferman, Matthew R. Guthaus | 2026-06-02 | 下载 | We introduce HighTide, an evolving AI-assisted benchmark suite. Specifically, the contributions are: (i) a diverse open-source suite spanning multiple design languages and technology nodes, (ii) Bazel... |
| ACRONYM: Accelerated Approximate Nearest Neighbor Search in Memory for Dynamic Vector Databases | Md Mizanur Rahaman Nayan, Tianqi Zhang, Flavio Ponzina, Tajana Rosing, Azad J Naeemi | 2026-06-02 | 下载 | Vector database search with frequent updates is increasingly critical in applications such as retrieval augmented generation, recommendation systems, and large-scale embedding retrieval. |
| ZK-Flex: A Flexible and Scalable Framework for Accelerating Zero-Knowledge Proofs | Adiwena Putra, Cuong Manh Duong, Anh Quang Pham, Joo-Young Kim | 2026-06-02 | 下载 | Zero-knowledge proofs (ZKP) allows a prover to convince a verifier of computational correctness without revealing private data, ensuring both privacy and verifiability. |
| MOSAIC: Efficient Mixture-of-Agent Scheduling via Adaptive Aggregation and Inference Concurrency | Saptarshi Mitra, Yifan Zhang, Rachid Karami, Phyo Pyae Moe Aung, Nazmul Takbir, Sreetama Sarkar, Souvik Kundu, Sitao Huang | 2026-06-02 | 下载 | Mixture-of-Agents (MoA) systems improve reasoning accuracy by routing each query to multiple expert LLMs and aggregating their outputs. Efficiently executing this workload on limited GPU resources has... |
| Glass Box at Orbit: A Constitutional AI Verification Framework for Trustworthy Autonomous CubeSat Intelligence | Karthik Barma, Anil Sanneboyina, V C Premchand Yadav | 2026-06-02 | 下载 | The space industry is quietly building toward something nobody has fully reckoned with: orbital data centers running thousands of autonomous AI workloads with no human in the loop, 550 km above the Ea... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ACEAPEX: Parallel LZ77 Decoding via Encode-Time Absolute Offset Resolution | Yakiv Shavidze | 2026-06-02 | 下载 | LZ77-based codecs exhibit a fundamental sequential bottleneck in decoding: each back-reference depends on previously decompressed data, preventing multi-core scaling. |
| Notarized Agents: Receiver-Attested Confidential Receipts for AI Agent Actions | Juan Figuera | 2026-06-02 | 下载 | Current AI agent observability is structurally compromised: the entity producing the activity log is the same entity whose activity is being logged. |
| EvalStop: Using World Feedback to Detect and Correct Reward Overoptimization in Multi-Tenant RLHF Platforms | Guilin Zhang, Chuanyi Sun, Shahryar Sarkani, John M. Fossaceca | 2026-06-02 | 下载 | Cloud LLM fine-tuning platforms increasingly serve RLHF workloads, where a learned reward model is optimized as a proxy for human quality. As Gao et al. |
| UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing | Xinming Wei, Chao Jin, Tuo Dai, Yinmin Zhong, Shan Yu, Chengxu Yang, Bingyang Wu, Zili Zhang, Jing Mai, Qianchao Zhu, Zhouyang Li, Yuliang Liu, Guojie Luo | 2026-06-02 | 下载 | Large-scale expert parallelism (EP) is becoming pivotal for training and serving frontier MoE models, but it also amplifies device-level expert load imbalance into compute stragglers, token all-to-all... |
| NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference | Mubarak Adetunji Ojewale | 2026-06-02 | 下载 | Disaggregated LLM inference forces the KV cache to traverse the datacenter network before decoding begins, so transfer time enters directly into the Time to First Token (TTFT) budget. |
| E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments | Truong-Thanh Le, Amir Taherkordi, Hoang-Loc La, Frank Eliassen, Phuong Hoai Ha, Peiyuan Guan | 2026-06-02 | 下载 | Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the models themselves, practical deployment must address cost effic... |
| CADET: A Modular Platform for Evaluating Distributed Cooperative Autonomy in Connected Autonomous Vehicles | Pragya Sharma, Brian Wang, Mani Srivastava | 2026-06-02 | 下载 | Deep learning models are increasingly central to autonomous vehicle (AV) pipelines, yet their integration has traditionally followed a monolithic design where perception, planning, and control execute... |
| Fast TetraBFT: Optimizing Latency Where It Matters | Antonio J. Fernández-Pinto, Manuel Bravo, Gregory Chockler, Alexey Gotsman | 2026-06-02 | 下载 | Unauthenticated Byzantine consensus protocols achieve optimal failure resilience while relying only on authenticated point-to-point channels, not authenticated messages. |
| Deterministic Distance Approximation in MPC via Improved Hitting Sets | Kyungjin Cho, Michal Dory, Yannic Maus, Tijn de Vos | 2026-06-02 | 下载 | In this paper, we provide the first deterministic algorithms with sublogarithmic round complexity for spanners and approximate shortest paths in various MPC models. |
| SIGMA: A Versatile Streaming Graph Partitioner for Vertex- and Edge-Balanced Distributed GNN Training | Barbara Hoffmann, Shai Dorian Peretz, Adil Chhabra, Ahmet Kadir Yalcinkaya, Ruben Mayer, Christian Schulz | 2026-06-02 | 下载 | Distributed Graph Neural Network (GNN) training depends critically on how the underlying graph is partitioned across compute resources. Existing graph partitioners focus either on vertex partitioning ... |
| Demystifying Pipeline Parallelism: First Theory for PipeDream | Ivan Ilin, Peter Richtárik | 2026-06-02 | 下载 | Training modern machine learning models increasingly requires computation to be distributed across many accelerators. Data parallelism remains the default choice and is often paired with tensor-parall... |
| Predicting Lakehouse Performance in Clouds: An Empirical Exploration of Query Runtime Variance | James Nurdin, Wei Liu, Richard Mccreadie, Lauritz Thamsen | 2026-06-02 | 下载 | Data analytics increasingly runs on distributed lakehouse systems, where platform operators must optimise monetary, resource, and environmental costs. |
| BlobShuffle: Cost-Effective Repartitioning in Stream Processing Systems via Object Storage Exemplified with Kafka Streams | Sören Henning, Otmar Ertl, Adriano Vogel | 2026-06-02 | 下载 | Shuffling or repartitioning data streams is an essential operation of state-of-the-art stream processing frameworks to support stateful workloads in a large-scale, distributed setting. |
| OpenAgenet/OAN: Technical Architecture for Trust-Governed Agent Identity and Discovery | Jinliang Xu | 2026-06-02 | 下载 | This paper describes the technical architecture of OpenAgenet / OAN. OAN is a protocol-neutral trust layer for open Agent interconnection. It specifies the role architecture, identity objects, registr... |
| Libra: Efficient Resource Management for Agentic RL Post-Training | Kaiwen Chen, Xin Tan, Jingzong Li, Hong Xu | 2026-06-02 | 下载 | Reinforcement learning (RL) has become a standard post-training paradigm for large language models (LLMs), extending beyond preference alignment to complex reasoning and multi-turn agentic behaviors. |
| Brief Announcement: Generative Markov Model for Distributed Computing Systems | Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta | 2026-06-02 | 下载 | Emerging distributed computing paradigms, such as the computing continuum, are inherently heterogeneous, stochastic, and complex. Efficiently and effectively utilizing all available resources across t... |
| FOLD: Fuzzy Online Deduplication for Very Large Evolving Datasets via Approximate Nearest Neighbor Search | Nelson Bore, Pritish Mishra, Constantin Adam, Eyal de Lara, Oana Balmau | 2026-06-02 | 下载 | Fuzzy deduplication is key to constructing large language model training corpora. However, classic Locality-Sensitive Hashing pipelines scale poorly as corpora grow and are ill-suited to continuous in... |
| DriftSched: Adaptive QoS-Aware Scheduling under Runtime Token Drift for Multi-Tenant GPU Inference | Kathiravan Palaniappan | 2026-06-02 | 下载 | The rapid growth of large language model (LLM) inference services has increased the demand for efficient multi-tenant GPU scheduling. While modern inference runtimes such as vLLM improve throughput th... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Anycast Performance in Context | Eric Liang | 2026-06-02 | 下载 | IP anycast lets a service advertise one address from many physical sites, leaving BGP to map each client to a site. It is central to the DNS root server system, public resolvers, and some content deli... |
| NetKV: Network-Aware Decode Instance Selection for Disaggregated LLM Inference | Mubarak Adetunji Ojewale | 2026-06-02 | 下载 | Disaggregated LLM inference forces the KV cache to traverse the datacenter network before decoding begins, so transfer time enters directly into the Time to First Token (TTFT) budget. |
| AUGUSTE: Online-Learning dApp for Predictive URLLC Scheduling | Maxime Elkael, Michele Polese, Yunseong Lee, Koichiro Furueda, Tommaso Melodia | 2026-06-02 | 下载 | Ultra Reliable and Low Latency Communications (URLLC) was one of the main motivations behind 5G, with 3GPP advertising 1-10 ms latency targets for applications such as industrial automation, Vehicle-T... |
| Towards Intrusion Detection Systems for RPL-based IoT Networks using Foundation Models | Elias Lunderbye, Sourasekhar Banerjee, Christian Rohner, Andreas Johnsson | 2026-06-02 | 下载 | AI-based intrusion detection systems (IDS) have shown promise in detecting attacks on IoT systems. In this work, we explore the use of foundation models to detect and identify attacks, with a specific... |
| Throughput Optimization for Multi-AP IEEE P802.11bq Networks Based on Combinatorial Multi-Armed Bandits | Anshan Yuan, Mingqi Han, Xinghua Sun | 2026-06-02 | 下载 | This paper addresses distributed throughput optimization for dense multi-AP IEEE P802.11bq networks. We develop a packet-level model that jointly captures cross-link carrier-sense multiple access with... |
| When BBR Meets Live Streaming | Xu Yan, Tong Li, Bo Wu, Cheng Luo, Jiuxiang Zhu, Laizhong Cui | 2026-06-02 | 下载 | Recently, industrial pioneers like Amazon, Tencent, ByteDance, and Huawei have been adopting BBR as their congestion control algorithm for live-streaming applications, including TikTok Live. |
| Rain: RDMA-assisted In-Network Scheduling for Microsecond-scale Workloads | Zhihuang Ma, Xingming Cui, Xiaoliang Chen, Zuqing Zhu | 2026-06-02 | 下载 | Modern data center applications increasingly require microsecond-scale service time with strict tail latency requirements, which can hardly be realized with existing in-network task schedulers due to ... |
| Brief Announcement: Generative Markov Model for Distributed Computing Systems | Alfreds Lapkovskis, Ali Beikmohammadi, Sindri Magnússon, Praveen Kumar Donta | 2026-06-02 | 下载 | Emerging distributed computing paradigms, such as the computing continuum, are inherently heterogeneous, stochastic, and complex. Efficiently and effectively utilizing all available resources across t... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents | Yingqi Zhang | 2026-06-02 | 下载 | Large language model (LLM) agents are evolving from request-response assistants into long-running software actors: they maintain state across model calls, fork subtasks, wait for external events, requ... |