2026-07-02
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| HyNoC: A Hybrid Circuit-Switch/Wormhole Network-on-Chip for Distributed VLIW Computing on FPGA | Christophe Clienti | 2026-07-02 | 下载 | Network-on-Chip (NoC) architectures have become the standard interconnect fabric for many-core systems, yet most proposals face a fundamental trade-off between latency, area, and congestion management... |
| Probabilistic Memory for Trustworthy Edge Intelligence | Likai Pei, Jiahao Zheng, Xueji Zhao, Emilie Ye, Jianbo Liu, Hanqing Tao, Ming-Yen Lee, Ruiyang Qin, Yiyu Shi, Shimeng Yu, X. Sharon Hu, Ningyuan Cao | 2026-07-02 | 下载 | Probabilistic computation plays an important role in trustworthy edge intelligence to quantify uncertainty, enhance robustness, reconstruct data, and protect privacy, but its adoption is limited by th... |
| APEIRON: composing smart TDAQ systems for high energy physics experiments | Roberto Ammendola, Andrea Biagioni, Carlotta Chiarini, Andrea Ciardiello, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Pierpaolo Perticaroli, Cristian Rossi, Francesco Simula, Matteo Turisini, Piero Vicini | 2026-07-02 | 下载 | We present APEIRON, a distributed heterogeneous processing framework comprising both hardware architecture and software stack for multi-FPGA systems. |
| A 2048-spin bulk acoustic wave Ising machine for number partitioning and Sudoku | Venkatesh Vadde, Roman Ovcharov, Victor H. González, Roman Khymyn, Artem Litvinenko, Johan Åkerman | 2026-07-02 | 下载 | Optical coherent Ising machines based on time-multiplexing have demonstrated significant progress in terms of connectivity and spin scalability. |
| Approximate Attention Weighting for Sustainable FPGA-Based Vision Transformer Inference | Muhammad Usman, Muhammad Akmal Shafique, Shujaat Khan, Dorit Merhof | 2026-07-02 | 下载 | Vision Transformers have reshaped computer vision by using self-attention to capture global context across image regions. This makes them attractive for edge visual inspection and monitoring in applic... |
| 3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving | Jaehun Lee, In-Jun Jung, Joo-Young Kim | 2026-07-02 | 下载 | Large language model (LLM) serving increasingly combines prefill-decode (PD) disaggregation with tensor parallelism (TP) to support large models and long contexts. In conventional 2D/2. |
| MxGLUT: A Reconfigurable LUT-Centric Broadcast Dataflow Accelerator for Mixed-Precision GEMM | Weiyu Zhou, Chen Ding, Mingyuan Liu, Liangyu Gan, Yukun Feng, Hao Jia, Haoming Chu, Lirong Zheng, Ning Ma, Yuxiang Huan | 2026-07-02 | 下载 | Large language model (LLM) inference suffers from growing inefficiency across the prefill and decode phases, especially under weight-only quantization, where activations remain in FP8 while weights ar... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Remora: Scale-out Deterministic Execution for Smart Contracts | Zhengqing Liu, Alberto Sonnino, Igor Zablotchi, Eleftherios Kokoris Kogias, Marios Kogias | 2026-07-02 | 下载 | Modern blockchains rely on a modular architecture that decouples consensus from execution. Recent advances in consensus algorithms have shifted the bottleneck to the execution layer, which must determ... |
| HyNoC: A Hybrid Circuit-Switch/Wormhole Network-on-Chip for Distributed VLIW Computing on FPGA | Christophe Clienti | 2026-07-02 | 下载 | Network-on-Chip (NoC) architectures have become the standard interconnect fabric for many-core systems, yet most proposals face a fundamental trade-off between latency, area, and congestion management... |
| LLMoxie: Exploring Agentic AI for Scientific Software Development | Landung Setiawan, Anant Mittal, Cordero Core, Anshul Tambay, Carlos Garcia Jurado Suarez, David A. C. Beck, Andrew J. Connolly, Vani Mandava | 2026-07-02 | 下载 | In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise inference, a LiteLLM/MLflow control plane for authentication, budge... |
| FlintKV: A Fast Durable Storage Engine for Modern Databases | Sergey Egorov, Gregory Chockler, Brijesh Dongol, Dan O'Keeffe, Sadegh Keshavarzi | 2026-07-02 | 下载 | Byte-addressable non-volatile memory (NVM) offers an opportunity to rethink storage engine architectures. While recent NVM key-value stores achieve high throughput for ingestion and point lookups, the... |
| WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs | Mauricio Fadel Argerich, Jonathan Fürst, Marta Patiño-Martínez | 2026-07-02 | 下载 | Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to the most efficient G... |
| Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing Data | Daniel M. Jimenez-Gutierrez, Enrique Zuazua, Georgios Kellaris, Joaquin del Rio, Oleksii Sliusarenko, Xabi Uribe-Etxebarria | 2026-07-02 | 下载 | Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster response, operational security environments, infrastructure m... |
| Elasticity in Parallel Sparse Triangular Solve | Raphael S. Steiner, Christos K. Matzoros, Pál András Papp, Toni Böhnlein, A. N. Yzelman | 2026-07-02 | 下载 | We introduce stale synchronous parallel as a mode of execution in parallel sparse triangular linear system solve and present a general directed-acyclic-graph scheduler capable of producing such schedu... |
| Securing People and their Machines Against Major Faults | Ohad Eitan, Idit Keidar, Ehud Shapiro | 2026-07-02 | 下载 | We consider grassroots platforms -- distributed systems of agents consisting of people identified by self-chosen public keys and their machines (smartphones) -- and wish to make them secure against \e... |
| Cadence: Extreme Pipelining with Multiple Concurrent Proposers | Fatima Elsheimy, Mohammad Mussadiq Jalalzai, Tobias Klenze, Jovan Komatovic, Mike Setrin, Victor Shoup, Kushal Babel, Lioba Heimbach, Jason Milionis | 2026-07-02 | 下载 | We present Cadence, a Byzantine fault-tolerant multi-proposer consensus protocol with arbitrarily low block intervals, optimal resilience, and optimal fast-path latency. |
| Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of Lines | Bojie Li | 2026-07-02 | 下载 | Hardware accelerators now sit on the critical path of online serving. GPUs, FPGAs, and increasingly remote services such as hardware security modules, post-quantum KEMs, and inference servers. |
| Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving | Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj, Renee St. Amant, Victor Rühle | 2026-07-02 | 下载 | Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymmetry: under bursty, heavy-tailed workloads prefill... |
| Scalable and Distributed Silhouette Approximation | Ilie Sarpe, Federico Altieri, Andrea Pietracaprina, Geppino Pucci, Fabio Vandin | 2026-07-02 | 下载 | The silhouette is one of the most widely used measures to assess the quality of a -clustering of a dataset of elements. Its evaluation requires no information beyond the clustering assignment. |
| Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models | Xuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty | 2026-07-02 | 下载 | This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes various existing and novel parallelism techniques at... |
| Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference | Wenchen Han, Gingfung Matthew Yeung, Marco Barletta, William Toner, Amory Hoste, Adam Barker | 2026-07-02 | 下载 | Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems. |
| HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention | Chao Yuan, Pan Li, Yingnan Sun, Jing Liu | 2026-07-02 | 下载 | All-to-all based sequence parallelism methods execute communication and computation strictly in serial when processing medium-long sequences, resulting in hardware resource underutilization. |
| Exploiting Task-Based Parallelism for the Red-Black Gauss-Seidel Method on 2D Grids | Shiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter | 2026-07-02 | 下载 | Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques. |
| Arachne: Orchestrating Cascades for Efficient Text-to-Video Model Training | Peng Yu, Yuankai Fan, Yang Qiu, Tian Li, Bihuan Chen, Yin Chen, Qizhen Weng | 2026-07-02 | 下载 | The rising demand for AI-generated videos is fueled by advances in large-scale Text-to-Video (T2V) models, trained on extensive datasets of video clips spanning diverse resolutions and durations. |
| SCAPE: Accurate and Efficient LLM Training with Extreme Sparse Communication | Mingkai Zheng, Junlin Chen, Haotian Xie, Zhao Zhang | 2026-07-02 | 下载 | Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronization and parameter reco... |
| DeadPool: Resilient LLM Training with Hot-Swapping via Zero-Overhead Checkpoint | Haotian Xie, Junlin Chen, Mingkai Zheng, Lishan Yang, Zhao Zhang | 2026-07-02 | 下载 | State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack. |
| OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU Clusters | D. Balamurugan, Thomas W. Bush | 2026-07-02 | 下载 | Serving large language models (LLMs) on a shared, heterogeneous GPU cluster requires users and operators to select the GPU type, tensor-parallel degree, and precision before committing valuable node-h... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving | Jiaying Meng, Bojie Li | 2026-07-02 | 下载 | Real-time interaction models -- Moshi, MiniCPM-o, Qwen-Omni -- turn serving into a periodic real-time task: on every frame a session ingests streaming audio and must respond by a recurring wall-clock ... |
| Criticality-Based Guard Rail Validation for AI Agent Decisions in Autonomous Telecom Networks | Ravi Kant Sharma | 2026-07-02 | 下载 | The evolution toward fully autonomous telecommunications networks (Autonomous Network Levels 4-5) requires AI/ML agents to make real-time network decisions without human intervention. |
| CSI Simulation: Why Additive Noise Fails and How to Fix It | Aymen Bouferroum, Ildi Alla, Vincent Lenders, Valeria Loscri | 2026-07-02 | 下载 | Channel State Information (CSI) has become a widely used wireless channel sensing modality for applications such as indoor localization, activity recognition, and respiration monitoring. |
| Enabling Real-Time AI in O-RAN: Deploying andMeasuring AI Inside a Near-RT RIC xApp | Lawrence Obiuwevwi, Krzysztof J. Rechowicz, Sampath Jayarathna, Safdar Hussain Bouk, Fahmida Afrin, C. Nicolas Barati, Neda Moghim, Valentina Nanou, Muhammad Enayetur Rahman, Sachin Shetty | 2026-07-02 | 下载 | Open Radio Access Network (O-RAN) architectures introduce programmable Near-Real-Time RAN Intelligent Controllers (Near-RT RICs) that support closed-loop control through xApps at timescales from 10 ms... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Characterizing and Bridging the Diagnostic Gap in eBPF Verifier Rejections | Yusheng Zheng, Zhengjie Ji, Weichen Tao, Xiangyu Gao, Jianchang Su, Wei Zhang, Andi Quinn, Dan Williams | 2026-07-02 | 下载 | eBPF lets developers run custom programs inside the Linux kernel, where a verifier proves each program safe. However, when the verifier rejects a program, the unclear error makes repair challenging: t... |
| Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots | Ling Xu, Chuyu Han, Borui Li, Hao Wu, Shiqi Jiang, Ting Cao, Chuanyou Li, Sheng Zhong, Shuai Wang | 2026-07-02 | 下载 | Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, an... |
| Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of Lines | Bojie Li | 2026-07-02 | 下载 | Hardware accelerators now sit on the critical path of online serving. GPUs, FPGAs, and increasingly remote services such as hardware security modules, post-quantum KEMs, and inference servers. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Markovian Arrival Process Parameter Estimation of Quasi-birth-death Queueing Systems with Utilization Data | Chen Li, Junjun Zheng, Hiroyuki Okamura, Tadashi Dohi | 2026-07-02 | 下载 | Parameter estimation for queueing systems is commonly performed using inter-arrival times, waiting times, or queue-length observations. However, such detailed observations are often unavailable in pra... |