Skip to content

2026-07-02 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
HyNoC: A Hybrid Circuit-Switch/Wormhole Network-on-Chip for Distributed VLIW Computing on FPGAChristophe Clienti2026-07-02下载Network-on-Chip (NoC) architectures have become the standard interconnect fabric for many-core systems, yet most proposals face a fundamental trade-off between latency, area, and congestion management...
Probabilistic Memory for Trustworthy Edge IntelligenceLikai Pei, Jiahao Zheng, Xueji Zhao, Emilie Ye, Jianbo Liu, Hanqing Tao, Ming-Yen Lee, Ruiyang Qin, Yiyu Shi, Shimeng Yu, X. Sharon Hu, Ningyuan Cao2026-07-02下载Probabilistic computation plays an important role in trustworthy edge intelligence to quantify uncertainty, enhance robustness, reconstruct data, and protect privacy, but its adoption is limited by th...
APEIRON: composing smart TDAQ systems for high energy physics experimentsRoberto Ammendola, Andrea Biagioni, Carlotta Chiarini, Andrea Ciardiello, Paolo Cretaro, Ottorino Frezza, Francesca Lo Cicero, Alessandro Lonardo, Michele Martinelli, Pier Stanislao Paolucci, Pierpaolo Perticaroli, Cristian Rossi, Francesco Simula, Matteo Turisini, Piero Vicini2026-07-02下载We present APEIRON, a distributed heterogeneous processing framework comprising both hardware architecture and software stack for multi-FPGA systems.
A 2048-spin bulk acoustic wave Ising machine for number partitioning and SudokuVenkatesh Vadde, Roman Ovcharov, Victor H. González, Roman Khymyn, Artem Litvinenko, Johan Åkerman2026-07-02下载Optical coherent Ising machines based on time-multiplexing have demonstrated significant progress in terms of connectivity and spin scalability.
Approximate Attention Weighting for Sustainable FPGA-Based Vision Transformer InferenceMuhammad Usman, Muhammad Akmal Shafique, Shujaat Khan, Dorit Merhof2026-07-02下载Vision Transformers have reshaped computer vision by using self-attention to capture global context across image regions. This makes them attractive for edge visual inspection and monitoring in applic...
3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM ServingJaehun Lee, In-Jun Jung, Joo-Young Kim2026-07-02下载Large language model (LLM) serving increasingly combines prefill-decode (PD) disaggregation with tensor parallelism (TP) to support large models and long contexts. In conventional 2D/2.
MxGLUT: A Reconfigurable LUT-Centric Broadcast Dataflow Accelerator for Mixed-Precision GEMMWeiyu Zhou, Chen Ding, Mingyuan Liu, Liangyu Gan, Yukun Feng, Hao Jia, Haoming Chu, Lirong Zheng, Ning Ma, Yuxiang Huan2026-07-02下载Large language model (LLM) inference suffers from growing inefficiency across the prefill and decode phases, especially under weight-only quantization, where activations remain in FP8 while weights ar...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Remora: Scale-out Deterministic Execution for Smart ContractsZhengqing Liu, Alberto Sonnino, Igor Zablotchi, Eleftherios Kokoris Kogias, Marios Kogias2026-07-02下载Modern blockchains rely on a modular architecture that decouples consensus from execution. Recent advances in consensus algorithms have shifted the bottleneck to the execution layer, which must determ...
HyNoC: A Hybrid Circuit-Switch/Wormhole Network-on-Chip for Distributed VLIW Computing on FPGAChristophe Clienti2026-07-02下载Network-on-Chip (NoC) architectures have become the standard interconnect fabric for many-core systems, yet most proposals face a fundamental trade-off between latency, area, and congestion management...
LLMoxie: Exploring Agentic AI for Scientific Software DevelopmentLandung Setiawan, Anant Mittal, Cordero Core, Anshul Tambay, Carlos Garcia Jurado Suarez, David A. C. Beck, Andrew J. Connolly, Vani Mandava2026-07-02下载In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise inference, a LiteLLM/MLflow control plane for authentication, budge...
FlintKV: A Fast Durable Storage Engine for Modern DatabasesSergey Egorov, Gregory Chockler, Brijesh Dongol, Dan O'Keeffe, Sadegh Keshavarzi2026-07-02下载Byte-addressable non-volatile memory (NVM) offers an opportunity to rethink storage engine architectures. While recent NVM key-value stores achieve high throughput for ingestion and point lookups, the...
WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMsMauricio Fadel Argerich, Jonathan Fürst, Marta Patiño-Martínez2026-07-02下载Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requires matching specific LLMs to the most efficient G...
Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing DataDaniel M. Jimenez-Gutierrez, Enrique Zuazua, Georgios Kellaris, Joaquin del Rio, Oleksii Sliusarenko, Xabi Uribe-Etxebarria2026-07-02下载Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster response, operational security environments, infrastructure m...
Elasticity in Parallel Sparse Triangular SolveRaphael S. Steiner, Christos K. Matzoros, Pál András Papp, Toni Böhnlein, A. N. Yzelman2026-07-02下载We introduce stale synchronous parallel as a mode of execution in parallel sparse triangular linear system solve and present a general directed-acyclic-graph scheduler capable of producing such schedu...
Securing People and their Machines Against Major FaultsOhad Eitan, Idit Keidar, Ehud Shapiro2026-07-02下载We consider grassroots platforms -- distributed systems of agents consisting of people identified by self-chosen public keys and their machines (smartphones) -- and wish to make them secure against \e...
Cadence: Extreme Pipelining with Multiple Concurrent ProposersFatima Elsheimy, Mohammad Mussadiq Jalalzai, Tobias Klenze, Jovan Komatovic, Mike Setrin, Victor Shoup, Kushal Babel, Lioba Heimbach, Jason Milionis2026-07-02下载We present Cadence, a Byzantine fault-tolerant multi-proposer consensus protocol with arbitrarily low block intervals, optimal resilience, and optimal fast-path latency.
Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of LinesBojie Li2026-07-02下载Hardware accelerators now sit on the critical path of online serving. GPUs, FPGAs, and increasingly remote services such as hardware security modules, post-quantum KEMs, and inference servers.
Towards Load-Aware Prefill Deflection for Disaggregated LLM ServingShrikara Arun, Anjaly Parayil, Srikant Bharadwaj, Renee St. Amant, Victor Rühle2026-07-02下载Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymmetry: under bursty, heavy-tailed workloads prefill...
Scalable and Distributed Silhouette ApproximationIlie Sarpe, Federico Altieri, Andrea Pietracaprina, Geppino Pucci, Fabio Vandin2026-07-02下载The silhouette is one of the most widely used measures to assess the quality of a kk-clustering of a dataset of nn elements. Its evaluation requires no information beyond the clustering assignment.
Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts ModelsXuan-Phi Nguyen, Shrey Pandit, Yiran Zhao, Semih Yavuz, Silvio Savarese, Shafiq Joty2026-07-02下载This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes various existing and novel parallelism techniques at...
Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context InferenceWenchen Han, Gingfung Matthew Yeung, Marco Barletta, William Toner, Amory Hoste, Adam Barker2026-07-02下载Long-context inference is increasingly common in large language model (LLM) serving, driven by retrieval-augmented generation and agentic systems.
HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel AttentionChao Yuan, Pan Li, Yingnan Sun, Jing Liu2026-07-02下载All-to-all based sequence parallelism methods execute communication and computation strictly in serial when processing medium-long sequences, resulting in hardware resource underutilization.
Exploiting Task-Based Parallelism for the Red-Black Gauss-Seidel Method on 2D GridsShiting Long, Gustavo Ramirez-Hidalgo, Andreas Frommer, Dirk Pleiter2026-07-02下载Gauss-Seidel is a well-established iterative method for the solution of linear systems, and multicoloring has been widely used to increase parallelism in iterative solution techniques.
Arachne: Orchestrating Cascades for Efficient Text-to-Video Model TrainingPeng Yu, Yuankai Fan, Yang Qiu, Tian Li, Bihuan Chen, Yin Chen, Qizhen Weng2026-07-02下载The rising demand for AI-generated videos is fueled by advances in large-scale Text-to-Video (T2V) models, trained on extensive datasets of video clips spanning diverse resolutions and durations.
SCAPE: Accurate and Efficient LLM Training with Extreme Sparse CommunicationMingkai Zheng, Junlin Chen, Haotian Xie, Zhao Zhang2026-07-02下载Communication increasingly dominates the cost of Large Language Model (LLM) pre-training, especially under data-parallel and sharded training schemes, where gradient synchronization and parameter reco...
DeadPool: Resilient LLM Training with Hot-Swapping via Zero-Overhead CheckpointHaotian Xie, Junlin Chen, Mingkai Zheng, Lishan Yang, Zhao Zhang2026-07-02下载State-of-the-art large language model (LLM) training takes tens of thousands of graphics processing units (GPUs) for months and encounters failures across the software and hardware stack.
OmniPilot: An Uncertainty-Aware LLM Inference Advisor for Heterogeneous GPU ClustersD. Balamurugan, Thomas W. Bush2026-07-02下载Serving large language models (LLMs) on a shared, heterogeneous GPU cluster requires users and operators to select the GPU type, tensor-parallel degree, and precision before committing valuable node-h...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model ServingJiaying Meng, Bojie Li2026-07-02下载Real-time interaction models -- Moshi, MiniCPM-o, Qwen-Omni -- turn serving into a periodic real-time task: on every frame a session ingests streaming audio and must respond by a recurring wall-clock ...
Criticality-Based Guard Rail Validation for AI Agent Decisions in Autonomous Telecom NetworksRavi Kant Sharma2026-07-02下载The evolution toward fully autonomous telecommunications networks (Autonomous Network Levels 4-5) requires AI/ML agents to make real-time network decisions without human intervention.
CSI Simulation: Why Additive Noise Fails and How to Fix ItAymen Bouferroum, Ildi Alla, Vincent Lenders, Valeria Loscri2026-07-02下载Channel State Information (CSI) has become a widely used wireless channel sensing modality for applications such as indoor localization, activity recognition, and respiration monitoring.
Enabling Real-Time AI in O-RAN: Deploying andMeasuring AI Inside a Near-RT RIC xAppLawrence Obiuwevwi, Krzysztof J. Rechowicz, Sampath Jayarathna, Safdar Hussain Bouk, Fahmida Afrin, C. Nicolas Barati, Neda Moghim, Valentina Nanou, Muhammad Enayetur Rahman, Sachin Shetty2026-07-02下载Open Radio Access Network (O-RAN) architectures introduce programmable Near-Real-Time RAN Intelligent Controllers (Near-RT RICs) that support closed-loop control through xApps at timescales from 10 ms...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Characterizing and Bridging the Diagnostic Gap in eBPF Verifier RejectionsYusheng Zheng, Zhengjie Ji, Weichen Tao, Xiangyu Gao, Jianchang Su, Wei Zhang, Andi Quinn, Dan Williams2026-07-02下载eBPF lets developers run custom programs inside the Linux kernel, where a verifier proves each program safe. However, when the verifier rejects a program, the unclear error makes repair challenging: t...
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous RobotsLing Xu, Chuyu Han, Borui Li, Hao Wu, Shiqi Jiang, Ting Cao, Chuanyou Li, Sheng Zhong, Shuai Wang2026-07-02下载Embodied AI models now span vision-language-action (VLA) models and world-action models (WAMs), but practical deployment remains fragmented across model-specific Python stacks, backend assumptions, an...
Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of LinesBojie Li2026-07-02下载Hardware accelerators now sit on the critical path of online serving. GPUs, FPGAs, and increasingly remote services such as hardware security modules, post-quantum KEMs, and inference servers.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Markovian Arrival Process Parameter Estimation of Quasi-birth-death Queueing Systems with Utilization DataChen Li, Junjun Zheng, Hiroyuki Okamura, Tadashi Dohi2026-07-02下载Parameter estimation for queueing systems is commonly performed using inter-arrival times, waiting times, or queue-length observations. However, such detailed observations are often unavailable in pra...

基于 VitePress 构建 · 使用本地搜索查找论文