2026-08-03
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| LowRank-SSM: Hardware-Software Co-Design for Rank-Reduced Mamba Acceleration on FPGA | Haocheng Xu, Bhardwaj Bhat, Yu-an Chou, Zhiheng Chen, Leyao Han, Yifan Zhang, Ye Qiao, Saptarshi Mitra, Sitao Huang | 2026-08-03 | 下载 | State Space Models(SSMs) such as Mamba and Mamba-2 achieve linear-time autoregressive inference, making them attractive for latency-sensitive and resource-constrained deployment. |
| LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction Extension | Pingqing Zheng, Jiayin Qin, Fuqi Zhang, Zishen Wan, Shang Wu, Yu Cao, Caiwen Ding, Yang Katie Zhao | 2026-08-03 | 下载 | Domain-specific Instruction Set Architecture eXtensions (ISAX) are widely adopted in the RISC-V ecosystem to accelerate emerging workloads, but implementing and validating ISAXes across different core... |
| Oasis: Hiding the Cost of Querying Parquet Files in the Datapath | Jonas Dann, Luca Tagliavini, Gustavo Alonso | 2026-08-03 | 下载 | Cloud-native database systems disaggregate compute and storage resources to improve cost efficiency over traditional monolithic architectures through elasticity and resource pooling. |
| DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and Reorganization | Minnan Pei, Gang Li, Zeyu Zhu, Siting Wang, Junwen Si, Zhuoran Song, Yu Feng, Fangxin Liu, Xiaoyao Liang, Jian Cheng | 2026-08-03 | 下载 | 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: increasing the number of ... |
| LEAP: A Self-Supervised Per-Cycle Toggle Propagation Model Supports Fast, Transferable, and Early Analysis of Layout Power | Wenkai Li, Yuchao Wu, Ziyan Guo, Yao Lu, Wenji Fang, Mengming Li, Zhiyao Xie | 2026-08-03 | 下载 | Accurate power analysis is critical in VLSI design, as it directly impacts power optimization strategies. However, traditional approaches are often hindered by the substantial runtime required for per... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Diameter-Free Distributed Frequency Control for Graph Coloring in the CONGEST Model | Amit Nir, David Peleg | 2026-08-03 | 下载 | This paper presents two randomized proper-coloring algorithms that control color frequencies in the synchronous CONGEST model without paying a diameter-dependent coordination cost. |
| Distributed Algorithms for Near-Equitable Coloring | Amit Nir, David Peleg | 2026-08-03 | 下载 | For an -vertex graph of maximum degree Δ and diameter , an equitable (Δ+1)-coloring is a vertex coloring where the frequency of each color (namely, the number of vertices it colors) are all ... |
| Configurable and Hierarchical Allreduce | Valentino Guerrini, Ke Fan, Sidharth Kumar | 2026-08-03 | 下载 | MPI_Allreduce is among the most performance-critical collectives in large-scale scientific computing and distributed machine learning, yet the small- and medium-message regime remains challenging: lat... |
| AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies | Qiushi Lin, Chaojie Zhang, Íñigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic | 2026-08-03 | 下载 | The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, ... |
| Analyzing GPU Performance in Virtualized Environments: A~Case Study | Adel Belkhiri, Michel Dagenais | 2026-08-03 | 下载 | The graphics processing unit (GPU) plays a crucial role in boosting application performance and enhancing computational tasks. Thanks to its parallel architecture and energy efficiency, the GPU has be... |
| Greedy-Like Defective Coloring: Distributed Algorithms and Applications | Marc Fuchs, Fabian Kuhn | 2026-08-03 | 下载 | A -defective -coloring of a graph is a coloring of the nodes with colors such that every node has at most neighbors of the same color. |
| Epico: Long-Lived WebAssembly Components for High-Performance Serverless Stream Processing | Matteo Della Bartola, Valerio Besozzi, Patrizio Dazzi, Marco Danelutto | 2026-08-03 | 下载 | While serverless computing is popular, its dominant Function-as-a-Service (FaaS) model is ill-suited for stream processing because its stateless, centrally orchestrated functions cannot efficiently ha... |
| Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair Scheduling | Dayi Yao, Zijie Zhou | 2026-08-03 | 下载 | This paper studies a resource-allocation inefficiency in batched large language model (LLM) serving: heterogeneous requests that share a decode batch impose max-driven computational costs on one anoth... |
| Fast Discovery of Inclusion Dependencies with Desbordante | Alexander Smirnov, Anton Chizhov, Ilya Shchuckin, Nikita Bobrov, George Chernishev | 2026-08-03 | 下载 | Inclusion dependency is a relation between attributes of tables that indicates possible Primary Key-Foreign Key references. Automatic discovery of inclusion dependencies is a relevant problem for both... |
| DEFT: Joint Task Placement and DVFS for Energy-Efficient Multi-GPU Runtimes | Jing Chen, Miquel Pericas | 2026-08-03 | 下载 | Energy efficiency has become a first-order concern in modern high-performance computing systems, as it directly determines achievable throughput under fixed power budgets. |
| Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO | Ngoc Hung Nguyen, Bjorn Landfeldt | 2026-08-03 | 下载 | This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints. |
| TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving | Mengmeng Zhu, Yuxuan Sun, Wei Chen, Bo Ai | 2026-08-03 | 下载 | The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and re... |
| Diagnosing High-Performance BFT Consensus via Mixture Modeling of Block Time Distributions | Hongru He, Akihiro Fujihara | 2026-08-03 | 下载 | High-performance Byzantine Fault Tolerant (BFT) blockchains are designed to achieve high throughput and low latency, yet their observed block time distributions often reveal complex behaviors arising ... |
| Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling | Cunchen Hu, Liangliang Xu, Tian Liu, Min Lyu, Yongkun Li, Sa Wang, Shuo Quan, Yanan Yang, Wenda Tang, Yiduo Wang, Fu Yu, Jie Wu | 2026-08-03 | 下载 | Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. |
| FedJigsaw: Multi-Agent Collaborative Model Reassembly for Decentralized Heterogeneous Federated Learning | Jifeng Chen, Haibo Zhang, Yawen Chen | 2026-08-03 | 下载 | Model Heterogeneous Federated Learning (MHFL) addresses client-level resource heterogeneity by allowing each participant to train a personalized model architecture under a shared training objective. |
| HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model Serving | Yuning Zhang, Dong Yuan | 2026-08-03 | 下载 | Omni models unify text, speech, image, and multimodal reasoning in a single serving backend, but this unified deployment exposes a new scheduling problem. |
| LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing | Wen Zan, Jiaqi Zhang, Jianchao Tan, Hong Liu, Cunguang Wang, Xiang Li, Duyue Ma, Guanyu Wu, Yifan Lu, Fengcun Li, Yerui Sun, Peng Pei, Yuchen Xie, Xunliang Cai | 2026-08-03 | 下载 | DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive scoring ove... |
| Preserving Admission Responsibility in Multi-Tenant Large Language Model Prefix Caches | Zhiyu Wang, Rajkumar Buyya | 2026-08-03 | 下载 | Shared prefix caching turns Graphics Processing Unit (GPU) memory into persistent state shared across Large Language Model (LLM) tenants. A group that materializes new Key-Value (KV) blocks can force ... |
| PrefixPlace: Provable Prefix Key-Value Placement for Large Language Model Serving under Heterogeneous Compute and Transfer Costs | Zhiyu Wang, Rajkumar Buyya | 2026-08-03 | 下载 | Prefix Key-Value (KV) reuse avoids repeated prefill in Large Language Model (LLM) inference, but local misses require recomputation or replica fetches. |
| Bole: Efficient Tree Speculation for Hybrid-Attention Language Models | Li Wang, Yi Su, Xiabao Wu, Chiran You, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng, Fangxin Liu, Jie Zhang, Chen Tian, Chengying Huan | 2026-08-03 | 下载 | Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound. |
| Source-Bounded Exact Recovery over Docker's Logs API | Kelvin Amoaba | 2026-08-03 | 下载 | Docker can retain records that a collector misses before attachment or during downtime. A persisted read position does not by itself ensure recovery after lifecycle changes. |
| Meganeura: Portable GPU Training and Inference through Vulkan and Metal | Dzmitry Malyshau | 2026-08-03 | 下载 | Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries. Meganeura asks whether one compact native compiler can span both phases on consumer GPUs. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Age of Information in Non-Terrestrial Networks with Energy Harvesting | Fangming Zhao, Nikolaos Pappas, Shi Jin, Howard H. Yang | 2026-08-03 | 下载 | We analyze the timeliness of status-update delivery in a low Earth orbit (LEO) satellite-assisted energy-harvesting Internet of Things network using the Age of Information (AoI) metric. |
| Dynamic Modeling of Target Cell Location for Mobility Robustness Analysis in Cellular Networks: Technical Report | Kiichi Tokuyama | 2026-08-03 | 下载 | Mobility robustness optimization (MRO) requires an appropriate selection of handover (HO) parameters such as the time-to-trigger (TTT) and the offset margin to balance HO failures and ping-pong HOs. |
| In-Network Market Prediction Using Machine Learning and Limit Order Books | Xinpeng Hong, Changgang Zheng, Joshua Lilley, Stefan Zohren, Noa Zilberman | 2026-08-03 | 下载 | Machine learning is significantly transforming algorithmic trading, yet the requirement for rapid execution speeds persists. While both aspects aim to boost profitability, embedding advanced machine-l... |
| Broadcast Rate Limits in Wi-Fi: A Forgotten Bottleneck for Collaborative Edge LLM Inference | Liujianfu Wang, Yuyang Du, Shiqi Xu, Soung Chang Liew | 2026-08-03 | 下载 | LLM deployment is migrating from data centers to edge devices, where Mixture-of-Experts (MoE) models offer a promising path: sparse expert activation allows the model to be spread across multiple low-... |
| A Spatio-Temporal Model for Information Freshness in Massive Random Access | Andrea Munari, Alessandro Buratto, Federico Chiariotti, Leonardo Badia, Petar Popovski | 2026-08-03 | 下载 | Massive connectivity, a key building block of 5G, is expected to play an important role in the next generation of wireless systems, and its expected requirements are being revolutionized through the m... |
| TurboRetry: Mitigating Large-Scale QUIC Handshake Floods with Off-the-Shelf DPU Offloading | Jiahao Wu, Heng Pan, Kai Lv, Zhenyu Li, Yanbiao Li, Gaogang Xie | 2026-08-03 | 下载 | The modern transport protocol QUIC is designed to enhance network performance and security, but it remains vulnerable to handshake flooding attacks. |
| When Discovery Becomes a Storm: A ROS 2 Discovery Model for Wireless Robotic Networks | Yeonwoo Choi, Sanghoon Lee, Kyung-Joon Park | 2026-08-03 | 下载 | In Robot Operating System 2 (ROS 2), Data Distribution Service (DDS) participants must discover one another before exchanging data. In wireless environments, delayed or lost discovery messages cause r... |
| Measuring Post-Quantum TLS Deployment Across UK Internet Sectors | Konstantinos Loizou, Essam Ghadafi | 2026-08-03 | 下载 | Post-quantum cryptography (PQC) is becoming an important component of long-term trust in Internet-facing infrastructure. Publicly observable PQC support provides evidence of externally visible deploym... |
| Emulation vs Simulation: A Case Study from Congestion Control Algorithms in Low Earth Orbit Satellite Networks | Aiden Valentine, Mihai Mazilu, James Knowles, Ian Wakeman, George Parisis | 2026-08-03 | 下载 | Evaluating congestion control is inherently challenging because performance depends on the interaction between the congestion-control algorithm, transport stack, application behaviour, measurement pro... |
| Energy-Latency Trade-offs in O-RAN with Distributed Baseband Processing and AI Inference | Urooj Tariq, Rishu Raj, Shashi Raj Pandey, Merim Dzaferagic, Petar Popovski, Dan Kilper | 2026-08-03 | 下载 | The Open Radio Access Network (O-RAN) architecture introduces flexible functional splits and open interfaces that enable distributed and centralized deployment of baseband processing. |
| Cross-Layer Optimization and System-Level Design of Next-Generation Wireless Networks via Intelligent RAN Control | Maria Tsampazi | 2026-08-03 | 下载 | Recent years have seen the evolution of the traditional Radio Access Network (RAN) toward more open, programmable, disaggregated, and intelligent architectures, known as an Open RAN. |
| Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPO | Ngoc Hung Nguyen, Bjorn Landfeldt | 2026-08-03 | 下载 | This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints. |
| TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous Driving | Mengmeng Zhu, Yuxuan Sun, Wei Chen, Bo Ai | 2026-08-03 | 下载 | The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and re... |
| Predictive Exposure and Cryptographic Readiness: A Vendor-Neutral Framework, a, Bounded Multivocal Evidence Analysis, and Reproducible Synthetic Evaluation for SD-WAN Environments | Saeed Alam | 2026-08-03 | 下载 | SD-WAN teams often use static severity scores to decide what to fix first. These scores do not show live exploitation, network exposure, attack paths, business impact, or cryptographic migration risk. |
| CENTILE: A Telemetry Foundation Model Evaluated by the Decisions It Drives | Zifan Zhang, Zhichao Hou, Tingxiang Ji, Yuchen Liu | 2026-08-03 | 下载 | Modern computing and networking infrastructure emits telemetry continuously, yet operators convert it into decisions with a separate predictor per task, entity, and horizon. |
| On Topology's Role in ML Training Performance | Sarah McClure, Tegan Wilson, Brad Karp, Michael Mitzenmacher, Sylvia Ratnasamy, Scott Shenker, Minlan Yu | 2026-08-03 | 下载 | Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basic topologies: the fa... |
| LEO-Aware DRL Meta-Scheduler for 5G Non-Terrestrial Network Slicing | Víctor Vilchez, Tiago P. C. de Andrade, Edward Hinojosa, Edmundo Madeira, and Carlos A. Astudillo | 2026-08-03 | 下载 | The integration of Low Earth Orbit (LEO) Non-Terrestrial Networks (NTNs) into 5G and upcoming 6G architectures introduces various challenges, including severe propagation delays, ultra-high base stati... |
| LLM-Driven Automated Reward Design for Reinforcement Learning-Based Routing in LEO Satellite Networks | Walter P. Casas, Nelson L. S. da Fonseca, and Carlos A. Astudillo | 2026-08-03 | 下载 | Routing in Low Earth Orbit (LEO) satellite networks is challenging due to highly dynamic topologies and spatio-temporal network conditions. Reinforcement Learning (RL) has emerged as a promising appro... |
| Sensitivity-driven Adaptive Contention Window Optimization for IEEE 802.11 based V2I Networks | Aytül Bozkurt | 2026-08-03 | 下载 | In vehicle-to-infrastructure (V2I) communication the setting of IEEE 802.11 Distributed Coordination Function (DCF) parameters has a decisive bearing on performance, yet the literature seldom pins dow... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies | Qiushi Lin, Chaojie Zhang, Íñigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic | 2026-08-03 | 下载 | The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, ... |
| Mutate to Bypass: Autonomous Endpoint Evasion via Knowledge-Driven Multi-Agent Orchestration | Weifeng Yuan, Wenbo Guo, Qingyun Du, Jun Chen, Feng Dong, Haoyu Wang, Yang Liu | 2026-08-03 | 下载 | Public reports and open-source resources expose many EDR evasion techniques, but it remains unclear whether commercial Endpoint Detection and Response (EDR) systems can withstand these documented atta... |
| Source-Bounded Exact Recovery over Docker's Logs API | Kelvin Amoaba | 2026-08-03 | 下载 | Docker can retain records that a collector misses before attachment or during downtime. A persisted read position does not by itself ensure recovery after lifecycle changes. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Analyzing GPU Performance in Virtualized Environments: A~Case Study | Adel Belkhiri, Michel Dagenais | 2026-08-03 | 下载 | The graphics processing unit (GPU) plays a crucial role in boosting application performance and enhancing computational tasks. Thanks to its parallel architecture and energy efficiency, the GPU has be... |
| FastGFDs: Efficient Validation of Graph Functional Dependencies with Desbordante | Anton Chernikov, Yurii Litvinov, Kirill Smirnov, George Chernishev | 2026-08-03 | 下载 | Graph functional dependencies (GFD) are a recently-developed concept aimed at capturing both topological structures in graphs and functional dependencies between attributes. |
| Fast Discovery of Inclusion Dependencies with Desbordante | Alexander Smirnov, Anton Chizhov, Ilya Shchuckin, Nikita Bobrov, George Chernishev | 2026-08-03 | 下载 | Inclusion dependency is a relation between attributes of tables that indicates possible Primary Key-Foreign Key references. Automatic discovery of inclusion dependencies is a relevant problem for both... |
| HiResNets: Native Full-HD Video Recognition with Foveal Residual Streams | Shivani Mall, Swarnim Jain, Joao F. Henriques | 2026-08-03 | 下载 | Much of the recent progress in image and video recognition has come at the cost of memory: larger models, increased resolution, and longer temporal contexts. |
| TELLER: Non-intrusive Cross-Layer Root-Cause Analysis for LLM Inference | Ruilin Xu, Junyi Li, Pengfei Chen, Zongxuan Xie | 2026-08-03 | 下载 | Large language model (LLM) inference has evolved from an offline workload into a continuously operated software service, yet root-cause analysis remains difficult because a single request spans the in... |
| Diagnosing High-Performance BFT Consensus via Mixture Modeling of Block Time Distributions | Hongru He, Akihiro Fujihara | 2026-08-03 | 下载 | High-performance Byzantine Fault Tolerant (BFT) blockchains are designed to achieve high throughput and low latency, yet their observed block time distributions often reveal complex behaviors arising ... |
| Sensitivity-driven Adaptive Contention Window Optimization for IEEE 802.11 based V2I Networks | Aytül Bozkurt | 2026-08-03 | 下载 | In vehicle-to-infrastructure (V2I) communication the setting of IEEE 802.11 Distributed Coordination Function (DCF) parameters has a decisive bearing on performance, yet the literature seldom pins dow... |