Skip to content

2026-08-03 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
LowRank-SSM: Hardware-Software Co-Design for Rank-Reduced Mamba Acceleration on FPGAHaocheng Xu, Bhardwaj Bhat, Yu-an Chou, Zhiheng Chen, Leyao Han, Yifan Zhang, Ye Qiao, Saptarshi Mitra, Sitao Huang2026-08-03下载State Space Models(SSMs) such as Mamba and Mamba-2 achieve linear-time autoregressive inference, making them attractive for latency-sensitive and resource-constrained deployment.
LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction ExtensionPingqing Zheng, Jiayin Qin, Fuqi Zhang, Zishen Wan, Shang Wu, Yu Cao, Caiwen Ding, Yang Katie Zhao2026-08-03下载Domain-specific Instruction Set Architecture eXtensions (ISAX) are widely adopted in the RISC-V ecosystem to accelerate emerging workloads, but implementing and validating ISAXes across different core...
Oasis: Hiding the Cost of Querying Parquet Files in the DatapathJonas Dann, Luca Tagliavini, Gustavo Alonso2026-08-03下载Cloud-native database systems disaggregate compute and storage resources to improve cost efficiency over traditional monolithic architectures through elasticity and resource pooling.
DeGS: A Scalable 3DGS Architecture via Decoupled Workload Parsing and ReorganizationMinnan Pei, Gang Li, Zeyu Zhu, Siting Wang, Junwen Si, Zhuoran Song, Yu Feng, Fangxin Liu, Xiaoyao Liang, Jian Cheng2026-08-03下载3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time novel view synthesis, yet existing 3DGS accelerators suffer from poor architectural scalability: increasing the number of ...
LEAP: A Self-Supervised Per-Cycle Toggle Propagation Model Supports Fast, Transferable, and Early Analysis of Layout PowerWenkai Li, Yuchao Wu, Ziyan Guo, Yao Lu, Wenji Fang, Mengming Li, Zhiyao Xie2026-08-03下载Accurate power analysis is critical in VLSI design, as it directly impacts power optimization strategies. However, traditional approaches are often hindered by the substantial runtime required for per...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Diameter-Free Distributed Frequency Control for Graph Coloring in the CONGEST ModelAmit Nir, David Peleg2026-08-03下载This paper presents two randomized proper-coloring algorithms that control color frequencies in the synchronous CONGEST model without paying a diameter-dependent coordination cost.
Distributed Algorithms for Near-Equitable ColoringAmit Nir, David Peleg2026-08-03下载For an nn-vertex graph of maximum degree Δ and diameter DD, an equitable (Δ+1)-coloring is a vertex coloring where the frequency of each color (namely, the number of vertices it colors) are all ...
Configurable and Hierarchical AllreduceValentino Guerrini, Ke Fan, Sidharth Kumar2026-08-03下载MPI_Allreduce is among the most performance-critical collectives in large-scale scientific computing and distributed machine learning, yet the small- and medium-message regime remains challenging: lat...
AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane PoliciesQiushi Lin, Chaojie Zhang, Íñigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic2026-08-03下载The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, ...
Analyzing GPU Performance in Virtualized Environments: A~Case StudyAdel Belkhiri, Michel Dagenais2026-08-03下载The graphics processing unit (GPU) plays a crucial role in boosting application performance and enhancing computational tasks. Thanks to its parallel architecture and energy efficiency, the GPU has be...
Greedy-Like Defective Coloring: Distributed Algorithms and ApplicationsMarc Fuchs, Fabian Kuhn2026-08-03下载A dd-defective cc-coloring of a graph G=(V,E)G=(V,E) is a coloring of the nodes VV with cc colors such that every node has at most dd neighbors of the same color.
Epico: Long-Lived WebAssembly Components for High-Performance Serverless Stream ProcessingMatteo Della Bartola, Valerio Besozzi, Patrizio Dazzi, Marco Danelutto2026-08-03下载While serverless computing is popular, its dominant Function-as-a-Service (FaaS) model is ill-suited for stream processing because its stateless, centrally orchestrated functions cannot efficiently ha...
Efficiency and Cost Alignment in Batched LLM Serving via Resource-Fair SchedulingDayi Yao, Zijie Zhou2026-08-03下载This paper studies a resource-allocation inefficiency in batched large language model (LLM) serving: heterogeneous requests that share a decode batch impose max-driven computational costs on one anoth...
Fast Discovery of Inclusion Dependencies with DesbordanteAlexander Smirnov, Anton Chizhov, Ilya Shchuckin, Nikita Bobrov, George Chernishev2026-08-03下载Inclusion dependency is a relation between attributes of tables that indicates possible Primary Key-Foreign Key references. Automatic discovery of inclusion dependencies is a relevant problem for both...
DEFT: Joint Task Placement and DVFS for Energy-Efficient Multi-GPU RuntimesJing Chen, Miquel Pericas2026-08-03下载Energy efficiency has become a first-order concern in modern high-performance computing systems, as it directly determines achievable throughput under fixed power budgets.
Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPONgoc Hung Nguyen, Bjorn Landfeldt2026-08-03下载This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints.
TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous DrivingMengmeng Zhu, Yuxuan Sun, Wei Chen, Bo Ai2026-08-03下载The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and re...
Diagnosing High-Performance BFT Consensus via Mixture Modeling of Block Time DistributionsHongru He, Akihiro Fujihara2026-08-03下载High-performance Byzantine Fault Tolerant (BFT) blockchains are designed to achieve high throughput and low latency, yet their observed block time distributions often reveal complex behaviors arising ...
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency ScalingCunchen Hu, Liangliang Xu, Tian Liu, Min Lyu, Yongkun Li, Sa Wang, Shuo Quan, Yanan Yang, Wenda Tang, Yiduo Wang, Fu Yu, Jie Wu2026-08-03下载Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption.
FedJigsaw: Multi-Agent Collaborative Model Reassembly for Decentralized Heterogeneous Federated LearningJifeng Chen, Haibo Zhang, Yawen Chen2026-08-03下载Model Heterogeneous Federated Learning (MHFL) addresses client-level resource heterogeneity by allowing each participant to train a personalized model architecture under a shared training objective.
HorizonServe: Coordinating Request Scheduling with GPU Sharing for Omni-Model ServingYuning Zhang, Dong Yuan2026-08-03下载Omni models unify text, speech, image, and multimodal reasoning in a single serving backend, but this unified deployment exposes a new scheduling problem.
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer IndexingWen Zan, Jiaqi Zhang, Jianchao Tan, Hong Liu, Cunguang Wang, Xiang Li, Duyue Ma, Guanyu Wu, Yifan Lu, Fengcun Li, Yerui Sun, Peng Pei, Yuchen Xie, Xunliang Cai2026-08-03下载DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrained by the indexer's expensive O(L2)O(L^2) scoring ove...
Preserving Admission Responsibility in Multi-Tenant Large Language Model Prefix CachesZhiyu Wang, Rajkumar Buyya2026-08-03下载Shared prefix caching turns Graphics Processing Unit (GPU) memory into persistent state shared across Large Language Model (LLM) tenants. A group that materializes new Key-Value (KV) blocks can force ...
PrefixPlace: Provable Prefix Key-Value Placement for Large Language Model Serving under Heterogeneous Compute and Transfer CostsZhiyu Wang, Rajkumar Buyya2026-08-03下载Prefix Key-Value (KV) reuse avoids repeated prefill in Large Language Model (LLM) inference, but local misses require recomputation or replica fetches.
Bole: Efficient Tree Speculation for Hybrid-Attention Language ModelsLi Wang, Yi Su, Xiabao Wu, Chiran You, Yongchao Liu, Zhan Qiu, Juelu Zhang, Jiajun Zheng, Fangxin Liu, Jie Zhang, Chen Tian, Chengying Huan2026-08-03下载Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound.
Source-Bounded Exact Recovery over Docker's Logs APIKelvin Amoaba2026-08-03下载Docker can retain records that a collector misses before attachment or during downtime. A persisted read position does not by itself ensure recovery after lifecycle changes.
Meganeura: Portable GPU Training and Inference through Vulkan and MetalDzmitry Malyshau2026-08-03下载Training and deployed inference often cross export, conversion, and platform-specific runtime boundaries. Meganeura asks whether one compact native compiler can span both phases on consumer GPUs.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Age of Information in Non-Terrestrial Networks with Energy HarvestingFangming Zhao, Nikolaos Pappas, Shi Jin, Howard H. Yang2026-08-03下载We analyze the timeliness of status-update delivery in a low Earth orbit (LEO) satellite-assisted energy-harvesting Internet of Things network using the Age of Information (AoI) metric.
Dynamic Modeling of Target Cell Location for Mobility Robustness Analysis in Cellular Networks: Technical ReportKiichi Tokuyama2026-08-03下载Mobility robustness optimization (MRO) requires an appropriate selection of handover (HO) parameters such as the time-to-trigger (TTT) and the offset margin to balance HO failures and ping-pong HOs.
In-Network Market Prediction Using Machine Learning and Limit Order BooksXinpeng Hong, Changgang Zheng, Joshua Lilley, Stefan Zohren, Noa Zilberman2026-08-03下载Machine learning is significantly transforming algorithmic trading, yet the requirement for rapid execution speeds persists. While both aspects aim to boost profitability, embedding advanced machine-l...
Broadcast Rate Limits in Wi-Fi: A Forgotten Bottleneck for Collaborative Edge LLM InferenceLiujianfu Wang, Yuyang Du, Shiqi Xu, Soung Chang Liew2026-08-03下载LLM deployment is migrating from data centers to edge devices, where Mixture-of-Experts (MoE) models offer a promising path: sparse expert activation allows the model to be spread across multiple low-...
A Spatio-Temporal Model for Information Freshness in Massive Random AccessAndrea Munari, Alessandro Buratto, Federico Chiariotti, Leonardo Badia, Petar Popovski2026-08-03下载Massive connectivity, a key building block of 5G, is expected to play an important role in the next generation of wireless systems, and its expected requirements are being revolutionized through the m...
TurboRetry: Mitigating Large-Scale QUIC Handshake Floods with Off-the-Shelf DPU OffloadingJiahao Wu, Heng Pan, Kai Lv, Zhenyu Li, Yanbiao Li, Gaogang Xie2026-08-03下载The modern transport protocol QUIC is designed to enhance network performance and security, but it remains vulnerable to handshake flooding attacks.
When Discovery Becomes a Storm: A ROS 2 Discovery Model for Wireless Robotic NetworksYeonwoo Choi, Sanghoon Lee, Kyung-Joon Park2026-08-03下载In Robot Operating System 2 (ROS 2), Data Distribution Service (DDS) participants must discover one another before exchanging data. In wireless environments, delayed or lost discovery messages cause r...
Measuring Post-Quantum TLS Deployment Across UK Internet SectorsKonstantinos Loizou, Essam Ghadafi2026-08-03下载Post-quantum cryptography (PQC) is becoming an important component of long-term trust in Internet-facing infrastructure. Publicly observable PQC support provides evidence of externally visible deploym...
Emulation vs Simulation: A Case Study from Congestion Control Algorithms in Low Earth Orbit Satellite NetworksAiden Valentine, Mihai Mazilu, James Knowles, Ian Wakeman, George Parisis2026-08-03下载Evaluating congestion control is inherently challenging because performance depends on the interaction between the congestion-control algorithm, transport stack, application behaviour, measurement pro...
Energy-Latency Trade-offs in O-RAN with Distributed Baseband Processing and AI InferenceUrooj Tariq, Rishu Raj, Shashi Raj Pandey, Merim Dzaferagic, Petar Popovski, Dan Kilper2026-08-03下载The Open Radio Access Network (O-RAN) architecture introduces flexible functional splits and open interfaces that enable distributed and centralized deployment of baseband processing.
Cross-Layer Optimization and System-Level Design of Next-Generation Wireless Networks via Intelligent RAN ControlMaria Tsampazi2026-08-03下载Recent years have seen the evolution of the traditional Radio Access Network (RAN) toward more open, programmable, disaggregated, and intelligent architectures, known as an Open RAN.
Learning-Based Collaborative MEC for LLM Inference with Soft-Deadline Awareness via Transformer-Enhanced PPONgoc Hung Nguyen, Bjorn Landfeldt2026-08-03下载This paper investigates collaborative mobile edge computing (MEC) servers for large language model (LLM) inference under soft deadline constraints.
TALSC: Timeliness-Aware Large-Small VLM Collaboration for Infrastructure-Assisted Autonomous DrivingMengmeng Zhu, Yuxuan Sun, Wei Chen, Bo Ai2026-08-03下载The deployment of Vision-Language Models (VLMs) in autonomous driving (AD) systems is constrained by on-board computing power, restricting vehicles to small VLMs (SVLMs) with limited perception and re...
Predictive Exposure and Cryptographic Readiness: A Vendor-Neutral Framework, a, Bounded Multivocal Evidence Analysis, and Reproducible Synthetic Evaluation for SD-WAN EnvironmentsSaeed Alam2026-08-03下载SD-WAN teams often use static severity scores to decide what to fix first. These scores do not show live exploitation, network exposure, attack paths, business impact, or cryptographic migration risk.
CENTILE: A Telemetry Foundation Model Evaluated by the Decisions It DrivesZifan Zhang, Zhichao Hou, Tingxiang Ji, Yuchen Liu2026-08-03下载Modern computing and networking infrastructure emits telemetry continuously, yet operators convert it into decisions with a separate predictor per task, entity, and horizon.
On Topology's Role in ML Training PerformanceSarah McClure, Tegan Wilson, Brad Karp, Michael Mitzenmacher, Sylvia Ratnasamy, Scott Shenker, Minlan Yu2026-08-03下载Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basic topologies: the fa...
LEO-Aware DRL Meta-Scheduler for 5G Non-Terrestrial Network SlicingVíctor Vilchez, Tiago P. C. de Andrade, Edward Hinojosa, Edmundo Madeira, and Carlos A. Astudillo2026-08-03下载The integration of Low Earth Orbit (LEO) Non-Terrestrial Networks (NTNs) into 5G and upcoming 6G architectures introduces various challenges, including severe propagation delays, ultra-high base stati...
LLM-Driven Automated Reward Design for Reinforcement Learning-Based Routing in LEO Satellite NetworksWalter P. Casas, Nelson L. S. da Fonseca, and Carlos A. Astudillo2026-08-03下载Routing in Low Earth Orbit (LEO) satellite networks is challenging due to highly dynamic topologies and spatio-temporal network conditions. Reinforcement Learning (RL) has emerged as a promising appro...
Sensitivity-driven Adaptive Contention Window Optimization for IEEE 802.11 based V2I NetworksAytül Bozkurt2026-08-03下载In vehicle-to-infrastructure (V2I) communication the setting of IEEE 802.11 Distributed Coordination Function (DCF) parameters has a decisive bearing on performance, yet the literature seldom pins dow...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane PoliciesQiushi Lin, Chaojie Zhang, Íñigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic2026-08-03下载The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent, ...
Mutate to Bypass: Autonomous Endpoint Evasion via Knowledge-Driven Multi-Agent OrchestrationWeifeng Yuan, Wenbo Guo, Qingyun Du, Jun Chen, Feng Dong, Haoyu Wang, Yang Liu2026-08-03下载Public reports and open-source resources expose many EDR evasion techniques, but it remains unclear whether commercial Endpoint Detection and Response (EDR) systems can withstand these documented atta...
Source-Bounded Exact Recovery over Docker's Logs APIKelvin Amoaba2026-08-03下载Docker can retain records that a collector misses before attachment or during downtime. A persisted read position does not by itself ensure recovery after lifecycle changes.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Analyzing GPU Performance in Virtualized Environments: A~Case StudyAdel Belkhiri, Michel Dagenais2026-08-03下载The graphics processing unit (GPU) plays a crucial role in boosting application performance and enhancing computational tasks. Thanks to its parallel architecture and energy efficiency, the GPU has be...
FastGFDs: Efficient Validation of Graph Functional Dependencies with DesbordanteAnton Chernikov, Yurii Litvinov, Kirill Smirnov, George Chernishev2026-08-03下载Graph functional dependencies (GFD) are a recently-developed concept aimed at capturing both topological structures in graphs and functional dependencies between attributes.
Fast Discovery of Inclusion Dependencies with DesbordanteAlexander Smirnov, Anton Chizhov, Ilya Shchuckin, Nikita Bobrov, George Chernishev2026-08-03下载Inclusion dependency is a relation between attributes of tables that indicates possible Primary Key-Foreign Key references. Automatic discovery of inclusion dependencies is a relevant problem for both...
HiResNets: Native Full-HD Video Recognition with Foveal Residual StreamsShivani Mall, Swarnim Jain, Joao F. Henriques2026-08-03下载Much of the recent progress in image and video recognition has come at the cost of memory: larger models, increased resolution, and longer temporal contexts.
TELLER: Non-intrusive Cross-Layer Root-Cause Analysis for LLM InferenceRuilin Xu, Junyi Li, Pengfei Chen, Zongxuan Xie2026-08-03下载Large language model (LLM) inference has evolved from an offline workload into a continuously operated software service, yet root-cause analysis remains difficult because a single request spans the in...
Diagnosing High-Performance BFT Consensus via Mixture Modeling of Block Time DistributionsHongru He, Akihiro Fujihara2026-08-03下载High-performance Byzantine Fault Tolerant (BFT) blockchains are designed to achieve high throughput and low latency, yet their observed block time distributions often reveal complex behaviors arising ...
Sensitivity-driven Adaptive Contention Window Optimization for IEEE 802.11 based V2I NetworksAytül Bozkurt2026-08-03下载In vehicle-to-infrastructure (V2I) communication the setting of IEEE 802.11 Distributed Coordination Function (DCF) parameters has a decisive bearing on performance, yet the literature seldom pins dow...

基于 VitePress 构建 · 使用本地搜索查找论文