Skip to content

2026-09-20 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
SPLASH: Co-Designing Sparse Attention with High-Bandwidth Flash for Efficient Long-Context InferenceAditya Anirudh Jonnalagadda, Agasthi Haputhanthri, Pranav Dangi, Rohan Juneja, Wenshuo Yue, Aritra Bagchi, Bin Gao, Tulika Mitra2026-09-20下载The key-value (KV) cache has become the dominant consumer of memory in large language model (LLM) serving systems as context lengths, concurrency, and request lifetimes grow.
VSpector: Specification-Driven Bug Detection for RISC-V CPUsTianyu Jia, Zhaoyang Yu, Yuanliang Chen, Wei You, Jianjun Huang, Bin Liang2026-09-20下载Detecting RTL design bugs in open-source RISC-V CPU implementations is critical for ensuring system reliability. Traditional detection approaches inherently rely on predefined artifacts.
WaveletECO: A Closed-Loop Physical ECO Platform and a Specialized Local Language ModelGuoxiang Xu, Guozhen Ji, Zijian Luo, Zhengrui Chen, Qi Sun, Cheng Zhuo2026-09-20下载Engineering change order (ECO) is an important step in repairing timing and electrical violations during the late stages of chip design. Existing Agentic EDA methods primarily focus on tool invocation...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Adaptive Determinantal Client Scheduling in Federated LearningWen Xu, Ben Liang, Gary Boudreau, Hamza Sokun2026-09-20下载Scheduling clients for model training is critical in federated learning due to both data and system heterogeneity. Most previous works focus on the quality of the scheduled clients to achieve faster c...
SAGE: Optimal-Stopping Peer Selection for Decentralised Federated LearningKe Xiao, Qiyuan Wang, Christos Anagnostopoulos2026-09-20下载Decentralised federated learning replaces server aggregation with peer-to-peer model exchange, making collaborator selection a local decision under uncertainty.
TriFleetRCA: On-Premise LLM Root Cause Analysis for KubernetesRohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary2026-09-20下载Root cause analysis at a remote site is slow: evidence is scattered across pod logs, Kubernetes events and cluster-level objects, and many operators cannot send production logs to a hosted model at al...
Explicit State and Resource Contracts for Low-Precision Pipeline Parallel Training under Captured GraphsGenlang Chen, Junyi Zhu2026-09-20下载CUDA Graphs eliminate launch overheads by replaying tensor operations over static virtual addresses. However, FP8 pipeline training continuously alters the scaling states, microbatches, and deferred b...
Conflicting Pattern Formation by Teams of Anonymous, Fully Disoriented RobotsAnimesh Maiti, Prakhar Shukla, Subhash Bhagat2026-09-20下载Two groups of autonomous, anonymous, and oblivious mobile robots are deployed in the two-dimensional Euclidean plane, each assigned a distinct task.
Economical and efficient big data sharing with i-CloudThepparit Banditwattanawong, Masawee Masdisornchote, Putchong Uthayopas2026-09-20下载Big data can be hosted on cloud and being shared distributedly through cloud services in an unprecedented volume, variety and velocity. This causes not only cloud network congestions and delayed cloud...
Co-occurrence Patterns of LoRA Adapters in Production Diffusion Model Inference ServicesTao Zhang, Bin Liao, Tao Zhou, Yanping Liu2026-09-20下载Low-rank adaptation (LoRA) has become a key technology for serving large-scale personalized large language models and diffusion models in the cloud.
Accurate Distributed Tracing for Large-Scale AI Infrastructure: Time Synchronization as a Foundation for Reliable ObservabilityHesham Elbakoury, Ankur Sharma2026-09-20下载Distributed tracing in large-scale AI infrastructure fails silently when clock accuracy is insufficient: causal events are misordered, fault attribution is corrupted, and performance diagnoses are unr...
Accurate Simulation of Distributed Training Jobs with Network Contention ModelingYeonho Yoo, Hyunho Lee, Hyunmok Choi, Chuck Yoo, Gyeongsik Yang2026-09-20下载Trace-driven simulation is widely used to evaluate distributed training (DT) jobs in GPU clusters, but existing simulators either ignore network contention or approximate it with a fixed penalty.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Remote IoT Source Monitoring with Delayed FeedbackAndrea Munari2026-09-20下载Remote source monitoring is a key use case for Internet of things (IoT) applications, calling on efficient communications protocols to ensure timely data delivery.
SemDHT: Certified Semantic Discovery for Peer-to-Peer Agent Networks over Exact-Key DHTsTaotao Wang, Chonghe Zhao, Shengli Zhang, Soung Chang Liew2026-09-20下载Agents may need capabilities exposed through external agent endpoints or service APIs. When a requester is not already bound to a provider, it must discover advertised capabilities matching its task a...
Feature Suppression and Differential Privacy for Residential Traffic Classification: A Two-Home Federated StudyMárton Pál Lipcsey-Magyar, Adrian Pekar2026-09-20下载Residential traffic classification supports service management, but learning across homes must account for heterogeneous traffic and privacy constraints.
Protocol-Flexible Custom NFC for Wire-Free Wearable Sensor NetworksRiku Maeda, Akihito Noda2026-09-20下载This study presents a custom near-field communication (NFC) system with software-defined protocol implementation for wearable sensors distributed across the body.
Accurate Simulation of Distributed Training Jobs with Network Contention ModelingYeonho Yoo, Hyunho Lee, Hyunmok Choi, Chuck Yoo, Gyeongsik Yang2026-09-20下载Trace-driven simulation is widely used to evaluate distributed training (DT) jobs in GPU clusters, but existing simulators either ignore network contention or approximate it with a fixed penalty.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
When the Agent Becomes the Kernel: A Systematization of Security on the Path to AI-Native Operating SystemsLi Zhang, Yang Sun, Jie Shi2026-09-20下载Large language model agents are now privileged principals that take consequential actions: editing code repositories, operating inboxes, completing purchases.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM WorkflowsVivek Kumar Singh, Preeti Priyam, Gautam Bhowmick2026-09-20下载Every node in a multi-agent large language model (LLM) workflow retrieves context from memory and injects it into its prompt, where those injected tokens are billed as input tokens at the same per-tok...

基于 VitePress 构建 · 使用本地搜索查找论文