Skip to content

2026-09-26 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Bandwidth, Not FLOPS: FFT Kernels, Matrix Units and SAR Imaging on Apple M6Mohamed Amine Bergach2026-09-26下载The fast Fourier transform (FFT) underlies radar, imaging and scientific computing. A classic rule for fast GPU FFTs is to compute the largest block that fits on chip and compose larger transforms fro...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Tessera: Demand-Driven KV Cache Management for Retrieval-Augmented LLM ServingFei Fang, Chung-Hsiang Lo, Yi Liu, Yifan Hua, Chen Qian2026-09-26下载RAG and retrieval-based agent memory both inject retrieved content into LLM prompts, as document chunks and recalled memory records, respectively.
Simple and Fast Signature-Free Blockchain ConsensusGiuliano Losa, Xuechao Wang, Zhuolun Xiang, Qianyu Yu2026-09-26下载Signature-free protocols avoid the cost of post-quantum signatures. We present two simple signature-free blockchain consensus protocols for eventual synchrony with optimal good-case commit latency (th...
Analyzing 10 Petabit/s Network Data with Accelerated Associative (Token) ArraysJeremy Kepner, Hayden Jananthan, LaToya Anderson, William Arcand, David Bestor, William Bergeron, Chansup Byun, Alex Bonn, Daniel Burrill, Vijay Gadepally, Michael Houle, Matthew Hubbell, Michael Jones, Piotr Luszczek, Peter Michaleas, Lauren Milechin, Julie Mullen, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee, Alex Pentland2026-09-26下载As networks expand and become an ever more critical infrastructure to modern society the need to analyze these networks with the highest regard for privacy is essential to ensure their proper function...
Time Semantics and Liveness Artifacts in Adversarial Consensus SimulationNasit S Sony, Xianzhong Ding2026-09-26下载Adversarial consensus simulators require not only a model of message delivery but also a model of time. For timeout-sensitive protocols, coupling protocol-time progression to scheduler activity can al...
ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU WorkstationsJ. Paul Liu, Uthpala Herath, Andrew Petersen2026-09-26下载Traditional scientific computing requires researchers to translate computational intent into environment configuration, resource requests, and executable jobs, then diagnose failures from scheduler st...
Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market DataVincent Maciejewski2026-09-26下载HFT systems are conventionally built as a single-threaded event loop, on the rule that every thread hop adds latency. We test that rule against a measurement study of more than a year of CME market da...
Over-the-Air Federated Learning in Heterogeneous Mobile Wireless NetworksMing Xiang, Nicolò Michelusi, Yonina C. Eldar, Lili Su2026-09-26下载Over-the-air computation has emerged as a scalable and efficient solution for deploying federated learning algorithms in wireless networks by exploiting waveform superposition for simultaneous model a...
Agentic Network Traffic MonitoringManuel Tsoukatos, Hayden Jananthan, Jeremy Kepner2026-09-26下载As the use of agentic artificial intelligence increases in nearly every industry, there exists a widening attack surface. It is necessary to monitor agents to ensure that agents are acting in a way th...
Near-Optimal Distributed Domination in Planar GraphsWojciech Wawrzyniak2026-09-26下载We give a deterministic (8+ε)(8+\varepsilon)-approximation for minimum dominating set on planar graphs in a constant number of rounds of the LOCAL model, for every \varepsilon\>0.
Ask Without Telling: Local SLMs Consult Cloud LLMs Without Revealing Task IntentYanmeng Wang, Yunxuan Li, Shilong Fan, Yuhan Zheng, Tsung-Hui Chang2026-09-26下载As local small language models (SLMs) increasingly collaborate with more capable cloud large language models (LLMs), a natural privacy question arises: Can a local SLM obtain cloud LLM guidance while ...
Trapped by Their Own Rollouts: Understanding Aggregation--Rollout Feedback in Federated On-Policy DistillationJinqian Chen, Jihua Zhu, Chang Liu2026-09-26下载On-policy distillation (OPD) is a promising approach to language-model adaptation, aligning teacher supervision with the student's own generated trajectories.
SCLATE: a Substrate for Continual-Learning Agent Training and EvaluationYoungmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang, Manjot Bilkhu2026-09-26下载Continual-learning agents are systems of models, harnesses, and memory operating over long multi-session horizons. Evaluating and training them requires interleaving tasks with agent-side events such ...
GPUPHOT: A Python Framework for High-Performance GPU-Accelerated Photometry and Distributed Astronomical Data ReductionSamuel Lemes-Perera, Miguel R. Alarcon, Miquel Serra-Ricart, Pino Caballero-Gil2026-09-26下载We present GPUPHOT, an open-source Python framework for GPU-accelerated real-time photometry and astrometry of astronomical CCD and scientific CMOS images.
AgentReplay: Token-Wise Trace Replay Is Essential for Fair Serving System Performance BenchmarkingZaifeng Pan, Michael Wang, Chris Wu, Zhengding Hu, Xinwei Qiang, Zhongkai Yu, Yufei Ding2026-09-26下载LLM-based agents execute multi-turn workflows with interleaved model inference and tool calls, making efficient serving increasingly important.
RR-Evict: Fine-Grained Prefix Cache Eviction beyond LRU for Agentic LLM ServingZaifeng Pan, Chris Wu, Zhengding Hu, Xinwei Qiang, Zhongkai Yu, Yufei Ding2026-09-26下载LLM-based agents execute long-horizon tasks through repeated model calls interleaved with tool execution and user interaction. As each call extends the history accumulated in previous turns, prefix ca...
Federated Subspace Guided Vision-Language-Action Policy Distillation for Non-IID Multi-Robot ManipulationBiprodip Pal, Kaushik Roy, Yanming Zhu, Brendan Tidd, Alan Wee-Chung Liew, Peyman Moghadam2026-09-26下载Federated learning offers a natural way for multiple robots to jointly improve manipulation policies without requiring centralized access to training demonstrations.
No-Restart Elasticity in an Adaptive Runtime System for Cloud-Native HPCAditya Bhosale, Laxmikant Kale2026-09-26下载Exploiting discounted spot instances for HPC requires an application to change its resource allocation at runtime, shrinking ahead of an interruption and expanding onto replacement capacity.
SparSP: Exploiting Communication Sparsity for Sequence-Parallel Video DiTsDesen Sun, Xinrui Zhong, Yuke Wang, Sihang Liu2026-09-26下载Diffusion Transformers have become the dominant architecture for video generation. Their substantial computational cost motivates scaling inference across multi-GPU servers, yet efficient scaling rema...
Federated 3D Gaussian Splatting for Large-Scale Scene Reconstruction at Wireless EdgeGuanlin Wu, Chao Hu, Pu Chen, Juyong Zhang, Han Hu, Shuguang Cui, Jie Xu2026-09-26下载Three-dimensional (3D) Gaussian splatting (3D-GS) has emerged as a promising technique for large-scale scene reconstruction due to its high rendering efficiency and fidelity.
ThreadShift: Transparent Thread-Level Offloading on Transient Cloud Resources Using MPKsAntoine Murat, Clément Burgelin, Rachid Guerraoui2026-09-26下载Transient cloud resources offer significant cost savings, but their unpredictability makes them hard to use for applications that cannot be safely restarted after reclamation.
REBASE: Device-Cloud Experience Coherence for GUI Agents Across App UpdatesBeining Wu, Jun Huang, Yanxiao Zhao2026-09-26下载A graphical user interface (GUI) agent that ships on a phone runs a small model and reuses experience: action paths recorded on earlier runs and cached from the cloud.
Empowering Hybrid Attention Models on NPUsYinyuan Zhang, Daliang Xu, Xiaolong Huang, Wangsong Yin, Yun Ma, Mengwei Xu, Gang Huang2026-09-26下载Hybrid attention models have emerged as a crucial architecture for Large Language Models (LLMs) (e.g., the Qwen3.5 and Kimi series). Their memory and computational efficiency make them highly attracti...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Analyzing 10 Petabit/s Network Data with Accelerated Associative (Token) ArraysJeremy Kepner, Hayden Jananthan, LaToya Anderson, William Arcand, David Bestor, William Bergeron, Chansup Byun, Alex Bonn, Daniel Burrill, Vijay Gadepally, Michael Houle, Matthew Hubbell, Michael Jones, Piotr Luszczek, Peter Michaleas, Lauren Milechin, Julie Mullen, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee, Alex Pentland2026-09-26下载As networks expand and become an ever more critical infrastructure to modern society the need to analyze these networks with the highest regard for privacy is essential to ensure their proper function...
Adaptive and Resilient Dual-Layer Resource Slicing for Hovering Aerial Backhaul NetworksChuan-Chi Lai, Jen-Hsiang Li2026-09-26下载This paper investigates adaptive and resilient dual-layer resource slicing in hovering aerial agent (HAA)-assisted backhaul networks for heterogeneous 5G/6G services, including enhanced mobile broadba...
Agentic Network Traffic MonitoringManuel Tsoukatos, Hayden Jananthan, Jeremy Kepner2026-09-26下载As the use of agentic artificial intelligence increases in nearly every industry, there exists a widening attack surface. It is necessary to monitor agents to ensure that agents are acting in a way th...
Equilibrium Joining Strategies for Queues in Two-Phase Random EnvironmentKonstantin Avrachenkov, Uri Yechiali2026-09-26下载We study equilibrium joining strategies in an M/M/1-type queueing system with strategic customers operating in a two-phase random environment described as a continuous-time Markov process.
Toward Agentic Optical Networks: A Vision of LLM Agent-Driven Autonomous Lifecycle ManagementYao Zhang, Shengnan Li, Yuchen Song, Yidi Wang, Yue Pang, Wenbin Chen, Xiaotian Jiang, Xiao Luo, Meixia Fu, Min Zhang, Yongli Zhao, Shanguo Huang, Alan Pak Tao Lau, Danshi Wang2026-09-26下载As optical networks continue to expand in scale, complexity, and service diversity, the implementation of automation has become essential for ensuring agility, efficiency, and reliability in lifecycle...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
ThreadShift: Transparent Thread-Level Offloading on Transient Cloud Resources Using MPKsAntoine Murat, Clément Burgelin, Rachid Guerraoui2026-09-26下载Transient cloud resources offer significant cost savings, but their unpredictability makes them hard to use for applications that cannot be safely restarted after reclamation.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Analyzing 10 Petabit/s Network Data with Accelerated Associative (Token) ArraysJeremy Kepner, Hayden Jananthan, LaToya Anderson, William Arcand, David Bestor, William Bergeron, Chansup Byun, Alex Bonn, Daniel Burrill, Vijay Gadepally, Michael Houle, Matthew Hubbell, Michael Jones, Piotr Luszczek, Peter Michaleas, Lauren Milechin, Julie Mullen, Andrew Prout, Albert Reuther, Antonio Rosa, Charles Yee, Alex Pentland2026-09-26下载As networks expand and become an ever more critical infrastructure to modern society the need to analyze these networks with the highest regard for privacy is essential to ensure their proper function...
Packets, Transactions and Queues: Design Principles for HFT Systems from a Measurement Study of CME Market DataVincent Maciejewski2026-09-26下载HFT systems are conventionally built as a single-threaded event loop, on the rule that every thread hop adds latency. We test that rule against a measurement study of more than a year of CME market da...
Change the Product, Keep the Parameters: Associative Algebra Layers for TransformersIlya Koziev, Ivan Oseledets2026-09-26下载Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper produc...
FA-Bench: A Benchmark for Word-Level and Phone-Level Forced-Alignment and ASR Timestamps Under Clean and Noisy ConditionsWei Chu, Yuanzhe Dong, Ke Tan, Dong Han, Yichao Zhou, Ruchao Fan, Bingshen Mu, Jingbei Li, Vishwas Shetty, Sarthak Bisht, Ziyue Qiu, Massa Baali, Rita Singh, Bhisha Raj2026-09-26下载Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so thei...
Bandwidth, Not FLOPS: FFT Kernels, Matrix Units and SAR Imaging on Apple M6Mohamed Amine Bergach2026-09-26下载The fast Fourier transform (FFT) underlies radar, imaging and scientific computing. A classic rule for fast GPU FFTs is to compute the largest block that fits on chip and compose larger transforms fro...
SparSP: Exploiting Communication Sparsity for Sequence-Parallel Video DiTsDesen Sun, Xinrui Zhong, Yuke Wang, Sihang Liu2026-09-26下载Diffusion Transformers have become the dominant architecture for video generation. Their substantial computational cost motivates scaling inference across multi-GPU servers, yet efficient scaling rema...

基于 VitePress 构建 · 使用本地搜索查找论文