2026-09-16
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Splyce: SIMD Vectorization of Sparse Coiteration | Kabilan Mahathevan, Poorna Gunathilaka, Kirshanthan Sundararajah | 2026-09-16 | 下载 | Sparse tensor contractions are bottlenecked by sparse-sparse coiteration loops that resist standard loop vectorization. We present Splyce, an auto-vectorization framework in MLIR that overcomes this t... |
| Do AI Agents Understand Computer Architecture? | Ambika Sharan, Grigory Chirkov, Soheil Abbasloo | 2026-09-16 | 下载 | Agents are increasingly asked to design hardware, and increasingly reported to succeed. Such reports establish that a design improved; they cannot establish why. |
| Rosetta: Automating First-Principles Performance Modeling Using Multi-Agent LLMs | Karthikeyan Sankaralingam | 2026-09-16 | 下载 | Analytical performance models --- derivations of throughput or speedup from hardware parameters --- make claims independently verifiable and expose binding constraints, yet rarely accompany architectu... |
| Analog Pin Directionality as an Exfiltration Attack Surface in Mixed-Signal ICs | Ramana Ranganatham, Chirag Adiga, Michael Zuzak, Tejasvi Das | 2026-09-16 | 下载 | Mixed-signal SoCs rely on nominally input-only analog pins to acquire off-chip signals, but the directionality of these interfaces is generally treated as a functional property rather than explicitly ... |
| Locus: A Framework for Exploring and Optimizing Point Addition Hardware for Zero-Knowledge Proofs | Gaurav Kuwar, Alhad Daftardar, Jianqiao Mo, Siddharth Garg, Brandon Reagen | 2026-09-16 | 下载 | Zero-Knowledge Proofs (ZKPs) are critical for privacy-preserving and verifiable computation, but their cryptographic primitives impose high computational overheads. |
| Quantifying the Effect of HCLs on a Fixed-Microarchitecture MXFP4 Accelerator | Daniele Passaretti, Sajjad Tamimi, Nicola Dall'Ora | 2026-09-16 | 下载 | Hardware Construction Languages (HCLs) aim to improve hardware design productivity while generating register-transfer-level (RTL) circuits without changing the designer's microarchitecture. |
| HBFlex: A Flexible Memory System for Bridging Fine-Grained LLM States and Coarse-Grained HBF Parallel Execution | Shuzhang Zhong, Weikai Xu, Yifan Zhou, Tongbin Zhao, Tenghao Zhao, Yifei Kang, Cunyin Chang, Shu Li, Guangyu Sun, Meng Li | 2026-09-16 | 下载 | Large language models (LLMs) require increasing memory capacity to accommodate growing model weights and KV caches. High-Bandwidth Flash (HBF) offers high memory density and aggregate read bandwidth t... |
| Automated Instruction Encoding Synthesis for Modern GPU ISA Compression | Mingyuan Ma, Hu He | 2026-09-16 | 下载 | Modern GPU kernels increasingly stress the instruction supply path, while fixed instruction containers can leave substantial footprint slack. This paper presents an automated encoding-synthesis framew... |
| MeshKV: A Network-on-Chip KV Cache Fabric for Scalable Transformer Decoding Accelerators | Dong Liu, Yanxuan Yu | 2026-09-16 | 下载 | Autoregressive transformer decoding is constrained by irregular key-value (KV) cache movement on tiled accelerators. Prior compression and DRAM-placement systems still concentrate traffic on centraliz... |
| Epic: Efficient Programming Paradigm for In-Storage Computing | Yuyue Wang, Zhenyu Zhang, Glenn Reinman, Huaicheng Li | 2026-09-16 | 下载 | In-storage computing (ISC) reduces host--storage data movement by executing computation inside computational storage devices (CSDs). For multi-stage applications, realizing these benefits requires coo... |
| VeriBugBench: An Empirically Grounded Framework for Constructing Verilog RTL Debugging Benchmarks | Xiankai Meng, Kejian Feng, Xinlin Zhao, Zhuo Zhang, Yan Lei, Xiaoguang Mao, Jiang Wu | 2026-09-16 | 下载 | RTL source-level debugging research requires benchmark artifacts that provide faulty designs together with precise change locations, executable test stimuli, and reproducible configurations. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Reputation as Community Memory for the Agentic Web | Ryan Chard, Gus Ellerm, Alexander Brace, Alok Kamatar, Suman Raj, Ian Foster, Kyle Chard | 2026-09-16 | 下载 | Agents can now externalize experience into memory, consolidating historical traces into semantic knowledge and procedural shortcuts that persist between sessions. |
| Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling | Mobina Kashaniyan, Ali Jannesari | 2026-09-16 | 下载 | Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number ... |
| Replication-Aware Placement of Functions and Data in the Edge-Cloud Continuum | Dario d'Abate, Matteo Cenzato, Matteo Briscini, Arianna Dragoni, Alessandro Margara | 2026-09-16 | 下载 | Function-as-a-Service (FaaS) has emerged as the prominent programming model for the edge-cloud continuum. FaaS inherently decouples stateless functions from their persistent state. |
| Ermes: a Stateful Serverless Platform for the Edge-to-Cloud Continuum | Matteo Cenzato, Dario d'Abate, Arianna Dragoni, Giacomo Orsenigo, Luca Tosetti, Matteo Briscini, Alessandro Margara | 2026-09-16 | 下载 | Function-as-a-Service (FaaS) is a widely adopted paradigm to simplify application deployment across the edge-to-cloud continuum. However, its stateless nature forces functions to retrieve their state ... |
| Fluid Notarization: Verifiable Evolution of Concurrently Edited Structured Documents | Amos Brocco, Giuliano Gremlich, Roberto Guidi | 2026-09-16 | 下载 | Traditional blockchain-based document notarization follows a snapshot-oriented model in which each document revision is represented as an independent state anchored on-chain through a cryptographic re... |
| Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It | Yipeng Liu, Yingqiang Zhang, Feifei Li, Huanchen Zhang | 2026-09-16 | 下载 | An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time. Serving systems decide whether that cache stays, leaves, or comes back by gue... |
| A Distributed Computing Framework for Satellite Swarms | Ezra Fielding, Clement Demazure, Guthemberg Silvestre, Felipe Alves Suana, Philippe Quéinnec | 2026-09-16 | 下载 | The rise of large satellite constellations and Distributed Space Systems (DSS) demands generalized frameworks that enable fault-tolerant, autonomous distributed space applications. |
| Vigil: Accountable Liveness against Selective Silence | Jiawei Cheng, Huiping Sun, Rui Zhou, Jinjue Zhou, Zhong Chen | 2026-09-16 | 下载 | BFT accountability is well understood for safety violations, and recent work attributes global liveness violations; \emph{recipient-selective} silence remains unresolved. |
| RayOrch: Programming and Executing Lineage-Controlled Multi-Grain Dataflows for Foundation-Model Data Preparation | Xiaochen Ma, Zimo Meng, Junzhu Liang, Youhe Jiang, Yue Cheng, Hao Liang, Bohan Zeng, Dengchun Li, Lu Ma, Zhengyang Zhao, Zhen Hao Wong, Runming He, Meiyi Qiang, Jiangtao Guan, Binhang Yuan, Wentao Zhang | 2026-09-16 | 下载 | Preparing high quality training data for foundation models requires scalable pipelines that transform heterogeneous documents and videos into structured records. |
| AUPE: Collaborative byzantine fault-tolerant peer-sampling | Augusta Mukam, Joachim Bruneau-Queyreix, Laurent Réveillère | 2026-09-16 | 下载 | Peer sampling is a crucial primitive in distributed systems, used to manage overlays and disseminate information in large-scale scenarios such as permissionless blockchain systems. |
| COMPASS-ABS: Reducing Fragmentation in Shared GPU Clusters for Deep Learning Training Workloads | Yukai Zhou, Hongfan Wu | 2026-09-16 | 下载 | With the rapid advancement of deep learning technology, shared GPU clusters receive an increasing number of deep learning training (DLT) jobs. |
| PatchyBFT: Automating Diversification of Fault-Tolerant Systems using LLMs | Arne Vogel, Christian Berger, Rüdiger Kapitza | 2026-09-16 | 下载 | Fault-tolerant agreement protocols fail if replicas share a common flaw that simultaneously affects more replicas than the tolerable threshold. |
| From Pixels to Semantics: Edge AI for UAV-Based Critical Infrastructure Inspection | Reza Farahani, Naser Hossein Motlagh, Zoha Azimi, Christian Timmerer, Lorenzo Carnevale, Sasu Tarkoma, Schahram Dustdar | 2026-09-16 | 下载 | Critical infrastructure assets such as bridges, tunnels, dams, and power line networks require timely and scalable inspection. While conventional manual inspection remains costly and hazardous, unmann... |
| GeoMesh: Workload-Balanced and Sign-Compressed Geo-Distributed LLM Training | Changyong Shin, Jaerim Park, Minchul Kang, Younghun Go, Zhixiong Niu, Yongqiang Xiong, Gyeongsik Yang, Chuck Yoo | 2026-09-16 | 下载 | Large language models are increasingly trained on GPUs distributed across multiple regions, but geo-distributed training is challenging in practice. |
| Zero-I/O Fault Recovery for Sharded Deep Learning via Dynamic Framework Dependency Rebinding | Genlang Chen, Junyi Zhu | 2026-09-16 | 下载 | Distributed model training at scale is frequently interrupted by transient network failures, conventionally forcing cluster managers to abort all processes and roll back to the latest checkpoint. |
| Token Latency Fairness: Performance Isolation for Multi-Tenant LLM Serving | Dev Bali, Soujanya Ponnapalli, Yichuan Wang, Natacha Crooks, Scott Shenker, Matei Zaharia | 2026-09-16 | 下载 | LLM serving is typically offered as a shared, multi-tenant service, where high-demand workloads from one client can cause latency SLO violations for others. |
| SSD-LLaMA: SSD-Native Inference for Trillion-Parameter MoE at 1+ Token/s on a Consumer PC | Fangzhou Liang, Yibin Shen, Jianmin Hu, Jiayang Xu, Hanchi Gao, Minxian Xu, Zili Meng | 2026-09-16 | 下载 | Frontier open-weight language models increasingly use Mixture-of-Experts (MoE) architectures to expand model capacity while activating only a small subset of experts per token. |
| vidax: A Unified JAX Framework for Video Generative Models on Accelerator Meshes | Congyue Deng | 2026-09-16 | 下载 | Open-source video generative models ship almost exclusively as PyTorch/CUDA reference implementations. This leaves Cloud TPU pods without a production-ready inference path, despite offering large, cos... |
| Towards Training Private LLMs: Exploring Fine-Tuning Language Models on Apple Silicon with RDMA over Thunderbolt | En-Ming Huang, Yao-Ting Hsieh, Hsiang-Yu Tsou, Mu-Chi Chen, Shih-Hao Hung, H. T. Kung | 2026-09-16 | 下载 | Private large language model (LLM) fine-tuning is increasingly important for organizations that need to adapt models using sensitive data, but it often exceeds the memory capacity of commodity datacen... |
| Large Language Model based air quality monitoring and localized alert generation | Ricardo Vieira, Luis Tavares, Kaylane Lima, Lucas De Souza, Arthur Poggy, João Lima, Vitor Pinheiro, Markus Endler | 2026-09-16 | 下载 | Poor indoor air quality can cause up to five times more direct health problems to occupants than outdoor air. In particular, it may cause headaches, fatigue, eye/throat irritation, and long-time expos... |
| ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference | Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao, Murali Annavaram, Salman Avestimehr | 2026-09-16 | 下载 | Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention an... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| When Agents Look Like Beacons: NIDS Evasion by Model Context Protocol Traffic | Muhammad Abdullah Sohail | 2026-09-16 | 下载 | The Model Context Protocol (MCP) standardizes communication between autonomous Artificial Intelligence (AI) agents and remote tools over Streamable HTTP. |
| Taming the Agentic RAN: Stability-Guaranteed Arbitration of Autonomous AI Agents in O-RAN | Seyed Bagher Hashemi Natanzi, Bo Tang | 2026-09-16 | 下载 | The O-RAN control plane is becoming agentic: autonomous AI agents, deployed as rApps by different vendors, independently close control loops over shared radio resources. |
| A Distributed Computing Framework for Satellite Swarms | Ezra Fielding, Clement Demazure, Guthemberg Silvestre, Felipe Alves Suana, Philippe Quéinnec | 2026-09-16 | 下载 | The rise of large satellite constellations and Distributed Space Systems (DSS) demands generalized frameworks that enable fault-tolerant, autonomous distributed space applications. |
| Reliability-Guided Trusted Repeater Node Selection in QKD-Enabled Metro Optical Networks | Arup Kumar Marik, Basabdatta Palit, Sadananda Behera | 2026-09-16 | 下载 | Quantum Key Distribution (QKD) in optical networks can provide information-theoretic security but is limited by its operability range, thereby requiring repeater nodes for extended coverage. |
| Toward Composable Network Digital Twins: A Subgraph-Based Latency Prediction Study | Shenjia Ding, David Flynn, Paul Harvey | 2026-09-16 | 下载 | Modern networks must support changing topologies, configurations, and performance objectives, motivating fast and reliable performance estimation. |
| Netkit: Specializing Linux Packet Delivery for Container Networks | Daniel Borkmann, Paul Chaignon | 2026-09-16 | 下载 | Cloud-native microservices architectures rely on network namespaces for isolation, with the overhead of container communications remaining a critical performance bottleneck. |
| MoQSplat: Adaptive Progressive Streaming of 3D Gaussian Splatting via MoQ | Emanuele Artioli, Mohammadreza Ghafari, Md Tariqul Islam, Farzad Tashtarian, Christian Rothenberg, Christian Timmerer | 2026-09-16 | 下载 | 3D Gaussian Splatting (3DGS) enables photorealistic novel view synthesis, but transmitting gigabyte-scale scene data remains challenging for immersive applications. |
| Jamming Detection in 5G/6G Networks: From O-RAN Concept to OCUDU Deployment | Marcin Hoffmann, Lukasz Kulacz, Osama Baldo, Marcin Pakula, Balaji Raghothaman | 2026-09-16 | 下载 | RF jamming poses a severe threat to 5G/6G networks, increasing packet latency being especially disruption to mission-critical URLLC services. This paper introduces a proactive Jamming Detection xApp (... |
| PentestChain: A Cost-Aware, MCP-Orchestrated Framework for Automated Penetration Testing with Free-Tier LLMs | Rushabh Vipulkumar Patel, Dipo Dunsin, Mohammed Almaiah, Mohamed Chahine Ghanem | 2026-09-16 | 下载 | AI-driven penetration testing has been demonstrated with premium frontier models such as GPT-4, but the per-engagement token cost makes continuous, automated testing unaffordable for the smaller organ... |
| SemABR: Measuring Video Semantic Fidelity with Multimodal LLMs for Adaptive Bitrate Streaming | Shiqi Xu, Soung Chang Liew, Yuyang Du | 2026-09-16 | 下载 | Conventional video metrics such as PSNR, SSIM, and VMAF measure visual distortion or perceptual quality, but they do not directly capture semantic preservation: whether compression retains a video's o... |
| The Operable Pareto Front: Distilling Offline Search into Run-Time Control for Multi-Objective UAV Edge-Computing Scheduling | Qiao Liao, Zhiyong Feng, Bin Wu, Guodong Fan | 2026-09-16 | 下载 | A UAV mobile edge computing (MEC) fleet trades energy against delay, and its schedules form a Pareto front; we call a scheduler operable when the fleet can be asked for any point on that front at run ... |
| Human Exposure to Non-Ionizing Radiation from Indoor Distributed Antenna System: Shopping Mall Measurement Analysis | Júlia da L. A. Silva, Vicente A. de Sousa,, Marcio E. C. Rodrigues, Fred Sizenando Rossiter Pinheiro, Gutembergue Soares da Silva, Halysson B. Mendonça, Ricardo Q. de F. H. Silva, João V. L. da Silva, Fernanda E. S. Galdino, Vitor F. C. de Carvalho, Lucas I. C. Medeiros | 2026-09-16 | 下载 | It is crucial to monitor the levels of Non-Ionizing Radiation (NIR) to which the general population may be exposed and compare them to the limits defined in the current standards, in view of the rapid... |
| Open-source emulation-based test environment to settle O-RAN-compliant trials | Ramon Fontes, Allan Martins, Vicente Sousa, Kaio Dantas, Lucas Medeiros, Pedro Alves, Marcelo Fernandes, Iago Rego, Eduardo Aranha, Vinícius Filho, Mateus Goldbarg, Wysterlanya Barros, Roger Immich, Augusto V. Neto | 2026-09-16 | 下载 | Experimental tools are a key factor in both academic and industrial research communities to create design evaluations of new networking technologies that involve troubleshooting or changing the planni... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Ask the Tool, Don't Guess: Agent Tool Calls Hold Their Progress, and the Serving System Should Read It | Yipeng Liu, Yingqiang Zhang, Feifei Li, Huanchen Zhang | 2026-09-16 | 下载 | An agentic request spends substantial wall-clock time waiting for tools, and its KV cache holds GPU memory the whole time. Serving systems decide whether that cache stays, leaves, or comes back by gue... |
| Netkit: Specializing Linux Packet Delivery for Container Networks | Daniel Borkmann, Paul Chaignon | 2026-09-16 | 下载 | Cloud-native microservices architectures rely on network namespaces for isolation, with the overhead of container communications remaining a critical performance bottleneck. |
| Position: It is Time to Virtualize Foundation Models with a Self-evolving Operating System Layer | Suparna Bhattacharya, Tarun Kumar, Cong Xu, Satish Kumar Mopur, Jiahao Li, Ashish Mishra, Aalap Tripathy, Annmary Justine Koomthanam, Martin Foltin, Ian Foster | 2026-09-16 | 下载 | AI applications have shifted from single, monolithic foundation models (FM) to compound agentic systems. Yet today's stacks remain fragmented: even as protocols (e.g. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling | Mobina Kashaniyan, Ali Jannesari | 2026-09-16 | 下载 | Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number ... |
| Splyce: SIMD Vectorization of Sparse Coiteration | Kabilan Mahathevan, Poorna Gunathilaka, Kirshanthan Sundararajah | 2026-09-16 | 下载 | Sparse tensor contractions are bottlenecked by sparse-sparse coiteration loops that resist standard loop vectorization. We present Splyce, an auto-vectorization framework in MLIR that overcomes this t... |
| Rosetta: Automating First-Principles Performance Modeling Using Multi-Agent LLMs | Karthikeyan Sankaralingam | 2026-09-16 | 下载 | Analytical performance models --- derivations of throughput or speedup from hardware parameters --- make claims independently verifiable and expose binding constraints, yet rarely accompany architectu... |