2026-09-17
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Can Agents Design Better Chips with a Higher Level Abstraction? | Zijian Ding, Yang Zou, Yizhou Sun, Jason Cong | 2026-09-17 | 下载 | Large Language Model (LLM) agents are increasingly being explored for chip design, but most existing approaches operate directly at RTL. We ask whether agents can design better chips by leveraging hig... |
| A Multi-Engine Dataflow for MoE Decoding on Scratchpad-Based Tensor Accelerators | Bin Ma, Wenjie Fan, Dong Li | 2026-09-17 | 下载 | Mixture-of-Experts (MoE) decoding on scratchpad-based tensor accelerators (STA) is dominated by moving expert weights while the compute engines sit idle. |
| RISC-V and machine learning: a survey | Shriman Keshri, Apparna Singh, Chinmaya Kumar Palo, Shreya Adya, Subhankar Mishra | 2026-09-17 | 下载 | The intersection of open-source processor architectures and machine learning is driving the demand for customizable, efficient, and accessible hardware. |
| Evaluating Positive Feedback Adiabatic Logic in 16nm FinFET with a Realistic Power-Clock | Franciszek Łukowski, Maciej Pyrzowski, Aida Todri-Sanial | 2026-09-17 | 下载 | Adiabatic logic reuses the energy stored on load capacitances through quasi-reversible switching, enabling a lower minimum energy consumption than conventional static CMOS. |
| Evaluation of Power-Clock Waveforms for Positive Feedback Adiabatic Logic in 16 nm FinFET Technology | Maciej Szymon Pyrzowski, Franciszek Łukowski, Aida Todri-Sanial | 2026-09-17 | 下载 | Adiabatic logic can recover part of the energy stored on load capacitances through quasi-reversible switching, but its waveform-optimized operation in FinFET technology and at multi-GHz frequencies re... |
| High-frequency Multispeculative Multiply-Accumulation Unit for Fused Posit Arithmetic | Mario Alonso, Miguel Ángel Sacristán, Guillermo Botella, Alberto A. Del Barrio | 2026-09-17 | 下载 | Posit arithmetic offers a compelling alternative to the IEEE 754 floating-point standard, providing enhanced accuracy. Its fused multiply-accumulate operations avoid intermediate rounding, ensuring ex... |
| MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration | Yuan Liao, Jae-sun Seo | 2026-09-17 | 下载 | The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. |
| Quantum computers will not be that different: A blueprint for quantum computer architecture at scale | Torsten Hoefler, Matthias Troyer | 2026-09-17 | 下载 | Quantum computers are technologically novel and unusual, but at system scale they should be engineered using many of the same principles that govern classical heterogeneous accelerators. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Cloud-Side Transactional Orchestration Framework for Resource-Constrained Embedded Systems | Pravin Nagare, Aditya Sabbineni, Preetam Dedu, Willison Lopes | 2026-09-17 | 下载 | As digital commerce ecosystems expand into low-end consumer electronics (CE), hardware constraints-specifically limited CPU duty cycles and volatile heap fragmentation-become significant bottlenecks f... |
| SensorWF: A FAIR Generalizable Workflow Framework for Scientific Time-Series Analysis | Logan Luna, Joseph Rigo, Kellan Shew, Raul Alejandro Vargas-Acosta | 2026-09-17 | 下载 | Scientific sensor data is foundational across disciplines including spacecraft engineering, clinical medicine, and atmospheric science. In each context, pipelines are constructed to ingest raw archive... |
| DLB: Distributed Load Balancing at Scale for Generative AI Inference | Santiago R. Balseiro, Bartek Wydrowski, Sameer Agarwal, David Applegate, Aaron Archer, Soheil Hassas Yeganeh, Alex Iriza, Bobby Kleinberg, Balasubramanian Sivan, Pranav Vaish, Oscar Zegarra, Wenxin Zhang, Vahab Mirrokni, Amin Vahdat | 2026-09-17 | 下载 | The reliance on scarce and expensive accelerators such as GPUs and TPUs in modern datacenters places unprecedented demands on backend infrastructure. |
| How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach? | Gaurav Agarwal, Ashish Garg, Isha Singhal | 2026-09-17 | 下载 | Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a frontier model produces correct kernels for 91. |
| FedeRage: Provably Convergent Agnostic Federated Learning under General Client Drift | Herlock Rahimi, Dionysis Kalogerias | 2026-09-17 | 下载 | Federated learning (FL) enables collaborative model training without sharing raw data, but its performance degrades under non-IID data and stochastic client participation. |
| Efficient Non-Uniform Quantum Hermite Transform through Adaptive Sampling | Nitay Mayo, Aryeh Lev Zabokritskiy | 2026-09-17 | 下载 | On the span of the first oscillator modes, Gauss--Hermite quadrature gives an exact change of basis between mode coefficients and weighted position space samples. |
| PixelFlow: Token-Level Workload Management for Efficient Distributed DiT Serving | Zhexiang Zhang, Minchen Yu, Yifan Sun, Xu Bai, Xingliang Yuan, Adel N. Toosi | 2026-09-17 | 下载 | Online image generation with Diffusion Transformers (DiTs) must meet latency service-level objectives (SLOs) while using GPU resources efficiently. |
| Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated Learning | Simon Süwer, Julian Klemm, Elisa Acitelli, Mathieu Almeida, Lucia Altucci, Zsolt Bagyura, Michelangela Barbieri, Zsolt-Zoltán Bedő, Rosaria Benedetti, Béla Bihari, Csongor Csalóka, Lucia Dicunta, Stanislav Ehrlich, Bjoern M. Eskofier, Sándor-József Fejér, Georg Fröwis, Walter Hötzendorfer, Alexandra Kautzky-Willer, Jens Johann Georg Lohmann, Marianna Maranghi, Lorenzo Marconi, Rudolf Mayer, Wouter Leonard Megchelenbrink, Monika Moga, Adham Mottalib, Sanjeev Mehta, Madeleine Müller, Thomas Nyström, Balázs-Attila Orbán, Paul O'Toole, Giuseppe Paolisso, Paolo Parini, Matteo Pedrelli, Enrico Petrillo, Philipp Poindl, Niklas Probul, Anastasia Pustozerova, Tanja Šarčević, Lukas Weilguny, Jan Baumbach, Andreas Maier | 2026-09-17 | 下载 | Federated learning enables collaborative training without sharing patient-level data, but most studies remain simulations. Based on five requirements derived from the literature, we analyzed 14 FL fra... |
| A Kubernetes-Native Request Router for Quality-Aware Inference Serving in the Computing Continuum | Ignjat Karanovic, Pantelis A. Frangoudis, Ivan Čilić, Ivana Podnar Žarko, Schahram Dustdar | 2026-09-17 | 下载 | We introduce Adaptive Score-based Routing Balancer (ASRB), a dynamic, score-based request routing mechanism for Kubernetes-based service deployments over the computing continuum. |
| Accelerating Sharded Data Parallelism at Scale with Federated Learning | Gianluca Mittone, Marco Aldinucci | 2026-09-17 | 下载 | The symbiotic scaling of artificial intelligence models and high-performance computing systems continually creates algorithmic challenges in their convergence. |
| Competition, Collusion, and Corruption: The Spectrum of MEV Attacks on DAG-Based BFT Consensus Protocols | Iliya Mirzaei, Heer Patel, Chenyuan Wu, Mohammad Javad Amiri | 2026-09-17 | 下载 | Byzantine Fault-Tolerant (BFT) protocols guarantee safety and liveness despite the malicious failure of nodes. However, they do not prevent adversarial manipulation of transaction order, where the ord... |
| FedeRICo: Federated Region-Influenced Coupling for Traffic Flow Prediction | Fermin Orozco, Man Luo, Johan Wahlström | 2026-09-17 | 下载 | Urban traffic forecasting often relies on information distributed across stakeholders who may be unable to share raw data due to privacy or commercial constraints, motivating federated spatial-tempora... |
| XIR: A Framework for Interoperability across Cross-Chain Protocols Based on a Verifiable Intermediate Representation | Yushen Li, Linpeng Jia, Jiaying Feng, Ziliang Liao, Yi Sun | 2026-09-17 | 下载 | Cross-chain protocols enable applications to exchange messages across blockchains. Under point-to-point configurations, communication depends on a direct connection between the source and destination ... |
| Distributed Edge Inference: an Experimental Study on Multiview Detection | Gianluca Mittone, Giulio Malenza, Marco Aldinucci, Robert Birke | 2026-09-17 | 下载 | Computing is evolving rapidly to cater to the increasing demand for sophisticated services, and Cloud computing lays a solid foundation for flexible on-demand provisioning. |
| P-GADMM: Parallel Group-Based ADMM for Asynchronous Optimization in Heterogeneous Edge Networks | Gaiguo Wei, Qingying Zhang, Heqiang Wang, Yu Zhang, Xiaoxiong Zhong | 2026-09-17 | 下载 | The Alternating Direction Method of Multipliers (ADMM) is widely used for distributed optimization, but its synchronous implementation can suffer from efficiency loss in heterogeneous edge networks, w... |
| Efficiently Distributed Federated Learning | Gianluca Mittone, Robert Birke, Marco Aldinucci | 2026-09-17 | 下载 | Federated Learning (FL) is experiencing a substantial research interest, with many frameworks being developed to allow practitioners to build federations easily and quickly. |
| VERA: Reinforcement Learning for Dynamic Memory Scaling of HPC Workloads in Kubernetes | Ade Pramono, Jie Ren, Ivy Peng | 2026-09-17 | 下载 | Memory over-provisioning results in resource underutilization when HPC workloads run on Kubernetes. The default Vertical Pod Autoscaler (VPA) cannot anticipate phase-driven memory spikes for first-run... |
| The Life of a Token: from Words to Bits on the Wire | Davide Avesani, Pengwenlong Gu, Sotiris Skaperas, Stefano Secci | 2026-09-17 | 下载 | Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. |
| Xronos: Heterogeneity-Aware Tensor Parallelism for Collaborative LLM Fine-Tuning on Edge CPUs | Wonmi Choi, Sunjae Park, Dohyeok Kwon, Zhixiong Niu, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang | 2026-09-17 | 下载 | Collaborative fine-tuning on edge devices adapts large language models to domain-specific data while keeping each device's data local. State-of-the-art (SOTA) collaborative fine-tuning techniques are ... |
| Hopper: Bounded-Memory Collaborative Debiasing for Byzantine-Tolerant Peer Sampling | Augusta Mukam, Joachim Bruneau-Queyreix, Laurent Reveillère | 2026-09-17 | 下载 | Byzantine-tolerant peer sampling relies on continuously refreshed views, yet an adversary can bias the identifier streams used to construct them. |
| Safe Exploration of Arbitrary Dynamic Dangerous Networks | Caterina Feletti, Paola Flocchini, Giuseppe Prencipe, Nicola Santoro | 2026-09-17 | 下载 | Given a team of agents on the nodes of a graph-based network, the exploration problem requires each node to be visited by at least one agent. In the classical distributed setting of static networks, a... |
| Sketching the Error, Not the Product: Post Hoc Fault Recovery for Half Precision GPU Matrix Multiplication | Pranav Napolean, Vikas Srivastava, Napolean Periathambi | 2026-09-17 | 下载 | Silent data corruption (SDC) from defective accelerators now interrupts large scale training, yet deployed mitigations act on whole nodes. Algorithm based fault tolerance (ABFT) for a single GEMM has ... |
| Syndrome Decoding for Silent Data Corruption in Quantized Integer GPU Arithmetic | Pranav Napolean, Vikas Srivastava, Napolean Periathambi | 2026-09-17 | 下载 | Quantized neural network inference runs integer matrix multiplications on GPU tensor cores, and the INT32 accumulators inside those cores have neither parity nor ECC. |
| Detecting Soft Errors in Parallel Software with LLM-tuned Instruction Duplication | Yafan Huang, Guanpeng Li | 2026-09-17 | 下载 | We propose PaRID (PaRallel Instruction Duplication), a software-directed soft error detection framework that requires only compile-time effort for multithreading parallel programs. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Cloud-Side Transactional Orchestration Framework for Resource-Constrained Embedded Systems | Pravin Nagare, Aditya Sabbineni, Preetam Dedu, Willison Lopes | 2026-09-17 | 下载 | As digital commerce ecosystems expand into low-end consumer electronics (CE), hardware constraints-specifically limited CPU duty cycles and volatile heap fragmentation-become significant bottlenecks f... |
| NetInspector: Measuring and Improving LLM Capabilities for Reliable Intent-Based Networking Policy Generation | Yuxuan Zhang, Hongxin Hu, Guofei Gu | 2026-09-17 | 下载 | Modern networks are large in scale and heterogeneous in configuration, making manual policy management increasingly impractical. Intent-Based Networking (IBN) addresses this by automating the translat... |
| GeoRIS: Geofencing With Reconfigurable Intelligent Surfaces | André Gomes, Arthur S. de Sena, Luiz A. DaSilva, Jacek Kibilda | 2026-09-17 | 下载 | Geofencing refers to controlling the availability of wireless services within a network perimeter. In this paper, we study how RIS can be used to achieve geofencing in outdoor-to-indoor network scenar... |
| RUN-O-RAN: An O-RAN-Native Architecture Enabling Cooperative Uplink Localization | Viola Bernazzoli, Alberto Ceresoli, Ilario Filippini | 2026-09-17 | 下载 | Accurate positioning is increasingly required in indoor and dense urban environments; nonetheless, satellite-based systems are not always available, and standardized 5G localization solutions remain d... |
| NS3Learn: Transferring 5G NR Mode-2 Reception Realism from ns-3 to the Veins/SUMO Stack for Connected-Vehicle Safety Assessment | Rasheed Bello, Arthur Mukwaya, Gurcan Comert, Varghese Vaidyan, Vijay Bendigeri, Anthony Dontoh, Jagruti Sahoo, Judith Mwakalonge | 2026-09-17 | 下载 | Connected-vehicle safety evaluations rely on coupled traffic and network simulations, but standard channel models ignore radio resource competition in 5G NR sidelink Mode-2, reporting unrealistically ... |
| Value-Based Massive Access through Goal-Oriented Irregular Repetition Slotted ALOHA | Pietro Talli, Andrea Munari, Federico Mason, Federico Chiariotti, Andrea Zanella | 2026-09-17 | 下载 | The goal-oriented communication paradigm is poised to enable novel real-time applications by easing the burden on communication networks while still delivering task-relevant information. |
| STR-Agent: An LLM-Driven Agent for QoS-Aware Routing in LEO Satellite Networks | Bowen Lu, Mugen Peng, Yaohua Sun, Hongyu Wang, Kerui Guo, Wenjia Xu | 2026-09-17 | 下载 | LEO satellite networks feature dynamic topologies, time-varying links, and diverse service requirements, which make conventional routing schemes difficult to support fine-grained quality-of-service (Q... |
| The Life of a Token: from Words to Bits on the Wire | Davide Avesani, Pengwenlong Gu, Sotiris Skaperas, Stefano Secci | 2026-09-17 | 下载 | Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. |
| Trigger Timing, Deadline Readiness, and Event-Aligned Accounting for Dynamic Ad Insertion | Prashant Chaudhary, Kapil Khandelwal | 2026-09-17 | 下载 | Dynamic ad insertion comparisons can conflate trigger, reach, readiness, playback, billability and measurement even when the accounting is arithmetically correct. |
| PyStream: Enhancing Video Streaming Evaluation | Samuel Radler, Leon Prüller, Emanuele Artioli, Farzad Tashtarian, Christian Timmerer | 2026-09-17 | 下载 | As streaming services become more commonplace, analyzing their behavior effectively under different network conditions is crucial. This is normally quite expensive, requiring multiple players with dif... |
| Reachability, Not Observation: Containing Systems Whose Wiring Changes | Yoshiaki Takashita | 2026-09-17 | 下载 | Containment decisions -- where to put a firewall, which links to monitor, what a program may reach -- are computed from an observed structure, and observation is a snapshot. |
| Quantum computers will not be that different: A blueprint for quantum computer architecture at scale | Torsten Hoefler, Matthias Troyer | 2026-09-17 | 下载 | Quantum computers are technologically novel and unusual, but at system scale they should be engineered using many of the same principles that govern classical heterogeneous accelerators. |
| A Bi-Objective Routing Framework for Hybrid Terrestrial-Satellite Quantum Networks | Yashpreet Khambay, Nitish K. Panigrahy | 2026-09-17 | 下载 | Hybrid terrestrial-satellite quantum networks combine terrestrial fiber infrastructure with free-space links to enable long-distance entanglement distribution. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| An Approximate Queueing Model of LLM Inference Serving for SLO-Driven Autoscaling | Vishakha Ramani, Asser N. Tantawi | 2026-09-17 | 下载 | Performance models of LLM servers support both latency evaluation and the design of controllers for autoscaling against service level objectives (SLOs) and for inference optimization. |
| Scaling Fourier-Based Sparse Matrix Analysis on GPUs | Ruifeng Zhang, Sai Krishna Teja Varma Manthena, Jiajia Li, Xipeng Shen | 2026-09-17 | 下载 | Sparse computations are important workloads in applications such as scientific computing, graph neural networks (GNNs), and machine learning. While many sparse operations can benefit from modern GPUs,... |
| Accelerating Sharded Data Parallelism at Scale with Federated Learning | Gianluca Mittone, Marco Aldinucci | 2026-09-17 | 下载 | The symbiotic scaling of artificial intelligence models and high-performance computing systems continually creates algorithmic challenges in their convergence. |
| Efficiently Distributed Federated Learning | Gianluca Mittone, Robert Birke, Marco Aldinucci | 2026-09-17 | 下载 | Federated Learning (FL) is experiencing a substantial research interest, with many frameworks being developed to allow practitioners to build federations easily and quickly. |
| Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics | Wonmi Choi, Minuk Park, Zhixiong Niu, Yongqiang Xiong, Chuck Yoo, Gyeongsik Yang | 2026-09-17 | 下载 | LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. |
| The Life of a Token: from Words to Bits on the Wire | Davide Avesani, Pengwenlong Gu, Sotiris Skaperas, Stefano Secci | 2026-09-17 | 下载 | Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks. |
| Trigger Timing, Deadline Readiness, and Event-Aligned Accounting for Dynamic Ad Insertion | Prashant Chaudhary, Kapil Khandelwal | 2026-09-17 | 下载 | Dynamic ad insertion comparisons can conflate trigger, reach, readiness, playback, billability and measurement even when the accounting is arithmetically correct. |
| Whittle index approach to multi-server scheduling with convex delay costs and impatient customers | Samuli Aalto | 2026-09-17 | 下载 | We consider the dynamic scheduling problem in a multi-class M/G/N + M queue with convex delay costs and impatient customers that have exponential abandonment times. |
| Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image Generation | Daeun Kim, Junwha Hong, Changhun Oh, Yoonsung Kim, Yoonhyeong Lee, Jongse Park | 2026-09-17 | 下载 | Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures. |
| PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM Serving | Omkar Shewale, Deepak Kumar, Divakar Kumar Yadav | 2026-09-17 | 下载 | Repeated prompt prefixes are increasingly common in LLM serving workloads, appearing in system prompts, templated retrieval-augmented generation pipelines, agent frameworks, and multi-turn conversatio... |
| A Bi-Objective Routing Framework for Hybrid Terrestrial-Satellite Quantum Networks | Yashpreet Khambay, Nitish K. Panigrahy | 2026-09-17 | 下载 | Hybrid terrestrial-satellite quantum networks combine terrestrial fiber infrastructure with free-space links to enable long-distance entanglement distribution. |