Skip to content

2026-09-17 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Can Agents Design Better Chips with a Higher Level Abstraction?Zijian Ding, Yang Zou, Yizhou Sun, Jason Cong2026-09-17下载Large Language Model (LLM) agents are increasingly being explored for chip design, but most existing approaches operate directly at RTL. We ask whether agents can design better chips by leveraging hig...
A Multi-Engine Dataflow for MoE Decoding on Scratchpad-Based Tensor AcceleratorsBin Ma, Wenjie Fan, Dong Li2026-09-17下载Mixture-of-Experts (MoE) decoding on scratchpad-based tensor accelerators (STA) is dominated by moving expert weights while the compute engines sit idle.
RISC-V and machine learning: a surveyShriman Keshri, Apparna Singh, Chinmaya Kumar Palo, Shreya Adya, Subhankar Mishra2026-09-17下载The intersection of open-source processor architectures and machine learning is driving the demand for customizable, efficient, and accessible hardware.
Evaluating Positive Feedback Adiabatic Logic in 16nm FinFET with a Realistic Power-ClockFranciszek Łukowski, Maciej Pyrzowski, Aida Todri-Sanial2026-09-17下载Adiabatic logic reuses the energy stored on load capacitances through quasi-reversible switching, enabling a lower minimum energy consumption than conventional static CMOS.
Evaluation of Power-Clock Waveforms for Positive Feedback Adiabatic Logic in 16 nm FinFET TechnologyMaciej Szymon Pyrzowski, Franciszek Łukowski, Aida Todri-Sanial2026-09-17下载Adiabatic logic can recover part of the energy stored on load capacitances through quasi-reversible switching, but its waveform-optimized operation in FinFET technology and at multi-GHz frequencies re...
High-frequency Multispeculative Multiply-Accumulation Unit for Fused Posit ArithmeticMario Alonso, Miguel Ángel Sacristán, Guillermo Botella, Alberto A. Del Barrio2026-09-17下载Posit arithmetic offers a compelling alternative to the IEEE 754 floating-point standard, providing enhanced accuracy. Its fused multiply-accumulate operations avoid intermediate rounding, ensuring ex...
MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model AccelerationYuan Liao, Jae-sun Seo2026-09-17下载The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization.
Quantum computers will not be that different: A blueprint for quantum computer architecture at scaleTorsten Hoefler, Matthias Troyer2026-09-17下载Quantum computers are technologically novel and unusual, but at system scale they should be engineered using many of the same principles that govern classical heterogeneous accelerators.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Cloud-Side Transactional Orchestration Framework for Resource-Constrained Embedded SystemsPravin Nagare, Aditya Sabbineni, Preetam Dedu, Willison Lopes2026-09-17下载As digital commerce ecosystems expand into low-end consumer electronics (CE), hardware constraints-specifically limited CPU duty cycles and volatile heap fragmentation-become significant bottlenecks f...
SensorWF: A FAIR Generalizable Workflow Framework for Scientific Time-Series AnalysisLogan Luna, Joseph Rigo, Kellan Shew, Raul Alejandro Vargas-Acosta2026-09-17下载Scientific sensor data is foundational across disciplines including spacecraft engineering, clinical medicine, and atmospheric science. In each context, pipelines are constructed to ingest raw archive...
DLB: Distributed Load Balancing at Scale for Generative AI InferenceSantiago R. Balseiro, Bartek Wydrowski, Sameer Agarwal, David Applegate, Aaron Archer, Soheil Hassas Yeganeh, Alex Iriza, Bobby Kleinberg, Balasubramanian Sivan, Pranav Vaish, Oscar Zegarra, Wenxin Zhang, Vahab Mirrokni, Amin Vahdat2026-09-17下载The reliance on scarce and expensive accelerators such as GPUs and TPUs in modern datacenters places unprecedented demands on backend infrastructure.
How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?Gaurav Agarwal, Ashish Garg, Isha Singhal2026-09-17下载Language models can now write GPU kernels that outperform PyTorch. We evaluate five model configurations on KernelBench level 1 and find that a frontier model produces correct kernels for 91.
FedeRage: Provably Convergent Agnostic Federated Learning under General Client DriftHerlock Rahimi, Dionysis Kalogerias2026-09-17下载Federated learning (FL) enables collaborative model training without sharing raw data, but its performance degrades under non-IID data and stochastic client participation.
Efficient Non-Uniform Quantum Hermite Transform through Adaptive SamplingNitay Mayo, Aryeh Lev Zabokritskiy2026-09-17下载On the span of the first NN oscillator modes, Gauss--Hermite quadrature gives an exact change of basis between mode coefficients and NN weighted position space samples.
PixelFlow: Token-Level Workload Management for Efficient Distributed DiT ServingZhexiang Zhang, Minchen Yu, Yifan Sun, Xu Bai, Xingliang Yuan, Adel N. Toosi2026-09-17下载Online image generation with Diffusion Transformers (DiTs) must meet latency service-level objectives (SLOs) while using GPU resources efficiently.
Multi-center Medical Data Mining with FL-Net - A One-stop Shop for Federated LearningSimon Süwer, Julian Klemm, Elisa Acitelli, Mathieu Almeida, Lucia Altucci, Zsolt Bagyura, Michelangela Barbieri, Zsolt-Zoltán Bedő, Rosaria Benedetti, Béla Bihari, Csongor Csalóka, Lucia Dicunta, Stanislav Ehrlich, Bjoern M. Eskofier, Sándor-József Fejér, Georg Fröwis, Walter Hötzendorfer, Alexandra Kautzky-Willer, Jens Johann Georg Lohmann, Marianna Maranghi, Lorenzo Marconi, Rudolf Mayer, Wouter Leonard Megchelenbrink, Monika Moga, Adham Mottalib, Sanjeev Mehta, Madeleine Müller, Thomas Nyström, Balázs-Attila Orbán, Paul O'Toole, Giuseppe Paolisso, Paolo Parini, Matteo Pedrelli, Enrico Petrillo, Philipp Poindl, Niklas Probul, Anastasia Pustozerova, Tanja Šarčević, Lukas Weilguny, Jan Baumbach, Andreas Maier2026-09-17下载Federated learning enables collaborative training without sharing patient-level data, but most studies remain simulations. Based on five requirements derived from the literature, we analyzed 14 FL fra...
A Kubernetes-Native Request Router for Quality-Aware Inference Serving in the Computing ContinuumIgnjat Karanovic, Pantelis A. Frangoudis, Ivan Čilić, Ivana Podnar Žarko, Schahram Dustdar2026-09-17下载We introduce Adaptive Score-based Routing Balancer (ASRB), a dynamic, score-based request routing mechanism for Kubernetes-based service deployments over the computing continuum.
Accelerating Sharded Data Parallelism at Scale with Federated LearningGianluca Mittone, Marco Aldinucci2026-09-17下载The symbiotic scaling of artificial intelligence models and high-performance computing systems continually creates algorithmic challenges in their convergence.
Competition, Collusion, and Corruption: The Spectrum of MEV Attacks on DAG-Based BFT Consensus ProtocolsIliya Mirzaei, Heer Patel, Chenyuan Wu, Mohammad Javad Amiri2026-09-17下载Byzantine Fault-Tolerant (BFT) protocols guarantee safety and liveness despite the malicious failure of nodes. However, they do not prevent adversarial manipulation of transaction order, where the ord...
FedeRICo: Federated Region-Influenced Coupling for Traffic Flow PredictionFermin Orozco, Man Luo, Johan Wahlström2026-09-17下载Urban traffic forecasting often relies on information distributed across stakeholders who may be unable to share raw data due to privacy or commercial constraints, motivating federated spatial-tempora...
XIR: A Framework for Interoperability across Cross-Chain Protocols Based on a Verifiable Intermediate RepresentationYushen Li, Linpeng Jia, Jiaying Feng, Ziliang Liao, Yi Sun2026-09-17下载Cross-chain protocols enable applications to exchange messages across blockchains. Under point-to-point configurations, communication depends on a direct connection between the source and destination ...
Distributed Edge Inference: an Experimental Study on Multiview DetectionGianluca Mittone, Giulio Malenza, Marco Aldinucci, Robert Birke2026-09-17下载Computing is evolving rapidly to cater to the increasing demand for sophisticated services, and Cloud computing lays a solid foundation for flexible on-demand provisioning.
P-GADMM: Parallel Group-Based ADMM for Asynchronous Optimization in Heterogeneous Edge NetworksGaiguo Wei, Qingying Zhang, Heqiang Wang, Yu Zhang, Xiaoxiong Zhong2026-09-17下载The Alternating Direction Method of Multipliers (ADMM) is widely used for distributed optimization, but its synchronous implementation can suffer from efficiency loss in heterogeneous edge networks, w...
Efficiently Distributed Federated LearningGianluca Mittone, Robert Birke, Marco Aldinucci2026-09-17下载Federated Learning (FL) is experiencing a substantial research interest, with many frameworks being developed to allow practitioners to build federations easily and quickly.
VERA: Reinforcement Learning for Dynamic Memory Scaling of HPC Workloads in KubernetesAde Pramono, Jie Ren, Ivy Peng2026-09-17下载Memory over-provisioning results in resource underutilization when HPC workloads run on Kubernetes. The default Vertical Pod Autoscaler (VPA) cannot anticipate phase-driven memory spikes for first-run...
The Life of a Token: from Words to Bits on the WireDavide Avesani, Pengwenlong Gu, Sotiris Skaperas, Stefano Secci2026-09-17下载Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks.
Xronos: Heterogeneity-Aware Tensor Parallelism for Collaborative LLM Fine-Tuning on Edge CPUsWonmi Choi, Sunjae Park, Dohyeok Kwon, Zhixiong Niu, Yeonho Yoo, Chuck Yoo, Gyeongsik Yang2026-09-17下载Collaborative fine-tuning on edge devices adapts large language models to domain-specific data while keeping each device's data local. State-of-the-art (SOTA) collaborative fine-tuning techniques are ...
Hopper: Bounded-Memory Collaborative Debiasing for Byzantine-Tolerant Peer SamplingAugusta Mukam, Joachim Bruneau-Queyreix, Laurent Reveillère2026-09-17下载Byzantine-tolerant peer sampling relies on continuously refreshed views, yet an adversary can bias the identifier streams used to construct them.
Safe Exploration of Arbitrary Dynamic Dangerous NetworksCaterina Feletti, Paola Flocchini, Giuseppe Prencipe, Nicola Santoro2026-09-17下载Given a team of agents on the nodes of a graph-based network, the exploration problem requires each node to be visited by at least one agent. In the classical distributed setting of static networks, a...
Sketching the Error, Not the Product: Post Hoc Fault Recovery for Half Precision GPU Matrix MultiplicationPranav Napolean, Vikas Srivastava, Napolean Periathambi2026-09-17下载Silent data corruption (SDC) from defective accelerators now interrupts large scale training, yet deployed mitigations act on whole nodes. Algorithm based fault tolerance (ABFT) for a single GEMM has ...
Syndrome Decoding for Silent Data Corruption in Quantized Integer GPU ArithmeticPranav Napolean, Vikas Srivastava, Napolean Periathambi2026-09-17下载Quantized neural network inference runs integer matrix multiplications on GPU tensor cores, and the INT32 accumulators inside those cores have neither parity nor ECC.
Detecting Soft Errors in Parallel Software with LLM-tuned Instruction DuplicationYafan Huang, Guanpeng Li2026-09-17下载We propose PaRID (PaRallel Instruction Duplication), a software-directed soft error detection framework that requires only compile-time effort for multithreading parallel programs.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Cloud-Side Transactional Orchestration Framework for Resource-Constrained Embedded SystemsPravin Nagare, Aditya Sabbineni, Preetam Dedu, Willison Lopes2026-09-17下载As digital commerce ecosystems expand into low-end consumer electronics (CE), hardware constraints-specifically limited CPU duty cycles and volatile heap fragmentation-become significant bottlenecks f...
NetInspector: Measuring and Improving LLM Capabilities for Reliable Intent-Based Networking Policy GenerationYuxuan Zhang, Hongxin Hu, Guofei Gu2026-09-17下载Modern networks are large in scale and heterogeneous in configuration, making manual policy management increasingly impractical. Intent-Based Networking (IBN) addresses this by automating the translat...
GeoRIS: Geofencing With Reconfigurable Intelligent SurfacesAndré Gomes, Arthur S. de Sena, Luiz A. DaSilva, Jacek Kibilda2026-09-17下载Geofencing refers to controlling the availability of wireless services within a network perimeter. In this paper, we study how RIS can be used to achieve geofencing in outdoor-to-indoor network scenar...
RUN-O-RAN: An O-RAN-Native Architecture Enabling Cooperative Uplink LocalizationViola Bernazzoli, Alberto Ceresoli, Ilario Filippini2026-09-17下载Accurate positioning is increasingly required in indoor and dense urban environments; nonetheless, satellite-based systems are not always available, and standardized 5G localization solutions remain d...
NS3Learn: Transferring 5G NR Mode-2 Reception Realism from ns-3 to the Veins/SUMO Stack for Connected-Vehicle Safety AssessmentRasheed Bello, Arthur Mukwaya, Gurcan Comert, Varghese Vaidyan, Vijay Bendigeri, Anthony Dontoh, Jagruti Sahoo, Judith Mwakalonge2026-09-17下载Connected-vehicle safety evaluations rely on coupled traffic and network simulations, but standard channel models ignore radio resource competition in 5G NR sidelink Mode-2, reporting unrealistically ...
Value-Based Massive Access through Goal-Oriented Irregular Repetition Slotted ALOHAPietro Talli, Andrea Munari, Federico Mason, Federico Chiariotti, Andrea Zanella2026-09-17下载The goal-oriented communication paradigm is poised to enable novel real-time applications by easing the burden on communication networks while still delivering task-relevant information.
STR-Agent: An LLM-Driven Agent for QoS-Aware Routing in LEO Satellite NetworksBowen Lu, Mugen Peng, Yaohua Sun, Hongyu Wang, Kerui Guo, Wenjia Xu2026-09-17下载LEO satellite networks feature dynamic topologies, time-varying links, and diverse service requirements, which make conventional routing schemes difficult to support fine-grained quality-of-service (Q...
The Life of a Token: from Words to Bits on the WireDavide Avesani, Pengwenlong Gu, Sotiris Skaperas, Stefano Secci2026-09-17下载Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks.
Trigger Timing, Deadline Readiness, and Event-Aligned Accounting for Dynamic Ad InsertionPrashant Chaudhary, Kapil Khandelwal2026-09-17下载Dynamic ad insertion comparisons can conflate trigger, reach, readiness, playback, billability and measurement even when the accounting is arithmetically correct.
PyStream: Enhancing Video Streaming EvaluationSamuel Radler, Leon Prüller, Emanuele Artioli, Farzad Tashtarian, Christian Timmerer2026-09-17下载As streaming services become more commonplace, analyzing their behavior effectively under different network conditions is crucial. This is normally quite expensive, requiring multiple players with dif...
Reachability, Not Observation: Containing Systems Whose Wiring ChangesYoshiaki Takashita2026-09-17下载Containment decisions -- where to put a firewall, which links to monitor, what a program may reach -- are computed from an observed structure, and observation is a snapshot.
Quantum computers will not be that different: A blueprint for quantum computer architecture at scaleTorsten Hoefler, Matthias Troyer2026-09-17下载Quantum computers are technologically novel and unusual, but at system scale they should be engineered using many of the same principles that govern classical heterogeneous accelerators.
A Bi-Objective Routing Framework for Hybrid Terrestrial-Satellite Quantum NetworksYashpreet Khambay, Nitish K. Panigrahy2026-09-17下载Hybrid terrestrial-satellite quantum networks combine terrestrial fiber infrastructure with free-space links to enable long-distance entanglement distribution.

cs.PF - Performance ​

标题作者发布日期PDF摘要
An Approximate Queueing Model of LLM Inference Serving for SLO-Driven AutoscalingVishakha Ramani, Asser N. Tantawi2026-09-17下载Performance models of LLM servers support both latency evaluation and the design of controllers for autoscaling against service level objectives (SLOs) and for inference optimization.
Scaling Fourier-Based Sparse Matrix Analysis on GPUsRuifeng Zhang, Sai Krishna Teja Varma Manthena, Jiajia Li, Xipeng Shen2026-09-17下载Sparse computations are important workloads in applications such as scientific computing, graph neural networks (GNNs), and machine learning. While many sparse operations can benefit from modern GPUs,...
Accelerating Sharded Data Parallelism at Scale with Federated LearningGianluca Mittone, Marco Aldinucci2026-09-17下载The symbiotic scaling of artificial intelligence models and high-performance computing systems continually creates algorithmic challenges in their convergence.
Efficiently Distributed Federated LearningGianluca Mittone, Robert Birke, Marco Aldinucci2026-09-17下载Federated Learning (FL) is experiencing a substantial research interest, with many frameworks being developed to allow practitioners to build federations easily and quickly.
Not All AI Agents Are Equal: Characterizing Resource and Performance DynamicsWonmi Choi, Minuk Park, Zhixiong Niu, Yongqiang Xiong, Chuck Yoo, Gyeongsik Yang2026-09-17下载LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers.
The Life of a Token: from Words to Bits on the WireDavide Avesani, Pengwenlong Gu, Sotiris Skaperas, Stefano Secci2026-09-17下载Large Language Models (LLMs) transform vast collections of unstructured text into semantic patterns used for language generation and reasoning tasks.
Trigger Timing, Deadline Readiness, and Event-Aligned Accounting for Dynamic Ad InsertionPrashant Chaudhary, Kapil Khandelwal2026-09-17下载Dynamic ad insertion comparisons can conflate trigger, reach, readiness, playback, billability and measurement even when the accounting is arithmetically correct.
Whittle index approach to multi-server scheduling with convex delay costs and impatient customersSamuli Aalto2026-09-17下载We consider the dynamic scheduling problem in a multi-class M/G/N + M queue with convex delay costs and impatient customers that have exponential abandonment times.
Understanding and Exploiting Diagonal Attention Sparsity in Autoregressive Image GenerationDaeun Kim, Junwha Hong, Changhun Oh, Yoonsung Kim, Yoonhyeong Lee, Jongse Park2026-09-17下载Autoregressive image generation has emerged as a paradigm for multimodal AI systems due to its compatibility with transformer-based LLM serving infrastructures.
PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM ServingOmkar Shewale, Deepak Kumar, Divakar Kumar Yadav2026-09-17下载Repeated prompt prefixes are increasingly common in LLM serving workloads, appearing in system prompts, templated retrieval-augmented generation pipelines, agent frameworks, and multi-turn conversatio...
A Bi-Objective Routing Framework for Hybrid Terrestrial-Satellite Quantum NetworksYashpreet Khambay, Nitish K. Panigrahy2026-09-17下载Hybrid terrestrial-satellite quantum networks combine terrestrial fiber infrastructure with free-space links to enable long-distance entanglement distribution.

基于 VitePress 构建 · 使用本地搜索查找论文