2026-08-10
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Monophonic Audio Synthesizer Using FPGAs | Michael Smith, D. G. Perera | 2026-08-10 | 下载 | Signal synthesis is used in every aspect of the electronics world, where sinusoidal waveforms are used to perform functions such as clocking, signal transmission, feedback controls, and other applicat... |
| ArchAgent v2: A Case Study with the Data Prefetching Championship | Abraham Gonzalez, Raghav Gupta, Akanksha Jain, Hanna Alam, Alexander Novikov, Po-Sen Huang, Matej Balog, Marvin Eisenberger, Sergey Shirobokov, Ngân Vũ, Hank Levy, Borivoje Nikolić, Sagar Karandikar, Martin Dixon, Parthasarathy Ranganathan | 2026-08-10 | 下载 | Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search sp... |
| FSGen: Agile Fused and Sparse Accelerator Generator with Accurate Power Model for LLM Applications | Jay Zhe-An Mok, Qijun Zhang, Zhiyao Xie | 2026-08-10 | 下载 | With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains. |
| SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple Silicon | Tianhong Xu, Saion K. Roy, Ruyi Ding, Aidong Adam Ding, Yunsi Fei | 2026-08-10 | 下载 | Modern heterogeneous System-on-Chip designs integrate CPU cores and a GPU that share a last-level cache (LLC) or system-level cache (SLC). This sharing exposes a new cross-domain attack surface, and e... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| What Actually Serializes GPU LZ77 Decode: Three Decoders, Three Mechanisms, and an Encode-Time Lever That Removes the Last One | Yakiv Shavidze | 2026-08-10 | 下载 | The sequential part of GPU LZ77 decode is not where the field assumes it is. Across three decoder architectures on an H100 we measure that parse, not copy, holds 64-72% of device-resident decode time;... |
| SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client Ranks | Yue Xia, Tayyebeh Jahani-Nezhad, Mayank Bakshi, Rawad Bitar | 2026-08-10 | 下载 | We consider federated parameter efficient fine-tuning of large neural networks with low-rank adaptation (LoRA,~Hu et al.\ 2022). Combining LoRA with federated PEFT introduces challenges absent from ei... |
| Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4 | Matt J. Borowski, Blazej Osinski | 2026-08-10 | 下载 | High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matrix loads with ldmatrix, and matrix multiply-accumulate operati... |
| Thread Scaling of Hexaly on the TDVRPTW across Two Model Encodings. An Experimental Report: External-Function Serialization, Slice-Count Choice, and the Time-Sliced Thread Ladder | Florian Rascoussier | 2026-08-10 | 下载 | Thread scaling in Hexaly depends first on how the Time-Dependent Vehicle Routing Problem with Time Windows is modeled. We compare two Python encodings: one evaluates continuous travel-time functions t... |
| Certified Split Windows for Parallel Lexing: Recovering Boundaries Where No Byte Certifies | Nicklas Nidhögg | 2026-08-10 | 下载 | A certified split point lets a parallel lexer cut unlexed input at a single byte with the serial token stream provably preserved, but several conventional token sets in the predecessor's controlled st... |
| Defining Decentralization: An Ontological Perspective | Jakub Kacper Szeląg, Aydin Abadi, Mohammad Naseri | 2026-08-10 | 下载 | Decentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as security, distributed computing, artificial intelligence, clo... |
| Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach | Xinyi Xu, Bingnan Xiao, Shuang Qin, Gang Feng, Tony Q. S. Quek | 2026-08-10 | 下载 | Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., and , providing an efficient way to fine-tune large models in federated learning ... |
| SpSYRK: Half the Work in Distributed Sparse Matrix Multiplication | Thomas McFarland, Julian Bellavita, Giulia Guidi | 2026-08-10 | 下载 | The symmetric rank- update (SYRK), \C = \A\A^\top, computes the dot product between each pair of rows of \A, producing the Gram matrix \C. |
| A Preliminary Study on Simultaneous Coscheduling for Discrete GPU vs. Fused GPU | Poorna Gunathilaka, Nabayan Chaudhury, Kirshanthan Sundararajah, Wu-chun Feng | 2026-08-10 | 下载 | CPU-GPU coscheduling enables simultaneous execution of an application across both processing units, but its efficiency depends on workload partitioning and memory architecture. |
| ptimal Distributed Maximum Flow Approximation in Undirected Planar Graphs | Yaseen Abd-Elhaleem, Michal Dory, Oren Weimann | 2026-08-10 | 下载 | Persistent efforts in recent years have been devoted to devising distributed algorithms for fundamental optimization problems in planar graphs. |
| Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching | Kristian Schwethelm, Daniel Rueckert, Georgios Kaissis | 2026-08-10 | 下载 | A main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more ... |
| How Accurately Can the Energy Use of Spark Applications Be Estimated Based on Resource Utilisation? | Youssef Moawad, Kathleen West, Vasilis Bountris, Philipp Thamm, Yehia Elkhatib, Lauritz Thamsen | 2026-08-10 | 下载 | Distributed batch data processing applications are widely executed on cloud-based resources where restricted user access to node-level hardware energy counters hinders transparent sustainability accou... |
| Beyond the Limits: Flexible and Congestion-Aware Cluster Scheduling for the Cloud | Oliver Larsson, Thijs Metsch, Cristian Klein, Erik Elmroth | 2026-08-10 | 下载 | Workload scheduling in cloud environments often relies on simplistic assumptions about application resource needs and hardware utilization. Overlooking application-level performance objectives and har... |
| UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on Edge | Tianhao Jiang, Hang Gu, Teng Wang, Qianyu Cheng, ZhenDong Zheng, Cheng Tang, Qiyue Su, Wenqi Lou, Lei Gong, Chao Wang, Xi Li, Xuehai Zhou | 2026-08-10 | 下载 | Edge LLM inference combines sparsity and low-bit quantization to meet device memory, latency, and power limits. Yet quantization shrinks weight payloads without proportionally reducing sparse metadata... |
| FEAST: Federated Shared-Space Training for Resource-Heterogeneous Clients | Bostan Khan, Masoud Daneshtalab | 2026-08-10 | 下载 | Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per deployment limit is costly. |
| Track me if you can: Ephemeral coin tracing | Ignacio Amores-Sesar, Christian Cachin, Rohit Chatterjee, Luiza Soezima, François-Xavier Wicht, Michelle Yeo | 2026-08-10 | 下载 | Privacy-preserving payment systems are well understood, yet their adoption in regulated settings, such as central bank digital currencies (CBDCs), institutional stablecoins, and other compliant paymen... |
| FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning | Radwan Selo, Majid Kundroo, Taehong Kim | 2026-08-10 | 下载 | Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL faces significant challenges due to data heterogeneity, partic... |
| FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning | Van Truong Vo, Khoa Nguyen, Taehong Kim | 2026-08-10 | 下载 | Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learning (DFL). |
| Beyond Fast Contractions: Attenuation and Recovery of Matrix-Engine Speedups in High-Order Finite Elements | Yinuo Wang, Lin Gan, Tianqi Mao, Zeyu Song, Wubing Wan, Jiayu Fu, Zekun Yin, Yuyang Jin, Xiaohui Duan, Wei Xue, Guangwen Yang | 2026-08-10 | 下载 | Modern processors increasingly provide matrix engines whose peak arithmetic throughput greatly exceeds conventional SIMD, but scientific applications rarely realize this advantage end to end. |
| A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume Metric | Gunika Verma, Aashutosh A, Pooja Srinivas, Yogesh Simmhan, Ayush Choure, Harshit Shah, Mayukh Das, Prashant Sasatte, Chetan Bansal, Abhijit Pai, Suraj Dixit, Achint Agrawal | 2026-08-10 | 下载 | Large-scale managed cloud databases leverage sophisticated load Packing and Migration (PAM) algorithms, which provide the efficiencies necessary for running these services at scale on cloud resources. |
| SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key Normalization | Gyudong Kim, Wonjun Han, Young Geun Kim | 2026-08-10 | 下载 | Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). However, under Tensor Parallelism (TP), layerwise QK-Norm introduces additional cr... |
| GPU-Accelerated Conic Quadratic Programming with Local Linear Convergence under Strict Complementarity | Hongpei Li, Yicheng Huang, Huikang Liu, Dongdong Ge, Yinyu Ye | 2026-08-10 | 下载 | We present PDHCG-CQP, a GPU-accelerated first-order solver for large-scale conic convex quadratic programming. PDHCG-CQP supports affine constraints and Cartesian products of nonnegative, second-order... |
| Universal Rendezvous of Anonymous Agents with Footprints | Bibhuti Das | 2026-08-10 | 下载 | Deterministic rendezvous for two anonymous mobile agents starting simultaneously from two distinct nodes of an anonymous connected graph and navigating synchronously in the graph requires that they me... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| ChronoSSM: Training for Temporally Aware Representations in Autoregressive State Space Models | Adrien Schoen, Nachiketa Ratnakar Patil, Arjun Bhagoji, Francesco Bronzino | 2026-08-10 | 下载 | Modern sequence models, from Transformers to State Space Models, have enabled powerful generative modeling across diverse domains, yet they are typically trained to predict what happens while treating... |
| A Bird's-Eye View on Security Considerations in RFCs | Jukka Ruohonen, Qusai Ramadan | 2026-08-10 | 下载 | Request for comments (RFCs) are Internet standards, memorandums, and related technical documents about core Internet protocols made via and released by the Internet Engineering Task Force (IETF). |
| Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable Cells | Emma Takács, Mátyás Hajós, Ádám Juniki, Ádám Fischer, Zoltán Komáromi, Kristóf Abai, Dániel Horváth, Sándor Máthé, Konstantinos Kousias, Bence Tipary | 2026-08-10 | 下载 | Human-Robot Collaboration (HRC) plays a vital role in dynamic, high mix, low volume industrial scenarios such as remanufacturing, which frequently face workcell rearrangements. |
| Abstractions for Network Intelligence: A Reference Architecture for AI at the Wireless Edge | Salil Reddy, Haohuang Wen, Ness Shroff, Venki Ramaswamy, Zhiqiang Lin, Elisa Bertino, Jim Kurose, Anish Arora | 2026-08-10 | 下载 | Networks are increasingly adopting AI as are AI applications leveraging networks. Awareness sharing between networks and AI applications promises to unlock higher levels of network utilization and app... |
| A Semantic Communication Approach to Fiducial Marker Processing in 5G-Enabled Edge SLAM | Boris Radovanovic, Vukan Ninkovic, Katarina Vidojevic, Buda Bajic Papuga, Dejan Vukobratovic | 2026-08-10 | 下载 | Autonomous robots increasingly rely on edge computing to offload computationally intensive perception tasks while maintaining real-time operation over 5G networks. |
| Enabling Beyond-Visual-Line-of-Sight Drones Operation over Open RAN 5G Networks with Slicing | Pau Baguer, Esteban Municio, Gines Garcia-Aviles, Xavier Costa-Pérez | 2026-08-10 | 下载 | Among the foretold claims of the transition from 5G to 6G, Beyond-Visual-Line-of-Sight (BVLoS) drone operation has emerged as a prominent Internet-of-Robots enabler. |
| Quantum-Classical Coexistence Network Tomography | Xuchuang Wang, Joseph C. Chapman, Aneesh Ramaswamy, Matheus Guedes de Andrade, Yu-Zhen Janice Chen, Joseph M. Lukens, Gayane Vardoyan, Don Towsley | 2026-08-10 | 下载 | Quantum-classical coexistence networks (QCNs) share optical fiber between quantum and classical signals via wavelength-division multiplexing, offering a practical path to quantum communication over ex... |
| White paper: A perspective on civilian-to-defence research transfer to SDD | Rute C. Sofia, Daniel Mendez, Simon Barner, Hao Shen, Julian Woermann, Andrea Stocco, Axel von Arnim, Holger Pfeifer, Alexander Pretschner | 2026-08-10 | 下载 | Military capability is increasingly determined by software. Yet defence platforms are procured on decade-long timescales, while the software and AI models they carry must evolve in days or hours. |
| Automated Synthesis of Deterministic Cross-Domain Interfaces | Konstantinos Christodoulopoulos, Antonis Selentis-Boulntadakis | 2026-08-10 | 下载 | Deterministic networking spans heterogeneous domains. At each boundary, two domains must agree on an assume--guarantee contract: what traffic the client may inject, and the QoS the carrier will hold f... |
| SparsePilot: Belief-Guided Network Planning under Sparse Wireless Measurements | Xuanhao Luo, Jiayuan Huang, Longyu Zhou, Mingzhe Chen, Yuchen Liu | 2026-08-10 | 下载 | Unmanned aerial vehicles (UAVs) have emerged as a promising solution for on-demand wireless coverage planning in urban environments. Existing learning-based UAV control methods, however, typically rel... |
| Active Hemispherical Metasurface Transmitarray Antenna for Wide-Angle 3D Beam Steering and Target Tracking | Somayeh Komeylian, Christopher Paolini | 2026-08-10 | 下载 | A stacked multilayer hemispherical graphene-based transmitarray antenna operating at 250 GHz is developed to achieve wide-angle 3D beam steering. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| What Actually Serializes GPU LZ77 Decode: Three Decoders, Three Mechanisms, and an Encode-Time Lever That Removes the Last One | Yakiv Shavidze | 2026-08-10 | 下载 | The sequential part of GPU LZ77 decode is not where the field assumes it is. Across three decoder architectures on an H100 we measure that parse, not copy, holds 64-72% of device-resident decode time;... |
| Machine Shape and Hierarchical Blocking: A Mathematics of Arrays Formalization, with an Open Problem in Hierarchical Shape Occupancy | Lenore M Mullin | 2026-08-10 | 下载 | A companion empirical study found that dense matrix multiplication block sizes calibrated on Apple M1 Pro correspond to two cache-t formulas that mispredict badly on a dierent chip's known cache sizes... |
| Performance and Cost-Aware Cache Provisioning | Ridwanul Tanvir, George Kesidis | 2026-08-10 | 下载 | While traditional cache policy evaluations fix capacity - often at 0.1% of the dataset - and measure the resulting hit rate, practical edge-cloud deployments require balancing both storage and computa... |