Skip to content

2026-08-10 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Monophonic Audio Synthesizer Using FPGAsMichael Smith, D. G. Perera2026-08-10下载Signal synthesis is used in every aspect of the electronics world, where sinusoidal waveforms are used to perform functions such as clocking, signal transmission, feedback controls, and other applicat...
ArchAgent v2: A Case Study with the Data Prefetching ChampionshipAbraham Gonzalez, Raghav Gupta, Akanksha Jain, Hanna Alam, Alexander Novikov, Po-Sen Huang, Matej Balog, Marvin Eisenberger, Sergey Shirobokov, Ngân Vũ, Hank Levy, Borivoje Nikolić, Sagar Karandikar, Martin Dixon, Parthasarathy Ranganathan2026-08-10下载Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search sp...
FSGen: Agile Fused and Sparse Accelerator Generator with Accurate Power Model for LLM ApplicationsJay Zhe-An Mok, Qijun Zhang, Zhiyao Xie2026-08-10下载With the growing demand of artificial intelligence (AI) applications, large language models (LLMs) have become important workloads in many domains.
SLAC: Access-Driven CPU-to-GPU Side-channel Attacks via System-Level Cache on Apple SiliconTianhong Xu, Saion K. Roy, Ruyi Ding, Aidong Adam Ding, Yunsi Fei2026-08-10下载Modern heterogeneous System-on-Chip designs integrate CPU cores and a GPU that share a last-level cache (LLC) or system-level cache (SLC). This sharing exposes a new cross-domain attack surface, and e...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
What Actually Serializes GPU LZ77 Decode: Three Decoders, Three Mechanisms, and an Encode-Time Lever That Removes the Last OneYakiv Shavidze2026-08-10下载The sequential part of GPU LZ77 decode is not where the field assumes it is. Across three decoder architectures on an H100 we measure that parse, not copy, holds 64-72% of device-resident decode time;...
SeFoRA: Sketch-Aggregated Federated Low-Rank Adaptation with Heterogeneous Client RanksYue Xia, Tayyebeh Jahani-Nezhad, Mayank Bakshi, Rawad Bitar2026-08-10下载We consider federated parameter efficient fine-tuning of large neural networks with low-rank adaptation (LoRA,~Hu et al.\ 2022). Combining LoRA with federated PEFT introduces challenges absent from ei...
Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4Matt J. Borowski, Blazej Osinski2026-08-10下载High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matrix loads with ldmatrix, and matrix multiply-accumulate operati...
Thread Scaling of Hexaly on the TDVRPTW across Two Model Encodings. An Experimental Report: External-Function Serialization, Slice-Count Choice, and the Time-Sliced Thread LadderFlorian Rascoussier2026-08-10下载Thread scaling in Hexaly depends first on how the Time-Dependent Vehicle Routing Problem with Time Windows is modeled. We compare two Python encodings: one evaluates continuous travel-time functions t...
Certified Split Windows for Parallel Lexing: Recovering Boundaries Where No Byte CertifiesNicklas Nidhögg2026-08-10下载A certified split point lets a parallel lexer cut unlexed input at a single byte with the serial token stream provably preserved, but several conventional token sets in the predecessor's controlled st...
Defining Decentralization: An Ontological PerspectiveJakub Kacper Szeląg, Aydin Abadi, Mohammad Naseri2026-08-10下载Decentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as security, distributed computing, artificial intelligence, clo...
Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive ApproachXinyi Xu, Bingnan Xiao, Shuang Qin, Gang Feng, Tony Q. S. Quek2026-08-10下载Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., AA and BB, providing an efficient way to fine-tune large models in federated learning ...
SpSYRK: Half the Work in Distributed Sparse Matrix MultiplicationThomas McFarland, Julian Bellavita, Giulia Guidi2026-08-10下载The symmetric rank-kk update (SYRK), \C = \A\A^\top, computes the dot product between each pair of rows of \A, producing the Gram matrix \C.
A Preliminary Study on Simultaneous Coscheduling for Discrete GPU vs. Fused GPUPoorna Gunathilaka, Nabayan Chaudhury, Kirshanthan Sundararajah, Wu-chun Feng2026-08-10下载CPU-GPU coscheduling enables simultaneous execution of an application across both processing units, but its efficiency depends on workload partitioning and memory architecture.
O~\tilde{\text{O}}ptimal Distributed Maximum Flow Approximation in Undirected Planar GraphsYaseen Abd-Elhaleem, Michal Dory, Oren Weimann2026-08-10下载Persistent efforts in recent years have been devoted to devising distributed algorithms for fundamental optimization problems in planar graphs.
Depth-adaptive Inference of Looped Language Models via Continuous Depth BatchingKristian Schwethelm, Daniel Rueckert, Georgios Kaissis2026-08-10下载A main promise of looped language models (LMs) is depth-adaptive inference. By iterating a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more ...
How Accurately Can the Energy Use of Spark Applications Be Estimated Based on Resource Utilisation?Youssef Moawad, Kathleen West, Vasilis Bountris, Philipp Thamm, Yehia Elkhatib, Lauritz Thamsen2026-08-10下载Distributed batch data processing applications are widely executed on cloud-based resources where restricted user access to node-level hardware energy counters hinders transparent sustainability accou...
Beyond the Limits: Flexible and Congestion-Aware Cluster Scheduling for the CloudOliver Larsson, Thijs Metsch, Cristian Klein, Erik Elmroth2026-08-10下载Workload scheduling in cloud environments often relies on simplistic assumptions about application resource needs and hardware utilization. Overlooking application-level performance objectives and har...
UnionSparse: An Index-Efficient Sparsity Framework for Low-Bit Sparse LLM Inference on EdgeTianhao Jiang, Hang Gu, Teng Wang, Qianyu Cheng, ZhenDong Zheng, Cheng Tang, Qiyue Su, Wenqi Lou, Lei Gong, Chao Wang, Xi Li, Xuehai Zhou2026-08-10下载Edge LLM inference combines sparsity and low-bit quantization to meet device memory, latency, and power limits. Yet quantization shrinks weight payloads without proportionally reducing sparse metadata...
FEAST: Federated Shared-Space Training for Resource-Heterogeneous ClientsBostan Khan, Masoud Daneshtalab2026-08-10下载Federated learning (FL) must serve devices with varying computational capabilities. A fixed model cannot suit all devices, while training one model per deployment limit is costly.
Track me if you can: Ephemeral coin tracingIgnacio Amores-Sesar, Christian Cachin, Rohit Chatterjee, Luiza Soezima, François-Xavier Wicht, Michelle Yeo2026-08-10下载Privacy-preserving payment systems are well understood, yet their adoption in regulated settings, such as central bank digital currencies (CBDCs), institutional stablecoins, and other compliant paymen...
FedTVD: Balancing Data Quality and Quantity for Robust Federated LearningRadwan Selo, Majid Kundroo, Taehong Kim2026-08-10下载Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL faces significant challenges due to data heterogeneity, partic...
FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learningVan Truong Vo, Khoa Nguyen, Taehong Kim2026-08-10下载Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learning (DFL).
Beyond Fast Contractions: Attenuation and Recovery of Matrix-Engine Speedups in High-Order Finite ElementsYinuo Wang, Lin Gan, Tianqi Mao, Zeyu Song, Wubing Wan, Jiayu Fu, Zekun Yin, Yuyang Jin, Xiaohui Duan, Wei Xue, Guangwen Yang2026-08-10下载Modern processors increasingly provide matrix engines whose peak arithmetic throughput greatly exceeds conventional SIMD, but scientific applications rarely realize this advantage end to end.
A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume MetricGunika Verma, Aashutosh A, Pooja Srinivas, Yogesh Simmhan, Ayush Choure, Harshit Shah, Mayukh Das, Prashant Sasatte, Chetan Bansal, Abhijit Pai, Suraj Dixit, Achint Agrawal2026-08-10下载Large-scale managed cloud databases leverage sophisticated load Packing and Migration (PAM) algorithms, which provide the efficiencies necessary for running these services at scale on cloud resources.
SwiftQK: Fast and Communication-Efficient Tensor Parallelism for Query-Key NormalizationGyudong Kim, Wonjun Han, Young Geun Kim2026-08-10下载Query-Key Normalization (QK-Norm) improves the training stability and quality of modern Large Language Models (LLMs). However, under Tensor Parallelism (TP), layerwise QK-Norm introduces additional cr...
GPU-Accelerated Conic Quadratic Programming with Local Linear Convergence under Strict ComplementarityHongpei Li, Yicheng Huang, Huikang Liu, Dongdong Ge, Yinyu Ye2026-08-10下载We present PDHCG-CQP, a GPU-accelerated first-order solver for large-scale conic convex quadratic programming. PDHCG-CQP supports affine constraints and Cartesian products of nonnegative, second-order...
Universal Rendezvous of Anonymous Agents with FootprintsBibhuti Das2026-08-10下载Deterministic rendezvous for two anonymous mobile agents starting simultaneously from two distinct nodes of an anonymous connected graph and navigating synchronously in the graph requires that they me...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
ChronoSSM: Training for Temporally Aware Representations in Autoregressive State Space ModelsAdrien Schoen, Nachiketa Ratnakar Patil, Arjun Bhagoji, Francesco Bronzino2026-08-10下载Modern sequence models, from Transformers to State Space Models, have enabled powerful generative modeling across diverse domains, yet they are typically trained to predict what happens while treating...
A Bird's-Eye View on Security Considerations in RFCsJukka Ruohonen, Qusai Ramadan2026-08-10下载Request for comments (RFCs) are Internet standards, memorandums, and related technical documents about core Internet protocols made via and released by the Internet Engineering Task Force (IETF).
Removing Infrastructure Barriers in Human-Robot Collaboration Through Wireless Reconfigurable CellsEmma Takács, Mátyás Hajós, Ádám Juniki, Ádám Fischer, Zoltán Komáromi, Kristóf Abai, Dániel Horváth, Sándor Máthé, Konstantinos Kousias, Bence Tipary2026-08-10下载Human-Robot Collaboration (HRC) plays a vital role in dynamic, high mix, low volume industrial scenarios such as remanufacturing, which frequently face workcell rearrangements.
Abstractions for Network Intelligence: A Reference Architecture for AI at the Wireless EdgeSalil Reddy, Haohuang Wen, Ness Shroff, Venki Ramaswamy, Zhiqiang Lin, Elisa Bertino, Jim Kurose, Anish Arora2026-08-10下载Networks are increasingly adopting AI as are AI applications leveraging networks. Awareness sharing between networks and AI applications promises to unlock higher levels of network utilization and app...
A Semantic Communication Approach to Fiducial Marker Processing in 5G-Enabled Edge SLAMBoris Radovanovic, Vukan Ninkovic, Katarina Vidojevic, Buda Bajic Papuga, Dejan Vukobratovic2026-08-10下载Autonomous robots increasingly rely on edge computing to offload computationally intensive perception tasks while maintaining real-time operation over 5G networks.
Enabling Beyond-Visual-Line-of-Sight Drones Operation over Open RAN 5G Networks with SlicingPau Baguer, Esteban Municio, Gines Garcia-Aviles, Xavier Costa-Pérez2026-08-10下载Among the foretold claims of the transition from 5G to 6G, Beyond-Visual-Line-of-Sight (BVLoS) drone operation has emerged as a prominent Internet-of-Robots enabler.
Quantum-Classical Coexistence Network TomographyXuchuang Wang, Joseph C. Chapman, Aneesh Ramaswamy, Matheus Guedes de Andrade, Yu-Zhen Janice Chen, Joseph M. Lukens, Gayane Vardoyan, Don Towsley2026-08-10下载Quantum-classical coexistence networks (QCNs) share optical fiber between quantum and classical signals via wavelength-division multiplexing, offering a practical path to quantum communication over ex...
White paper: A perspective on civilian-to-defence research transfer to SDDRute C. Sofia, Daniel Mendez, Simon Barner, Hao Shen, Julian Woermann, Andrea Stocco, Axel von Arnim, Holger Pfeifer, Alexander Pretschner2026-08-10下载Military capability is increasingly determined by software. Yet defence platforms are procured on decade-long timescales, while the software and AI models they carry must evolve in days or hours.
Automated Synthesis of Deterministic Cross-Domain InterfacesKonstantinos Christodoulopoulos, Antonis Selentis-Boulntadakis2026-08-10下载Deterministic networking spans heterogeneous domains. At each boundary, two domains must agree on an assume--guarantee contract: what traffic the client may inject, and the QoS the carrier will hold f...
SparsePilot: Belief-Guided Network Planning under Sparse Wireless MeasurementsXuanhao Luo, Jiayuan Huang, Longyu Zhou, Mingzhe Chen, Yuchen Liu2026-08-10下载Unmanned aerial vehicles (UAVs) have emerged as a promising solution for on-demand wireless coverage planning in urban environments. Existing learning-based UAV control methods, however, typically rel...
Active Hemispherical Metasurface Transmitarray Antenna for Wide-Angle 3D Beam Steering and Target TrackingSomayeh Komeylian, Christopher Paolini2026-08-10下载A stacked multilayer hemispherical graphene-based transmitarray antenna operating at 250 GHz is developed to achieve wide-angle 3D beam steering.

cs.PF - Performance ​

标题作者发布日期PDF摘要
What Actually Serializes GPU LZ77 Decode: Three Decoders, Three Mechanisms, and an Encode-Time Lever That Removes the Last OneYakiv Shavidze2026-08-10下载The sequential part of GPU LZ77 decode is not where the field assumes it is. Across three decoder architectures on an H100 we measure that parse, not copy, holds 64-72% of device-resident decode time;...
Machine Shape and Hierarchical Blocking: A Mathematics of Arrays Formalization, with an Open Problem in Hierarchical Shape OccupancyLenore M Mullin2026-08-10下载A companion empirical study found that dense matrix multiplication block sizes calibrated on Apple M1 Pro correspond to two cache-t formulas that mispredict badly on a dierent chip's known cache sizes...
Performance and Cost-Aware Cache ProvisioningRidwanul Tanvir, George Kesidis2026-08-10下载While traditional cache policy evaluations fix capacity - often at 0.1% of the dataset - and measure the resulting hit rate, practical edge-cloud deployments require balancing both storage and computa...

基于 VitePress 构建 · 使用本地搜索查找论文