2026-07-20
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| PIP-NTT: Towards a Scalable Memory-Parallelized Accelerator for Iterative NTT in PQC | Malik Imran, Ayesha Khalid, Ciara Rafferty, Safiullah Khan, Muhammad Rashid, Maire O'Neill | 2026-07-20 | 下载 | The iterative forward and inverse number theoretic transform (NTT) is a key component in lattice-based post-quantum cryptography (PQC), typically implemented using Cooley-Tukey and Gentleman-Sande but... |
| A Progressive Approach to Synthesizable RTL Design Generation Using LLMs | Xiangfei Kong, Tasnim Tabassum, Marwan Abdelwahab, Hao Zheng | 2026-07-20 | 下载 | Large language models can generate register-transfer-level (RTL) designs directly from natural language specifications. Their failures, however, arise mostly from understanding rather than coding \cit... |
| Empowering On-Device Model Adaptation with an Edge AI Inference Accelerator | Mateusz Piechocki, Alessandro Capotondi, Marek Kraft | 2026-07-20 | 下载 | On-device model adaptation is essential to enable lifelong personalization on resource-constrained hardware, but compute, power, and memory limitations of such devices make end-to-end backpropagation ... |
| Hardware Mechanisms to Dynamically Throttle AI Performance | Haiyue Ma, Lauren Malek, Joseph Forzani, David Wentzlaff | 2026-07-20 | 下载 | As more capable AI models are increasingly integrated into critical computer systems, the lack of control over AI intent motivates safety mechanisms. |
| Fixed Point Exploration For CV-QKD IR QC-MET-LDPC Toward Hardware Implementation | Guilherme Vergne de Oliveira, Mauro Queiroz Nooblath Neto, Micael Andrade Dias, Francisco Revson Fernandes Pereira, Francisco Marcos de Assis, Valéria Loureiro da Silva, Nelson Alves Ferreira | 2026-07-20 | 下载 | High-speed LDPC decoding is a major bottleneck in CV-QKD and motivates hardware acceleration with fixed-point arithmetic. This work compares SPA, MSA, and NMS under a unified low-SNR fixed-point frame... |
| SEAM-V: A Hybrid-Decoupled RISC-V Vector Processor with Backend-Visible EP Context for Sustained Vector Throughput | Weiying Wang, Zhiwei Zhang | 2026-07-20 | 下载 | Data-parallel workloads in deep learning and scientific computing continue to drive demand for higher processor throughput, energy efficiency, and scalability. |
| New Number Formats for FFT IP Cores in Optical OFDM Transceivers | Lukas Krupp, Gustavo Magalhães Gomes De Souza, Sani Nassif, Norbert Wehn | 2026-07-20 | 下载 | Many state-of-the-art DSP implementations use fixed-point arithmetic due to its reduced hardware complexity and high throughput compared to conventional floating-point arithmetic. |
| PRISM: Sensitivity-Aware PolynoMial PRuning for EffIcient Neural Network Encryption | Sahaj Majavdia, Mahdi Taheri | 2026-07-20 | 下载 | Structured pruning is essential for making neural network inference feasible under homomorphic encryption (HE), yet its impact on model reliability has remained unexplored. |
| Cross-Domain Acceleration of Open Modification Search: From Commodity Platforms to Emerging Memory and Storage Devices | Sumukh Pinge, Chang Eun Song, Po-Kai Hsu, Zheyu Li, Ashkan Moradifirouzabadi, Yanru Chen, Xiangjin Wu, Wei-Chen Chen, Eric Pop, Shimeng Yu, H. -S. Philip Wong, Tajana Rosing, Mingu Kang | 2026-07-20 | 下载 | Open modification search (OMS) in mass spectrometry (MS) is a data-intensive workload whose performance is dominantly limited by reference data movement rather than computation. |
| D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector Adaptation | Chang Eun Song, Sumukh Pinge, Tianqi Zhang, Sung Eun Kim, Tajana S. Rosing, Mingu Kang | 2026-07-20 | 下载 | Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases. |
| Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows | Jinyuan Deng, Zhengrui Chen, Xufeng Wei, Tianyu Xing, Chenyi Wen, Cheng Zhuo | 2026-07-20 | 下载 | LLM-driven agent systems have emerged as a promising paradigm for electronic design automation (EDA), demonstrating strong potential for automating complex design workflows. |
| Isolation Failure From Shared Storage: Characterizing and Exploiting Page-Cache SCA Leakage Across Containers and VMs | Alon Abudraham, Xingyu Chen, Itamar Levi, Ari Trachtenberg | 2026-07-20 | 下载 | Modern cloud platforms increasingly combine strong software isolation mechanisms with shared hardware resources to improve performance and resource efficiency. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| What Governs Decode Throughput in Absolute-Offset GPU LZ77? A Work-Granularity Mechanism and an Encode-Time Min-Match-Length Lever | Yakiv Shavidze | 2026-07-20 | 下载 | The ACEAPEX line of work established a lossless LZ77 format whose back-references are absolute output positions, giving parallel, compressed-resident GPU decode with sub-millisecond region seek. |
| uSTM: A Lightweight and Efficient STM Supporting General Types and Deferred Aborts | Zachary Kent, Guy Blelloch, André Costa | 2026-07-20 | 下载 | Software Transactional Memory (STM) systems allow developers to more easily exploit multicore architectures by wrapping arbitrary sequential code in transactions that are executed concurrently. |
| HyMCache: A KV Cache Framework for Multi-Turn LLM Serving with CXL-Hybrid Memory | Hakbeom Jang, Inho Song, Sam H. Noh, Jongryool Kim | 2026-07-20 | 下载 | Long-context, multi-turn, and agentic LLM workloads increasingly reuse previously processed context, making KV-cache reuse essential for reducing redundant computation. |
| ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent Kernels | Bingyang Wu, Chao Jin, Zili Zhang, Xinming Wei, Yinmin Zhong, Ruidong Zhu, Chengxu Yang, Xin Jin, Yuliang Liu | 2026-07-20 | 下载 | LLMs scale Mixture-of-Experts (MoE) parameters for superior intelligence, but massive weights and dynamic computation impede efficient serving. |
| A Curvature-Aware Rank-Adaptive Distributed Augmented-Lagrangian Solver for Large-Scale SDPs | Hongpei Li, Huikang Liu, Dongdong Ge, Yinyu Ye | 2026-07-20 | 下载 | We present CARDAL (Curvature-Aware Rank-Adaptive Distributed Augmented Lagrangian), a distributed multi-GPU solver for large-scale semidefinite programs (SDPs) based on a rank-adaptive Burer-Monteiro ... |
| AutoEncoder-Compressed Parallel Split Learning for Pre-trained Model Fine-Tuning | Bas Meuwissen, Vasileios Tsouvalas, Nirvana Meratnia | 2026-07-20 | 下载 | Distributed Fine-Tuning (DFT) of large-scale Foundation Models (FMs) on resource-constrained edge devices is limited by local compute constraints and communication overhead. |
| Entanglement geometry separates circuit cutting, classical hardness, and trainability | Maria Gragera Garces, Sabina Drăgoi, Lirandë Pira | 2026-07-20 | 下载 | Circuit cutting promises to scale quantum computations beyond current hardware, but variational quantum advantage also requires low cutting overhead, classical hardness, and trainability. |
| Mobius Learning: Cyclic Depth Folding in Transformers | Tongtian Zhu | 2026-07-20 | 下载 | Transformer-based language models organize computation along an ordered depth axis, where shallow and deep blocks often develop distinct representational roles. |
| Byzantine Fault-Tolerant Post-Quantum Distributed Quorum Signatures | Quentin Kniep, Jakub Sliwinski, Roger Wattenhofer | 2026-07-20 | 下载 | Threshold, aggregate, and multi-signatures -- which we collectively call quorum signatures -- certify that a quorum of nodes endorsed a statement, with a certificate as small as a single signature. |
| A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix | Changzheng Ma | 2026-07-20 | 下载 | Multi-head Latent Attention (MLA) ships two implementations in Megatron-Core: an explicit form used for training and an absorbed form -- which slashes collective communication by gathering only the co... |
| PRISM: Sensitivity-Aware PolynoMial PRuning for EffIcient Neural Network Encryption | Sahaj Majavdia, Mahdi Taheri | 2026-07-20 | 下载 | Structured pruning is essential for making neural network inference feasible under homomorphic encryption (HE), yet its impact on model reliability has remained unexplored. |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Uplink SRS-Based Real-Time Indoor Localization System over OpenAirInterface | Ping-Yu Hsieh, Chieh-Chun Chen, Navid Nikaein, Ray-Guang Cheng | 2026-07-20 | 下载 | Indoor localization is one of the important services for future 5G-Advanced and 6G systems. This paper presents an uplink Sounding Reference Signal (SRS)-based real-time indoor localization system imp... |
| Evaluating Power Control Strategies for UORA in IEEE 802.11be Systems with Capture Effect | Kuan-Chin Li, Ting-Wei Hung, Lain-Chyr Hwang, Pengwenlong Gu, Ray-Guang Cheng | 2026-07-20 | 下载 | Uplink OFDMA-based random access (UORA) is a new channel access mechanism that supports uplink multiuser access in the new generation WiFi systems. |
| Quality over Quantity: Value-Driven Distributed Congestion Control for the Collective Perception Service | Tengfei Lyu, Florian A. Schiegg, Md Noor-A-Rahim, Dirk Pesch, Aisling O'Driscoll | 2026-07-20 | 下载 | While the Collective Perception Service (CPS) enables the exchange of sensor information among Intelligent Transport System Stations (ITS-S'), frequent transmission of Collective Perception Messages (... |
| Cost-Aware Uplink MPQUIC Scheduling via Multi-Objective Bayesian Optimization | Thanh Trung Nguyen, Thanh Le, Phi Le Nguyen, Kien Nguyen | 2026-07-20 | 下载 | Multipath QUIC (MPQUIC) enables simultaneous uplink transmission over heterogeneous access networks such as Wi-Fi and LTE, improving reliability and performance. |
| AI Agent Communications in AI-Native 6G Network: Status, Challenges and Opportunities | Qiang Duan | 2026-07-20 | 下载 | The rapid development of agentic AI and multi-agent systems is establishing AI agent communication as a fundamental requirement for the future Internet. |
| ClouDens: Operational Context-Aware Anomaly Detection for Large-scale Cloud System Monitoring | Thu T. H. Doan, Mohammad Saiful Islam, Andriy Miranskyy, Ngoc-Thanh Nguyen, Rogardt Heldal, Patrizio Pelliccione | 2026-07-20 | 下载 | With the rapid growth of cloud computing infrastructures in scale and complexity, network monitoring for Large-scale Cloud Systems (LCSs) has become increasingly challenging, requiring automated and r... |
| Human Grounded Evaluation of Large Language Models for Optical Network Automation | Kiarash Rezaei, Omran Ayoub, Paolo Monti, Carlos Natalino | 2026-07-20 | 下载 | Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families. |
| Enhanced Dynamic Beamwidth Selection-based THz MAC Protocol for Wireless Data Center Networks | Muhammad Absaruddin, Saim Ghafoor, Mubashir Husain Rehmani | 2026-07-20 | 下载 | Terahertz (THz) wireless communication offers a promising alternative to traditional wired links in data centres (DCs), enabling ultra-high data rates, low latency, and greater scalability. |
| PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication Networks | Wen Qiu, Zhiqiang He, Wei Zhao, Hiroshi Masui | 2026-07-20 | 下载 | Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state une... |
| Beyond Car Sharing: Uncertainty-Aware Pooling of Vehicular Compute at the Network Edge | Wellington Lobato, Nadjib Achir, Aline Carneiro Viana | 2026-07-20 | 下载 | Connected vehicles increasingly embed AI accelerators, offering a substantial yet volatile source of supplemental compute near the network edge. |
| Mobile Network Control with a World Model | Maxime Bouton, Ioanna Mitsioni, Simon Lindståhl, Jaeseong Jeong | 2026-07-20 | 下载 | The increasing complexity of mobile networks necessitates intelligent and dynamic control strategies for efficient, energy-conserving management. |
| Token Communications (TokCom): A Unified AI-Native Communication Framework | Yaru Fu, Liang Ji, Sabita Maharjan, Tony Q. S. Quek | 2026-07-20 | 下载 | As artificial intelligence (AI) evolves from static perception to generative reasoning and autonomous agency, the fundamental principles of wireless communications are undergoing a paradigm shift. |
| Self-Directed Spectrum Allocation Framework for Integrated TN-NTN 6G Networks | Vaskar Chakma, Wooyeol Choi | 2026-07-20 | 下载 | This paper proposes a self-adaptive channel assignment framework based on Q-learning, where agents learn optimal policies by observing network load, interference conditions, and temporal traffic dynam... |
| Diverge-Merge Formation and MAC Control in Structured Airspace | Kai Xiong, Xingyu Wu, Ba Zhang, Li Wei, Min Zeng, Supeng Leng | 2026-07-20 | 下载 | The rapid scaling of advanced air mobility (AAM) makes corridor-based structured airspace a promising infrastructure for high-density unmanned aerial vehicle (UAV) traffic. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization | Alex Mathai, Shobini Iyer, Aleksandr Nogikh, Petros Maniatis, Franjo Ivancic, Junfeng Yang, Baishakhi Ray | 2026-07-20 | 下载 | Coding agents are increasingly used to accelerate code generation in many downstream tasks, such as fixing bugs, building applications, and prototyping. |
| SuperPass: Fast-Tracking Blocking Threads to Mitigate Priority Inversion on Mobile Devices | Lei Li, Yu Liang, Riwei Pan, Youcheng Sun, Nan Guan, Tei-Wei Kuo, Chun Jason Xue | 2026-07-20 | 下载 | Priority inversion occurs when a high-priority thread is delayed by a lower-priority one. Although well studied in real-time systems, its impact in general-purpose OSes (e.g. |
| Isolation Failure From Shared Storage: Characterizing and Exploiting Page-Cache SCA Leakage Across Containers and VMs | Alon Abudraham, Xingyu Chen, Itamar Levi, Ari Trachtenberg | 2026-07-20 | 下载 | Modern cloud platforms increasingly combine strong software isolation mechanisms with shared hardware resources to improve performance and resource efficiency. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| What Governs Decode Throughput in Absolute-Offset GPU LZ77? A Work-Granularity Mechanism and an Encode-Time Min-Match-Length Lever | Yakiv Shavidze | 2026-07-20 | 下载 | The ACEAPEX line of work established a lossless LZ77 format whose back-references are absolute output positions, giving parallel, compressed-resident GPU decode with sub-millisecond region seek. |
| A Taxonomy of Distance Metrics for Time-Sensitive Importance Splitting: Timer Bounds, Resampling, and the Global Age | Gabriel Dengler, Carlos E. Budde, Laura Carnevali | 2026-07-20 | 下载 | Importance splitting (ISPLIT) evaluates the probabilities of rare events in non-Markovian models. It requires a heuristic importance function (IFUN) that estimates the distance to the target. |
| Cross-Domain Acceleration of Open Modification Search: From Commodity Platforms to Emerging Memory and Storage Devices | Sumukh Pinge, Chang Eun Song, Po-Kai Hsu, Zheyu Li, Ashkan Moradifirouzabadi, Yanru Chen, Xiangjin Wu, Wei-Chen Chen, Eric Pop, Shimeng Yu, H. -S. Philip Wong, Tajana Rosing, Mingu Kang | 2026-07-20 | 下载 | Open modification search (OMS) in mass spectrometry (MS) is a data-intensive workload whose performance is dominantly limited by reference data movement rather than computation. |
| SALT: Salience-Aware Lexical Trie for Long-Context Compression | Oteo Mamo, Hyunjin Yi, Joydhriti Choudhury, Shangqian Gao, Weikuan Yu | 2026-07-20 | 下载 | As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottlenecks in inference systems. |