Skip to content

2026-07-20 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
PIP-NTT: Towards a Scalable Memory-Parallelized Accelerator for Iterative NTT in PQCMalik Imran, Ayesha Khalid, Ciara Rafferty, Safiullah Khan, Muhammad Rashid, Maire O'Neill2026-07-20下载The iterative forward and inverse number theoretic transform (NTT) is a key component in lattice-based post-quantum cryptography (PQC), typically implemented using Cooley-Tukey and Gentleman-Sande but...
A Progressive Approach to Synthesizable RTL Design Generation Using LLMsXiangfei Kong, Tasnim Tabassum, Marwan Abdelwahab, Hao Zheng2026-07-20下载Large language models can generate register-transfer-level (RTL) designs directly from natural language specifications. Their failures, however, arise mostly from understanding rather than coding \cit...
Empowering On-Device Model Adaptation with an Edge AI Inference AcceleratorMateusz Piechocki, Alessandro Capotondi, Marek Kraft2026-07-20下载On-device model adaptation is essential to enable lifelong personalization on resource-constrained hardware, but compute, power, and memory limitations of such devices make end-to-end backpropagation ...
Hardware Mechanisms to Dynamically Throttle AI PerformanceHaiyue Ma, Lauren Malek, Joseph Forzani, David Wentzlaff2026-07-20下载As more capable AI models are increasingly integrated into critical computer systems, the lack of control over AI intent motivates safety mechanisms.
Fixed Point Exploration For CV-QKD IR QC-MET-LDPC Toward Hardware ImplementationGuilherme Vergne de Oliveira, Mauro Queiroz Nooblath Neto, Micael Andrade Dias, Francisco Revson Fernandes Pereira, Francisco Marcos de Assis, Valéria Loureiro da Silva, Nelson Alves Ferreira2026-07-20下载High-speed LDPC decoding is a major bottleneck in CV-QKD and motivates hardware acceleration with fixed-point arithmetic. This work compares SPA, MSA, and NMS under a unified low-SNR fixed-point frame...
SEAM-V: A Hybrid-Decoupled RISC-V Vector Processor with Backend-Visible EP Context for Sustained Vector ThroughputWeiying Wang, Zhiwei Zhang2026-07-20下载Data-parallel workloads in deep learning and scientific computing continue to drive demand for higher processor throughput, energy efficiency, and scalability.
New Number Formats for FFT IP Cores in Optical OFDM TransceiversLukas Krupp, Gustavo Magalhães Gomes De Souza, Sani Nassif, Norbert Wehn2026-07-20下载Many state-of-the-art DSP implementations use fixed-point arithmetic due to its reduced hardware complexity and high throughput compared to conventional floating-point arithmetic.
PRISM: Sensitivity-Aware PolynoMial PRuning for EffIcient Neural Network EncryptionSahaj Majavdia, Mahdi Taheri2026-07-20下载Structured pruning is essential for making neural network inference feasible under homomorphic encryption (HE), yet its impact on model reliability has remained unexplored.
Cross-Domain Acceleration of Open Modification Search: From Commodity Platforms to Emerging Memory and Storage DevicesSumukh Pinge, Chang Eun Song, Po-Kai Hsu, Zheyu Li, Ashkan Moradifirouzabadi, Yanru Chen, Xiangjin Wu, Wei-Chen Chen, Eric Pop, Shimeng Yu, H. -S. Philip Wong, Tajana Rosing, Mingu Kang2026-07-20下载Open modification search (OMS) in mass spectrometry (MS) is a data-intensive workload whose performance is dominantly limited by reference data movement rather than computation.
D-NOVA: In-Storage Retrieval Accelerator via Dual-Bound 3D NAND-Optimized Similarity Search with Vector AdaptationChang Eun Song, Sumukh Pinge, Tianqi Zhang, Sung Eun Kim, Tajana S. Rosing, Mingu Kang2026-07-20下载Retrieval-Augmented Generation (RAG) enhances the factual grounding of large language model (LLM) inference by retrieving relevant information from external knowledge bases.
Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA WorkflowsJinyuan Deng, Zhengrui Chen, Xufeng Wei, Tianyu Xing, Chenyi Wen, Cheng Zhuo2026-07-20下载LLM-driven agent systems have emerged as a promising paradigm for electronic design automation (EDA), demonstrating strong potential for automating complex design workflows.
Isolation Failure From Shared Storage: Characterizing and Exploiting Page-Cache SCA Leakage Across Containers and VMsAlon Abudraham, Xingyu Chen, Itamar Levi, Ari Trachtenberg2026-07-20下载Modern cloud platforms increasingly combine strong software isolation mechanisms with shared hardware resources to improve performance and resource efficiency.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
What Governs Decode Throughput in Absolute-Offset GPU LZ77? A Work-Granularity Mechanism and an Encode-Time Min-Match-Length LeverYakiv Shavidze2026-07-20下载The ACEAPEX line of work established a lossless LZ77 format whose back-references are absolute output positions, giving parallel, compressed-resident GPU decode with sub-millisecond region seek.
uSTM: A Lightweight and Efficient STM Supporting General Types and Deferred AbortsZachary Kent, Guy Blelloch, André Costa2026-07-20下载Software Transactional Memory (STM) systems allow developers to more easily exploit multicore architectures by wrapping arbitrary sequential code in transactions that are executed concurrently.
HyMCache: A KV Cache Framework for Multi-Turn LLM Serving with CXL-Hybrid MemoryHakbeom Jang, Inho Song, Sam H. Noh, Jongryool Kim2026-07-20下载Long-context, multi-turn, and agentic LLM workloads increasingly reuse previously processed context, making KV-cache reuse essential for reducing redundant computation.
ExpertPlex: A High-Goodput Disaggregated Serving System for MoE LLMs with Adaptive Persistent KernelsBingyang Wu, Chao Jin, Zili Zhang, Xinming Wei, Yinmin Zhong, Ruidong Zhu, Chengxu Yang, Xin Jin, Yuliang Liu2026-07-20下载LLMs scale Mixture-of-Experts (MoE) parameters for superior intelligence, but massive weights and dynamic computation impede efficient serving.
A Curvature-Aware Rank-Adaptive Distributed Augmented-Lagrangian Solver for Large-Scale SDPsHongpei Li, Huikang Liu, Dongdong Ge, Yinyu Ye2026-07-20下载We present CARDAL (Curvature-Aware Rank-Adaptive Distributed Augmented Lagrangian), a distributed multi-GPU solver for large-scale semidefinite programs (SDPs) based on a rank-adaptive Burer-Monteiro ...
AutoEncoder-Compressed Parallel Split Learning for Pre-trained Model Fine-TuningBas Meuwissen, Vasileios Tsouvalas, Nirvana Meratnia2026-07-20下载Distributed Fine-Tuning (DFT) of large-scale Foundation Models (FMs) on resource-constrained edge devices is limited by local compute constraints and communication overhead.
Entanglement geometry separates circuit cutting, classical hardness, and trainabilityMaria Gragera Garces, Sabina Drăgoi, Lirandë Pira2026-07-20下载Circuit cutting promises to scale quantum computations beyond current hardware, but variational quantum advantage also requires low cutting overhead, classical hardness, and trainability.
Mobius Learning: Cyclic Depth Folding in TransformersTongtian Zhu2026-07-20下载Transformer-based language models organize computation along an ordered depth axis, where shallow and deep blocks often develop distinct representational roles.
Byzantine Fault-Tolerant Post-Quantum Distributed Quorum SignaturesQuentin Kniep, Jakub Sliwinski, Roger Wattenhofer2026-07-20下载Threshold, aggregate, and multi-signatures -- which we collectively call quorum signatures -- certify that a quorum of nodes endorsed a statement, with a certificate as small as a single signature.
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient FixChangzheng Ma2026-07-20下载Multi-head Latent Attention (MLA) ships two implementations in Megatron-Core: an explicit form used for training and an absorbed form -- which slashes collective communication by gathering only the co...
PRISM: Sensitivity-Aware PolynoMial PRuning for EffIcient Neural Network EncryptionSahaj Majavdia, Mahdi Taheri2026-07-20下载Structured pruning is essential for making neural network inference feasible under homomorphic encryption (HE), yet its impact on model reliability has remained unexplored.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Uplink SRS-Based Real-Time Indoor Localization System over OpenAirInterfacePing-Yu Hsieh, Chieh-Chun Chen, Navid Nikaein, Ray-Guang Cheng2026-07-20下载Indoor localization is one of the important services for future 5G-Advanced and 6G systems. This paper presents an uplink Sounding Reference Signal (SRS)-based real-time indoor localization system imp...
Evaluating Power Control Strategies for UORA in IEEE 802.11be Systems with Capture EffectKuan-Chin Li, Ting-Wei Hung, Lain-Chyr Hwang, Pengwenlong Gu, Ray-Guang Cheng2026-07-20下载Uplink OFDMA-based random access (UORA) is a new channel access mechanism that supports uplink multiuser access in the new generation WiFi systems.
Quality over Quantity: Value-Driven Distributed Congestion Control for the Collective Perception ServiceTengfei Lyu, Florian A. Schiegg, Md Noor-A-Rahim, Dirk Pesch, Aisling O'Driscoll2026-07-20下载While the Collective Perception Service (CPS) enables the exchange of sensor information among Intelligent Transport System Stations (ITS-S'), frequent transmission of Collective Perception Messages (...
Cost-Aware Uplink MPQUIC Scheduling via Multi-Objective Bayesian OptimizationThanh Trung Nguyen, Thanh Le, Phi Le Nguyen, Kien Nguyen2026-07-20下载Multipath QUIC (MPQUIC) enables simultaneous uplink transmission over heterogeneous access networks such as Wi-Fi and LTE, improving reliability and performance.
AI Agent Communications in AI-Native 6G Network: Status, Challenges and OpportunitiesQiang Duan2026-07-20下载The rapid development of agentic AI and multi-agent systems is establishing AI agent communication as a fundamental requirement for the future Internet.
ClouDens: Operational Context-Aware Anomaly Detection for Large-scale Cloud System MonitoringThu T. H. Doan, Mohammad Saiful Islam, Andriy Miranskyy, Ngoc-Thanh Nguyen, Rogardt Heldal, Patrizio Pelliccione2026-07-20下载With the rapid growth of cloud computing infrastructures in scale and complexity, network monitoring for Large-scale Cloud Systems (LCSs) has become increasingly challenging, requiring automated and r...
Human Grounded Evaluation of Large Language Models for Optical Network AutomationKiarash Rezaei, Omran Ayoub, Paolo Monti, Carlos Natalino2026-07-20下载Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substantially across LLM families.
Enhanced Dynamic Beamwidth Selection-based THz MAC Protocol for Wireless Data Center NetworksMuhammad Absaruddin, Saim Ghafoor, Mubashir Husain Rehmani2026-07-20下载Terahertz (THz) wireless communication offers a promising alternative to traditional wired links in data centres (DCs), enabling ultra-high data rates, low latency, and greater scalability.
PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication NetworksWen Qiu, Zhiqiang He, Wei Zhao, Hiroshi Masui2026-07-20下载Most reinforcement learning controllers for these networks assume stationary conditions, and the few that handle change react to the external environment while leaving the network's internal state une...
Beyond Car Sharing: Uncertainty-Aware Pooling of Vehicular Compute at the Network EdgeWellington Lobato, Nadjib Achir, Aline Carneiro Viana2026-07-20下载Connected vehicles increasingly embed AI accelerators, offering a substantial yet volatile source of supplemental compute near the network edge.
Mobile Network Control with a World ModelMaxime Bouton, Ioanna Mitsioni, Simon Lindståhl, Jaeseong Jeong2026-07-20下载The increasing complexity of mobile networks necessitates intelligent and dynamic control strategies for efficient, energy-conserving management.
Token Communications (TokCom): A Unified AI-Native Communication FrameworkYaru Fu, Liang Ji, Sabita Maharjan, Tony Q. S. Quek2026-07-20下载As artificial intelligence (AI) evolves from static perception to generative reasoning and autonomous agency, the fundamental principles of wireless communications are undergoing a paradigm shift.
Self-Directed Spectrum Allocation Framework for Integrated TN-NTN 6G NetworksVaskar Chakma, Wooyeol Choi2026-07-20下载This paper proposes a self-adaptive channel assignment framework based on Q-learning, where agents learn optimal policies by observing network load, interference conditions, and temporal traffic dynam...
Diverge-Merge Formation and MAC Control in Structured AirspaceKai Xiong, Xingyu Wu, Ba Zhang, Li Wei, Min Zeng, Supeng Leng2026-07-20下载The rapid scaling of advanced air mobility (AAM) makes corridor-based structured airspace a promising infrastructure for high-density unmanned aerial vehicle (UAV) traffic.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory MinimizationAlex Mathai, Shobini Iyer, Aleksandr Nogikh, Petros Maniatis, Franjo Ivancic, Junfeng Yang, Baishakhi Ray2026-07-20下载Coding agents are increasingly used to accelerate code generation in many downstream tasks, such as fixing bugs, building applications, and prototyping.
SuperPass: Fast-Tracking Blocking Threads to Mitigate Priority Inversion on Mobile DevicesLei Li, Yu Liang, Riwei Pan, Youcheng Sun, Nan Guan, Tei-Wei Kuo, Chun Jason Xue2026-07-20下载Priority inversion occurs when a high-priority thread is delayed by a lower-priority one. Although well studied in real-time systems, its impact in general-purpose OSes (e.g.
Isolation Failure From Shared Storage: Characterizing and Exploiting Page-Cache SCA Leakage Across Containers and VMsAlon Abudraham, Xingyu Chen, Itamar Levi, Ari Trachtenberg2026-07-20下载Modern cloud platforms increasingly combine strong software isolation mechanisms with shared hardware resources to improve performance and resource efficiency.

cs.PF - Performance ​

标题作者发布日期PDF摘要
What Governs Decode Throughput in Absolute-Offset GPU LZ77? A Work-Granularity Mechanism and an Encode-Time Min-Match-Length LeverYakiv Shavidze2026-07-20下载The ACEAPEX line of work established a lossless LZ77 format whose back-references are absolute output positions, giving parallel, compressed-resident GPU decode with sub-millisecond region seek.
A Taxonomy of Distance Metrics for Time-Sensitive Importance Splitting: Timer Bounds, Resampling, and the Global AgeGabriel Dengler, Carlos E. Budde, Laura Carnevali2026-07-20下载Importance splitting (ISPLIT) evaluates the probabilities of rare events in non-Markovian models. It requires a heuristic importance function (IFUN) that estimates the distance to the target.
Cross-Domain Acceleration of Open Modification Search: From Commodity Platforms to Emerging Memory and Storage DevicesSumukh Pinge, Chang Eun Song, Po-Kai Hsu, Zheyu Li, Ashkan Moradifirouzabadi, Yanru Chen, Xiangjin Wu, Wei-Chen Chen, Eric Pop, Shimeng Yu, H. -S. Philip Wong, Tajana Rosing, Mingu Kang2026-07-20下载Open modification search (OMS) in mass spectrometry (MS) is a data-intensive workload whose performance is dominantly limited by reference data movement rather than computation.
SALT: Salience-Aware Lexical Trie for Long-Context CompressionOteo Mamo, Hyunjin Yi, Joydhriti Choudhury, Shangqian Gao, Weikuan Yu2026-07-20下载As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottlenecks in inference systems.

基于 VitePress 构建 · 使用本地搜索查找论文