Skip to content

2026-06-07 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Accuracy-Configurable Floating-Point Multiplier Design for SRAM-Based Compute-in-MemoryYiqi Zhou, Junhao Lu, Jiale Yu, Zhuo Xu, Yang He, Yue Yuan, Shan Shen, Daying Sun2026-06-07下载Digital Compute-in-Memory (DCiM) reduces data movement and has become a promising solution for energy-efficient edge AI. However, most existing DCiM frameworks still primarily target integer or fixed-...
Programming Domain-Specific FPGA Hardblocks from HLS: An RTL Blackbox ApproachRuthwik Reddy Sunketa, Jeevesh Choudhury, Aman Arora2026-06-07下载Domain-specific Field Programmable Gate Array (FPGA) architectures increasingly integrate specialized hardblocks, such as Tensor Slices, to accelerate artificial intelligence and machine learning work...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
A Low-Latency Semantic State Estimator using Latent Predictive Learning for Dynamic Network Monitoring and OrchestrationHari Madhukumar, Haiyuan Li, Xiaolan Liu, Andy Corston-Petrie, Dimitra Simeonidou2026-06-07下载Closed-loop network monitoring and orchestration increasingly require semantic interpretations of live telemetry beyond raw counter collection.
Parallel SMT Solving via Dynamic Partitioning, Core-Guided Pruning, and Online Backbone DetectionIlana Shapiro, Sorin Lerner, Nikolaj Bjørner2026-06-07下载Exploiting parallelism in modern CPU architectures remains a longstanding challenge in optimizing SMT solvers. We introduce a novel parallel framework that dynamically builds a binary partition tree o...
Aperon Technical Report: Hierarchical No-Pointer Tangent-Local Search for High-Dimensional Approximate Nearest NeighborsYong Fu2026-06-07下载We present HNTL (Hierarchical No-pointer Tangent-Local), the core vector indexing and candidate generation framework of the Aperon vector memory system. Proximity graphs (e.g.
APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute RebalancingHong Guo, Nianhui Guo, Weixing Wang, Jona Otholt, Christoph Meinel, Haojin Yang2026-06-07下载W4A4 quantization promises full utilization of INT4 Tensor Cores, yet group dequantization overhead on CUDA Cores has driven existing systems to mixed-precision fallbacks.
SpectrumKV: Per-Token Mixed-Precision KV Cache Transfer for Prefill-Decode Disaggregated LLM ServingYang Pengju2026-06-07下载Prefill-decode (PD) disaggregation decouples prompt processing from token generation, but it also turns the key-value (KV) cache into a network payload.
Auditable Graph-Guided Root Cause Analysis for Kubernetes IncidentsAnastasiia Kuvshinova, Seungmin Jin2026-06-07下载Kubernetes incidents are diagnosed reliably only when a root-cause system's reported gains come from incident evidence rather than scenario-specific shortcuts.
Unifying von-Neumann HPC and Neuromorphic Acceleration via the EBRAINS Research Infrastructure: A Framework for High-Performance WorkflowsKrishna Kant Singh, Charl Linssen, Eric Müller, Eleni Mathioulaki, Wouter Klijn, Lena Oden2026-06-07下载Modern scientific workflows increasingly span diverse computing architectures, yet executing a single computational model across disparate systems often forces researchers to maintain fragmented, site...
FlashCP: Load-Balanced Communication-Efficient Context Parallelism for LLM TrainingZheng Wang, Eric Liu, Linan Jiang, Zhongkai Yu, Zaifeng Pan, Yue Guan, Yuke Wang, Yufei Ding2026-06-07下载Context parallelism (CP) is essential for training large-scale, long-context language models, as it partitions sequences to reduce memory overhead.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
SCOPE: A Syndrome-Driven Control Plane for QEC-Enabled Quantum NetworksXiaojie Fan, Zian Wang, Ashutosh Tiwari, Himanshu Gupta2026-06-07下载As quantum networks evolve from experimental testbeds to fault-tolerant systems, the primary performance metric shifts from physical link fidelity to end-to-end logical error rate.
Systems-Level Planning and Coordination of Truck-Drone Collaborative Delivery NetworksDidem Cicek, Burak Kantarci2026-06-07下载Urban last-mile parcel delivery increasingly relies on heterogeneous fleets whose performance depends on timely coordination, reliable communication, and scalable control.
Block coordinate descent for joint delay-energy optimization in multi-hop D2D networksKai-Xiang Hu, Jacek Gondzio, Caixia Kou2026-06-07下载In multi-hop device-to-device (D2D) networks, the optimization of network-level metrics is particularly difficult due to the tight coupling between network-layer routing and physical-layer resource al...

cs.PF - Performance ​

标题作者发布日期PDF摘要
An Empirical Comparison of General Context-Free ParsersHuan Vo, Danushka Liyanage, Hong Jin Kang, Sasha Rubin, Rahul Gopinath2026-06-07下载Parsing underpins a vast range of software engineering tasks, from compilers and static analyzers to language servers and fuzz testing tools. Yet most parsers deployed in practice are deterministic (L...

基于 VitePress 构建 · 使用本地搜索查找论文