Skip to content

2026-06-18 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generationSelin Yildirim, Yingbing Huang, Deming Chen2026-06-18下载We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), which jointly optimizes neural network architectures and their h...
Memory-Centric Computing: Security Benefits and Challenges of Processing-in-DRAMIsmail Emir Yuksel, F. Nisa Bostanci, Ataberk Olgun, Onur Mutlu2026-06-18下载Today's computing systems are processor-centric: they require frequent data movement between processing elements (e.g., CPU) and main memory (DRAM), leading to significant inefficiencies in performanc...
ExSpike: A General Full-Event Neuromorphic Architecture for Exploiting Irregular Sparsity with Event CompressionYuehai Chen, Farhad Merchant2026-06-18下载Spiking neural networks (SNNs) promise energy-efficient computing due to their sparse spatio-temporal activity. However, effectively translating such irregular sparsity into practical performance and ...
Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the EdgeChanda Gupta, Sanidhya Bhatia, Shaurya Priyadarshi, Himani Panwar, Rishad Shafik, Sudip Roy2026-06-18下载Tsetlin Machine (TM) is a logic-based machine learning approach that relies on simple bitwise operations and finite-state automata, which makes it attractive for edge AI deployments.
Design and Evaluation of Energy-Efficient Whisper Dot-Product Kernel Offloading on a CGLA ArchitectureTakuto Ando, Yu Eto, Ayumu Takeuchi, Yasuhiko Nakashima2026-06-18下载In this paper, we implement and evaluate Whisper dot-product kernel offloading on IMAX, a programmable Coarse-Grained Linear Arrays (CGLAs) architecture. Whisper-tiny.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Multiword Arithmetic and Parallel ComputingJan Verschelde2026-06-18下载In many applications, the precision by the available hardware arithmetic is insufficient to guarantee accurate results. Multiword arithmetic is a special type of multiprecision arithmetic where a mult...
Process-Reward Tactic Evolution for Long-Horizon Bioinformatics WorkflowsLingzhi Yang, Yubo Fan, Song Wu, Gilchan Park2026-06-18下载LLM agents can write code and call tools, but reliable bioinformatics work requires long-horizon interaction with workflow software, typed data objects, provenance, and biological checks.
Memory-Centric Computing: Security Benefits and Challenges of Processing-in-DRAMIsmail Emir Yuksel, F. Nisa Bostanci, Ataberk Olgun, Onur Mutlu2026-06-18下载Today's computing systems are processor-centric: they require frequent data movement between processing elements (e.g., CPU) and main memory (DRAM), leading to significant inefficiencies in performanc...
Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI ServingLiang Su2026-06-18下载Mainstream LLM serving systems reuse prefix work mainly through paged or radix key-value (KV) caches. This is highly effective for high-throughput, high-concurrency serving, but it manages only one po...
Sovereign Execution Broker: Enforcing Certificate-Bound Authority in Agentic Control PlanesJun He, Deying Yu2026-06-18下载Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes.
Coarse Solvers for Exascale Solution of Poisson ProblemsThilina Ratnayaka, Paul Fischer, Luke Olson2026-06-18下载We present a two-level Schwarz method as an alternative to Algebraic Multigrid method(AMG) used as the last level (coarse) solver of the p-multigrid pMG preconditioner for pressure Poisson equation re...
ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU ClustersJiasheng Zhou, Longbin Zeng, Clavis Chen, Ruiming Lu, Qinwei Yang, Leyi Ye, Ray Ying, Key Zhang2026-06-18下载Large-scale LLM training requires always-on, fine-grained observability for effective performance diagnosis at scale. Coarse resource monitors alone cannot localize root causes, and fine-grained profi...
Quantum ring all-reduce: communication and privacy advantages for distributed learningMaría Gragera Garcés, Lirandë Pira2026-06-18下载Machine learning models have scaled to unprecedented sizes, making training across distributed devices the de facto standard in the field. In this work, we explore how quantum communications can make ...
The Correctness Illusion in LLM-Generated GPU KernelsDipankar Sarkar2026-06-18下载Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks. The number of inputs varies between benchmarks.
Online Dynamic Batching with Formal Guarantees for LLM TrainingDian Li, Zekun Wang, Yaoru Wang, Jiahong Yan2026-06-18下载Modern LLM training breaks a core assumption behind offline batch samplers: the true training cost of a sample is only observable after preprocessing, augmentation, templating, tokenization, and multi...
The Bi-Channel Networking Paradigm for Database Systems in the CloudGeorg Kreuzmayr, Muhammad El-Hindi, Benjamin Wagner, Tobias Ziegler, Viktor Leis2026-06-18下载When network links were slow, cloud and distributed database systems could rely on generic kernel abstractions and treat network communication as a black box.
EVM Workloads in the Wild: Evidence for Multi-Dimensional Gas Metering, State Growth, Delayed Execution, and ParallelismLioba Heimbach, Kushal Babel, Jason Milionis2026-06-18下载Gas metering on EVM-compatible blockchains assumes that execution conditions are stable: that the resource mix is constant enough to justify collapsing execution costs into a single scalar with fixed ...
Multi-Orientation Edge-Minimum Repair for Non-Redundant Fault-Tolerant Broadcasting in Dense Eisenstein--Jacobi NetworksBader Albader2026-06-18下载Dense Eisenstein--Jacobi (EJ) networks are degree-six algebraic interconnection networks whose finite quotient geometry is naturally represented by a hexagonal axial-coordinate ball.
Fault-Tolerant Shared-Relay Communication in Circulant Interconnection NetworksBader Albader, Galal Hassan, Mohamed R. Al-Mulla2026-06-18下载Circulant interconnection networks provide symmetric addressing, compact generator descriptions, and uniform local connectivity. This paper maps a degree--redundancy landscape for a fault-tolerant two...
Certified Euclidean-Residue Minimal-Alignment Switch Decompositions for Three Edge-Disjoint Hamiltonian Cycles in Eisenstein--Jacobi NetworksBader Albader2026-06-18下载Eisenstein--Jacobi (EJ) networks are degree-six quotient-lattice interconnection networks. For a generator α=a+bρ, let N=a2+ab+b2N=a^2+ab+b^2 and d=gcd(a,b)d=\gcd(a,b).
SAC: Disaggregated KV Cache System for Sparse Attention LLMs with CXLRuiyang Ma, Teng Ma, Junru Li, Hantian Zha, Xuchun Shang, Qingda Hu, Zheng Liu, Xinjun Yang, Tao Ma, Guojie Luo2026-06-18下载The scaling of LLMs toward long-context inference has shifted the primary serving system bottleneck from computation to memory capacity. Traditional solutions for dense attention models rely on RDMA-b...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Can Quantum Receiver Beat the SIC Limit in Multiple Access Networks?Shiqian Guo, Huaiyu Dai, Jianqing Liu2026-06-18下载Successive interference cancellation (SIC) is an important technique for 5G/B5G wireless receivers to resolve interfering signals from multiple users.
Frequency Lock EncodingJoshua Montierth, Michael Rice, Philip Lundrigan2026-06-18下载Modern wireless systems are designed with excess synchronization bandwidth to ensure reliable operation under worst-case conditions. This paper presents a protocol-agnostic secondary signaling layer t...
Implicit Semantic-Aware Communication Based on Hypergraph ReasoningYiwei Liao, Shurui Tu, Yong Xiao, Yingyu Li, Guangming Shi2026-06-18下载Semantic-aware communication has emerged as a transformative paradigm for next-generation communication systems, shifting the fundamental goal from transmitting bit-level symbols to reliably recoverin...
Multi-Orientation Edge-Minimum Repair for Non-Redundant Fault-Tolerant Broadcasting in Dense Eisenstein--Jacobi NetworksBader Albader2026-06-18下载Dense Eisenstein--Jacobi (EJ) networks are degree-six algebraic interconnection networks whose finite quotient geometry is naturally represented by a hexagonal axial-coordinate ball.
Fault-Tolerant Shared-Relay Communication in Circulant Interconnection NetworksBader Albader, Galal Hassan, Mohamed R. Al-Mulla2026-06-18下载Circulant interconnection networks provide symmetric addressing, compact generator descriptions, and uniform local connectivity. This paper maps a degree--redundancy landscape for a fault-tolerant two...
Certified Euclidean-Residue Minimal-Alignment Switch Decompositions for Three Edge-Disjoint Hamiltonian Cycles in Eisenstein--Jacobi NetworksBader Albader2026-06-18下载Eisenstein--Jacobi (EJ) networks are degree-six quotient-lattice interconnection networks. For a generator α=a+bρ, let N=a2+ab+b2N=a^2+ab+b^2 and d=gcd(a,b)d=\gcd(a,b).

cs.PF - Performance ​

标题作者发布日期PDF摘要
UltraQuant: 4-bit KV Caching for Context-Heavy AgentsInesh Chakrabarti, David Limpus, Aditi Ghai Rana, Bowen Bao, Spandan Tiwari, Thiago Crepaldi, Ashish Sirasao2026-06-18下载Context-heavy agents place unusual pressure on the key-value (KV) cache: long prefixes are reused across many short turns, while concurrency determines whether the serving system can keep GPUs utilize...
Randomized Sketching is Robust to Low-Precision Rounding on GPUsAryaman Jeendgar, Clément Flint, Hartwig Anzt2026-06-18下载Randomized sketching is a core primitive in randomized numerical linear algebra. On modern hardware architectures, in particular on GPUs, the performance of sparse sketches is limited by memory traffi...

基于 VitePress 构建 · 使用本地搜索查找论文