2026-06-18
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation | Selin Yildirim, Yingbing Huang, Deming Chen | 2026-06-18 | 下载 | We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), which jointly optimizes neural network architectures and their h... |
| Memory-Centric Computing: Security Benefits and Challenges of Processing-in-DRAM | Ismail Emir Yuksel, F. Nisa Bostanci, Ataberk Olgun, Onur Mutlu | 2026-06-18 | 下载 | Today's computing systems are processor-centric: they require frequent data movement between processing elements (e.g., CPU) and main memory (DRAM), leading to significant inefficiencies in performanc... |
| ExSpike: A General Full-Event Neuromorphic Architecture for Exploiting Irregular Sparsity with Event Compression | Yuehai Chen, Farhad Merchant | 2026-06-18 | 下载 | Spiking neural networks (SNNs) promise energy-efficient computing due to their sparse spatio-temporal activity. However, effectively translating such irregular sparsity into practical performance and ... |
| Low-Energy Reduced RISC-V Instruction Subset Processor for Tsetlin Machine Inference at the Edge | Chanda Gupta, Sanidhya Bhatia, Shaurya Priyadarshi, Himani Panwar, Rishad Shafik, Sudip Roy | 2026-06-18 | 下载 | Tsetlin Machine (TM) is a logic-based machine learning approach that relies on simple bitwise operations and finite-state automata, which makes it attractive for edge AI deployments. |
| Design and Evaluation of Energy-Efficient Whisper Dot-Product Kernel Offloading on a CGLA Architecture | Takuto Ando, Yu Eto, Ayumu Takeuchi, Yasuhiko Nakashima | 2026-06-18 | 下载 | In this paper, we implement and evaluate Whisper dot-product kernel offloading on IMAX, a programmable Coarse-Grained Linear Arrays (CGLAs) architecture. Whisper-tiny. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Multiword Arithmetic and Parallel Computing | Jan Verschelde | 2026-06-18 | 下载 | In many applications, the precision by the available hardware arithmetic is insufficient to guarantee accurate results. Multiword arithmetic is a special type of multiprecision arithmetic where a mult... |
| Process-Reward Tactic Evolution for Long-Horizon Bioinformatics Workflows | Lingzhi Yang, Yubo Fan, Song Wu, Gilchan Park | 2026-06-18 | 下载 | LLM agents can write code and call tools, but reliable bioinformatics work requires long-horizon interaction with workflow software, typed data objects, provenance, and biological checks. |
| Memory-Centric Computing: Security Benefits and Challenges of Processing-in-DRAM | Ismail Emir Yuksel, F. Nisa Bostanci, Ataberk Olgun, Onur Mutlu | 2026-06-18 | 下载 | Today's computing systems are processor-centric: they require frequent data movement between processing elements (e.g., CPU) and main memory (DRAM), leading to significant inefficiencies in performanc... |
| Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving | Liang Su | 2026-06-18 | 下载 | Mainstream LLM serving systems reuse prefix work mainly through paged or radix key-value (KV) caches. This is highly effective for high-throughput, high-concurrency serving, but it manages only one po... |
| Sovereign Execution Broker: Enforcing Certificate-Bound Authority in Agentic Control Planes | Jun He, Deying Yu | 2026-06-18 | 下载 | Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes. |
| Coarse Solvers for Exascale Solution of Poisson Problems | Thilina Ratnayaka, Paul Fischer, Luke Olson | 2026-06-18 | 下载 | We present a two-level Schwarz method as an alternative to Algebraic Multigrid method(AMG) used as the last level (coarse) solver of the p-multigrid pMG preconditioner for pressure Poisson equation re... |
| ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU Clusters | Jiasheng Zhou, Longbin Zeng, Clavis Chen, Ruiming Lu, Qinwei Yang, Leyi Ye, Ray Ying, Key Zhang | 2026-06-18 | 下载 | Large-scale LLM training requires always-on, fine-grained observability for effective performance diagnosis at scale. Coarse resource monitors alone cannot localize root causes, and fine-grained profi... |
| Quantum ring all-reduce: communication and privacy advantages for distributed learning | María Gragera Garcés, Lirandë Pira | 2026-06-18 | 下载 | Machine learning models have scaled to unprecedented sizes, making training across distributed devices the de facto standard in the field. In this work, we explore how quantum communications can make ... |
| The Correctness Illusion in LLM-Generated GPU Kernels | Dipankar Sarkar | 2026-06-18 | 下载 | Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks. The number of inputs varies between benchmarks. |
| Online Dynamic Batching with Formal Guarantees for LLM Training | Dian Li, Zekun Wang, Yaoru Wang, Jiahong Yan | 2026-06-18 | 下载 | Modern LLM training breaks a core assumption behind offline batch samplers: the true training cost of a sample is only observable after preprocessing, augmentation, templating, tokenization, and multi... |
| The Bi-Channel Networking Paradigm for Database Systems in the Cloud | Georg Kreuzmayr, Muhammad El-Hindi, Benjamin Wagner, Tobias Ziegler, Viktor Leis | 2026-06-18 | 下载 | When network links were slow, cloud and distributed database systems could rely on generic kernel abstractions and treat network communication as a black box. |
| EVM Workloads in the Wild: Evidence for Multi-Dimensional Gas Metering, State Growth, Delayed Execution, and Parallelism | Lioba Heimbach, Kushal Babel, Jason Milionis | 2026-06-18 | 下载 | Gas metering on EVM-compatible blockchains assumes that execution conditions are stable: that the resource mix is constant enough to justify collapsing execution costs into a single scalar with fixed ... |
| Multi-Orientation Edge-Minimum Repair for Non-Redundant Fault-Tolerant Broadcasting in Dense Eisenstein--Jacobi Networks | Bader Albader | 2026-06-18 | 下载 | Dense Eisenstein--Jacobi (EJ) networks are degree-six algebraic interconnection networks whose finite quotient geometry is naturally represented by a hexagonal axial-coordinate ball. |
| Fault-Tolerant Shared-Relay Communication in Circulant Interconnection Networks | Bader Albader, Galal Hassan, Mohamed R. Al-Mulla | 2026-06-18 | 下载 | Circulant interconnection networks provide symmetric addressing, compact generator descriptions, and uniform local connectivity. This paper maps a degree--redundancy landscape for a fault-tolerant two... |
| Certified Euclidean-Residue Minimal-Alignment Switch Decompositions for Three Edge-Disjoint Hamiltonian Cycles in Eisenstein--Jacobi Networks | Bader Albader | 2026-06-18 | 下载 | Eisenstein--Jacobi (EJ) networks are degree-six quotient-lattice interconnection networks. For a generator α=a+bρ, let and . |
| SAC: Disaggregated KV Cache System for Sparse Attention LLMs with CXL | Ruiyang Ma, Teng Ma, Junru Li, Hantian Zha, Xuchun Shang, Qingda Hu, Zheng Liu, Xinjun Yang, Tao Ma, Guojie Luo | 2026-06-18 | 下载 | The scaling of LLMs toward long-context inference has shifted the primary serving system bottleneck from computation to memory capacity. Traditional solutions for dense attention models rely on RDMA-b... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Can Quantum Receiver Beat the SIC Limit in Multiple Access Networks? | Shiqian Guo, Huaiyu Dai, Jianqing Liu | 2026-06-18 | 下载 | Successive interference cancellation (SIC) is an important technique for 5G/B5G wireless receivers to resolve interfering signals from multiple users. |
| Frequency Lock Encoding | Joshua Montierth, Michael Rice, Philip Lundrigan | 2026-06-18 | 下载 | Modern wireless systems are designed with excess synchronization bandwidth to ensure reliable operation under worst-case conditions. This paper presents a protocol-agnostic secondary signaling layer t... |
| Implicit Semantic-Aware Communication Based on Hypergraph Reasoning | Yiwei Liao, Shurui Tu, Yong Xiao, Yingyu Li, Guangming Shi | 2026-06-18 | 下载 | Semantic-aware communication has emerged as a transformative paradigm for next-generation communication systems, shifting the fundamental goal from transmitting bit-level symbols to reliably recoverin... |
| Multi-Orientation Edge-Minimum Repair for Non-Redundant Fault-Tolerant Broadcasting in Dense Eisenstein--Jacobi Networks | Bader Albader | 2026-06-18 | 下载 | Dense Eisenstein--Jacobi (EJ) networks are degree-six algebraic interconnection networks whose finite quotient geometry is naturally represented by a hexagonal axial-coordinate ball. |
| Fault-Tolerant Shared-Relay Communication in Circulant Interconnection Networks | Bader Albader, Galal Hassan, Mohamed R. Al-Mulla | 2026-06-18 | 下载 | Circulant interconnection networks provide symmetric addressing, compact generator descriptions, and uniform local connectivity. This paper maps a degree--redundancy landscape for a fault-tolerant two... |
| Certified Euclidean-Residue Minimal-Alignment Switch Decompositions for Three Edge-Disjoint Hamiltonian Cycles in Eisenstein--Jacobi Networks | Bader Albader | 2026-06-18 | 下载 | Eisenstein--Jacobi (EJ) networks are degree-six quotient-lattice interconnection networks. For a generator α=a+bρ, let and . |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| UltraQuant: 4-bit KV Caching for Context-Heavy Agents | Inesh Chakrabarti, David Limpus, Aditi Ghai Rana, Bowen Bao, Spandan Tiwari, Thiago Crepaldi, Ashish Sirasao | 2026-06-18 | 下载 | Context-heavy agents place unusual pressure on the key-value (KV) cache: long prefixes are reused across many short turns, while concurrency determines whether the serving system can keep GPUs utilize... |
| Randomized Sketching is Robust to Low-Precision Rounding on GPUs | Aryaman Jeendgar, Clément Flint, Hartwig Anzt | 2026-06-18 | 下载 | Randomized sketching is a core primitive in randomized numerical linear algebra. On modern hardware architectures, in particular on GPUs, the performance of sparse sketches is limited by memory traffi... |