2026-06-15
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Energy-efficient codon optimization on thermodynamic hardware | Andraz Jelincic, Ross C. Walker | 2026-06-15 | 下载 | The growing energy demand for computation is becoming increasingly unsustainable. Thermodynamic computing, which harnesses physical thermal fluctuations as a computational resource rather than suppres... |
| PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM Agents for VLSI Physical Design | Qiufeng Li, Rongqian Chen, Quan Cheng, Chengxuan Wang, Sizhe Tang, Wuxi Li, Duo Ding, Chia-Tung Ho, Haoxing Ren, David Z. Pan, Tian Lan, Weidong Cao | 2026-06-15 | 下载 | Large Language Models and vision-language models have shown remarkable success in the front-end design of Very Large-Scale Integrated Circuits, yet their capabilities for VLSI physical design remain s... |
| From Compression to Deployment: Real-Time and Energy-Efficient FastGRNN on Ultra-Constrained Microcontrollers | Emre Can Kizilates | 2026-06-15 | 下载 | The dominant trajectory of modern machine learning has been to scale up: larger models, larger accelerators, larger memory budgets. Yet a multi-year global semiconductor supply constraint and the grow... |
| HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting | Alper Yıldırım | 2026-06-15 | 下载 | Simple linear and frequency-domain models remain surprisingly competitive in long-horizon time-series forecasting, and recent mechanistic evidence suggests that standard forecasting benchmarks may not... |
| Shift-Left High-Level Synthesis Verification via Knowledge-Augmented LLM Agent | Zhihan Xiao, Zhe Zhao, Luke Ztz Hu, Songping Mai | 2026-06-15 | 下载 | High-Level Synthesis (HLS) enables rapid hardware development by translating C/C++ programs into hardware implementations. Functional consistency verification between golden C specifications and HLS-o... |
| Neural dynamical systems on ferroelectric compute-in-memory for real-time forecasting | Keshava Katti, Adithya Selvakumar, Pratik Chaudhari, Deep Jariwala | 2026-06-15 | 下载 | Neural dynamical systems are expressive temporal predictors that capture continuous-time dynamics through fine-grained state updates. However, this sequential structure maps poorly onto digital hardwa... |
| Architecture Carbon Tool v3: Enabling Sustainability-aware Silicon System Design Exploration | Vincent T. Lee, Bilge Acun, Zachary Lewis, Carole-Jean Wu | 2026-06-15 | 下载 | As the carbon cost of manufacturing and operating semiconductor devices has come into sharper focus, sustainability has gradually emerged as a new system architecture design metric. |
| From the NYU Ultracomputer to Modern Exascale: A Historical and Architectural Survey of In-Network Computing and Scalable Synchronization | Lars Warren Ericson | 2026-06-15 | 下载 | This paper presents a historical and technical survey of the hardware architectures, interconnection networks, and synchronization primitives that have shaped massively parallel systems over the past ... |
| DataGuard: Guaranteeing Private Training in Systolic-array Based Accelerators | Pawan Kumar Sanjaya, Christina Giannoula, Nikhil Shreekumar, Ian Colbert, Alec Dewulf, Mehdi Saeedi, Ihab Amer, Gabor Sines, Nandita Vijaykumar | 2026-06-15 | 下载 | Differential privacy (DP) and federated learning (FL) have emerged as important privacy-preserving approaches when using sensitive data to train machine learning (ML) models. |
| TreeGRNG: Binary Tree Gaussian Random Number Generator for Efficient Probabilistic AI Hardware | Jonas Crols, Guilherme Paim, Shirui Zhao, Marian Verhelst | 2026-06-15 | 下载 | Bayesian Neural Networks (BNNs) offer opportunities for greatly enhancing the trustworthiness of conventional neural networks by monitoring the uncertainties in decision-making. |
| Towards Delta Aware Training: Efficient DNN Weight Storage for Resource-Constrained FPGAs | David Peter Federl, Lukas Einhaus, Andreas Erbslöh, Gregor Schiele | 2026-06-15 | 下载 | The deployment of embedded deep neural networks on resource-constrained field programmable gate arrays (FPGAs) is challenging due to limited memory and computational capacities. |
| NeuronFabric: A Software Reference Architecture for On-Chip Transformer Training with Local Adam | Evgeny Ukladchikov | 2026-06-15 | 下载 | Publicly documented accelerator architectures generally separate training computation from optimizer-state updates or rely on external memory and host orchestration. |
| MPX: A Unified Systolic Array for Matrix and Polynomial Multiplication | George Alexakis, Dimitrios Schoinianakis, Giorgos Dimitrakopoulos | 2026-06-15 | 下载 | Polynomial multiplication is a fundamental kernel in Fully Homomorphic Encryption (FHE) and post-quantum cryptography (PQC) and is commonly accelerated through Number Theoretic Transforms (NTTs). |
| Embedded Arena: Iterative Optimization via Hardware Feedback | Zhihan Zhang, Alexander Le Metzger, Jiuyang Lyu, Chun-Cheng Chang, Jiayi Shao, Yujia Liu, Emmanuel Azuh Mensah, Edward Wang, Kurtis Heimerl, Gregory D. Abowd, Shwetak Patel, Natasha Jaques, Vikram Iyer | 2026-06-15 | 下载 | Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints. |
| AIA: A 16nm Multicore SoC for Approximate Inference Acceleration Exploiting Non-normalized Knuth-Yao Sampling and Inter-Core Register Sharing | Shirui Zhao, Nimish Shah, Wannes Meert, Marian Verhelst | 2026-06-15 | 下载 | Probabilistic graphical models (PMs) are popular to empower machine learning with the ability of reasoning and decision-making. To perform approximate inference in PMs, sampling-based Markov Chain Mon... |
| When Proofs Meet Hardware: Comparing NTT and SumCheck in Zero-Knowledge Systems | Jianqiao Mo, Alhad Daftardar, Barath GaneshKumar, Kaiyue Guo, Hong Wang, Benedikt Bunz, Siddharth Garg, Brandon Reagen | 2026-06-15 | 下载 | In the ZKP community, it has long been discussed that the SumCheck protocol is asymptotically more efficient than the Number Theoretic Transform (NTT), requiring only arithmetic versus $O(N \lo... |
| AIA: A Customized Multi-core RISC-V SoC for Discrete Sampling Workloads in 16 nm | Shirui Zhao, Nimish Shah, Wannes Meert, Marian Verhelst | 2026-06-15 | 下载 | Probabilistic models (PMs) are essential in advancing machine learning capabilities, particularly in safety-critical applications involving reasoning and decision-making. |
| Beyond CPU-GPU Frequency: Memory-Clock and Tail Effects in Edge Inference Latency Estimation | Jaehoon Kang | 2026-06-15 | 下载 | Frequency-aware latency estimators enable deadline-aware DVFS for edge ML inference by modeling latency over CPU and GPU frequencies. We present a measurement study on an NVIDIA Jetson Orin Nano showi... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Space-Efficient Lock-Free Linear-Probing Hash Table | Hagit Attiya, Rotem Oshman, Noa Schiller | 2026-06-15 | 下载 | Linear probing is one of the simplest and most space-efficient approaches to hash table design, and is widely used in sequential settings due to its compact memory layout. |
| Local Fault Repair of Perfect Resource Placements in Eisenstein--Jacobi Networks | Bader Albader | 2026-06-15 | 下载 | Perfect resource placements in dense Eisenstein--Jacobi (EJ) networks partition the network into hexagonal radius- service cells. This paper studies local repair of such placements after resource f... |
| Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems | Sajjad Khan | 2026-06-15 | 下载 | Multi-agent LLM systems share state through memory stores, vector indices, and tool registries. We model such sharing as long-running read-generate-write operations under deterministic-generation sema... |
| Diagonal-Budgeted Trotterization for Efficient Quantum Hamiltonian Simulation | Srikar Chundury, Blake Burgstahler, Jiajia Li, In-Saeng Suh, Frank Mueller | 2026-06-15 | 下载 | Efficient classical simulation of quantum Hamiltonian dynamics is often bottlenecked by exponential state growth and the overhead of generic sparse linear algebra. |
| Re-Rooting-Based Fault-Tolerant Broadcasting in Dense Gaussian Networks | Bader Albader, Mohamed R. Al-Mulla, Galal Hassan | 2026-06-15 | 下载 | Dense Gaussian networks provide degree-4 interconnection topologies with small diameter and regular structure, making them suitable for efficient one-to-all broadcasting. |
| Single-Connection Mixed-Criticality Transport with CATS: Bounded Guarantees, Three Structural Limits, and a QUIC Escape | Syed Muhammad Aqdas Rizvi | 2026-06-15 | 下载 | Mixed-criticality applications, such as satellite terminals, industrial telemetry, embedded systems, tactical, and other constrained links, often multiplex a small, latency-critical message class and ... |
| Tangram: Hiding GPU Heterogeneity for Efficient LLM Parallelization | Yanda Tao, Pedro F. Silvestre, Marcel Wagenländer, Peter Pietzuch | 2026-06-15 | 下载 | The scale of LLM training jobs requires parallelization planning over large GPU clusters. Due to different GPU types and interconnects added over time, these GPU clusters are increasingly heterogeneou... |
| A Unified Constant-Time Switch Rule for Constructing Edge-Disjoint Hamiltonian Cycles in Gaussian Networks | Bader Albader | 2026-06-15 | 下载 | Gaussian networks are degree-four symmetric interconnection networks defined over residue classes of Gaussian integers. Earlier work showed that when the generator α=a+bi satisfies , th... |
| Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection | Markus Bujotzek, Dimitrios Bounias, Stefan Denner, Ralf Floca, Maximilian Fischer, Peter Neher, Klaus Maier-Hein | 2026-06-15 | 下载 | While federated learning (FL) enables collaborative medical image segmentation without centralizing sensitive data, real-world deployment is frequently complicated by cross-site label imperfections su... |
| CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agents | Shubham Tiwari, Tapan Chugh, Nash Rickert, Simon Peter, Ratul Mahajan, Haiying Shen | 2026-06-15 | 下载 | Coding agents are a fast-growing LLM application, executing as long-running closed-loop sessions in which LLM generations alternate with external tool calls. |
| From the NYU Ultracomputer to Modern Exascale: A Historical and Architectural Survey of In-Network Computing and Scalable Synchronization | Lars Warren Ericson | 2026-06-15 | 下载 | This paper presents a historical and technical survey of the hardware architectures, interconnection networks, and synchronization primitives that have shaped massively parallel systems over the past ... |
| Robust and Automated Reconfiguration of Byzantine Wide-Area Replication | Rowdy Chotkan, Bulat Nasrulin, Johan Pouwelse, Jérémie Decouchant | 2026-06-15 | 下载 | Distributed systems handle adversarial nodes through redundancy, which imposes a significant performance overhead. In blockchain systems, Byzantine fault-tolerant state-machine replication (BFT-SMR) i... |
| DRIFT: Risk-Constrained Diffusion with Imitation Priors for Mixed-Autonomy Traffic Generation | Yaoshen Yu, Minghui Liwang, Wenbo Zhu, Xinlei Yi, Yiguang Hong, Yuhan Su, Seyyedali Hosseinalipour | 2026-06-15 | 下载 | Future intelligent transportation systems are envisioned to evolve toward a long-term mixed-autonomy paradigm, where human-driven vehicles (HVs) and autonomous vehicles (AVs) coexist within highly cou... |
| Incentives and Evidence in Learned Service Orchestration | Syed Izhan Khilji, Alireza Furutanpey, Schahram Dustdar | 2026-06-15 | 下载 | Reinforcement learning for service orchestration has been the subject of sustained research for over a decade, yet it is not used in production at scale. |
| Generated, Parallel, Scalable? A Study of Agentic AI-Generated Julia Code on Supercomputers | Linus Bantel, Anna-Lena Roth, Jonas Posner, Dirk Pflüger | 2026-06-15 | 下载 | Julia is increasingly used in HPC as a single-language alternative to combining high-level scripting with low-level systems languages, but achieving scalable performance still requires expertise in pa... |
| SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions | Feiyang Chen, Haibo Chen | 2026-06-15 | 下载 | Modern CPUs increasingly integrate matrix extensions, such as Arm Scalable Matrix Extension (SME), that provide high-throughput matrix execution within the CPU. |
| Fractional Verkle Trees: A Hypertree Decomposition and Verified Proof Serialization Architecture for High-Performance Blockchain State Accumulators | Ekleen Kaur, Everton Fraga | 2026-06-15 | 下载 | Modern blockchain state management faces a critical scalability bottleneck: maintaining cryptographic commitments over hundreds of millions of entries becomes computationally prohibitive. |
| Tropical: Enhancing SLO Attainment in Disaggregated LLM Serving via SLO-Aware Multiplexing | Jinming Ma, Jiefei Chen, Xiuhong Li, Jiangfei Duan, Haojie Duanmu, Xingcheng Zhang, Chao Yang, Dahua Lin | 2026-06-15 | 下载 | To guarantee service quality in transformer based large language model (LLM) serving, it is essential to meet the latency constraints of both the prefill phase (measured by Time-to-First-Token, TTFT) ... |
| StorRep: Storage Research Experiment Patterns on Chameleon Cloud and Trovi | Ray A. O. Sinurat, Yuyang Huang, Nanqinqin Li, Mark Powers, Michael Sherman, Kate Keahey, Haryadi S. Gunawi | 2026-06-15 | 下载 | Storage experiments are vital to advancing storage research, but creating extensible and reproducible storage artifacts can be a challenging task. |
| did:crdt: Coordination-Free Decentralised Identifiers via Signed CRDTs | Hugo O'Connor, Claire Barnes | 2026-06-15 | 下载 | Existing Decentralised Identifier (DID) methods require coordination, an agreed global order of operations, to update a DID document: blockchain-anchored methods incur fees and latency; lightweight pe... |
| Efficient Data Availability Sampling via Coded Distributed Arrays | Dang Pham Minh, Hung Vuong Huu, Duc A. Tran | 2026-06-15 | 下载 | Data availability is a fundamental bottleneck in modern blockchain networks. Most blockchain systems rely on a full-replication model, which requires downloading of a full block to verify its availabi... |
| SwiftCache: Efficient LLM Serving for Multi-turn Conversations with Heterogeneous KV Cache Sharing | Jianmin Hu, Minxian Xu, Sa Wang, Chong Ma, Min Shen, Kejiang Ye, Lin Qu, Chengzhong Xu | 2026-06-15 | 下载 | Multi-turn conversation is a fundamental scenario in LLM applications, widely used in chatbots and AI agents. As the conversation evolves, historical tokens accumulate continuously. |
| Beyond CPU-GPU Frequency: Memory-Clock and Tail Effects in Edge Inference Latency Estimation | Jaehoon Kang | 2026-06-15 | 下载 | Frequency-aware latency estimators enable deadline-aware DVFS for edge ML inference by modeling latency over CPU and GPU frequencies. We present a measurement study on an NVIDIA Jetson Orin Nano showi... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes | Shengli Zhang, Deen Ma, Zibin Lin, Taotao Wang | 2026-06-15 | 下载 | Large language models have accelerated the transition from passive conversational assistants to autonomous agents that can understand goals, plan actions, invoke tools, and execute multi-step tasks. |
| Local Fault Repair of Perfect Resource Placements in Eisenstein--Jacobi Networks | Bader Albader | 2026-06-15 | 下载 | Perfect resource placements in dense Eisenstein--Jacobi (EJ) networks partition the network into hexagonal radius- service cells. This paper studies local repair of such placements after resource f... |
| Cache to the Future: A Distributed Webpage Archive for Internet Blackouts | Ross Evans, Diogo Barradas | 2026-06-15 | 下载 | Internet blackouts, occurring due to technological mishaps or intentional governmental action, prevent citizens from accessing the internet. Citizens in regions where internet blackouts are common hav... |
| Renderable Partial Representations for Dynamic Gaussian Splatting under Incomplete Delivery | Faruk Alpay, Levent Sarioglu, Yaser Hadri | 2026-06-15 | 下载 | Dynamic Gaussian compression is normally optimized for complete files or complete progressive prefixes, but interactive rendering encounters partial representations: some spatiotemporal regions are pr... |
| Re-Rooting-Based Fault-Tolerant Broadcasting in Dense Gaussian Networks | Bader Albader, Mohamed R. Al-Mulla, Galal Hassan | 2026-06-15 | 下载 | Dense Gaussian networks provide degree-4 interconnection topologies with small diameter and regular structure, making them suitable for efficient one-to-all broadcasting. |
| Di5Guise: 5G Privacy with vSIM | Shirin Ebadi, Zach Moolman, Eric Keller, Tamara Lehman | 2026-06-15 | 下载 | SIM cards have been the key building block of user authenticationand security in cellular networks. While they are meant to serve as privacy protecting elements in cellular communications, they can be... |
| Single-Connection Mixed-Criticality Transport with CATS: Bounded Guarantees, Three Structural Limits, and a QUIC Escape | Syed Muhammad Aqdas Rizvi | 2026-06-15 | 下载 | Mixed-criticality applications, such as satellite terminals, industrial telemetry, embedded systems, tactical, and other constrained links, often multiplex a small, latency-critical message class and ... |
| A Unified Constant-Time Switch Rule for Constructing Edge-Disjoint Hamiltonian Cycles in Gaussian Networks | Bader Albader | 2026-06-15 | 下载 | Gaussian networks are degree-four symmetric interconnection networks defined over residue classes of Gaussian integers. Earlier work showed that when the generator α=a+bi satisfies , th... |
| Robust and Automated Reconfiguration of Byzantine Wide-Area Replication | Rowdy Chotkan, Bulat Nasrulin, Johan Pouwelse, Jérémie Decouchant | 2026-06-15 | 下载 | Distributed systems handle adversarial nodes through redundancy, which imposes a significant performance overhead. In blockchain systems, Byzantine fault-tolerant state-machine replication (BFT-SMR) i... |
| Predictive Dynamic Scheduling for Deterministic Communications in Beyond 5G | Syed Morsleen Riaz, M. Carmen Lucas-Estañ, Baldomero Coll-Perales, Javier Gozalvez | 2026-06-15 | 下载 | Next generation wireless networks must sustain deterministic service levels to support emerging time-sensitive applications. The ability to guarantee bounded latencies depends on the efficient managem... |
| Impact of ADAS and V2X Penetration Rates on Cooperative Active Safety | M. C. Lucas-Estañ, B. Coll-Perales, M. I. Khan, J. Gozalvez, O. Altintas | 2026-06-15 | 下载 | A major driver of connected and automated driving is cooperative active safety. The effectiveness of cooperative safety applications depends on the ability of vehicles to detect traffic safety risks i... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Single-Connection Mixed-Criticality Transport with CATS: Bounded Guarantees, Three Structural Limits, and a QUIC Escape | Syed Muhammad Aqdas Rizvi | 2026-06-15 | 下载 | Mixed-criticality applications, such as satellite terminals, industrial telemetry, embedded systems, tactical, and other constrained links, often multiplex a small, latency-critical message class and ... |
| CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agents | Shubham Tiwari, Tapan Chugh, Nash Rickert, Simon Peter, Ratul Mahajan, Haiying Shen | 2026-06-15 | 下载 | Coding agents are a fast-growing LLM application, executing as long-running closed-loop sessions in which LLM generations alternate with external tool calls. |
| StorRep: Storage Research Experiment Patterns on Chameleon Cloud and Trovi | Ray A. O. Sinurat, Yuyang Huang, Nanqinqin Li, Mark Powers, Michael Sherman, Kate Keahey, Haryadi S. Gunawi | 2026-06-15 | 下载 | Storage experiments are vital to advancing storage research, but creating extensible and reproducible storage artifacts can be a challenging task. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| The Right Call for Software Benchmarking: Consistent Decisions in Stateful Environments | Gábor Melis | 2026-06-15 | 下载 | In the perpetual pursuit of performance, modern computing systems rely ever more on stateful mechanisms to accommodate the dynamics of workloads and physical environments, bolstering efficiency but co... |
| OpenGadget3 GPU solver tests | A. Ragagnin, G. S. Karademir, F. Groth, K. Dolag, L. M. Böss, T. Castro, N. Hariharan, M. Aiello, L. Tornatore | 2026-06-15 | 下载 | We present an in-depth evaluation of the scalability and accuracy of the GPU porting of the N-body code for hydrodynamic cosmological simulations \og. |
| Single-Connection Mixed-Criticality Transport with CATS: Bounded Guarantees, Three Structural Limits, and a QUIC Escape | Syed Muhammad Aqdas Rizvi | 2026-06-15 | 下载 | Mixed-criticality applications, such as satellite terminals, industrial telemetry, embedded systems, tactical, and other constrained links, often multiplex a small, latency-critical message class and ... |
| SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions | Feiyang Chen, Haibo Chen | 2026-06-15 | 下载 | Modern CPUs increasingly integrate matrix extensions, such as Arm Scalable Matrix Extension (SME), that provide high-throughput matrix execution within the CPU. |
| Fractional Verkle Trees: A Hypertree Decomposition and Verified Proof Serialization Architecture for High-Performance Blockchain State Accumulators | Ekleen Kaur, Everton Fraga | 2026-06-15 | 下载 | Modern blockchain state management faces a critical scalability bottleneck: maintaining cryptographic commitments over hundreds of millions of entries becomes computationally prohibitive. |
| Beyond CPU-GPU Frequency: Memory-Clock and Tail Effects in Edge Inference Latency Estimation | Jaehoon Kang | 2026-06-15 | 下载 | Frequency-aware latency estimators enable deadline-aware DVFS for edge ML inference by modeling latency over CPU and GPU frequencies. We present a measurement study on an NVIDIA Jetson Orin Nano showi... |