Skip to content

2026-06-15 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Energy-efficient codon optimization on thermodynamic hardwareAndraz Jelincic, Ross C. Walker2026-06-15下载The growing energy demand for computation is becoming increasingly unsustainable. Thermodynamic computing, which harnesses physical thermal fluctuations as a computational resource rather than suppres...
PDAGENT-BENCH: Characterizing, Grounding, and Architecting LLM Agents for VLSI Physical DesignQiufeng Li, Rongqian Chen, Quan Cheng, Chengxuan Wang, Sizhe Tang, Wuxi Li, Duo Ding, Chia-Tung Ho, Haoxing Ren, David Z. Pan, Tian Lan, Weidong Cao2026-06-15下载Large Language Models and vision-language models have shown remarkable success in the front-end design of Very Large-Scale Integrated Circuits, yet their capabilities for VLSI physical design remain s...
From Compression to Deployment: Real-Time and Energy-Efficient FastGRNN on Ultra-Constrained MicrocontrollersEmre Can Kizilates2026-06-15下载The dominant trajectory of modern machine learning has been to scale up: larger models, larger accelerators, larger memory budgets. Yet a multi-year global semiconductor supply constraint and the grow...
HAMON: Passive Optical Sequence Mixing for Long-Horizon ForecastingAlper Yıldırım2026-06-15下载Simple linear and frequency-domain models remain surprisingly competitive in long-horizon time-series forecasting, and recent mechanistic evidence suggests that standard forecasting benchmarks may not...
Shift-Left High-Level Synthesis Verification via Knowledge-Augmented LLM AgentZhihan Xiao, Zhe Zhao, Luke Ztz Hu, Songping Mai2026-06-15下载High-Level Synthesis (HLS) enables rapid hardware development by translating C/C++ programs into hardware implementations. Functional consistency verification between golden C specifications and HLS-o...
Neural dynamical systems on ferroelectric compute-in-memory for real-time forecastingKeshava Katti, Adithya Selvakumar, Pratik Chaudhari, Deep Jariwala2026-06-15下载Neural dynamical systems are expressive temporal predictors that capture continuous-time dynamics through fine-grained state updates. However, this sequential structure maps poorly onto digital hardwa...
Architecture Carbon Tool v3: Enabling Sustainability-aware Silicon System Design ExplorationVincent T. Lee, Bilge Acun, Zachary Lewis, Carole-Jean Wu2026-06-15下载As the carbon cost of manufacturing and operating semiconductor devices has come into sharper focus, sustainability has gradually emerged as a new system architecture design metric.
From the NYU Ultracomputer to Modern Exascale: A Historical and Architectural Survey of In-Network Computing and Scalable SynchronizationLars Warren Ericson2026-06-15下载This paper presents a historical and technical survey of the hardware architectures, interconnection networks, and synchronization primitives that have shaped massively parallel systems over the past ...
DataGuard: Guaranteeing Private Training in Systolic-array Based AcceleratorsPawan Kumar Sanjaya, Christina Giannoula, Nikhil Shreekumar, Ian Colbert, Alec Dewulf, Mehdi Saeedi, Ihab Amer, Gabor Sines, Nandita Vijaykumar2026-06-15下载Differential privacy (DP) and federated learning (FL) have emerged as important privacy-preserving approaches when using sensitive data to train machine learning (ML) models.
TreeGRNG: Binary Tree Gaussian Random Number Generator for Efficient Probabilistic AI HardwareJonas Crols, Guilherme Paim, Shirui Zhao, Marian Verhelst2026-06-15下载Bayesian Neural Networks (BNNs) offer opportunities for greatly enhancing the trustworthiness of conventional neural networks by monitoring the uncertainties in decision-making.
Towards Delta Aware Training: Efficient DNN Weight Storage for Resource-Constrained FPGAsDavid Peter Federl, Lukas Einhaus, Andreas Erbslöh, Gregor Schiele2026-06-15下载The deployment of embedded deep neural networks on resource-constrained field programmable gate arrays (FPGAs) is challenging due to limited memory and computational capacities.
NeuronFabric: A Software Reference Architecture for On-Chip Transformer Training with Local AdamEvgeny Ukladchikov2026-06-15下载Publicly documented accelerator architectures generally separate training computation from optimizer-state updates or rely on external memory and host orchestration.
MPX: A Unified Systolic Array for Matrix and Polynomial MultiplicationGeorge Alexakis, Dimitrios Schoinianakis, Giorgos Dimitrakopoulos2026-06-15下载Polynomial multiplication is a fundamental kernel in Fully Homomorphic Encryption (FHE) and post-quantum cryptography (PQC) and is commonly accelerated through Number Theoretic Transforms (NTTs).
Embedded Arena: Iterative Optimization via Hardware FeedbackZhihan Zhang, Alexander Le Metzger, Jiuyang Lyu, Chun-Cheng Chang, Jiayi Shao, Yujia Liu, Emmanuel Azuh Mensah, Edward Wang, Kurtis Heimerl, Gregory D. Abowd, Shwetak Patel, Natasha Jaques, Vikram Iyer2026-06-15下载Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints.
AIA: A 16nm Multicore SoC for Approximate Inference Acceleration Exploiting Non-normalized Knuth-Yao Sampling and Inter-Core Register SharingShirui Zhao, Nimish Shah, Wannes Meert, Marian Verhelst2026-06-15下载Probabilistic graphical models (PMs) are popular to empower machine learning with the ability of reasoning and decision-making. To perform approximate inference in PMs, sampling-based Markov Chain Mon...
When Proofs Meet Hardware: Comparing NTT and SumCheck in Zero-Knowledge SystemsJianqiao Mo, Alhad Daftardar, Barath GaneshKumar, Kaiyue Guo, Hong Wang, Benedikt Bunz, Siddharth Garg, Brandon Reagen2026-06-15下载In the ZKP community, it has long been discussed that the SumCheck protocol is asymptotically more efficient than the Number Theoretic Transform (NTT), requiring only O(N)O(N) arithmetic versus $O(N \lo...
AIA: A Customized Multi-core RISC-V SoC for Discrete Sampling Workloads in 16 nmShirui Zhao, Nimish Shah, Wannes Meert, Marian Verhelst2026-06-15下载Probabilistic models (PMs) are essential in advancing machine learning capabilities, particularly in safety-critical applications involving reasoning and decision-making.
Beyond CPU-GPU Frequency: Memory-Clock and Tail Effects in Edge Inference Latency EstimationJaehoon Kang2026-06-15下载Frequency-aware latency estimators enable deadline-aware DVFS for edge ML inference by modeling latency over CPU and GPU frequencies. We present a measurement study on an NVIDIA Jetson Orin Nano showi...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Space-Efficient Lock-Free Linear-Probing Hash TableHagit Attiya, Rotem Oshman, Noa Schiller2026-06-15下载Linear probing is one of the simplest and most space-efficient approaches to hash table design, and is widely used in sequential settings due to its compact memory layout.
Local Fault Repair of Perfect Resource Placements in Eisenstein--Jacobi NetworksBader Albader2026-06-15下载Perfect resource placements in dense Eisenstein--Jacobi (EJ) networks partition the network into hexagonal radius-tt service cells. This paper studies local repair of such placements after resource f...
Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model SystemsSajjad Khan2026-06-15下载Multi-agent LLM systems share state through memory stores, vector indices, and tool registries. We model such sharing as long-running read-generate-write operations under deterministic-generation sema...
Diagonal-Budgeted Trotterization for Efficient Quantum Hamiltonian SimulationSrikar Chundury, Blake Burgstahler, Jiajia Li, In-Saeng Suh, Frank Mueller2026-06-15下载Efficient classical simulation of quantum Hamiltonian dynamics is often bottlenecked by exponential state growth and the overhead of generic sparse linear algebra.
Re-Rooting-Based Fault-Tolerant Broadcasting in Dense Gaussian NetworksBader Albader, Mohamed R. Al-Mulla, Galal Hassan2026-06-15下载Dense Gaussian networks provide degree-4 interconnection topologies with small diameter and regular structure, making them suitable for efficient one-to-all broadcasting.
Single-Connection Mixed-Criticality Transport with CATS: Bounded Guarantees, Three Structural Limits, and a QUIC EscapeSyed Muhammad Aqdas Rizvi2026-06-15下载Mixed-criticality applications, such as satellite terminals, industrial telemetry, embedded systems, tactical, and other constrained links, often multiplex a small, latency-critical message class and ...
Tangram: Hiding GPU Heterogeneity for Efficient LLM ParallelizationYanda Tao, Pedro F. Silvestre, Marcel Wagenländer, Peter Pietzuch2026-06-15下载The scale of LLM training jobs requires parallelization planning over large GPU clusters. Due to different GPU types and interconnects added over time, these GPU clusters are increasingly heterogeneou...
A Unified Constant-Time Switch Rule for Constructing Edge-Disjoint Hamiltonian Cycles in Gaussian NetworksBader Albader2026-06-15下载Gaussian networks are degree-four symmetric interconnection networks defined over residue classes of Gaussian integers. Earlier work showed that when the generator α=a+bi satisfies gcd(a,b)=1\gcd(a,b)=1, th...
Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method SelectionMarkus Bujotzek, Dimitrios Bounias, Stefan Denner, Ralf Floca, Maximilian Fischer, Peter Neher, Klaus Maier-Hein2026-06-15下载While federated learning (FL) enables collaborative medical image segmentation without centralizing sensitive data, real-world deployment is frequently complicated by cross-site label imperfections su...
CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding AgentsShubham Tiwari, Tapan Chugh, Nash Rickert, Simon Peter, Ratul Mahajan, Haiying Shen2026-06-15下载Coding agents are a fast-growing LLM application, executing as long-running closed-loop sessions in which LLM generations alternate with external tool calls.
From the NYU Ultracomputer to Modern Exascale: A Historical and Architectural Survey of In-Network Computing and Scalable SynchronizationLars Warren Ericson2026-06-15下载This paper presents a historical and technical survey of the hardware architectures, interconnection networks, and synchronization primitives that have shaped massively parallel systems over the past ...
Robust and Automated Reconfiguration of Byzantine Wide-Area ReplicationRowdy Chotkan, Bulat Nasrulin, Johan Pouwelse, Jérémie Decouchant2026-06-15下载Distributed systems handle adversarial nodes through redundancy, which imposes a significant performance overhead. In blockchain systems, Byzantine fault-tolerant state-machine replication (BFT-SMR) i...
DRIFT: Risk-Constrained Diffusion with Imitation Priors for Mixed-Autonomy Traffic GenerationYaoshen Yu, Minghui Liwang, Wenbo Zhu, Xinlei Yi, Yiguang Hong, Yuhan Su, Seyyedali Hosseinalipour2026-06-15下载Future intelligent transportation systems are envisioned to evolve toward a long-term mixed-autonomy paradigm, where human-driven vehicles (HVs) and autonomous vehicles (AVs) coexist within highly cou...
Incentives and Evidence in Learned Service OrchestrationSyed Izhan Khilji, Alireza Furutanpey, Schahram Dustdar2026-06-15下载Reinforcement learning for service orchestration has been the subject of sustained research for over a decade, yet it is not used in production at scale.
Generated, Parallel, Scalable? A Study of Agentic AI-Generated Julia Code on SupercomputersLinus Bantel, Anna-Lena Roth, Jonas Posner, Dirk Pflüger2026-06-15下载Julia is increasingly used in HPC as a single-language alternative to combining high-level scripting with low-level systems languages, but achieving scalable performance still requires expertise in pa...
SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix ExtensionsFeiyang Chen, Haibo Chen2026-06-15下载Modern CPUs increasingly integrate matrix extensions, such as Arm Scalable Matrix Extension (SME), that provide high-throughput matrix execution within the CPU.
Fractional Verkle Trees: A Hypertree Decomposition and Verified Proof Serialization Architecture for High-Performance Blockchain State AccumulatorsEkleen Kaur, Everton Fraga2026-06-15下载Modern blockchain state management faces a critical scalability bottleneck: maintaining cryptographic commitments over hundreds of millions of entries becomes computationally prohibitive.
Tropical: Enhancing SLO Attainment in Disaggregated LLM Serving via SLO-Aware MultiplexingJinming Ma, Jiefei Chen, Xiuhong Li, Jiangfei Duan, Haojie Duanmu, Xingcheng Zhang, Chao Yang, Dahua Lin2026-06-15下载To guarantee service quality in transformer based large language model (LLM) serving, it is essential to meet the latency constraints of both the prefill phase (measured by Time-to-First-Token, TTFT) ...
StorRep: Storage Research Experiment Patterns on Chameleon Cloud and TroviRay A. O. Sinurat, Yuyang Huang, Nanqinqin Li, Mark Powers, Michael Sherman, Kate Keahey, Haryadi S. Gunawi2026-06-15下载Storage experiments are vital to advancing storage research, but creating extensible and reproducible storage artifacts can be a challenging task.
did:crdt: Coordination-Free Decentralised Identifiers via Signed CRDTsHugo O'Connor, Claire Barnes2026-06-15下载Existing Decentralised Identifier (DID) methods require coordination, an agreed global order of operations, to update a DID document: blockchain-anchored methods incur fees and latency; lightweight pe...
Efficient Data Availability Sampling via Coded Distributed ArraysDang Pham Minh, Hung Vuong Huu, Duc A. Tran2026-06-15下载Data availability is a fundamental bottleneck in modern blockchain networks. Most blockchain systems rely on a full-replication model, which requires downloading of a full block to verify its availabi...
SwiftCache: Efficient LLM Serving for Multi-turn Conversations with Heterogeneous KV Cache SharingJianmin Hu, Minxian Xu, Sa Wang, Chong Ma, Min Shen, Kejiang Ye, Lin Qu, Chengzhong Xu2026-06-15下载Multi-turn conversation is a fundamental scenario in LLM applications, widely used in chatbots and AI agents. As the conversation evolves, historical tokens accumulate continuously.
Beyond CPU-GPU Frequency: Memory-Clock and Tail Effects in Edge Inference Latency EstimationJaehoon Kang2026-06-15下载Frequency-aware latency estimators enable deadline-aware DVFS for edge ML inference by modeling latency over CPU and GPU frequencies. We present a measurement study on an NVIDIA Jetson Orin Nano showi...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and PrototypesShengli Zhang, Deen Ma, Zibin Lin, Taotao Wang2026-06-15下载Large language models have accelerated the transition from passive conversational assistants to autonomous agents that can understand goals, plan actions, invoke tools, and execute multi-step tasks.
Local Fault Repair of Perfect Resource Placements in Eisenstein--Jacobi NetworksBader Albader2026-06-15下载Perfect resource placements in dense Eisenstein--Jacobi (EJ) networks partition the network into hexagonal radius-tt service cells. This paper studies local repair of such placements after resource f...
Cache to the Future: A Distributed Webpage Archive for Internet BlackoutsRoss Evans, Diogo Barradas2026-06-15下载Internet blackouts, occurring due to technological mishaps or intentional governmental action, prevent citizens from accessing the internet. Citizens in regions where internet blackouts are common hav...
Renderable Partial Representations for Dynamic Gaussian Splatting under Incomplete DeliveryFaruk Alpay, Levent Sarioglu, Yaser Hadri2026-06-15下载Dynamic Gaussian compression is normally optimized for complete files or complete progressive prefixes, but interactive rendering encounters partial representations: some spatiotemporal regions are pr...
Re-Rooting-Based Fault-Tolerant Broadcasting in Dense Gaussian NetworksBader Albader, Mohamed R. Al-Mulla, Galal Hassan2026-06-15下载Dense Gaussian networks provide degree-4 interconnection topologies with small diameter and regular structure, making them suitable for efficient one-to-all broadcasting.
Di5Guise: 5G Privacy with vSIMShirin Ebadi, Zach Moolman, Eric Keller, Tamara Lehman2026-06-15下载SIM cards have been the key building block of user authenticationand security in cellular networks. While they are meant to serve as privacy protecting elements in cellular communications, they can be...
Single-Connection Mixed-Criticality Transport with CATS: Bounded Guarantees, Three Structural Limits, and a QUIC EscapeSyed Muhammad Aqdas Rizvi2026-06-15下载Mixed-criticality applications, such as satellite terminals, industrial telemetry, embedded systems, tactical, and other constrained links, often multiplex a small, latency-critical message class and ...
A Unified Constant-Time Switch Rule for Constructing Edge-Disjoint Hamiltonian Cycles in Gaussian NetworksBader Albader2026-06-15下载Gaussian networks are degree-four symmetric interconnection networks defined over residue classes of Gaussian integers. Earlier work showed that when the generator α=a+bi satisfies gcd(a,b)=1\gcd(a,b)=1, th...
Robust and Automated Reconfiguration of Byzantine Wide-Area ReplicationRowdy Chotkan, Bulat Nasrulin, Johan Pouwelse, Jérémie Decouchant2026-06-15下载Distributed systems handle adversarial nodes through redundancy, which imposes a significant performance overhead. In blockchain systems, Byzantine fault-tolerant state-machine replication (BFT-SMR) i...
Predictive Dynamic Scheduling for Deterministic Communications in Beyond 5GSyed Morsleen Riaz, M. Carmen Lucas-Estañ, Baldomero Coll-Perales, Javier Gozalvez2026-06-15下载Next generation wireless networks must sustain deterministic service levels to support emerging time-sensitive applications. The ability to guarantee bounded latencies depends on the efficient managem...
Impact of ADAS and V2X Penetration Rates on Cooperative Active SafetyM. C. Lucas-Estañ, B. Coll-Perales, M. I. Khan, J. Gozalvez, O. Altintas2026-06-15下载A major driver of connected and automated driving is cooperative active safety. The effectiveness of cooperative safety applications depends on the ability of vehicles to detect traffic safety risks i...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Single-Connection Mixed-Criticality Transport with CATS: Bounded Guarantees, Three Structural Limits, and a QUIC EscapeSyed Muhammad Aqdas Rizvi2026-06-15下载Mixed-criticality applications, such as satellite terminals, industrial telemetry, embedded systems, tactical, and other constrained links, often multiplex a small, latency-critical message class and ...
CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding AgentsShubham Tiwari, Tapan Chugh, Nash Rickert, Simon Peter, Ratul Mahajan, Haiying Shen2026-06-15下载Coding agents are a fast-growing LLM application, executing as long-running closed-loop sessions in which LLM generations alternate with external tool calls.
StorRep: Storage Research Experiment Patterns on Chameleon Cloud and TroviRay A. O. Sinurat, Yuyang Huang, Nanqinqin Li, Mark Powers, Michael Sherman, Kate Keahey, Haryadi S. Gunawi2026-06-15下载Storage experiments are vital to advancing storage research, but creating extensible and reproducible storage artifacts can be a challenging task.

cs.PF - Performance ​

标题作者发布日期PDF摘要
The Right Call for Software Benchmarking: Consistent Decisions in Stateful EnvironmentsGábor Melis2026-06-15下载In the perpetual pursuit of performance, modern computing systems rely ever more on stateful mechanisms to accommodate the dynamics of workloads and physical environments, bolstering efficiency but co...
OpenGadget3 GPU solver testsA. Ragagnin, G. S. Karademir, F. Groth, K. Dolag, L. M. Böss, T. Castro, N. Hariharan, M. Aiello, L. Tornatore2026-06-15下载We present an in-depth evaluation of the scalability and accuracy of the GPU porting of the N-body code for hydrodynamic cosmological simulations \og.
Single-Connection Mixed-Criticality Transport with CATS: Bounded Guarantees, Three Structural Limits, and a QUIC EscapeSyed Muhammad Aqdas Rizvi2026-06-15下载Mixed-criticality applications, such as satellite terminals, industrial telemetry, embedded systems, tactical, and other constrained links, often multiplex a small, latency-critical message class and ...
SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix ExtensionsFeiyang Chen, Haibo Chen2026-06-15下载Modern CPUs increasingly integrate matrix extensions, such as Arm Scalable Matrix Extension (SME), that provide high-throughput matrix execution within the CPU.
Fractional Verkle Trees: A Hypertree Decomposition and Verified Proof Serialization Architecture for High-Performance Blockchain State AccumulatorsEkleen Kaur, Everton Fraga2026-06-15下载Modern blockchain state management faces a critical scalability bottleneck: maintaining cryptographic commitments over hundreds of millions of entries becomes computationally prohibitive.
Beyond CPU-GPU Frequency: Memory-Clock and Tail Effects in Edge Inference Latency EstimationJaehoon Kang2026-06-15下载Frequency-aware latency estimators enable deadline-aware DVFS for edge ML inference by modeling latency over CPU and GPU frequencies. We present a measurement study on an NVIDIA Jetson Orin Nano showi...

基于 VitePress 构建 · 使用本地搜索查找论文