Skip to content

2026-08-25 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
Automotive HSMs - Architectural Challenges and Security ImplicationsKrishna Teja Medam, Austin Bruce2026-08-25下载Automotive electronic control units (ECUs) increasingly depend on hardware-rooted security to protect software integrity, authenticity, and lifecycle management in the presence of remote and physical ...
An Open-Source Benchmark Suite of 3D-IC TestcasesRohan Soni, Jooyeon Jeong, Alexander Graening, Anthony Foo, Richard Chen, Puneet Gupta2026-08-25下载The physical design community has benefited from standardized, publicly available benchmark suites, which have enabled reproducible evaluation and driven significant advances in 2D place-and-route alg...
FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference AccelerationGeraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Yüzügüler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu, Zhou Ke, Shai Bergman, Ji Zhang2026-08-25下载LLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput. This constraint is especially acute in single-accelerator and small-node inference systems, whe...
Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization LevelsAmir Taherin, Sana Taghipour Anvari, Charles Amante, Yixiao Chen, Ruben Noroian, Zlatan Feric, Nicolas Bohm Agostini, Pu Zhao, José Cano, Bin Ren, Yanzhi Wang, David Kaeli2026-08-25下载Edge LLM deployment is shaped by more than model size and precision: inference backend, hardware platform, memory traffic, and power management all affect latency and efficiency.
Maia 200: A Software Defined Dataflow System for Large-scale AI AccelerationSherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan, Jesse Benson, Ofer Dekel, Saurabh Dighe, Anupama Kurpad, Artour Levin, Matthew Mattina, George Petre, Cheng Tang, Yuan Yu, Li Zhang, Torsten Hoefler2026-08-25下载We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth.
Simthesizer: An Agent-Driven Simulation Framework for LLM Serving SystemsWonung Kim, Hyunmin Choi, Minsu Kim, Jaehong Cho, Yeongwook Kim, Jongse Park2026-08-25下载System-level simulation is an essential tool for exploring the rapidly expanding design space of LLM serving systems, where real deployments remain costly and often infeasible.
Thermal Tuning Overhead in Wafer-Scale Optical Interconnects for LLM MoE Training: A Cross-Layer Analysis and Ferroelectric-Based MitigationSeongwon Yoon, Pin-Jun Chen, Shimeng Yu2026-08-25下载The rapid scaling of large language models (LLMs), particularly mixture-of-experts (MoE) architectures, has intensified interconnect demands because expert-parallel execution is communication-intensiv...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Probabilistic Performance Analysis of Parallel Signature Search Strategies in Multi-Level Tree NetworksJingwei Li, Thomas G. Robertazzi2026-08-25下载Hierarchical distributed search, locating a data pattern, or signature, across a tree-structured collection of files, underlies distributed index traversal, deep packet inspection and sequence alignme...
FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference AccelerationGeraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Yüzügüler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu, Zhou Ke, Shai Bergman, Ji Zhang2026-08-25下载LLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput. This constraint is especially acute in single-accelerator and small-node inference systems, whe...
Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization LevelsAmir Taherin, Sana Taghipour Anvari, Charles Amante, Yixiao Chen, Ruben Noroian, Zlatan Feric, Nicolas Bohm Agostini, Pu Zhao, José Cano, Bin Ren, Yanzhi Wang, David Kaeli2026-08-25下载Edge LLM deployment is shaped by more than model size and precision: inference backend, hardware platform, memory traffic, and power management all affect latency and efficiency.
Next-generation O-RAN Edge: Energy-aware Joint Placement and Migration of Cloud-Native FunctionsNguyen Phuc Tran, Brigitte Jaumard, Oscar Delgado2026-08-25下载The transition toward Open Radio Access Networks (O-RANs) is reshaping how cellular infrastructure is deployed, managed, and optimized. This paper investigates the energy-aware joint placement and mig...
Maia 200: A Software Defined Dataflow System for Large-scale AI AccelerationSherry Xu, Marco Heddes, Jackson Peng, Tom Savell, Monica Tang, Prashant Ranjan, Jesse Benson, Ofer Dekel, Saurabh Dighe, Anupama Kurpad, Artour Levin, Matthew Mattina, George Petre, Cheng Tang, Yuan Yu, Li Zhang, Torsten Hoefler2026-08-25下载We introduce Maia 200, an advanced AI accelerator delivering high performance-10 145 Tflop/s FP4 and 5072 Tflop/s FP8 within a 750W TDP and 7 TB/s HBM bandwidth.
Asynchronous Verifiable Information Dispersal with Low Space and Communication ComplexityThomas Locher, Yvonne-Anne Pignolet2026-08-25下载The primary goal of a distributed storage system is to ensure that clients can both write and read data in a reliable and consistent manner, even in the presence of failures.
Scalable datacenter replication with mostly-synchronous consensus on hardwareDavide Rovelli, Philipp Berdesinski, Rodrigo Otoni, Patrick Eugster2026-08-25下载Consistent replication of data among distributed processes -- a task involving the well-known consensus problem -- is notoriously expensive and hard to scale, affecting especially datacenter services ...
SatDL: Jointly Optimizing Data Redistribution and Training for Satellite-Based Distributed LearningHao Wu, Kin Whye Chew, Yizhan Han, Han Li, Jingxian Wang2026-08-25下载Satellite-based distributed learning promises to train machine-learning models directly in orbit using massive, globally dispersed sensor data, thereby avoiding large-scale data downloads to ground se...
Mixed-Precision SEM-Based CFD Simulations on GPUs: A Taylor-Green Vortex caseYanxiang Chen, Manuel Münsch, Roman Iakymchuk2026-08-25下载Mixed precision is a promising approach for reducing the computational cost and energy consumption of Computation Fluid Dynamics (CFD) simulations, but its effectiveness depends strongly on where prec...
An HPC Approach to Accelerate Tensor DecompositionsMarkus Hellgren, Erna Begovic Kovac, Hans O. Karlsson, Roman Iakymchuk2026-08-25下载Quantum systems grow in complexity so rapidly that even modest models become difficult to simulate, creating a strong need for methods that can handle high-dimensional data, also known as tensors.
pigzpp: Fast, Parallel, Portable Compression for the Whole StackThamme Gowda2026-08-25下载pigz is a widely deployed parallel gzip utility, but its process-global mutable state means that it was not designed as a reentrant, directly embeddable library.
A Few Shared Random Bits Suffice for Constant-Round Almost Stable MatchingYi-jun Chang, Kushagra Chatterjee2026-08-25下载We show that almost stable matching can be solved in constant distributed rounds on general bipartite graphs G=(V,E)G=(V,E) using only a few shared random bits.
ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud RemovalBohan Zhang, Chenyu Xu, Yijie Mao, Yuanming Shi2026-08-25下载Low-earth-orbit (LEO) satellites enable high-resolution, large-scale Earth observation for applications such as disaster monitoring and environmental surveillance.
Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM CompressionMohammad Mozaffari2026-08-25下载Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) a...
A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPUAndrew James Amos2026-08-25下载A self-organising map turns a large corpus into a browsable two-dimensional atlas, but building one at MEDLINE scale has been impractical: the best-matching-unit (BMU) search that dominates training i...
Trust, but Verify: Rigorously Profiling Best-Effort High-Performance Computing for Digital EvolutionMatthew Andres Moreno, Santiago Rodriguez Papa, Charles Ofria, Luis Zaman, Emily Dolson2026-08-25下载Developments in high-performance computing (HPC) technology continue to drastically increase quantities of available processing power. In the context of digital evolution, this explosive growth offers...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
BGPay: An Incentive-Compatible Mechanism for BGP Hijack FilteringTomasz Sadowy, Constantine Doumanidis, Maria Apostolaki2026-08-25下载BGP hijacking remains a persistent threat as existing defenses, including RPKI/ROV suffer from a fundamental incentive misalignment: the networks best positioned to filter malicious announcements bear...
ECO-COMM: An Ultra Low-Latency Event Camera based Optical Communication SystemChengling Xu, Keigo Hirakawa, Feng Ye2026-08-25下载Ultralow-latency communication is critical for emerging next-generation applications such as XR, real-time control, and distributed sensing. We present ECO-COMM, an event-camera-based optical communic...
A Capability Broker for Workflow-Network QoS Coordination in B5G/6G Industrial ServicesQize Guo, Yan Chen, Taleb Tarik, Bjoern Riemer, Hemant Zope, Hao Yu2026-08-25下载Industrial services in beyond-fifth-generation (B5G) and sixth-generation (6G) networks are increasingly executed as multi-phase workflows whose quality-of-service (QoS) demands change across ordered ...
LEMONS: Leveraging Model-Based Techniques to Enable Non-Intrusive Semantic Enrichment in Wireless Sensor NetworksJan Novacek, Arthur Kühlwein, Sebastian Reiter, Alexander Viehl, Oliver Bringmann, Wolfgang Rosenstiel2026-08-25下载The paper presents an efficient approach to the semantic enrichment of measured sensor data in Wireless Sensor Networks (WSNs), by bridging techniques from Model-driven Software Development (MDSD) and...
WiCi: Wireless GPU Computing InfrastructureYibin Shen, Wei Li, Kaiqiang Xu, Zili Meng2026-08-25下载LLM inference applications are gaining significant traction. The demand for inference is growing exponentially, and the GPU usage of inference is increasingly surpassing that of training.
A Dynamic-Kernel/QPacket Executable for Quantum Repeater Chains in Q2NS/ns-3Adam Pearson, Marcello Caleffi, Angela Sara Cacciapuoti2026-08-25下载The Quantum Internet operates on entanglement, a non-local, non-copyable, stateful network resource, which motivates protocol organization beyond classical layering.
Centrality-Based Deployment of Queue Policies in Acyclic Multipath Routing NetworksMahima Gupta, Acquin Biju, Rijul Jain, Dipesh Sharma, Sreelakshmi Manjunath2026-08-25下载Excessive queueing delays constitute a significant impediment to latency-sensitive network applications. Although effective deployment of Active Queue Management (AQM) strategies has been proposed as ...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Analyzing and Reducing Search Quality Differences in Vector Similarity SearchSara Mahdizadeh Shahri, Martin Prammer, Jignesh M. Patel, Akshitha Sriraman2026-08-25下载Modern database services scalably search over large data collections via Approximate Nearest Neighbor Search, which improves search performance at the cost of search quality, measured by recall.

cs.PF - Performance ​

标题作者发布日期PDF摘要
Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization LevelsAmir Taherin, Sana Taghipour Anvari, Charles Amante, Yixiao Chen, Ruben Noroian, Zlatan Feric, Nicolas Bohm Agostini, Pu Zhao, José Cano, Bin Ren, Yanzhi Wang, David Kaeli2026-08-25下载Edge LLM deployment is shaped by more than model size and precision: inference backend, hardware platform, memory traffic, and power management all affect latency and efficiency.
Next-generation O-RAN Edge: Energy-aware Joint Placement and Migration of Cloud-Native FunctionsNguyen Phuc Tran, Brigitte Jaumard, Oscar Delgado2026-08-25下载The transition toward Open Radio Access Networks (O-RANs) is reshaping how cellular infrastructure is deployed, managed, and optimized. This paper investigates the energy-aware joint placement and mig...
ABSTRACTS: Amsterdam Benchmark Suite for the Time and Resource Analysis of Clifford+T SimulatorsMatthew Sutcliffe, John van de Wetering2026-08-25下载Recent years have seen a rapid growth in literature presenting new methods for simulating non-Clifford quantum circuits with classical hardware.
Compression Trinity: Exploring Sparsity, Quantization, and Low-Rank Approximations for LLM CompressionMohammad Mozaffari2026-08-25下载Prohibitive computational and environmental costs impede the scalable deployment of Large Language Models (LLMs). Traditional compression techniques (sparsity, quantization, low-rank approximations) a...
The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain ProblemElioth Sanabria2026-08-25下载Large language model providers are compute constrained, and their universal response to congestion is to degrade service: route queries to smaller models, cut reasoning effort, truncate context.

基于 VitePress 构建 · 使用本地搜索查找论文