Skip to content

2026-09-09 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
REACH: Controller-Managed Long-Span ECC for HBM AI InferenceRui Xie, Yunhua Fang, Asad Ul Haq, Linsen Ma, Sanchari Sen, Swagath Venkataramani, Liu Liu, Tong Zhang2026-09-09下载High-Bandwidth Memory (HBM) cost motivates stronger controller protection that can support a wider range of device error rates. Long-span error-correcting codes provide stronger protection at a compar...
PASCAL: A Phase-Aware Shared-Cache Model for Parallel ScansZhongchun Zhou, Chengtao Lai, Songtao Mao2026-09-09下载In modern AI Accelerators and GPGPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query tiles share the same K/V block, GEMM, where ...
CertiFlash: A Formal Verification Framework for Flash Translation Layers in Computational Solid State DrivesHarshita Gupta, Mayank Kabra, Rakesh Nadig, Nika Mansouri Ghiasi, Sahand Divsalar, F. Nisa Bostanci, Ataberk Olgun, Konstantinos Kanellopoulos, Jisung Park, Haiyu Mao, Abdullah Giray Yaglikci, Mohammad Sadrosadati, Onur Mutlu2026-09-09下载Data-intensive applications move large amounts of data from storage to the compute unit, incurring significant data movement overhead. Storage-centric computing reduces this overhead by moving computa...
SAGE: Semantic-Aware Geographic Error Recovery for AI Data MovementPatrick S. Y. Hung, Zitong Wang, Zekai Zhang, Yu Hin Chan, Shengzhe Lyu, Ray C. C. Cheung2026-09-09下载AI interconnects typically protect and replay packets uniformly, yet numerical bit faults differ sharply in consequence: a low-order mantissa flip may resemble quantization noise, while a high-signifi...
AutoTrans: AI-Assisted Automatic Translation of Security Assertions for RISC-V ProcessorsSharjeel Imtiaz, Uljana Reinsalu, Tara Ghasempouri2026-09-09下载Reusing a set of verified security assertions across RISC-V processor targets remains one of the most expensive bottlenecks in hardware security verification.
Analytic Gradients and Nonadiabatic Couplings for Device-Resident DMRG-QD-NEVPT2 Through Conical Intersections on a Consumer GPURubén Darío Guerrero2026-09-09下载Nonadiabatic dynamics through a conical intersection needs both static and dynamic correlation and, at every geometry, an excited-state gradient and interstate nonadiabatic coupling (NACME); analytic ...
HermiCache: Enclave-Aware Cache Replacement for Trusted Execution EnvironmentsOussama Elmnaouri, Pascal Cotret, Vianney Lapôtre, Loïc Lagadec2026-09-09下载Trusted Execution Environments (TEEs) protect enclave memory from untrusted software but remain vulnerable to cache-based side-channel attacks due to shared microarchitectural resources.
AMEND: Audited Margins Enable Nonblocking Drops in GPU-PIM LLM DecodingZuxiong Tan, Will Wei-Jen Wang, Wei Shao, Ali Karkehabadi, Houman Homayoun, Avesta Sasan2026-09-09下载Autoregressive large language model (LLM) decoding re-reads a growing key-value (KV) cache at every step, so long-context attention is bound by graphics processing unit (GPU) memory bandwidth.
HBFSim: Fast and Faithful Simulation of High-Bandwidth Flash Under Real GPU ExecutionYanpeng Hu, Yiwei Yang, Yuanwu Zhu, Yusheng Zheng, Wei Zhang, Andi Quinn2026-09-09下载Serving a large language model (LLM) is limited by memory capacity. High-Bandwidth Flash (HBF) stacks NAND flash inside the accelerator package, one tier below high-bandwidth memory (HBM); the specifi...
Minimal Deadlock-Free Routing for Degree-Six Triangular-Lattice Meshes and Tori with Two Forbidden TurnsZibo Diao, Rongxi Sun2026-09-09下载Degree-six triangular-lattice interconnection networks offer substantial minimal-path diversity, but their additional directions complicate deadlock-free routing under wormhole flow control.
A Fully Wave-Domain Wideband MU MIMO OFDM Transmitter via Stacked Intelligent MetasurfacesZheao Li, Jiancheng An, Chau Yuen2026-09-09下载This paper proposes an advanced realization principle for wideband multiuser multiple-input multiple-output orthogonal frequency-division multiplexing (MU-MIMO OFDM) transmitters, where the convention...
UNISON: A Co-Designed Near-Memory Scheduler of Session KV Residency for LLM AgentsFan He, Yan Li, Xiaoyang Zeng2026-09-09下载Large language models are increasingly composed into agent loops that plan, call tools, and resume the same task after each action. These loops press a shared memory hierarchy harder than conventional...
Differential Stochastic Simulated Annealing Processor for Fully Connected 2048-Spin OptimizationNaoya Onizawa, Md Mohaimenul Alam, Sean Smithson, Duckgyu Shin, Takahiro Hanyu2026-09-09下载A 2,048-spin fully connected annealing processor based on differential stochastic simulated annealing (DSSA) is presented as an architectural design in TSMC 28 nm CMOS with a 3 mm x 4 mm post-layout a...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Navigating Small-World Networks with Distance PredictionsLadan Kian, Ming Ming Tan, Dariusz Kowalski2026-09-09下载The small-world phenomenon was given an algorithmic foundation by Kleinberg, who showed that in an augmented kk-dimensional lattice a decentralized greedy algorithm delivers a message in $O(\log^2 n)...
ExaServe: Large-Scale LLM Serving on Exascale HPC SystemsWenyi Wang, Shu Shi, Yadu Babuji, Ian Foster, Kyle Chard2026-09-09下载Cloud-native LLM serving frameworks have made deployment routine in data centers, yet deploying them on leadership-class supercomputers remains an engineering challenge requiring scheduler integration...
Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM ServingHongjian Fan, Kevin Zhang, David Habinsky, Sean Dykstra2026-09-09下载We present a Kubernetes Dynamic Resource Allocation (DRA) driver that makes composable CXL memory a schedulable cluster resource, and evaluate the resulting shared-memory tier for cross-node KV-cache ...
PASCAL: A Phase-Aware Shared-Cache Model for Parallel ScansZhongchun Zhou, Chengtao Lai, Songtao Mao2026-09-09下载In modern AI Accelerators and GPGPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query tiles share the same K/V block, GEMM, where ...
Avatar: Toward Autonomous End-to-End Orchestration of Scientific Workflows using LLMsSuman Raj, Hai Duc Nguyen, Haochen Pan, Ryan Chard, Kyle Chard, Ian Foster2026-09-09下载Scientific workflow management (WMSs) systems automate execution, yet orchestrate using fixed, hand-tuned rules. LLM agents promise more autonomous orchestration, but it remains unclear where to intro...
Stencil Computation at the Intersection of AI and HPCTimothee Ewart, Mauricio Araya-Polo2026-09-09下载Tensor compilers such as TinyTC and OpenAI Triton were originally developed for AI workloads, but the same tiling and memory abstractions can be applied to implement efficient high-order stencils for ...
CEDD-optimizer: Enabling Cost-Efficient Dataset Distillation on Geographically Distributed Edge SystemsDai Liu, Eishi Arima, Martin Schulz2026-09-09下载Centralized learning is a fundamental paradigm in modern AI, in which data are collected from distributed edge devices and aggregated at a central host for model training.
Introvert Clustering for Distributed Graph AlgorithmsYi-Jun Chang, Nima Dolatabadi2026-09-09下载We introduce a graph decomposition primitive called introvert clustering, which strengthens standard low-diameter clustering by guaranteeing that every clustered vertex keeps at least a $\left(\frac12...
Decentralized network congestion control for DAG-based distributed ledger systemMayank Pandey, Rachit Agarwal, Sandeep Kumar Shukla, Nishchal Kumar Verma2026-09-09下载We propose a variable and behavior-based node-specific proof-of-work (PoW) model for a directed acyclic graph (DAG)-based distributed ledger technology (DLT) network to mitigate decentralized network ...
Epoch: Compiling Diffusion Blocks for Sparse MoE ServingJianian Zhu, Hang Wu, Yinghui Li, Haojie Wang, Ruixuan Li, Jidong Zhai2026-09-09下载Diffusion language models generate text by refining a fixed-size block of token positions through many forward passes, a loop that does not match the per-forward execution unit used by most LLM servin...
Breaking Fault Lines: Unifying TEE-Assisted BFT Consensus in Partially Trusted WorldsXiaoqing Wen, Tong Liu, Jianyu Niu, Jialin Li, Cong Wang, Yinqian Zhang, Chen Feng2026-09-09下载This paper revisits TEE-assisted BFT under a universal partial-TEE model, where an arbitrary subset of replicas execute inside TEEs while the remaining replicas operate without hardware trust guarante...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
HAPS-RIS or HAPS-Relay: Which Outperforms Under Impairments with NOMA in 6G NTN?Bilal Karaman, Faicel Khennoufa, Ilhan Basturk, Metin Ozturk, Ferdi Kara, Sezai Taskin, Halim Yanikomeroglu2026-09-09下载This paper investigates the performance of high-altitude platform station (HAPS)-assisted communication systems employing either reconfigurable intelligent surfaces (RIS) or relay stations (RS) under ...
HybridFLow: SDN-Orchestrated Client Partitioning for Hybrid Federated LearningOsama Abu Hamdan, Rabin Pandey, Hao Che, Engin Arslan, Md Arifuzzaman2026-09-09下载Cross-silo Federated Learning (FL) enables geographically distributed institutions to collaboratively train machine learning models without sharing raw data.
Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries?Tianzhu Zhang, Chih-Kai Huang, Meikang Qiu2026-09-09下载AI agents are increasingly involved in network automation, where they can initiate configuration changes through mediated operational interfaces and assess the resulting state.
Storage-Scalable Progressive Semantic Communication via Knowledge-Base ReuseHeng Zhu, Ye Liu, Kun Zhu, Feifei Song2026-09-09下载Existing knowledge-base-assisted semantic communication schemes commonly adopt either single knowledge-base quantization (SKBQ) or multi-knowledge-base residual quantization (MKBQ).
Decentralized network congestion control for DAG-based distributed ledger systemMayank Pandey, Rachit Agarwal, Sandeep Kumar Shukla, Nishchal Kumar Verma2026-09-09下载We propose a variable and behavior-based node-specific proof-of-work (PoW) model for a directed acyclic graph (DAG)-based distributed ledger technology (DLT) network to mitigate decentralized network ...
Can AI Agents Detect and Repair Artifact Drift in Network Experiments?Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan, Meikang Qiu2026-09-09下载In recent years, AI agents have evolved into capable assistants that carry out multi-step tasks in digital environments. The network systems community is beginning to explore these capabilities in ope...
Lightweight Zero Trust via Automotive SDNFriedrich Wiemer, Florian Wagner2026-09-09下载Zonal in-vehicle networks ship Ethernet, MACsec, and TSN, but treat the network itself as trusted: once configured at the factory, there is no standardized runtime way to easily revoke access, rotate ...
NEXUS-MI: Communication-Aware Federated Personalization for Gateway-Coordinated Motor-Imagery Brain-Computer InterfacesDaniel Adu Worae, Aarthy Nagarajan2026-09-09下载Electroencephalography (EEG)-based motor-imagery brain-computer interfaces (MI-BCIs) vary across subjects and sessions, complicating personalization from limited calibration data.
Minimal Deadlock-Free Routing for Degree-Six Triangular-Lattice Meshes and Tori with Two Forbidden TurnsZibo Diao, Rongxi Sun2026-09-09下载Degree-six triangular-lattice interconnection networks offer substantial minimal-path diversity, but their additional directions complicate deadlock-free routing under wormhole flow control.
Efficient Graph Neural Networks for Multicarrier Wideband Hybrid Beamforming OptimizationBeier Li, Mai Vu2026-09-09下载6G wireless technology is poised to adopt higher and wider frequency bands, leveraging highly directional beamforming. However, the vast bandwidths amplify the impact of beam squinting.
Contextual Bandit-Based Decomposition of Network Slice Requirements under Cumulative Resource Budget ConstraintsMasaki Kobayashi, Akito Suzuki, Ryoichi Kawahara, Masahiro Kobayashi2026-09-09下载End-to-end (E2E) network slices (NSs) are provisioned across multiple domains of the 5G network. In hierarchical NS management, a tenant submits a network slice request (NSR), which specifies E2E serv...
Automated Mobile Video Objective Testing SystemEric Petajan, Jonathan Lynam, Morey Antebi, Hessam Moeini, David Lindero, Lars Ernstrom, Gyanesh Patra, Szilveszter Nadas2026-09-09下载Applying QoE analysis to optimize usage of cellular spectrum is of high interest to mobile network operators. A key challenge is to be able to perform QoE measurement across very different types of ap...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Violet: Enabling Full Virtualization for M-mode RTOS on RISC-VTaro Kito, Ryosuke Yamamoto, Keisuke Horii, Hiroki Masuda, Koichi Mouri2026-09-09下载In embedded systems, complex configurations may be required, such as the simultaneous execution of a real-time operating system (RTOS) and a general-purpose operating system (GPOS), or the operation o...
PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency ScalingWeisi Yang, Stephen Xia2026-09-09下载Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency.

cs.PF - Performance ​

标题作者发布日期PDF摘要
PASCAL: A Phase-Aware Shared-Cache Model for Parallel ScansZhongchun Zhou, Chengtao Lai, Songtao Mao2026-09-09下载In modern AI Accelerators and GPGPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query tiles share the same K/V block, GEMM, where ...
Elastoformer: Enabling Dynamic Adaptivity via Elastic Model TransformationSudaksh Kalra, Dolly Sapra2026-09-09下载EdgeAI systems are increasingly employing computer vision applications to enable intelligent, on-device decision-making in real-time. However, these deployments face highly dynamic operational conditi...
Forward-Free LLM Depth Pruning via Weight RedundancyVincent-Daniel Yun, Woosang Lim2026-09-09下载Depth pruning reduces large language model (LLM) inference cost by removing complete Transformer blocks. Activation-based methods collect hidden states through forward passes on calibration data, whil...
PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency ScalingWeisi Yang, Stephen Xia2026-09-09下载Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency.

基于 VitePress 构建 · 使用本地搜索查找论文