2026-09-09
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| REACH: Controller-Managed Long-Span ECC for HBM AI Inference | Rui Xie, Yunhua Fang, Asad Ul Haq, Linsen Ma, Sanchari Sen, Swagath Venkataramani, Liu Liu, Tong Zhang | 2026-09-09 | 下载 | High-Bandwidth Memory (HBM) cost motivates stronger controller protection that can support a wider range of device error rates. Long-span error-correcting codes provide stronger protection at a compar... |
| PASCAL: A Phase-Aware Shared-Cache Model for Parallel Scans | Zhongchun Zhou, Chengtao Lai, Songtao Mao | 2026-09-09 | 下载 | In modern AI Accelerators and GPGPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query tiles share the same K/V block, GEMM, where ... |
| CertiFlash: A Formal Verification Framework for Flash Translation Layers in Computational Solid State Drives | Harshita Gupta, Mayank Kabra, Rakesh Nadig, Nika Mansouri Ghiasi, Sahand Divsalar, F. Nisa Bostanci, Ataberk Olgun, Konstantinos Kanellopoulos, Jisung Park, Haiyu Mao, Abdullah Giray Yaglikci, Mohammad Sadrosadati, Onur Mutlu | 2026-09-09 | 下载 | Data-intensive applications move large amounts of data from storage to the compute unit, incurring significant data movement overhead. Storage-centric computing reduces this overhead by moving computa... |
| SAGE: Semantic-Aware Geographic Error Recovery for AI Data Movement | Patrick S. Y. Hung, Zitong Wang, Zekai Zhang, Yu Hin Chan, Shengzhe Lyu, Ray C. C. Cheung | 2026-09-09 | 下载 | AI interconnects typically protect and replay packets uniformly, yet numerical bit faults differ sharply in consequence: a low-order mantissa flip may resemble quantization noise, while a high-signifi... |
| AutoTrans: AI-Assisted Automatic Translation of Security Assertions for RISC-V Processors | Sharjeel Imtiaz, Uljana Reinsalu, Tara Ghasempouri | 2026-09-09 | 下载 | Reusing a set of verified security assertions across RISC-V processor targets remains one of the most expensive bottlenecks in hardware security verification. |
| Analytic Gradients and Nonadiabatic Couplings for Device-Resident DMRG-QD-NEVPT2 Through Conical Intersections on a Consumer GPU | Rubén Darío Guerrero | 2026-09-09 | 下载 | Nonadiabatic dynamics through a conical intersection needs both static and dynamic correlation and, at every geometry, an excited-state gradient and interstate nonadiabatic coupling (NACME); analytic ... |
| HermiCache: Enclave-Aware Cache Replacement for Trusted Execution Environments | Oussama Elmnaouri, Pascal Cotret, Vianney Lapôtre, Loïc Lagadec | 2026-09-09 | 下载 | Trusted Execution Environments (TEEs) protect enclave memory from untrusted software but remain vulnerable to cache-based side-channel attacks due to shared microarchitectural resources. |
| AMEND: Audited Margins Enable Nonblocking Drops in GPU-PIM LLM Decoding | Zuxiong Tan, Will Wei-Jen Wang, Wei Shao, Ali Karkehabadi, Houman Homayoun, Avesta Sasan | 2026-09-09 | 下载 | Autoregressive large language model (LLM) decoding re-reads a growing key-value (KV) cache at every step, so long-context attention is bound by graphics processing unit (GPU) memory bandwidth. |
| HBFSim: Fast and Faithful Simulation of High-Bandwidth Flash Under Real GPU Execution | Yanpeng Hu, Yiwei Yang, Yuanwu Zhu, Yusheng Zheng, Wei Zhang, Andi Quinn | 2026-09-09 | 下载 | Serving a large language model (LLM) is limited by memory capacity. High-Bandwidth Flash (HBF) stacks NAND flash inside the accelerator package, one tier below high-bandwidth memory (HBM); the specifi... |
| Minimal Deadlock-Free Routing for Degree-Six Triangular-Lattice Meshes and Tori with Two Forbidden Turns | Zibo Diao, Rongxi Sun | 2026-09-09 | 下载 | Degree-six triangular-lattice interconnection networks offer substantial minimal-path diversity, but their additional directions complicate deadlock-free routing under wormhole flow control. |
| A Fully Wave-Domain Wideband MU MIMO OFDM Transmitter via Stacked Intelligent Metasurfaces | Zheao Li, Jiancheng An, Chau Yuen | 2026-09-09 | 下载 | This paper proposes an advanced realization principle for wideband multiuser multiple-input multiple-output orthogonal frequency-division multiplexing (MU-MIMO OFDM) transmitters, where the convention... |
| UNISON: A Co-Designed Near-Memory Scheduler of Session KV Residency for LLM Agents | Fan He, Yan Li, Xiaoyang Zeng | 2026-09-09 | 下载 | Large language models are increasingly composed into agent loops that plan, call tools, and resume the same task after each action. These loops press a shared memory hierarchy harder than conventional... |
| Differential Stochastic Simulated Annealing Processor for Fully Connected 2048-Spin Optimization | Naoya Onizawa, Md Mohaimenul Alam, Sean Smithson, Duckgyu Shin, Takahiro Hanyu | 2026-09-09 | 下载 | A 2,048-spin fully connected annealing processor based on differential stochastic simulated annealing (DSSA) is presented as an architectural design in TSMC 28 nm CMOS with a 3 mm x 4 mm post-layout a... |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Navigating Small-World Networks with Distance Predictions | Ladan Kian, Ming Ming Tan, Dariusz Kowalski | 2026-09-09 | 下载 | The small-world phenomenon was given an algorithmic foundation by Kleinberg, who showed that in an augmented -dimensional lattice a decentralized greedy algorithm delivers a message in $O(\log^2 n)... |
| ExaServe: Large-Scale LLM Serving on Exascale HPC Systems | Wenyi Wang, Shu Shi, Yadu Babuji, Ian Foster, Kyle Chard | 2026-09-09 | 下载 | Cloud-native LLM serving frameworks have made deployment routine in data centers, yet deploying them on leadership-class supercomputers remains an engineering challenge requiring scheduler integration... |
| Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving | Hongjian Fan, Kevin Zhang, David Habinsky, Sean Dykstra | 2026-09-09 | 下载 | We present a Kubernetes Dynamic Resource Allocation (DRA) driver that makes composable CXL memory a schedulable cluster resource, and evaluate the resulting shared-memory tier for cross-node KV-cache ... |
| PASCAL: A Phase-Aware Shared-Cache Model for Parallel Scans | Zhongchun Zhou, Chengtao Lai, Songtao Mao | 2026-09-09 | 下载 | In modern AI Accelerators and GPGPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query tiles share the same K/V block, GEMM, where ... |
| Avatar: Toward Autonomous End-to-End Orchestration of Scientific Workflows using LLMs | Suman Raj, Hai Duc Nguyen, Haochen Pan, Ryan Chard, Kyle Chard, Ian Foster | 2026-09-09 | 下载 | Scientific workflow management (WMSs) systems automate execution, yet orchestrate using fixed, hand-tuned rules. LLM agents promise more autonomous orchestration, but it remains unclear where to intro... |
| Stencil Computation at the Intersection of AI and HPC | Timothee Ewart, Mauricio Araya-Polo | 2026-09-09 | 下载 | Tensor compilers such as TinyTC and OpenAI Triton were originally developed for AI workloads, but the same tiling and memory abstractions can be applied to implement efficient high-order stencils for ... |
| CEDD-optimizer: Enabling Cost-Efficient Dataset Distillation on Geographically Distributed Edge Systems | Dai Liu, Eishi Arima, Martin Schulz | 2026-09-09 | 下载 | Centralized learning is a fundamental paradigm in modern AI, in which data are collected from distributed edge devices and aggregated at a central host for model training. |
| Introvert Clustering for Distributed Graph Algorithms | Yi-Jun Chang, Nima Dolatabadi | 2026-09-09 | 下载 | We introduce a graph decomposition primitive called introvert clustering, which strengthens standard low-diameter clustering by guaranteeing that every clustered vertex keeps at least a $\left(\frac12... |
| Decentralized network congestion control for DAG-based distributed ledger system | Mayank Pandey, Rachit Agarwal, Sandeep Kumar Shukla, Nishchal Kumar Verma | 2026-09-09 | 下载 | We propose a variable and behavior-based node-specific proof-of-work (PoW) model for a directed acyclic graph (DAG)-based distributed ledger technology (DLT) network to mitigate decentralized network ... |
| Epoch: Compiling Diffusion Blocks for Sparse MoE Serving | Jianian Zhu, Hang Wu, Yinghui Li, Haojie Wang, Ruixuan Li, Jidong Zhai | 2026-09-09 | 下载 | Diffusion language models generate text by refining a fixed-size block of token positions through many forward passes, a loop that does not match the per-forward execution unit used by most LLM servin... |
| Breaking Fault Lines: Unifying TEE-Assisted BFT Consensus in Partially Trusted Worlds | Xiaoqing Wen, Tong Liu, Jianyu Niu, Jialin Li, Cong Wang, Yinqian Zhang, Chen Feng | 2026-09-09 | 下载 | This paper revisits TEE-assisted BFT under a universal partial-TEE model, where an arbitrary subset of replicas execute inside TEEs while the remaining replicas operate without hardware trust guarante... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| HAPS-RIS or HAPS-Relay: Which Outperforms Under Impairments with NOMA in 6G NTN? | Bilal Karaman, Faicel Khennoufa, Ilhan Basturk, Metin Ozturk, Ferdi Kara, Sezai Taskin, Halim Yanikomeroglu | 2026-09-09 | 下载 | This paper investigates the performance of high-altitude platform station (HAPS)-assisted communication systems employing either reconfigurable intelligent surfaces (RIS) or relay stations (RS) under ... |
| HybridFLow: SDN-Orchestrated Client Partitioning for Hybrid Federated Learning | Osama Abu Hamdan, Rabin Pandey, Hao Che, Engin Arslan, Md Arifuzzaman | 2026-09-09 | 下载 | Cross-silo Federated Learning (FL) enables geographically distributed institutions to collaboratively train machine learning models without sharing raw data. |
| Can AI Agents Deliver Verifiable Network-Wide Outcomes Across Authority Boundaries? | Tianzhu Zhang, Chih-Kai Huang, Meikang Qiu | 2026-09-09 | 下载 | AI agents are increasingly involved in network automation, where they can initiate configuration changes through mediated operational interfaces and assess the resulting state. |
| Storage-Scalable Progressive Semantic Communication via Knowledge-Base Reuse | Heng Zhu, Ye Liu, Kun Zhu, Feifei Song | 2026-09-09 | 下载 | Existing knowledge-base-assisted semantic communication schemes commonly adopt either single knowledge-base quantization (SKBQ) or multi-knowledge-base residual quantization (MKBQ). |
| Decentralized network congestion control for DAG-based distributed ledger system | Mayank Pandey, Rachit Agarwal, Sandeep Kumar Shukla, Nishchal Kumar Verma | 2026-09-09 | 下载 | We propose a variable and behavior-based node-specific proof-of-work (PoW) model for a directed acyclic graph (DAG)-based distributed ledger technology (DLT) network to mitigate decentralized network ... |
| Can AI Agents Detect and Repair Artifact Drift in Network Experiments? | Tianzhu Zhang, Weichen Tao, Changgang Zheng, Yusheng Zheng, Long Chen, Xiaoyi Fan, Meikang Qiu | 2026-09-09 | 下载 | In recent years, AI agents have evolved into capable assistants that carry out multi-step tasks in digital environments. The network systems community is beginning to explore these capabilities in ope... |
| Lightweight Zero Trust via Automotive SDN | Friedrich Wiemer, Florian Wagner | 2026-09-09 | 下载 | Zonal in-vehicle networks ship Ethernet, MACsec, and TSN, but treat the network itself as trusted: once configured at the factory, there is no standardized runtime way to easily revoke access, rotate ... |
| NEXUS-MI: Communication-Aware Federated Personalization for Gateway-Coordinated Motor-Imagery Brain-Computer Interfaces | Daniel Adu Worae, Aarthy Nagarajan | 2026-09-09 | 下载 | Electroencephalography (EEG)-based motor-imagery brain-computer interfaces (MI-BCIs) vary across subjects and sessions, complicating personalization from limited calibration data. |
| Minimal Deadlock-Free Routing for Degree-Six Triangular-Lattice Meshes and Tori with Two Forbidden Turns | Zibo Diao, Rongxi Sun | 2026-09-09 | 下载 | Degree-six triangular-lattice interconnection networks offer substantial minimal-path diversity, but their additional directions complicate deadlock-free routing under wormhole flow control. |
| Efficient Graph Neural Networks for Multicarrier Wideband Hybrid Beamforming Optimization | Beier Li, Mai Vu | 2026-09-09 | 下载 | 6G wireless technology is poised to adopt higher and wider frequency bands, leveraging highly directional beamforming. However, the vast bandwidths amplify the impact of beam squinting. |
| Contextual Bandit-Based Decomposition of Network Slice Requirements under Cumulative Resource Budget Constraints | Masaki Kobayashi, Akito Suzuki, Ryoichi Kawahara, Masahiro Kobayashi | 2026-09-09 | 下载 | End-to-end (E2E) network slices (NSs) are provisioned across multiple domains of the 5G network. In hierarchical NS management, a tenant submits a network slice request (NSR), which specifies E2E serv... |
| Automated Mobile Video Objective Testing System | Eric Petajan, Jonathan Lynam, Morey Antebi, Hessam Moeini, David Lindero, Lars Ernstrom, Gyanesh Patra, Szilveszter Nadas | 2026-09-09 | 下载 | Applying QoE analysis to optimize usage of cellular spectrum is of high interest to mobile network operators. A key challenge is to be able to perform QoE measurement across very different types of ap... |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Violet: Enabling Full Virtualization for M-mode RTOS on RISC-V | Taro Kito, Ryosuke Yamamoto, Keisuke Horii, Hiroki Masuda, Koichi Mouri | 2026-09-09 | 下载 | In embedded systems, complex configurations may be required, such as the simultaneous execution of a real-time operating system (RTOS) and a general-purpose operating system (GPOS), or the operation o... |
| PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling | Weisi Yang, Stephen Xia | 2026-09-09 | 下载 | Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency. |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| PASCAL: A Phase-Aware Shared-Cache Model for Parallel Scans | Zhongchun Zhou, Chengtao Lai, Songtao Mao | 2026-09-09 | 下载 | In modern AI Accelerators and GPGPUs, many concurrent cores repeatedly access the same shared data. This pattern occurs in attention, where different query tiles share the same K/V block, GEMM, where ... |
| Elastoformer: Enabling Dynamic Adaptivity via Elastic Model Transformation | Sudaksh Kalra, Dolly Sapra | 2026-09-09 | 下载 | EdgeAI systems are increasingly employing computer vision applications to enable intelligent, on-device decision-making in real-time. However, these deployments face highly dynamic operational conditi... |
| Forward-Free LLM Depth Pruning via Weight Redundancy | Vincent-Daniel Yun, Woosang Lim | 2026-09-09 | 下载 | Depth pruning reduces large language model (LLM) inference cost by removing complete Transformer blocks. Activation-based methods collect hidden states through forward passes on calibration data, whil... |
| PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling | Weisi Yang, Stephen Xia | 2026-09-09 | 下载 | Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency. |