Skip to content

2026-07-28 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent CrossbarSolomon Micheal Serunjogi, Rachmad Vidya Wicaksana Putra, Ayat Taha, Muhammad Shafique, Mahmoud Rasras2026-07-28下载Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference.
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer InferenceBowen Wang, Chi Zhang, Diyou Shen, Renzo Andri, Navaneeth Kunhi Purayil, Luca Benini2026-07-28下载Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference for Transformer models.
Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator SchedulingJiaqi Yang, Jiayi Li, Yihan Fu, Hongxiao Zhao, Zhan Chen, Qiuping Wu, Yuchao Yang, Bonan Yan2026-07-28下载Prefill-decode disaggregation (PD) and roofline-based operator placement are common strategies for partitioning Large Language Model (LLM) inference across heterogeneous systems, but they are often in...
ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS DesignJingbo Zhang, Haoxiang Sun, Wenbo Wang, Wenbo Zhang2026-07-28下载This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes three contributions.

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
Incast-Free MoE Rate-Based SchedulingEvyatar Cohen, Jose Yallouz, Alexander Shpiner, Mark Silberstein, Sylvia Ratnasamy, Isaac Keslassy2026-07-28下载Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks.
The Fabric Is the Cluster Driver: Cross-Layer eBPF Policies for GPU-CXL FabricsYiwei Yang, Andi Quinn2026-07-28下载We present fabric_ext, an eBPF middleware compiler and runtime for extensible OS policies over GPU--CXL fabrics. fabric_ext lets one policy program execute across GPU hooks, driver/runtime hooks, DPU/...
Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism StudyLukas Stepanek2026-07-28下载Mixture-of-experts (MoE) inference first aligns routed tokens into padded expert blocks, then executes packed quantized matrix multiplication over those blocks.
ProFlow: RL-Driven and Performance-Aware Proactive Flow Placement in Datacenter NetworksSourya Saha, Md Nurul Absur, Saptarshi Debroy2026-07-28下载In datacenter fabrics composed of leaf and aggregation switches, competing flows may become co-located on shared aggregation switches, creating congestion that can significantly degrade protected flow...
MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent CrossbarSolomon Micheal Serunjogi, Rachmad Vidya Wicaksana Putra, Ayat Taha, Muhammad Shafique, Mahmoud Rasras2026-07-28下载Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference.
Hermes: Low Tail-Latency Via Prefix ConsensusAlejandro Ranchal-Pedrosa, Dakai Kang, Neil Giridharan, Dahlia Malkhi, Mohammad Sadoghi, Ben Marsh2026-07-28下载Leader-based BFT protocols finalize through their leaders: a view whose leader is crashed or slow finalizes nothing, and the timeout that ends it admits no good setting.
Massively parallel numerical simulations with JuliaSimon Candelaresi, Benedict Geihe, Marco Artiano, Lars Christmann, Valentin Churavy, Andrés Rueda-Ramírez, Hendrik Ranocha, Gregor J. Gassner, Michael Schlottke-Lakemper2026-07-28下载The Julia programming language aims to provide a modern approach to develop high-performance computing (HPC) applications. It tries to achieve this by combining a high-level, dynamic interface with ju...
PowerScale: Energy-Efficient Geo-Distributed Model Training with Federated Datacenter PowerTalha Mehboob, Zhe Xu, Michael Zink, David Irwin2026-07-28下载The power demands of large-scale AI training increasingly exceed the capacity of any single data center, making geo-distributed training across power-constrained sites a practical necessity.
Optimistic Verifiable Claims: A Blockchain Protocol for Conditionally Confidential Bidding in Decentralized ManufacturingMarko Corn, Nejc Rožman, Primož Podržaj2026-07-28下载Decentralized manufacturing faces a pre-contractual impasse: a Provider cannot price a service accurately without inspecting the design file, yet the Consumer cannot share that file without exposing i...
WASP: A Configurable Framework for Portable Stateful Serverless ApplicationsMatteo Cenzato, Dario d'Abate, Arianna Dragoni, Giacomo Orsenigo, Luca Tosetti, Alessandro Margara2026-07-28下载WebAssembly (WASM) is emerging as a lightweight alternative to containers for Function-as-a-Service (FaaS) across the edge-cloud continuum. However, existing WASM-based serverless platforms are tightl...
CW-Ghost: Search-Free Granularity Selection for Helper-Thread Prefetching via Capacity WindowsYa Zhang, Tong Lei, Yao Chen, Yonggang Che, Chuanfu Xu, Haozhong Qiu, Yusong Tan2026-07-28下载Helper-thread prefetching hides the latency of irregular memory accesses by executing address dependency chains ahead of the main thread. However, its effectiveness depends on the range of future iter...
QCOEM: Quantum Cloud Orchestration with Evolutionary Multi-Objective OptimizationTam N. Pham, Hoa T. Nguyen, Quan Le-Trung2026-07-28下载Quantum cloud platforms need to dynamically orchestrate workloads across heterogeneous quantum computation backends whose noise profiles, qubit topologies, and queues vary over time.
Specula: Scaling formal specifications for autonomous model checking of system codeQian Cheng, Saad Mohammad Rafid Pial, Ruize Tang, Yiming Su, Emilie Ma, Finn Hackett, Ivan Beschastnikh, Yu Huang, Tianyin Xu2026-07-28下载Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding...

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Pramana: A Composable, Domain-Specific Backend for Empirical Networking ResearchJaber Daneshamooz, Eugene Vuong, Alagappan Ramanathan, Manni Moghimi, Haarika Manda, Satyam Kumar, Snithik Thode, Satyandra Guthula, Sylee Beltiukov, Dongsu Han, Tarun Mangla, Sangeetha Abdu Jyothi, Walter Willinger, Arpit Gupta2026-07-28下载Networking research advances by turning hypotheses into empirical evidence, so accelerating it means reducing the lag between ideation (synthesizing a hypothesis) and generating the data that tests it...
Incast-Free MoE Rate-Based SchedulingEvyatar Cohen, Jose Yallouz, Alexander Shpiner, Mark Silberstein, Sylvia Ratnasamy, Isaac Keslassy2026-07-28下载Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks.
Round Trip Time: A Benign Signal or an Indirect Window into Datacenter Workloads?Sourya Saha, Md Nurul Absur, Saptarshi Debroy2026-07-28下载Multi-tenant datacenter networks increasingly rely on shared leaf-spine fabrics, where traffic from multiple tenants traverses common network resources.
ProFlow: RL-Driven and Performance-Aware Proactive Flow Placement in Datacenter NetworksSourya Saha, Md Nurul Absur, Saptarshi Debroy2026-07-28下载In datacenter fabrics composed of leaf and aggregation switches, competing flows may become co-located on shared aggregation switches, creating congestion that can significantly degrade protected flow...
MAC-Gyver: Open, Programmable, Scheduling for AI-RAN 6G SystemsMaxime Elkael, Reshma Prasad, Tamerlan Aghayev, Salvatore D'Oro, Michele Polese, Tommaso Melodia2026-07-28下载Cellular networks are integrating Artificial Intelli- gence (AI) into radio access network control. The MAC scheduler is a promising target because it allocates a limited resource, spectrum, at every ...
Untangling Co-Drift: Proactive Multi-Intent Failure Prediction and Root-Cause Disambiguation for Self-Driving NetworksMd. Kamrul Hossain, Walid Aljoby2026-07-28下载The vision of self-driving networks that monitor, reason, and act upon themselves with minimal human intervention relies on tightly coupled monitoring, analytics, and actuation functions.
Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous NetworksRavi Kant Sharma, Ashutosh Uttam, Ajay Kumar2026-07-28下载Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust...
C-RE-ACT: Causal RE-ACTing Agent for O-RAN Forensic TriagePau Baguer, J. Xavier Salvat Lozano, Gines Garcia-Aviles, Xavier Costa-Pérez2026-07-28下载The shift to O-RAN architectures marks a turning point in cellular security, where increased openness and modularity directly translate into a broader attack surface.
Performance Evaluation of RF-powered IoT in Rural Areas: The Wireless Power Digital DivideHao Lin, Mustafa A. Kishk, Mohamed-Slim Alouini2026-07-28下载Bridging the digital divide is one of the goals of mobile networks in the future, and further building IoT networks in rural areas is a feasible solution.
Revisiting-Aware In-Orbit Edge Computing for Earth ObservationZehua Sun, Tao Ni, Kaiyan Cui, Weitao Xu, Jingxian Wang2026-07-28下载Typically, Earth observation satellites follow a rule of revisiting cycle to periodically pass over the same area of the Earth at regular intervals, which is jointly determined by their orbital proper...
The Model in the Middle: Toward AI-Native Real-Time CommunicationZiqian Liu, Minghao Li, Yiming Qiu2026-07-28下载Full-duplex omni models are transforming human--AI interaction from turn-based exchanges into continuous multimodal conversations in which speaking, listening, and reasoning unfold concurrently.
WALoMA: A Multitask Wireless Foundation Model via Adaptive Low-Rank Masked AutoencodersMadi Makin, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil2026-07-28下载This paper proposes a multitask wireless foundation model via adaptive low-rank masked autoencoders (WALoMA), a unified multi-task foundation model for sixth-generation (6G) wireless physical layer ar...
Robust Unsupervised Network Intrusion Detection via Federated Learning with Selective Aggregation under Anomalous Sample ContaminationShohei Kamiguchi, Takayuki Nishio2026-07-28下载Network intrusion detection systems (NIDS) have become essential for Internet of Things (IoT) environments, as malware targeting IoT devices continues to evolve in sophistication.

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Specula: Scaling formal specifications for autonomous model checking of system codeQian Cheng, Saad Mohammad Rafid Pial, Ruize Tang, Yiming Su, Emilie Ma, Finn Hackett, Ivan Beschastnikh, Yu Huang, Tianyin Xu2026-07-28下载Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding...

cs.PF - Performance ​

标题作者发布日期PDF摘要
Massively parallel numerical simulations with JuliaSimon Candelaresi, Benedict Geihe, Marco Artiano, Lars Christmann, Valentin Churavy, Andrés Rueda-Ramírez, Hendrik Ranocha, Gregor J. Gassner, Michael Schlottke-Lakemper2026-07-28下载The Julia programming language aims to provide a modern approach to develop high-performance computing (HPC) applications. It tries to achieve this by combining a high-level, dynamic interface with ju...
Bridging Compute- and Data-Optimal PretrainingTian Qin, Kimia Hamidieh, David Alvarez-Melis2026-07-28下载Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of hi...

基于 VitePress 构建 · 使用本地搜索查找论文