2026-07-28
cs.AR - Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar | Solomon Micheal Serunjogi, Rachmad Vidya Wicaksana Putra, Ayat Taha, Muhammad Shafique, Mahmoud Rasras | 2026-07-28 | 下载 | Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference. |
| At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference | Bowen Wang, Chi Zhang, Diyou Shen, Renzo Andri, Navaneeth Kunhi Purayil, Luca Benini | 2026-07-28 | 下载 | Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of inference for Transformer models. |
| Beyond Prefill-Decode Disaggregation: Dissecting LLM Inference for Heterogeneous Platforms via Dynamic Operator Scheduling | Jiaqi Yang, Jiayi Li, Yihan Fu, Hongxiao Zhao, Zhan Chen, Qiuping Wu, Yuchao Yang, Bonan Yan | 2026-07-28 | 下载 | Prefill-decode disaggregation (PD) and roofline-based operator placement are common strategies for partitioning Large Language Model (LLM) inference across heterogeneous systems, but they are often in... |
| ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design | Jingbo Zhang, Haoxiang Sun, Wenbo Wang, Wenbo Zhang | 2026-07-28 | 下载 | This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The workflow makes three contributions. |
cs.DC - Distributed, Parallel, and Cluster Computing
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Incast-Free MoE Rate-Based Scheduling | Evyatar Cohen, Jose Yallouz, Alexander Shpiner, Mark Silberstein, Sylvia Ratnasamy, Isaac Keslassy | 2026-07-28 | 下载 | Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks. |
| The Fabric Is the Cluster Driver: Cross-Layer eBPF Policies for GPU-CXL Fabrics | Yiwei Yang, Andi Quinn | 2026-07-28 | 下载 | We present fabric_ext, an eBPF middleware compiler and runtime for extensible OS policies over GPU--CXL fabrics. fabric_ext lets one policy program execute across GPU hooks, driver/runtime hooks, DPU/... |
| Route-Block Membership Selects Packed-AWQ Arithmetic: A Controlled Single-Fixture Mechanism Study | Lukas Stepanek | 2026-07-28 | 下载 | Mixture-of-experts (MoE) inference first aligns routed tokens into padded expert blocks, then executes packed quantized matrix multiplication over those blocks. |
| ProFlow: RL-Driven and Performance-Aware Proactive Flow Placement in Datacenter Networks | Sourya Saha, Md Nurul Absur, Saptarshi Debroy | 2026-07-28 | 下载 | In datacenter fabrics composed of leaf and aggregation switches, competing flows may become co-located on shared aggregation switches, creating congestion that can significantly degrade protected flow... |
| MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar | Solomon Micheal Serunjogi, Rachmad Vidya Wicaksana Putra, Ayat Taha, Muhammad Shafique, Mahmoud Rasras | 2026-07-28 | 下载 | Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over electronic accelerators for expediting Transformer inference. |
| Hermes: Low Tail-Latency Via Prefix Consensus | Alejandro Ranchal-Pedrosa, Dakai Kang, Neil Giridharan, Dahlia Malkhi, Mohammad Sadoghi, Ben Marsh | 2026-07-28 | 下载 | Leader-based BFT protocols finalize through their leaders: a view whose leader is crashed or slow finalizes nothing, and the timeout that ends it admits no good setting. |
| Massively parallel numerical simulations with Julia | Simon Candelaresi, Benedict Geihe, Marco Artiano, Lars Christmann, Valentin Churavy, Andrés Rueda-Ramírez, Hendrik Ranocha, Gregor J. Gassner, Michael Schlottke-Lakemper | 2026-07-28 | 下载 | The Julia programming language aims to provide a modern approach to develop high-performance computing (HPC) applications. It tries to achieve this by combining a high-level, dynamic interface with ju... |
| PowerScale: Energy-Efficient Geo-Distributed Model Training with Federated Datacenter Power | Talha Mehboob, Zhe Xu, Michael Zink, David Irwin | 2026-07-28 | 下载 | The power demands of large-scale AI training increasingly exceed the capacity of any single data center, making geo-distributed training across power-constrained sites a practical necessity. |
| Optimistic Verifiable Claims: A Blockchain Protocol for Conditionally Confidential Bidding in Decentralized Manufacturing | Marko Corn, Nejc Rožman, Primož Podržaj | 2026-07-28 | 下载 | Decentralized manufacturing faces a pre-contractual impasse: a Provider cannot price a service accurately without inspecting the design file, yet the Consumer cannot share that file without exposing i... |
| WASP: A Configurable Framework for Portable Stateful Serverless Applications | Matteo Cenzato, Dario d'Abate, Arianna Dragoni, Giacomo Orsenigo, Luca Tosetti, Alessandro Margara | 2026-07-28 | 下载 | WebAssembly (WASM) is emerging as a lightweight alternative to containers for Function-as-a-Service (FaaS) across the edge-cloud continuum. However, existing WASM-based serverless platforms are tightl... |
| CW-Ghost: Search-Free Granularity Selection for Helper-Thread Prefetching via Capacity Windows | Ya Zhang, Tong Lei, Yao Chen, Yonggang Che, Chuanfu Xu, Haozhong Qiu, Yusong Tan | 2026-07-28 | 下载 | Helper-thread prefetching hides the latency of irregular memory accesses by executing address dependency chains ahead of the main thread. However, its effectiveness depends on the range of future iter... |
| QCOEM: Quantum Cloud Orchestration with Evolutionary Multi-Objective Optimization | Tam N. Pham, Hoa T. Nguyen, Quan Le-Trung | 2026-07-28 | 下载 | Quantum cloud platforms need to dynamically orchestrate workloads across heterogeneous quantum computation backends whose noise profiles, qubit topologies, and queues vary over time. |
| Specula: Scaling formal specifications for autonomous model checking of system code | Qian Cheng, Saad Mohammad Rafid Pial, Ruize Tang, Yiming Su, Emilie Ma, Finn Hackett, Ivan Beschastnikh, Yu Huang, Tianyin Xu | 2026-07-28 | 下载 | Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding... |
cs.NI - Networking and Internet Architecture
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research | Jaber Daneshamooz, Eugene Vuong, Alagappan Ramanathan, Manni Moghimi, Haarika Manda, Satyam Kumar, Snithik Thode, Satyandra Guthula, Sylee Beltiukov, Dongsu Han, Tarun Mangla, Sangeetha Abdu Jyothi, Walter Willinger, Arpit Gupta | 2026-07-28 | 下载 | Networking research advances by turning hypotheses into empirical evidence, so accelerating it means reducing the lag between ideation (synthesizing a hypothesis) and generating the data that tests it... |
| Incast-Free MoE Rate-Based Scheduling | Evyatar Cohen, Jose Yallouz, Alexander Shpiner, Mark Silberstein, Sylvia Ratnasamy, Isaac Keslassy | 2026-07-28 | 下载 | Mixture of Experts (MoE) architectures have become key to large language models; however, their typical round-robin (RR) scheduling introduces significant bottlenecks. |
| Round Trip Time: A Benign Signal or an Indirect Window into Datacenter Workloads? | Sourya Saha, Md Nurul Absur, Saptarshi Debroy | 2026-07-28 | 下载 | Multi-tenant datacenter networks increasingly rely on shared leaf-spine fabrics, where traffic from multiple tenants traverses common network resources. |
| ProFlow: RL-Driven and Performance-Aware Proactive Flow Placement in Datacenter Networks | Sourya Saha, Md Nurul Absur, Saptarshi Debroy | 2026-07-28 | 下载 | In datacenter fabrics composed of leaf and aggregation switches, competing flows may become co-located on shared aggregation switches, creating congestion that can significantly degrade protected flow... |
| MAC-Gyver: Open, Programmable, Scheduling for AI-RAN 6G Systems | Maxime Elkael, Reshma Prasad, Tamerlan Aghayev, Salvatore D'Oro, Michele Polese, Tommaso Melodia | 2026-07-28 | 下载 | Cellular networks are integrating Artificial Intelli- gence (AI) into radio access network control. The MAC scheduler is a promising target because it allocates a limited resource, spectrum, at every ... |
| Untangling Co-Drift: Proactive Multi-Intent Failure Prediction and Root-Cause Disambiguation for Self-Driving Networks | Md. Kamrul Hossain, Walid Aljoby | 2026-07-28 | 下载 | The vision of self-driving networks that monitor, reason, and act upon themselves with minimal human intervention relies on tightly coupled monitoring, analytics, and actuation functions. |
| Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks | Ravi Kant Sharma, Ashutosh Uttam, Ajay Kumar | 2026-07-28 | 下载 | Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management standards lack a standardized mechanism for cross-vendor trust... |
| C-RE-ACT: Causal RE-ACTing Agent for O-RAN Forensic Triage | Pau Baguer, J. Xavier Salvat Lozano, Gines Garcia-Aviles, Xavier Costa-Pérez | 2026-07-28 | 下载 | The shift to O-RAN architectures marks a turning point in cellular security, where increased openness and modularity directly translate into a broader attack surface. |
| Performance Evaluation of RF-powered IoT in Rural Areas: The Wireless Power Digital Divide | Hao Lin, Mustafa A. Kishk, Mohamed-Slim Alouini | 2026-07-28 | 下载 | Bridging the digital divide is one of the goals of mobile networks in the future, and further building IoT networks in rural areas is a feasible solution. |
| Revisiting-Aware In-Orbit Edge Computing for Earth Observation | Zehua Sun, Tao Ni, Kaiyan Cui, Weitao Xu, Jingxian Wang | 2026-07-28 | 下载 | Typically, Earth observation satellites follow a rule of revisiting cycle to periodically pass over the same area of the Earth at regular intervals, which is jointly determined by their orbital proper... |
| The Model in the Middle: Toward AI-Native Real-Time Communication | Ziqian Liu, Minghao Li, Yiming Qiu | 2026-07-28 | 下载 | Full-duplex omni models are transforming human--AI interaction from turn-based exchanges into continuous multimodal conversations in which speaking, listening, and reasoning unfold concurrently. |
| WALoMA: A Multitask Wireless Foundation Model via Adaptive Low-Rank Masked Autoencoders | Madi Makin, Asmaa Abdallah, Abdulkadir Celik, Ahmed M. Eltawil | 2026-07-28 | 下载 | This paper proposes a multitask wireless foundation model via adaptive low-rank masked autoencoders (WALoMA), a unified multi-task foundation model for sixth-generation (6G) wireless physical layer ar... |
| Robust Unsupervised Network Intrusion Detection via Federated Learning with Selective Aggregation under Anomalous Sample Contamination | Shohei Kamiguchi, Takayuki Nishio | 2026-07-28 | 下载 | Network intrusion detection systems (NIDS) have become essential for Internet of Things (IoT) environments, as malware targeting IoT devices continues to evolve in sophistication. |
cs.OS - Operating Systems
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Specula: Scaling formal specifications for autonomous model checking of system code | Qian Cheng, Saad Mohammad Rafid Pial, Ruize Tang, Yiming Su, Emilie Ma, Finn Hackett, Ivan Beschastnikh, Yu Huang, Tianyin Xu | 2026-07-28 | 下载 | Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the specifications for highly effective model checking and bug finding... |
cs.PF - Performance
| 标题 | 作者 | 发布日期 | 摘要 | |
|---|---|---|---|---|
| Massively parallel numerical simulations with Julia | Simon Candelaresi, Benedict Geihe, Marco Artiano, Lars Christmann, Valentin Churavy, Andrés Rueda-Ramírez, Hendrik Ranocha, Gregor J. Gassner, Michael Schlottke-Lakemper | 2026-07-28 | 下载 | The Julia programming language aims to provide a modern approach to develop high-performance computing (HPC) applications. It tries to achieve this by combining a high-level, dynamic interface with ju... |
| Bridging Compute- and Data-Optimal Pretraining | Tian Qin, Kimia Hamidieh, David Alvarez-Melis | 2026-07-28 | 下载 | Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of hi... |