Skip to content

2026-04-30 ​

cs.AR - Architecture ​

标题作者发布日期PDF摘要
DPU or GPU for Accelerating Neural Networks Inference -- Why not both? Split CNN InferenceAli Emre Oztas, Mahir Demir, James Garside, Mikel Luj'an2026-04-30下载Video and image streaming on edge devices requires low latency. To address this, Neural Networks (NNs) are widely used, and prior work mainly focuses on accelerating them with single hardware units su...
I hope we don't do to trust what advertising has done to loveJade Alglave2026-04-30下载Advertising uses love to sell stuff, like nylons. It also uses the word "love" in trivialising ways -- do you "love" your oven? When I hear about trust in the context of AI, especially agentic, I hope...
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow ArchitecturesMuhammad Ihsan Al Hafiz, Artur Podobas2026-04-30下载Spiking neural networks (SNNs) are a promising paradigm for energy-efficient event-driven computation, but large-scale SNN execution remains challenging because sparse spike communication and synchron...
Affinity Tailor: Dynamic Locality-Aware Scheduling at ScaleJin Xin Ng, Ori Livneh, Richard O'Grady, Josh Don, Peng Ding, Samuel Grossman, Luis Otero, Chris Kennelly, David Lo, Carlos Villavieja2026-04-30下载Modern large multicore systems often run multiple workloads that share CPUs under schedulers such as Linux CFS. To keep CPUs busy, these schedulers load-balance runnable work, causing each workload to...
AME-PIM: Can Memory be Your Next Tensor Accelerator?Emanuele Venieri, Simone Manoni, Alberto Florian, Jaehyun Park, Kyomin Sohn, Andrea Bartolini2026-04-30下载High Bandwidth Memory with Processing-in-Memory (HBM-PIM) offers an opportunity to reduce data movement by executing computation directly inside memory, but current commercial platforms expose limited...
RuC: HDL-Agnostic Rule Completion Benchmark GenerationArnau Ayguadé Domingo, Miquel Alberti-Binimelis, Cristian Gutierrez-Gomez, Emanuele Parisi, Razine Moundir Ghorab, Miquel Moreto, Gokcen Kestor, Dario Garcia-Gasulla2026-04-30下载Large Language Models (LLMs) have rapidly improved in performance across code-related tasks, making their integration into Register Transfer Level (RTL) development increasingly attractive.
HAVEN: Hybrid Automated Verification ENgine for UVM Testbench Synthesis with LLMsChang-Chih Meng, Yu-Ren Lu, Guan-Yu Lin, Tsung Tai Yeh, Kai-Chiang Wu, I-Chen Wu2026-04-30下载Integrated Circuit (IC) verification consumes nearly 70% of the IC development cycle, and recent research leverages Large Language Models (LLMs) to automatically generate testbenches and reduce verifi...
CuLifter: Lifting GPU Binaries to Typed IRJisheng Zhao, Huanzhi Pu, Shinnung Jeong, Chihyo Ahn, Hyesoon Kim2026-04-30下载GPU compilers merge all data types into a single unified register file, erasing the type information that binary-analysis tools rely on. We show that type recovery from this untyped register file is t...
VitaLLM: A Versatile, Ultra-Compact Ternary LLM Accelerator with Dependency-Aware SchedulingZi-Wei Lin, Tian-Sheuan Chang2026-04-30下载Deploying Large Language Models (LLMs) on resource-constrained edge devices faces critical bottlenecks in memory bandwidth and power consumption. While ternary quantization (e.g., BitNet b1.
RCW-CIM: A Digital CIM-based LLM Accelerator with Read-Compute/WriteYan-Cheng Guo, Tian-Sheuan Chang, Jian-Wei Su2026-04-30下载Digital computing-in-memory (DCIM) has emerged as a promising solution for large language model (LLM) acceleration by minimizing data transfers between external DRAM and on-chip accelerators while mai...
Autoformalizing Memory Specifications with AgentsJan Ole Ernst, Dmitri Michelangelo Saberi, Derek Christ, Thomas Zimmermann, Rajath Salegame, Suhaas M. Bhat, Stanislav Levental, Thomas Dybdahl Ahle, Matthias Jung2026-04-30下载The primary goal of Design Verification (DV) is to ensure that a proposed chip design implementation (either in code, or physical form) exactly matches its specification and is free of functional erro...

cs.DC - Distributed, Parallel, and Cluster Computing ​

标题作者发布日期PDF摘要
2B2B or Not 2B2B: A Tale of Three Algorithms for Streaming: Covariance Estimation after Welford and Chan-Golub-LeVequeFelix Reichel2026-04-30下载We place three algorithms for computing the unbiased sample covariance matrix in streaming and distributed settings on a common algebraic, numerical, and statistical foundation.
Replication in Graph Partitioning and Scheduling ProblemsPál András Papp, Toni Böhnlein, A. N. Yzelman2026-04-30下载The efficient parallel execution of complex computations requires balancing the workload across processors while minimizing the communication between them.
Network Digital Untwinning: Towards Backward Optimization of Digital TwinsZifan Zhang, Dianwei Chen, Anjun Gao, Manhua Wang, Mingzhe Chen, Minghong Fang, Xianfeng Yang, Yuchen Liu2026-04-30下载Network digital twins (NDTs) are transforming network management by offering precise virtual replicas of physical network systems. However, their reliance on diverse and sensitive data introduces sign...
Akita: A High Usability Simulation Framework for Computer ArchitectureSabila Al Jannat, Ying Li, Mengyang He, Xuzhong Wang, Huizhi Zhao, Jingxiang Sun, Daoxuan Xu, Enze Xu, Yifan Sun2026-04-30下载Computer architecture simulation is essential for evaluating new designs without the need for costly tapeout. The community has developed dozens of valuable simulators that have enabled significant ar...
NeuroRing: Scaling Spiking Neural Networks via Multi-FPGA Bidirectional Ring Topologies and Stream-Dataflow ArchitecturesMuhammad Ihsan Al Hafiz, Artur Podobas2026-04-30下载Spiking neural networks (SNNs) are a promising paradigm for energy-efficient event-driven computation, but large-scale SNN execution remains challenging because sparse spike communication and synchron...
Characterizing Path-Independent Fees: A Route to Zero Impermanent Loss in CPMMsAndrey Voronin, Roman Vlasov, Vladimir Gorgadze, Andrey Seoev, Yury Yanovich2026-04-30下载Constant Product Market Makers use fees that are typically fixed proportions of trade size. When these fees are automatically reinvested into the pool, as in Uniswap~V2 and some designs of Uniswap V4,...
From Impermanent Loss to Sustainable Gain: Quantifying Profitability Zones for Liquidity Providers on DEXIgnat Melnikov, Roman Vlasov, Vladimir Gorgadze, Andrey Seoev, Yury Yanovich2026-04-30下载Decentralized Finance (DeFi) is a rapidly evolving segment of blockchain technology that enables a transformative approach to financial services through Web3 applications.
Exploring Sparse Matrix Multiplication Kernels on the Cerebras CS-3Milan Shah, Sheng Di, Michela Becchi2026-04-30下载In recent years, novel AI accelerators have emerged as promising alternatives to GPU for AI model training and inference tasks. One such accelerator, the Cerebras CS-3, achieves strong performance on ...
Distributed Santa Claus via Global RoundingTijn de Vos, Leo Wennmann, Malte Baumecker, Yannic Maus, Florian Schager2026-04-30下载In this paper, we consider the Santa Claus problem in the CONGEST model. This NP-hard problem can be modeled as a bipartite graph of children and gifts where an edge indicates that a child desires a g...
The Origins of MEV: Systematic Attribution of Arbitrage Opportunity Creation at ScaleAndrei Seoev, Dmitry Belousov, Anastasiia Smirnova, Ksenia Kurinova, Aleksei Smirnov, Denis Fedyanin, Yury Yanovich2026-04-30下载Maximal Extractable Value (MEV) represents billions of dollars in extracted value that fundamentally shapes blockchain network dynamics and participant incentives.
Affinity Tailor: Dynamic Locality-Aware Scheduling at ScaleJin Xin Ng, Ori Livneh, Richard O'Grady, Josh Don, Peng Ding, Samuel Grossman, Luis Otero, Chris Kennelly, David Lo, Carlos Villavieja2026-04-30下载Modern large multicore systems often run multiple workloads that share CPUs under schedulers such as Linux CFS. To keep CPUs busy, these schedulers load-balance runnable work, causing each workload to...
AnTi-MiCS: Analytical Framework for Bounding Time in Embedded Mixed-Criticality SystemsBehnaz Ranjbar, Akash Kumar2026-04-30下载In Mixed-Criticality (MC) systems, although the high Worst-Case Execution Time (WCET) serves as a conservative upper bound representing the task's maximum execution time under all conditions, obtainin...
AI Inference as Relocatable Electricity Demand: A Latency-Constrained Energy-Geography FrameworkXubin Luo, Cheng Yang2026-04-30下载AI inference is becoming a persistent and geographically distributed source of electricity demand. Unlike many traditional electrical loads, inference workloads can sometimes be executed away from the...
ZipCCL: Efficient Lossless Data Compression of Communication Collectives for Accelerating LLM TrainingWenxiang Lin, Xinglin Pan, Ruibo Fan, Shaohuai Shi, Xiaowen Chu2026-04-30下载Communication has emerged as a critical bottleneck in the distributed training of large language models (LLMs). While numerous approaches have been proposed to reduce communication overhead, the poten...
Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and CertificationBehnaz Ranjbar, Kirankumar Raveendiran, Sudeep Pasricha, Samarjit Chakraborty, Cecilia Carbonelli, Akash Kumar2026-04-30下载The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by escalating system complexity, hardware-software ...
Monadic Presburger Predicates have Robust Population ProtocolsPhilipp Czerner, Javier Esparza, Vincent Fischer, Roland Guttenberg, Julian Pins, Simon Reilich2026-04-30下载Population protocols are a model of distributed computation in which a collection of indistinguishable finite-state agents interact randomly in pairs to decide a predicate of their initial configurati...
Back to the Future: Rethinking Endorsement in Order-Execute BlockchainsRongji Huang, Yifeng Ye, Gerui Wang, Mingchao Wan, Yuxing Duan, Jingjing Zhang, Guangtao Xue, Shengyun Liu2026-04-30下载Due to regulatory compliance and governance management, modern (permissioned) blockchains require flexible endorsement, which allows the endorsement policy for each contract or state object to be indi...
Lightweight Tamper-Evident Log Integrity Verification for IoT Edge Environments: A Merkle Tree Pipeline with Adaptive ChunkingMuhammet Anil Yagiz, Fahrettin Horasan, Ahmet Hasim Yurttakal2026-04-30下载Integrity of audit logs produced by Internet of Things (IoT) devices is a prerequisite for post-incident forensics, regulatory compliance, and operational accountability.
A Study on the Performance of Distributed Training of Data-driven CFD SimulationsSergio Iserte, Alejandro González-Barberá, Paloma Barreda, Krzysztof Rojek2026-04-30下载Data-driven methods for computer simulations are blooming in many scientific areas. The traditional approach to simulating physical behaviors relies on solving partial differential equations (PDE).
Towards the Democratization and Standardization of Dynamic Resources with MPI SpawningSergio Iserte, Iker Martín-Alvarez, Krzystof Rojek, José I. Aliaga, Maribel Castillo, Antonio J. Peña2026-04-30下载This paper presents an efficient tool for managing dynamic resources in production high-performance computing (HPC) settings, focusing on flexibility, adaptability, and user-friendliness.

cs.NI - Networking and Internet Architecture ​

标题作者发布日期PDF摘要
Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM ServingJunsun Choi, Sam Son, Sunjin Choi, Hansung Kim, Yakun Sophia Shao, Scott Shenker, Sylvia Ratnasamy, Borivoje Nikolic2026-04-30下载Mixture-of-experts (MoE) architectures have turned LLM serving into a cluster-scale workload in which communication consumes a considerable portion of LLM serving runtime.
Fidelity-Guaranteed Entanglement Routing with Distributed Purification PlanningAnthony Gatti, Anoosha Fayyaz, Prashant Krishnamurthy, Kaushik P. Seshadreesan, Amy Babay2026-04-30下载Many quantum-network applications require end-to-end Bell pairs whose fidelity exceeds a request-specific threshold, but existing entanglement routing algorithms either optimize only throughput withou...
A Multi-Perspective Study of the Internet Shutdown in IranAli Sadeghi Jahromi, Jason Jaskolka2026-04-30下载Iran conducted two nationwide Internet shutdowns in January and March 2026, the latter ongoing at the time of writing and the longest documented Iranian disruption.
RouteProfile: Elucidating the Design Space of LLM Profiles for RoutingJingjun Xu, Hongji Pu, Tao Feng, Haozhen Zhang, Jiaxuan You, Ge Liu2026-04-30下载As the large language model (LLM) ecosystem expands, individual models exhibit varying capabilities across queries, benchmarks, and domains, motivating the development of LLM routing.
Network Digital Untwinning: Towards Backward Optimization of Digital TwinsZifan Zhang, Dianwei Chen, Anjun Gao, Manhua Wang, Mingzhe Chen, Minghong Fang, Xianfeng Yang, Yuchen Liu2026-04-30下载Network digital twins (NDTs) are transforming network management by offering precise virtual replicas of physical network systems. However, their reliance on diverse and sensitive data introduces sign...
DeGenTWeb: A First Look at LLM-dominant WebsitesSichang Steven He, Calvin Ardi, Ramesh Govindan, Harsha V. Madhyastha2026-04-30下载Many recent news reports have claimed that content generated by large language models (LLMs) is taking over the web. However, these claims are typically not based on a representative sample of the web...
A MEC-Based Optimization Framework for Dynamic Inductive ChargingEmre Akıskalıoğlu, Mustafa Atmaca, Lorenzo Ghiro, Giovanni Perin, Renato Lo Cigno2026-04-30下载Range anxiety and long recharging times remain critical barriers to electric vehicle adoption. Dynamic Inductive Charging (DIC) offers a compelling solution by enabling wireless power transfer while d...
NetSatBench: A Distributed LEO Constellation Emulator with an SRv6 Case StudyAndrea Detti, Shahram Dadras, Giuseppe Tropea2026-04-30下载NetSatBench is a distributed emulation platform for evaluating communication protocols and application workloads over large-scale LEO satellite systems.
Libra: Accelerating Socket I/O via Programmable Selective Data CopyingKairui Zhou, Shengkai Lin, Wei Zhang, Shizhen Zhao2026-04-30下载Layer-7 (L7) proxies are critical to modern cloud-native systems, yet their performance is increasingly bottlenecked by copying entire payloads across the kernel-user boundary.
LZn : Robust LoRa Frame Synchronization Under Frame Collisions and Ultra-Low SNR ConditionsJosé Álamos, Thomas C. Schmidt, Matthias Wählisch2026-04-30下载LoRa has become a widely adopted wireless modulation scheme in LPWANs due to its low cost, long range, and minimal transmission power. However, collisions between frames of the same spreading factor -...
Multi-Connectivity for UAVs: A Measurement Study of Integrating Cellular, Aerial Mesh, and LEO Satellite LinksAygun Baltaci, Irshad A. Meer, Mustafa Ozger, Cicek Cavdar, Dominic Schupke2026-04-30下载Future uncrewed aerial vehicle (UAV) systems increasingly combine heterogeneous communication technologies, such as low-latency aerial mesh, terrestrial cellular, and satellite links, to improve robus...
Unified 5G-IoT Framework with CAMARA Gateways and SDN FederationZihan Jia, Ze Wang, Chen Chen, Ziren Xiao, Fung Po Tso2026-04-30下载The convergence of 5G and IoT enables fully connected, intelligent environments, but it faces challenges from the fragmentation of public/private 5G networks and the heterogeneity of IoT networks.
ReVo: A Cross-Layer Reliable Volumetric Videoconferencing SystemAnkur Aditya, Diptyaroop Maji, Lingdong Wang, Bhavya Ramakrishna, Ramesh Sitaraman, Prashant Shenoy2026-04-30下载Volumetric videoconferencing enables immersive six Degrees of Freedom interactions by jointly transmitting visual appearance and 3D geometry. However, delivering volumetric video over today's networks...

cs.OS - Operating Systems ​

标题作者发布日期PDF摘要
Crab: A Semantics-Aware Checkpoint/Restore Runtime for Agent SandboxesTianyuan Wu, Chaokun Chang, Lunxi Cao, Wei Gao, Wei Wang2026-04-30下载Autonomous agents act through sandboxed containers and microVMs whose state spans filesystems, processes, and runtime artifacts. Checkpoint and restore (C/R) of this state is needed for fault toleranc...
Affinity Tailor: Dynamic Locality-Aware Scheduling at ScaleJin Xin Ng, Ori Livneh, Richard O'Grady, Josh Don, Peng Ding, Samuel Grossman, Luis Otero, Chris Kennelly, David Lo, Carlos Villavieja2026-04-30下载Modern large multicore systems often run multiple workloads that share CPUs under schedulers such as Linux CFS. To keep CPUs busy, these schedulers load-balance runnable work, causing each workload to...
treVM: Tiny Rust Embedded Virtual Machines with WASM on Variable Resource-Constrained HardwareAntoine Lavandier, Bastien Buil, Chrystel Gaber, Emmanuel Baccelli2026-04-30下载Software stacks embedded on microcontroller-based hardware typically provide rudimentary APIs programmed in C/C++, basic connectivity and, sometimes, a firmware update mechanism.

基于 VitePress 构建 · 使用本地搜索查找论文