Skip to main navigation Skip to search Skip to main content

Offline inverse reinforcement learning for joint optimization of energy costs and demand charge in industrial PV-battery load systems

  • Yulong Hu
  • , Sen Li*
  • *Corresponding author for this work

Research output: Contribution to journalJournal Articlepeer-review

Abstract

Industrial electricity bills are typically composed of two major components: the energy charge, which is based on the total accumulated energy consumption over a billing period (e.g., one month), and the demand charge, which depends on the highest peak power observed during the same period. Consequently, the joint optimization of energy costs (through energy arbitrage) and demand charges (through peak shaving) is crucial for effective cost management in industrial PV-battery load systems. However, this task remains fundamentally challenging due to the volatility of renewable generation and load, the complex temporal dependencies introduced by peak demand charges, and the competing objectives between immediate cost savings and long-term peak reduction—rendering existing model-based and data-driven energy management approaches inadequate for real-world applications. To tackle these challenges, this paper formulates the problem as a soft Markov Decision Process (MDP) and proposes a novel Offline Inverse Reinforcement Learning (OIRL) framework based on a dual reward-policy iterative optimization mechanism. Our approach introduces an innovative synthesis of contrastive reward learning—leveraging both expert demonstrations and on-policy trajectory rollouts—with conservative soft Q-learning optimization. This architecture enables accurate reconstruction of implicit reward structures through comparative analysis of expert and agent behaviors, while ensuring stable policy improvement via regularized value function updates with pessimistic value initialization. Extensive experiments using real-world data from our industrial partner in China demonstrate that OIRL achieves substantial energy arbitrage and peak shaving improvement compared to state-of-the-art reinforcement learning baselines in energy management. Furthermore, the framework maintains robust performance across diverse operating conditions, establishing a new paradigm for intelligent control of industrial PV-battery load systems.

Original languageEnglish
Article number127416
JournalApplied Energy
Volume408
Early online date19 Jan 2026
DOIs
Publication statusPublished - 1 Apr 2026

Bibliographical note

Publisher Copyright:
© 2026 Elsevier Ltd

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 7 - Affordable and Clean Energy
    SDG 7 Affordable and Clean Energy
  2. SDG 9 - Industry, Innovation, and Infrastructure
    SDG 9 Industry, Innovation, and Infrastructure

Keywords

  • PV-battery load systems
  • Demand charge
  • Peak shaving
  • Inverse reinforcement learning
  • Offline reinforcement learning regularization

Fingerprint

Dive into the research topics of 'Offline inverse reinforcement learning for joint optimization of energy costs and demand charge in industrial PV-battery load systems'. Together they form a unique fingerprint.

Cite this