An accurate data driven, state of charge (SoC) estimation strategy incorporating a Panasonic 18650 lithium-ion battery's (LiB's) primary parameters i.e. voltage, current and temperature is presented in this paper. The mechanism is based on reinforcement learning (RL) along with an ensemble least square boosting (LSB) algorithm and a terminal three point interpolation stage. A twin-delayed deep deterministic neural network-based policy gradient agent (TD3NN) is deployed during the RL training phase, along with a versatile reward function that not only focuses on the accuracy but also the smoothness of the estimated curve. The reward function is designed based on an absolute negative error term and additional bonuses to ensure smoothness and a descending SoC discharge pattern. The LSB algorithm incorporates the number of tree splits and learning cycles to improve the overall accuracy of the proposed approach. The data consists of 55 discharge cycles; a fivefold validation is performed by dividing the data into training and test at 80:20 ratio five times to ensure unbiased results. The numerical values of the performance metrics are reported and analyzed. The minimum root mean square error (RMSE) is reported to be 1.003% and 0.8704% for a single term absolute negative reward function and a multi-term reward function with added bonuses, respectively. These results are not only satisfactory in comparison with the existing literature but also validate the efficacy of the multi-term reward function approach.

Reinforcement Learning based State of Charge Estimation of a Lithium Ion Battery / Ali, S., Bianchi, V., De Munari, I.. - 2026:(2026), pp. 43-47. (2026 IEEE International Workshop on Metrology for Automotive (IEEE MetroAutomotive 2026) Brescia, Italy June 17-19, 2026) [10.1109/MetroAutomotive69354.2026.11644679].

Reinforcement Learning based State of Charge Estimation of a Lithium Ion Battery

Ali Sadia;Bianchi Valentina;De Munari Ilaria
2026-01-01

Abstract

An accurate data driven, state of charge (SoC) estimation strategy incorporating a Panasonic 18650 lithium-ion battery's (LiB's) primary parameters i.e. voltage, current and temperature is presented in this paper. The mechanism is based on reinforcement learning (RL) along with an ensemble least square boosting (LSB) algorithm and a terminal three point interpolation stage. A twin-delayed deep deterministic neural network-based policy gradient agent (TD3NN) is deployed during the RL training phase, along with a versatile reward function that not only focuses on the accuracy but also the smoothness of the estimated curve. The reward function is designed based on an absolute negative error term and additional bonuses to ensure smoothness and a descending SoC discharge pattern. The LSB algorithm incorporates the number of tree splits and learning cycles to improve the overall accuracy of the proposed approach. The data consists of 55 discharge cycles; a fivefold validation is performed by dividing the data into training and test at 80:20 ratio five times to ensure unbiased results. The numerical values of the performance metrics are reported and analyzed. The minimum root mean square error (RMSE) is reported to be 1.003% and 0.8704% for a single term absolute negative reward function and a multi-term reward function with added bonuses, respectively. These results are not only satisfactory in comparison with the existing literature but also validate the efficacy of the multi-term reward function approach.
2026
9798331551285
Reinforcement Learning based State of Charge Estimation of a Lithium Ion Battery / Ali, S., Bianchi, V., De Munari, I.. - 2026:(2026), pp. 43-47. (2026 IEEE International Workshop on Metrology for Automotive (IEEE MetroAutomotive 2026) Brescia, Italy June 17-19, 2026) [10.1109/MetroAutomotive69354.2026.11644679].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11381/3071458
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact