A bankruptcy prediction model is essential for stakeholders to avoid losses and market imbalances arising from the misallocation of resources. In the literature, several methods have been proposed for bankruptcy prediction. However, conventional statistical techniques and financial ratios have limitations in correctly identifying financially distressed firms in a reasonable time frame. In this study, we compare several machine learning approaches to define a bankruptcy prediction model using raw accounting data from financial statements. In particular, we used a dataset of 1’826’157 financial statements from 532’255 active and 76’464 bankrupt Italian firms (i.e., the last three financial statements for each firm), considering both financial ratios and the raw accounting data as input variables. To overcome the constraints of imbalance between the two classes of bankrupt and active firms, we used rebalancing techniques. We found that Random Forest is the best-performing model, exhibiting an accuracy of 98% which is an increase of 13.95% compared to the use of the financial ratios. Moreover, relative variable importance analysis shows that equity/total debts ratio and short-term debts are the most essential bankrupt predictors.
High-dimensional Data from Financial Statements for a Bankruptcy Prediction Model / Gabrielli, G., Melioli, A., Bertini, F.. - (2023), pp. 1-7. (International Workshop on Big Data Analytics in Finance and Commerce ) [10.1109/ICDEW58674.2023.00005].
High-dimensional Data from Financial Statements for a Bankruptcy Prediction Model
Gabrielli, Gianluca;Bertini, Flavio
2023-01-01
Abstract
A bankruptcy prediction model is essential for stakeholders to avoid losses and market imbalances arising from the misallocation of resources. In the literature, several methods have been proposed for bankruptcy prediction. However, conventional statistical techniques and financial ratios have limitations in correctly identifying financially distressed firms in a reasonable time frame. In this study, we compare several machine learning approaches to define a bankruptcy prediction model using raw accounting data from financial statements. In particular, we used a dataset of 1’826’157 financial statements from 532’255 active and 76’464 bankrupt Italian firms (i.e., the last three financial statements for each firm), considering both financial ratios and the raw accounting data as input variables. To overcome the constraints of imbalance between the two classes of bankrupt and active firms, we used rebalancing techniques. We found that Random Forest is the best-performing model, exhibiting an accuracy of 98% which is an increase of 13.95% compared to the use of the financial ratios. Moreover, relative variable importance analysis shows that equity/total debts ratio and short-term debts are the most essential bankrupt predictors.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


