Machine Learning Prediction of Mohs Hardness from Atomic Descriptors: Evaluation on an Independent Synthetic-Crystal Dataset
Publication Date : Jul-22-2026
Author(s) :
Volume/Issue :
Abstract :
Predicting how hard a material is just from its chemical makeup is difficult, because atomic-level properties don’t translate simply into large-scale mechanical behavior. This study tested four machine learning models — Linear Regression, Ridge Regression, Random Forest, and Gradient Boosting — to evaluate how well they could predict Mohs hardness using eleven averaged atomic descriptors. The models were trained on 622 natural mineral compositions (training set) and then tested on an independent set of 52 synthetic crystals (test set) they had never seen before, which allowed evaluation of how well they generalize from natural to synthetic materials. Because Random Forest and Gradient Boosting involve randomness, each one was trained 100 times with different random seeds, and results are reported as an average with standard deviation. Random Forest performed best (RMSE = 0.94 ± 0.02; test-set R² = 0.62, 95% CI 0.59–0.64), cutting error by about 35% compared to Linear Regression (RMSE = 1.43, R² = 0.10); a paired bootstrap comparison confirmed this improvement was unlikely to be due to chance (95% CI for the RMSE difference: 0.20–0.82, which does not include zero). Gradient Boosting performed comparably well (RMSE = 1.02 ± 0.01; R² = 0.54 ± 0.01). A feature-removal analysis using Random Forest, repeated 10 times per descriptor, found that the atomic-number-to-mass ratio (Z/A) mattered most, followed by average ionization energy. Covalent radius turned out to matter less than earlier estimates suggested, and its effect wasn’t clearly different from normal run-to-run noise once repeated trials were used. Density-related descriptors barely mattered at all. Overall, these results point to a real link between electronic and bonding-related descriptors and Mohs hardness, but the model’s test-set R² of about 0.62 means it only accounts for roughly 62% of the variation in the test set, and it struggles more with very hard materials — showing the limits of predicting hardness from composition alone.
