Researchers combined classification, regression and sustainability screening to search a vast oxide-perovskite landscape for materials with band gaps suited to next-generation solar absorbers.

Paper: Predictive modeling and discovery of double perovskite oxides for photovoltaics. AI-generated conceptual image created using ChatGPT/OpenAI
Finding suitable materials for efficient and sustainable solar cells remains an important challenge. In a study published in the journal Scientific Reports, researchers investigated whether machine learning could accelerate the search for promising oxide perovskites for photovoltaic applications. Oxide perovskites are particularly interesting because their electronic and optical properties can be adjusted through changes in composition.
However, many known oxide perovskites are either insulating or have band gaps that are too large for effective solar energy conversion. The real challenge, therefore, is to identify compositions with moderate band gaps that could serve as light-absorbing materials. Conventional discovery methods make this difficult.
Experimental trial and error requires considerable time and resources, while Density Functional Theory (DFT) calculations become computationally expensive when thousands of possible compositions need to be screened. As a result, many potentially useful oxide perovskites may remain unexplored.
A faster, more systematic screening approach is needed to narrow down the large number of candidates and identify promising materials before conducting detailed calculations and experiments.
Multi-Phase ML Framework
To identify promising double perovskite oxides for photovoltaic applications, the researchers developed a machine-learning framework for predicting electronic band gaps. They started from a previously generated database of 551,696 charge-neutral and geometrically formable oxide perovskites, screened using criteria such as the Goldschmidt tolerance and octahedral factors.
From this pool, 5,450 compounds with available DFT-calculated band-gap values using the Local Density Approximation (LDA) were selected for model development. After compounds containing organic A-site cations were excluded, 515,452 structures were retained as the candidate pool for model inference. Although LDA is known to underestimate band gaps, this limitation was considered during the analysis and later used when defining the photovoltaic screening window.
Since complete structural information was unavailable for many compounds, the researchers represented each material using composition-based feature vectorization rather than computationally expensive structure-based descriptors. The Oliynyk framework was chosen because it captures a broad range of meaningful elemental properties.
The resulting features were then refined in three stages. First, low-variance and highly correlated features were removed to reduce redundancy. Distance correlation and mutual information were subsequently used to identify features associated with band-gap values, including nonlinear relationships. Finally, LightGBM feature importance was used to remove variables with no importance to the model.
The resulting optimized feature set was then used to train and evaluate both classification and regression models for band-gap prediction.
Band Gap Prediction & Screening
The machine-learning framework demonstrated strong performance on test data in identifying oxide perovskites predicted to have suitable band gaps for photovoltaic applications. Two classification models were first developed using band-gap thresholds of 0.5 and 2.0 eV. Among the models tested, XGBoost performed best, achieving test accuracies of 96.6% and 95.7%.
The 2.0 eV classifier showed particularly high confidence. For the 2.0 eV model, 73.8% of the 36,229 structures had confidence levels above 95%, while more than 81% exceeded 90%. Confidence was lower for the 0.5 eV classifier, with 19.7% exceeding 95% confidence and 28.7% exceeding 90% confidence. The authors cautioned, however, that these test results mainly reflect performance within the chemically related dataset and may represent an upper bound when applied to more chemically distinct compositions.
Materials that passed the classification stage were then evaluated using regression models to more precisely predict their band gaps. Several train-test ratios were examined, and an 80:20 split was ultimately selected because it provided a good balance between training performance and reliable testing, with only small variations in mean absolute error. A weighted-voting ensemble combining Extra Trees, SVR, CatBoost, and LightGBM produced the best overall regression performance, with a mean absolute error of about 0.209 eV.
This screening identified 15,966 structures with predicted band gaps between 1.28 and 1.62 eV, an uncertainty-adjusted range based on the band gaps of high-efficiency perovskite solar cells. Because the LDA-based training data tend to underestimate band gaps, the lower boundary was extended from 1.48 to 1.28 eV while the upper boundary remained at 1.62 eV. This target was derived from predominantly halide-perovskite device performance rather than from demonstrated high-efficiency oxide-perovskite cells. Additional filters were then applied to account for structural formability, charge neutrality, toxicity, elemental availability, supply-chain concerns, and manufacturability. After this screening, only 38 promising oxide compositions remained.
Importantly, these candidates contained neither lead nor cadmium, reducing some of the toxicity concerns associated with existing photovoltaic materials. The approach also predicted band gaps close to 1.5 eV for candidates such as Ba2GeSnO6 and K2MoSnO6, comparable to the band gap of CdTe. Overall, the results narrow a very large materials space to a small group of candidates that deserve more detailed computational and experimental investigation.
Candidates Require Further Validation
In summary, the study developed a multi-phase machine learning framework to predict electronic band gaps of oxide perovskite structures, effectively prioritizing promising candidates for solar PV applications.
From an initial pool of over half a million screened structures, successive classification, regression, and structural screening narrowed the candidate space substantially before sustainability and manufacturability criteria reduced it to the final shortlist.
Through the subsequent application of filters for structural formability, charge neutrality, supply-chain sustainability, toxicity, elemental scarcity, and manufacturability, the researchers ultimately identified a refined set of 38 lead- and cadmium-free oxide compositions.
This work demonstrates the potential of data-driven machine learning frameworks for rapid property prediction and large-scale material filtering, providing a prioritized starting point for higher-fidelity calculations and experimental validation rather than establishing the 38 compositions as deployment-ready photovoltaic materials. Thermodynamic and dynamic stability, absorption and charge-transport properties, defect tolerance, synthesis feasibility, and device efficiency remain to be evaluated.