Open Access
ARTICLE
CALPHAD-Informed MAP Priors for Cold-Start Composition-Space Partitioning in Active Alloy Design
1 College of Artificial Intelligence, Tianjin University of Science and Technology, Tianjin, China
2 China Nuclear Power Engineering Co., Ltd., Beijing, China
3 School of Artificial Intelligence, Xiangyang Polytechnic University, Xiangyang, China
4 Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China
5 School of Automation, Southeast University, Nanjing, China
6 School of Materials Science and Engineering, Tianjin University, Tianjin, China
* Corresponding Authors: Xinjie Yao. Email: ; Zhoupeng Guo. Email:
; Dahai Xia. Email:
Computers, Materials & Continua 2026, 89(2), 15 https://doi.org/10.32604/cmc.2026.086475
Received 31 May 2026; Accepted 27 July 2026; Issue published 15 September 2026
Abstract
Cold-start alloy-design campaigns often have too few labeled compositions to reliably locate phase boundaries for tree-structured composition-space Gaussian process regression (TCGPR). We study a controlled way to incorporate external CALPHAD-like boundary information into this partitioning step. The proposed MP-TCGPR method adds a Gaussian MAP penalty centered on a thermodynamic boundary estimate and uses an adaptive width to reduce prior influence as node-level data accumulate. The revised theory distinguishes asymptotic convergence from convergence rate: a fixed-width prior is also asymptotically negligible under local regularity, whereas the adaptive schedule accelerates finite-sample prior release. Across 900 one-dimensional simulation runs spanning six Al-alloy cross-sections, MP-TCGPR reduces normalized Hausdorff distance at from to . Additional fixed-width MAP, CALPHAD-only, misspecification, , two-dimensional, and active-learning diagnostics show that most cold-start gain comes from a sufficiently accurate thermodynamic prior, while downstream AL benefits are modest in the tested budgets. The method is therefore presented as a reproducible cold-start partitioning strategy whose reliability depends on prior calibration, not as validation of a specific production CALPHAD database.Keywords
Accelerated alloy design often uses active learning (AL) to select the next composition or simulation that is expected to be most informative [1–3]; Bayesian optimization and uncertainty-comparison studies provide related guidance for sequential materials modeling [4], while multi-fidelity approaches [5] offer further extensions for alloy design. Recent reviews on active learning for materials design [6] and Bayesian optimization [7] provide complementary perspectives. A common difficulty is that property landscapes can change abruptly across thermodynamic phase boundaries; a surrogate trained across such boundaries may average incompatible regimes and mislead acquisition. Tree-structured composition-space Gaussian process regression (TCGPR) addresses this by recursively partitioning composition space and fitting region-specific GPs, drawing on Bayesian CART and treed GP ideas [8,9]. However, at cold start, the number of labeled compositions may be too small to locate a split threshold reliably. Broader reviews and demonstrations of machine learning for materials and alloys show that data efficiency and extrapolation remain central challenges for composition-space design [10,11]; failed-experiment learning offers another illustration of data-efficient materials discovery [12].
CALPHAD thermodynamic calculations provide external information about phase boundaries before new experiments are run [13,14]; uncertainty-aware CALPHAD and database-development studies provide related context [15–17]. This makes CALPHAD a natural source of prior information for partitioning, but using it requires caution. Commercial databases do not always provide calibrated Gaussian uncertainty for a specific boundary, and database errors can be systematic, temperature-dependent, and non-Gaussian. Accordingly, this paper studies a controlled methodological question rather than claiming validation of a particular production database: when a thermodynamic boundary estimate and an assumed uncertainty are available, can they stabilize a cold-start TCGPR split, and how sensitive is the result to prior bias?
We introduce MP-TCGPR, a MAP-prior variant of TCGPR that adds a Gaussian penalty centered on a CALPHAD-like boundary estimate to the split-threshold objective. The method also includes an adaptive prior-width schedule,
The contributions of the revised manuscript are:
1. A MAP split objective that injects thermodynamic boundary information into the TCGPR threshold score while leaving the rest of the tree and GP machinery unchanged.
2. A corrected local bias–variance analysis showing how prior bias can make MAP worse than MLE, together with explicit regularity limitations for full recursive trees.
3. New baselines separating CALPHAD-only, fixed-width MAP, and adaptive-width MAP, with paired statistical tests and confidence intervals.
4. Reproducible coordinate scaling from physical mol% to normalized split coordinates for six Al-alloy cross-sections, including the projected quaternary Al-Mg-Zn-Si case.
5. Additional misspecification,
Multi-fidelity Bayesian optimization approaches further extend these capabilities to scenarios with limited high-fidelity data.
2.1 Tree-Structured Composition-Space Gaussian Process Regression (TCGPR)
TCGPR denotes the composition-space tree-partitioning framework studied here. It builds on Bayesian CART and treed Gaussian-process ideas [8,9] for fitting composition-property surfaces that contain sharp discontinuities at thermodynamic phase boundaries. A single Gaussian process (GP) across a multi-phase composition space tends to smooth over such discontinuities [18], whereas TCGPR recursively partitions the composition domain with axis-aligned split thresholds and fits an independent GP within each resulting region. Recent advances in composition-space Gaussian process regression for alloy design have been widely reported.
Let
where
2.2 CALPHAD Boundary Information
The CALPHAD (Calculation of Phase Diagrams) methodology [13,14,19] uses parametric Gibbs-energy descriptions fitted to thermochemical and phase-equilibrium measurements to predict phase boundaries in alloy systems [15,17,20]. Production databases and software such as Thermo-Calc and PANDAT are widely used for Al-alloy design [21,22], but they do not universally provide a calibrated, Gaussian, per-boundary uncertainty estimate as a standard output. Uncertainty quantification in CALPHAD database development has attracted increasing attention [23], and the integration of CALPHAD with machine learning has been recently reviewed [24]. Recent aluminum-alloy and phase-prediction examples provide related materials-informatics context [25–27]. Therefore, in this paper
This distinction is important because the experiments below do not query a production thermodynamic database. They ask a narrower statistical question: if an external thermodynamic model supplies a boundary estimate
2.3 Relation to Existing Bayesian and Physics-Informed Methods
MP-TCGPR is related to, but narrower than, several existing research lines. Bayesian CART and Bayesian treed Gaussian-process models place priors on tree structures and region-specific response surfaces [8,9], whereas the present work keeps the TCGPR tree-search machinery fixed and adds an externally supplied thermodynamic prior only to the split-threshold score. Physics-informed Gaussian-process and hybrid thermodynamics–machine-learning workflows often encode physical knowledge through kernels, mean functions, constraints, uncertainty quantification, or auxiliary thermodynamic features [17]; MP-TCGPR instead uses a CALPHAD-like boundary estimate as a local MAP regularizer for composition-space partitioning. Classical convex-hull and phase-field methods address phase stability or microstructure evolution from different physical assumptions [28], and they are complementary baselines for future end-to-end materials benchmarks rather than direct replacements for the cold-start threshold estimator studied here. The methodological novelty should therefore be read as an incremental, reproducible modification to TCGPR plus explicit prior-bias diagnostics, not as a general new Bayesian-CALPHAD framework.
2.4 Bias–Variance View of MAP Thresholds
Near a locally regular split optimum, the MLE threshold can be approximated as
If the prior center has bias
Thus MAP improves over MLE only when the variance reduction exceeds the squared bias introduced by the prior. Equivalently, for a fixed
Eq. (2) also corrects an important asymptotic point. If
The proposed MP-TCGPR workflow is summarized in Fig. 1.

Figure 1: The MP-TCGPR pipeline. Four sequential stages: (1) A CALPHAD thermodynamic database supplies a phase-boundary estimate
MP-TCGPR modifies only the split-threshold scoring step of TCGPR. For a candidate split dimension
The quadratic penalty is evaluated in the same normalized coordinate system as the tree split. All other operations—candidate dimension search, grid search over
The adaptive schedule is
where
Proposition 1 (Local prior-weight decay): Assume that a fixed tree node has a unique interior threshold optimum, that the composite GP log marginal likelihood is locally quadratic in
This proposition is a local approximation, not a proof of convergence of the full recursive tree. Changing split dimensions, imbalanced leaf sizes, sparse samples near a boundary, multimodal likelihoods, and finite grid resolution can all violate the smooth single-threshold assumptions. The empirical convergence and fixed-width baseline experiments below therefore serve as checks on the practical magnitude of these effects.
3.3 Coordinate Scaling and Prior Calibration
All split coordinates are normalized to
The reported Hausdorff distances are dimensionless normalized distances unless explicitly stated otherwise. Table 1 gives the physical ranges and the normalized equivalent of a 0.05 mol% prior width for every tested system.

Because production CALPHAD uncertainty estimates were not available for these simulations,
We implement the TCGPR-style tree-partitioning baseline with explicit settings added for reproducibility, following the Bayesian CART and treed-GP modeling principles cited above. Each node evaluates all composition coordinates as candidate split dimensions and then optimizes the threshold by grid search over 100 uniformly spaced values on the normalized interval
4.2 Synthetic Property and Prior Generation
The one-dimensional experiments use a piecewise-GP data-generating process. For each system, the normalized boundary
The simulated CALPHAD estimate is
with system-specific
4.3 Alloy Systems and Dimensionality
The six test cases are Al-Zn-Mg, Al-Mg-Si, Al-Cu-Mg, Al-Zn-Cu, Al-Mg-Zn-Si, and Al-Cu-Li. Al-Mg-Zn-Si is correctly treated as a projected quaternary cross-section, not a ternary system. The one-dimensional experiments correspond to fixed-temperature composition cross-sections through two-phase regions; Table 1 gives the axis, temperature, physical range, and normalization used for each. The two-dimensional diagnostic uses a normalized slice with a linear boundary
4.4 Evaluation Metrics and Statistics
The primary boundary metric is normalized Hausdorff distance (HD), lower being better. For one-dimensional thresholds we additionally store false-positive and false-negative region fractions and intersection-over-union in the source data. For the two-dimensional slice we report both boundary HD and region F1. MP-better rate is the fraction of paired runs where the MAP-prior method has lower HD than MLE. Confidence intervals for mean HD use normal 95% intervals over runs; confidence intervals for MP-better rates use approximate binomial 95% intervals. Because TCGPR and MAP-prior variants are evaluated on matched systems, seeds, and labeled sets, significance is assessed with one-sided Wilcoxon signed-rank tests on paired HD reductions. The text reports paired median and interquartile range (IQR) summaries and states whether results remain significant after Bonferroni correction for the family of baseline comparisons.
5.1 Cold-Start Partition Accuracy
Fig. 2 illustrates the mechanism schematically: without the MAP prior (panel A), split estimates scatter widely; with the prior (panel B), the search is anchored near the thermodynamic estimate.

Figure 2: Schematic of cold-start partition behavior at
Table 2 and Fig. 3 report the one-dimensional partition benchmark over 900 runs (6 systems


Figure 3: Mean Hausdorff distance vs. cold-start size
Because HD captures only the maximum boundary discrepancy, we also record run-level false-positive and false-negative phase-region fractions in the benchmark outputs. For a one-dimensional threshold, these fractions are the one-sided interval lengths
5.2 Fixed-Width MAP and CALPHAD-Only Baselines
Table 3 isolates the incremental contribution of the adaptive schedule by adding the reviewer-requested baselines: pure MLE, CALPHAD-only, fixed-width MAP with

Fig. 4 is reinterpreted as an empirical optimizer-gap study rather than a claim of exact statistical convergence at finite

Figure 4: Empirical MAP–MLE grid-search gap as
5.4 Prior-Misspecification Failure Envelope
As shown in Fig. 5, the MP-better rate declines gradually as bias increases.

Figure 5: Robustness to prior bias: MP-better rate vs. bias
Table 4 reproduces the original deterministic-bias grid. The most important revision is interpretive: the below-threshold cell at

The extended grid in Fig. 6 uses 240 paired runs per cell and biases up to

Figure 6: Extended prior-bias failure envelope with uncertainty. Error bars show approximate 95% binomial confidence intervals over 240 paired runs per cell. The dashed line denotes the 80% practical-utility threshold. Performance declines smoothly beyond
5.5 Sensitivity and Higher-Dimensional Check
Table 5 reports the original prior-width sweep, and Table 6 adds the requested


Table 7 adds a two-dimensional composition-slice experiment with a linear phase boundary. The adaptive prior reduces boundary HD from

The revised active-learning diagnostic reports both cumulative regret and best-found optimality gap over 40 runs per budget. It is a budget-level utility check rather than a full closed-loop acquisition benchmark; therefore it is used only to test whether the partition gains plausibly translate into downstream optimization gains. At a budget of 100 evaluations, adaptive MAP reduces mean cumulative regret from 0.0404 for MLE partitioning to 0.0292, while the best-found gap changes only from 0.0275 to 0.0258. These effects are directionally favorable but modest compared with the boundary HD improvement. We therefore frame the method primarily as a cold-start partition estimator; the claim that improved partitions reliably produce large downstream AL gains requires larger real-material studies.
The main finding is that thermodynamic priors substantially improve cold-start composition-space partitioning. MP-TCGPR reduces normalized HD by
Several limitations remain. First, the CALPHAD priors are simulated rather than queried from Thermo-Calc, PANDAT, or another production database; real database errors may be systematic, non-Gaussian, and composition- or temperature-dependent. Second, most experiments use one-dimensional cross-sections, although the new 2D slice shows that the MAP-prior idea can extend to a simple low-dimensional boundary parameterization. Third, the synthetic piecewise-GP generator is favorable to tree partitioning, so comparisons with linear models, SVM/RBF classifiers, convex-hull boundary methods, phase-field models, high-throughput screening workflows, and other Bayesian optimization or partitioning approaches remain necessary for a complete benchmark [33–35]. Extrapolation benchmarks and high-throughput materials platforms provide complementary evaluation settings [36–38]; broader materials-informatics perspectives also motivate such benchmarking [39]. Fourth, the computational cost remains small for axis-aligned thresholds (one quadratic penalty per candidate threshold), but richer 2D/3D boundary families would require more expensive optimization and stronger regularization.
Recent adjacent materials-informatics studies further illustrate the broader context in which such partition and prior-calibration problems arise, including ML-assisted aluminum-alloy design, quasi-phase prediction, interpolation–extrapolation trade-offs, polymer-material property prediction, high-throughput first-principles screening, and interface calculations [25–27]. Interpolation–extrapolation, polymer-material, high-throughput screening, and interface studies provide additional related examples [40]; first-principles interface calculations further illustrate this broader context. Overall, MP-TCGPR should be viewed as a lightweight mechanism for incorporating external thermodynamic boundary estimates into TCGPR when the uncertainty of those estimates is credible. It is most useful in the data-scarce regime and most risky when the thermodynamic prior is overconfident or biased. The revised experiments are intended to make these tradeoffs explicit rather than to overstate validation on real CALPHAD or DFT data.
This study presents MP-TCGPR, a lightweight MAP-prior extension of TCGPR that stabilizes cold-start composition-space partitioning when a thermodynamic boundary estimate is available. Theoretical analysis further shows that fixed-width MAP priors become asymptotically negligible under regularity conditions, while adaptive prior scheduling primarily serves as a finite-sample mechanism to accelerate prior release as node-level data accumulate. Controlled one-dimensional benchmarks demonstrate that MP-TCGPR achieves the largest improvement under limited initial budgets, and the CALPHAD-only and fixed-width MAP baselines further confirm that the gain mainly originates from incorporating a calibrated thermodynamic prior.
Additional diagnostics characterize the applicability and limitations of the proposed approach. Prior-bias experiments indicate that the method remains reliable when the prior center is reasonably close to the true boundary, typically within one prior standard deviation. Two-dimensional experiments further suggest that the MAP-prior formulation can extend beyond one-dimensional thresholds, although more complex boundary representations require additional investigation. The active-learning utility study shows consistent but moderate regret reduction, suggesting that the current contribution primarily lies in improving cold-start partition estimation rather than fully optimizing end-to-end alloy-design workflows.
Future studies will focus on three aspects: validating prior assumptions using production CALPHAD databases (e.g., Thermo-Calc and PANDAT) with realistic uncertainty sources, developing higher-dimensional boundary models with conflict-aware prior control, and integrating improved partitioning into complete closed-loop alloy-design campaigns with comprehensive evaluation of acquisition efficiency, GP calibration, optimization performance, and robustness under uncertain thermodynamic knowledge.
Acknowledgement: Not applicable.
Funding Statement: The authors received no specific funding for this study.
Author Contributions: Conceptualization, Haipeng Hu and Dahai Xia; methodology, Haipeng Hu and Xinjie Yao; software, Haipeng Hu and Junjie Zhu; validation, Junjie Zhu, Tao Hong and Zhoupeng Guo; formal analysis, Haipeng Hu and Xinjie Yao; investigation, Haipeng Hu, Junjie Zhu and Tao Hong; writing—original draft preparation, Haipeng Hu; writing—review and editing, Tao Hong and Junjie Zhu; visualization, Haipeng Hu and Tao Hong; supervision, Xinjie Yao, Zhoupeng Guo and Dahai Xia; project administration, Dahai Xia. All authors reviewed and approved the final version of the manuscript.
Availability of Data and Materials: Data and Material will be made available on reasonable request.
Ethics Approval: Not applicable.
Conflicts of Interest: The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| Abbreviation | Definition |
| AL | Active Learning |
| CALPHAD | Calculation of Phase Diagrams |
| DFT | Density Functional Theory |
| GP | Gaussian Process |
| HD | Hausdorff Distance |
| MAP | Maximum A Posteriori |
| MLE | Maximum Likelihood Estimation |
| MP-TCGPR | MAP-Prior Tree-structured Composition-space Gaussian Process Regression |
| RBF | Radial Basis Function |
| TCGPR | Tree-structured Composition-space Gaussian Process Regression |
References
1. Lookman T, Balachandran PV, Xue D, Yuan R. Active learning in materials science with emphasis on adaptive sampling using uncertainties for targeted design. npj Comp Mater. 2019;5(1):21. doi:10.1038/s41524-019-0153-8. [Google Scholar] [CrossRef]
2. Balachandran PV, Xue D, Theiler J, Hogden J, Lookman T. Adaptive strategies for materials design using uncertainties. Sci Rep. 2016;6(1):19660. doi:10.1038/srep19660. [Google Scholar] [CrossRef]
3. Zhang Y, Apley DW, Chen W. Bayesian optimization for materials design with mixed quantitative and qualitative variables. Sci Rep. 2020;10(1):4924. doi:10.1038/s41598-020-60652-9. [Google Scholar] [CrossRef]
4. Tran K, Neiswanger W, Yoon J, Zhang Q, Xing E, Ulissi ZW. Methods for comparing uncertainty quantifications for material property predictions. Mach Learn Sci Technol. 2020;1(2):025006. doi:10.1088/2632-2153/ab7e1a. [Google Scholar] [CrossRef]
5. Khatamsaz D, Arroyave R, Allaire DL. Asynchronous multi-information source Bayesian optimization. J Mech Des. 2024;146(10):101708. doi:10.1115/1.4065064. [Google Scholar] [CrossRef]
6. Zong B, Li J, Yuan T, Wang J, Yuan R. Recent progress on machine learning with limited materials data: using tools from data science and domain knowledge. J Mater. 2025;11(3):100916. [Google Scholar]
7. Greenhill S, Rana S, Gupta S, Vellanki P, Venkatesh S. Bayesian optimization for adaptive experimental design: a review. IEEE Access. 2020;8:13937–48. doi:10.1109/access.2020.2966228. [Google Scholar] [CrossRef]
8. Chipman HA, George EI, McCulloch RE. Bayesian CART model search. J Am Stat Assoc. 1998;93(443):935–48. doi:10.1080/01621459.1998.10473750. [Google Scholar] [CrossRef]
9. Gramacy RB, Lee HKH. Bayesian treed Gaussian process models with an application to computer modeling. J Am Stat Assoc. 2008;103(483):1119–30. doi:10.1198/016214508000000689. [Google Scholar] [CrossRef]
10. Wei J, Chu X, Sun XY, Xu K, Deng HX, Chen J, et al. Machine learning in materials science. InfoMat. 2019;1(3):338–58. doi:10.1002/inf2.12028. [Google Scholar] [CrossRef]
11. Hart GLW, Mueller T, Toher C, Curtarolo S. Machine learning for alloys. Nat Rev Mater. 2021;6(8):730–55. doi:10.1038/s41578-021-00340-w. [Google Scholar] [CrossRef]
12. Raccuglia P, Elbert KC, Adler PDF, Falk C, Wenny MB, Mollo A, et al. Machine-learning-assisted materials discovery using failed experiments. Nature. 2016;533:73–6. doi:10.1038/nature17439. [Google Scholar] [CrossRef]
13. Lukas HL, Fries SG, Sundman B. Computational thermodynamics: the CALPHAD method. Cambridge, UK: Cambridge University Press; 2007. [Google Scholar]
14. Dinsdale AT. SGTE data for pure elements. Calphad. 1991;15(4):317–425. [Google Scholar]
15. Bocklund B, Otis R, Egorov A, Obaied A, Roslyakova I, Liu ZK. ESPEI for efficient thermodynamic database development, modification, and uncertainty quantification: application to Cu–Mg—CORRIGENDUM. MRS Commun. 2020;10(4):702–2. [Google Scholar]
16. Shahmir H, Forghani F. High-throughput computational and machine-learning design of high-entropy alloys based on thermodynamic and empirical parameters: Al-Si-Cr-Fe-(Ni,Mn) systems. Intermetallics. 2026;195(5):109328. doi:10.1016/j.intermet.2026.109328. [Google Scholar] [CrossRef]
17. Ury N, Otis R, Ravi V. Generalized method of sensitivity analysis for uncertainty quantification in CALPHAD calculations. Calphad. 2022;79:102504. doi:10.1016/j.calphad.2022.102504. [Google Scholar] [CrossRef]
18. Rasmussen CE, Williams CKI. Gaussian processes for machine learning. Cambridge, MA, USA: MIT Press; 2006. [Google Scholar]
19. Chang YA, Chen S, Zhang F, Yan X, Xie F, Schmid-Fetzer R, et al. Phase diagram calculation: past, present and future. Prog Mater Sci. 2004;49(3–4):313–45. doi:10.1016/s0079-6425(03)00025-2. [Google Scholar] [CrossRef]
20. Li S, Liu D, Li S, Chen M. A materials discovery method considering the trade-off phenomenon in machine learning prediction capabilities between interpolation and extrapolation: case study on multi-objective Mg-Zn-Al alloy design. Comput Mater Contin. 2026;87(2):14. doi:10.32604/cmc.2026.075830. [Google Scholar] [CrossRef]
21. Kaur M, Randhawa P, Jaiswal J, Dubal D, Bulakhe RN, Balakrishnan D, et al. Data-driven materials science using machine learning and computational modeling. Comput Mater Contin. 2026;88(2):4. doi:10.32604/cmc.2026.079503. [Google Scholar] [CrossRef]
22. Chen SL, Daniel S, Zhang F, Chang YA, Yan XY, Xie FY, et al. The PANDAT software package and its applications. Calphad. 2002;26(2):175–88. doi:10.1016/s0364-5916(02)00034-2. [Google Scholar] [CrossRef]
23. Otis R. Uncertainty reduction and quantification in computational thermodynamics. Comput Mater Sci. 2022;212:111590. doi:10.1016/j.commatsci.2022.111590. [Google Scholar] [CrossRef]
24. Liu F, Xiao X, Huang L, Tan L, Liu Y. Design of NiCoCrAl eutectic high entropy alloys by combining machine learning with CALPHAD method. Mater Today Commun. 2022;30:103172. doi:10.1016/j.mtcomm.2022.103172. [Google Scholar] [CrossRef]
25. Wang H, Duan Z, Guo Q, Zhang Y, Zhao Y. Machine learning design of aluminum-lithium alloys with high strength. Comput Mater Contin. 2023;77(2):1393–409. doi:10.32604/cmc.2023.045871. [Google Scholar] [CrossRef]
26. Lin Z, Yang C. Artificial intelligence design of sustainable aluminum alloys: a review. Comput Mater Contin. 2025;86(2):1–33. doi:10.32604/cmc.2025.070735. [Google Scholar] [CrossRef]
27. Zhu C, Zhao B, Naranjo Villota JL, Gao Z, Feng L. Quasi-phase equilibrium prediction of multi-element alloys based on machine learning and deep learning. Comput Mater Contin. 2023;76(1):49–64. doi:10.32604/cmc.2023.036729. [Google Scholar] [CrossRef]
28. Chen LQ. Phase-field models for microstructure evolution. Annu Rev Mater Res. 2002;32(1):113–40. doi:10.1146/annurev.matsci.32.112001.132041. [Google Scholar] [CrossRef]
29. Berger JO, Moreno E, Pericchi LR, Bayarri MJ, Bernardo JM, Cano JA, et al. An overview of robust Bayesian analysis. Test. 1994;3(1):5–124. doi:10.1007/bf02562676. [Google Scholar] [CrossRef]
30. Grunwald P, van Ommen T. Inconsistency of Bayesian inference for misspecified linear models, and a proposal for repairing it. Bayesian Anal. 2017;12(4):1069–103. [Google Scholar]
31. Harris CR, Millman KJ, van der Walt SJ, Gommers R, Virtanen P, Cournapeau D, et al. Array programming with NumPy. Nature. 2020;585:357–62. [Google Scholar]
32. Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825–30. [Google Scholar]
33. Shahriari B, Swersky K, Wang Z, Adams RP, de Freitas N. Taking the human out of the loop: a review of Bayesian optimization. Proc IEEE. 2016;104(1):148–75. doi:10.1109/jproc.2015.2494218. [Google Scholar] [CrossRef]
34. Frazier PI. A tutorial on Bayesian optimization. arXiv:1807.02811. 2018. [Google Scholar]
35. Dunn A, Wang Q, Ganose A, Dopp D, Jain A. Benchmarking materials property prediction methods: the Matbench test set and Automatminer reference algorithm. npj Comput Mater. 2020;6:138.Erratum in: npj Comput Mater. 2020;6:159. [Google Scholar]
36. Meredig B, Antono E, Church C, Hutchinson M, Ling J, Paradiso S, et al. Can machine learning identify the next high-temperature superconductor? Examining extrapolation performance for materials discovery. Mol Syst Des Eng. 2018;3(5):819–25. [Google Scholar]
37. Curtarolo S, Hart GLW, Nardelli MB, Mingo N, Sanvito S, Levy O. The high-throughput highway to computational materials design. Nat Mater. 2013;12(3):191–201. doi:10.1038/nmat3568. [Google Scholar] [CrossRef]
38. Jain A, Ong SP, Hautier G, Chen W, Richards WD, Dacek S, et al. Commentary: the materials project: a materials genome approach to accelerating materials innovation. APL Mater. 2013;1:011002. [Google Scholar]
39. Agrawal A, Choudhary A. Perspective: materials informatics and big data: realization of the fourth paradigm of science in materials science. APL Mater. 2016;4(5):053208. [Google Scholar]
40. Mo W, Lu Q, Zheng X, Yang M, Zeng Y, Li K, et al. CALPHAD-based cross-system knowledge transfer for rapid discovery of high-performance Al-Mg–Zn alloys. npj Comp Mater. 2026;12(1):201. doi:10.1038/s41524-026-02073-2. [Google Scholar] [CrossRef]
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools