ARCHIVES
Year 2026 · Volume 5 · Issue 3
Quantifying Evaluation Bias in Machine Learning Models for Concrete Compressive Strength: Mix-Level Replication in the Yeh Benchmark and Its Differential Effect on Model Comparison
Published Online: September-December 2026
Pages: 141-148
Cite this article
↗ https://www.doi.org/10.59256/indjcst.20260503018Abstract
The 1030-record dataset compiled by Yeh is the de facto benchmark of the concrete strength-prediction literature, and coefficients of determination reported on it routinely exceed 0.90. This study shows that part of that reported accuracy reflects the internal structure of the dataset rather than predictive capability. The 1030 records describe only 427 distinct mix proportions: one batch is cast and then tested at several curing ages, so that 76.1 per cent of records share their composition with at least one other record that differs only in age. Under the random partitioning used almost universally in this literature, models are therefore evaluated on mixes already seen during training. Five estimators spanning a range of flexibility are re-evaluated under random and mix-grouped five-fold cross-validation, each repeated ten times. The mix-grouped protocol lowers R² for every estimator, but unevenly: the reduction is negligible for ordinary least squares (0.002) and k-nearest neighbours (0.004) and material for support-vector regression (0.029), random forest (0.043) and gradient boosting (0.054). The apparent advantage of gradient boosting over least squares is consequently overstated by about 19 per cent. The penalty is concentrated in the sparsely sampled high-strength region (60 MPa and above), whose records are also the most heavily replicated; there, the root-mean-square error of the two tree ensembles rises by 42 and 44 per cent. A controlled experiment on synthetic data with a known generative law and a tunable per-mix batch effect confirms the mechanism: the gap is zero without a batch effect and grows monotonically with it. The study concludes that leakage-aware evaluation, rather than synthetic augmentation, is the appropriate corrective, and provides a single executable notebook that reproduces every result.
Related Articles
2026
Artificial Intelligence in Learning and Teaching
2026
Admin Assist: An AI – Driven Configuration and Orchestration for Enterprise Application
2026
Enhancing Blood Group Identification using pigeon inspired optimization: An Innovative Approach
2026
Eco-Genius: Power Up Smart, Power Down Waste
2026
Crowd-Sourced Disaster Response and Rescue Assistant
2026
Unveiling Deepfake Detection Using Vision Transformers: A Survey and Experimental Study
Share Article
Or copy link
*Instagram doesn't support direct link sharing from web. Copy the link and share it in your Instagram story or post.