Data-Driven Prediction and Optimisation of Concrete Compressive Strength
Machine learning trained on 1,030 concrete samples to predict how strong a mix will be, then run in reverse to find the strongest recipe the data supports.
Predicts strength within 10 MPa on 88% of samples
Individual coursework project. All modelling, tuning and optimisation my own.

- 1,030
- concrete samples, 8 inputs
- 5.41 MPa
- RMSE, 89% of variance explained
- 88%
- of predictions within 10 MPa
- 28 days → seconds
- time to a strength estimate
Overview
How strong concrete turns out depends on the recipe: cement, water, sand, stone and additives like slag, fly ash and plasticiser. The rule of thumb engineers still use was worked out on simple mixes and copes badly with modern ones, where the ingredients interact.
The usual way to find out is to mix a batch and wait 28 days for it to cure, which is slow and expensive while you are still choosing a recipe.

What I did
- Worked through what each of the eight inputs physically does before modelling anything, so every later choice had a reason behind it.
- Trained and compared a random forest, Gaussian process regression and a neural network on the same split of 1,030 samples.
- Tuned the network with a cross validated randomised search over 250 configurations covering layer size, activation, solver, learning rate, batch size and regularisation.
- Judged models on parity plots, residual spread and the gap between cross validation and test error rather than one accuracy score.
- Searched the eight dimensional recipe space with differential evolution, then bounded it to real data once the unbounded answer proved to be extrapolation.

Methods
- Random forest, Gaussian process regression and a neural network on 1,030 samples
- Cross validated randomised search over 250 hyperparameter combinations
- Differential evolution over the mix design space with physical bounds
- Parity, residual, error histogram and learning curve diagnostics
- Response surface analysis on cement content and water to cement ratio
Key results
- 01
The neural network is accurate enough to design with. It explains 89% of the variation in strength at 5.41 MPa RMSE, within 10 MPa on 88% of samples, with errors sitting symmetric around zero.
- 02
It generalises rather than memorises. Test error of 29.28 MPa² against cross validation error of 25.37 MPa² is a 15% gap, which is small enough to trust on unseen mixes.
- 03
One model looked sensible and was useless. Gaussian process regression averaged around 31 MPa of error with predictions bunched at low values, so it was reported and rejected rather than quietly dropped.
- 04
Unbounded optimisation gave a fantasy answer. It proposed a mix 84% stronger than anything in the data. Bounded to observed ranges it lands near 82 MPa in about 30 seconds, on high binder, low water and extended curing.
Outcome
The result is a usable design tool. Give it a recipe and it returns a likely strength in seconds instead of 28 days, or ask it for the best recipe within sensible limits and it finds one inside a minute. That removes a lot of trial batches early on.