Prepare completed flights
Standardised columns and mixed data types, removed duplicates, handled missing values, and excluded cancelled and diverted flights. Created route, scheduled-hour, and scheduled-duration features.
PROJECT 06 / Machine Learning
Comparing three classifiers using pre-flight features to predict arrival delays over 15 minutes.
March 2026 – June 2026
Data preparation & logistic regression
Conceptual modelling workflow
THE PROJECT AT A GLANCE
The team compared logistic regression, a decision tree, and a random forest using the same pre-flight inputs.
01 / OBJECTIVE
Predict whether a completed flight will arrive more than 15 minutes late using information available before the flight outcome is known, and compare interpretable and non-linear classification methods.
02 / MY CONTRIBUTION
I was responsible for data cleaning, preprocessing, feature engineering, and the logistic regression model, including evaluation and interpretation. The source data contained 101,000 records; cleaning produced 93,380 completed flights for modelling.
This was a three-person assignment. My model was the logistic regression baseline; teammates developed the decision tree and random forest and contributed exploratory analysis and pipeline integration. The comparison below reports the team’s saved notebook results.
03 / TECHNICAL APPROACH
Standardised columns and mixed data types, removed duplicates, handled missing values, and excluded cancelled and diverted flights. Created route, scheduled-hour, and scheduled-duration features.
Selected 14 carrier, airport, route, calendar, distance, and scheduled-time predictors. Excluded actual arrival/departure delays and other outcome-related fields to reduce target leakage.
Used a scikit-learn Pipeline and ColumnTransformer for imputation, one-hot encoding, and numerical scaling, followed by class-weighted logistic regression. The team used an 80/20 stratified train/test split.
Compared accuracy, precision, recall, F1-score, ROC-AUC, and confusion matrices across three models. With only 19.05% of flights delayed, F1 and recall helped expose the trade-off between missed delays and false alerts.
04 / OUTCOMES
| Model | F1-score | ROC-AUC | Recall |
|---|---|---|---|
| Logistic regression | 0.3873 | 0.6727 | 0.6222 |
| Decision tree | 0.4968 | 0.7695 | 0.6306 |
| Random forest | 0.4590 | 0.7483 | 0.6269 |
Recorded on an 18,676-flight test set from an 80/20 stratified split. The decision tree had the highest F1-score and ROC-AUC in this run. These are saved evaluation results; the models have not been rerun for this portfolio.
The study used historical data and scheduled inputs. Live weather, congestion, and aircraft rotation were outside the model, and future use would require validation on later flights.
Want to talk about this project?
Get in touch