All projects

PROJECT 06 / Machine Learning

Flight Delay Prediction

Comparing three classifiers using pre-flight features to predict arrival delays over 15 minutes.

WHEN

March 2026 – June 2026

CONTRIBUTION

Data preparation & logistic regression

TECHNOLOGIES
Pythonpandasscikit-learn
PREDICTIVE MODELLING06
SOURCE RECORDS101,000Flight records
PREPARED DATASET93,380Completed flights
3 classifiers compared14 pre-flight predictors
> 15 min

Conceptual modelling workflow

THE PROJECT AT A GLANCE

From noisy flight records to a fair model comparison.

The team compared logistic regression, a decision tree, and a random forest using the same pre-flight inputs.

Completed flights
93,380
Pre-flight predictors
14
Delayed flights
19.05%

01 / OBJECTIVE

What the project set out to do

Predict whether a completed flight will arrive more than 15 minutes late using information available before the flight outcome is known, and compare interpretable and non-linear classification methods.

02 / MY CONTRIBUTION

My part in the work

I was responsible for data cleaning, preprocessing, feature engineering, and the logistic regression model, including evaluation and interpretation. The source data contained 101,000 records; cleaning produced 93,380 completed flights for modelling.

This was a three-person assignment. My model was the logistic regression baseline; teammates developed the decision tree and random forest and contributed exploratory analysis and pipeline integration. The comparison below reports the team’s saved notebook results.

03 / TECHNICAL APPROACH

How it came together

01

Prepare completed flights

Standardised columns and mixed data types, removed duplicates, handled missing values, and excluded cancelled and diverted flights. Created route, scheduled-hour, and scheduled-duration features.

02

Keep the inputs available before flight

Selected 14 carrier, airport, route, calendar, distance, and scheduled-time predictors. Excluded actual arrival/departure delays and other outcome-related fields to reduce target leakage.

03

Train a transparent baseline

Used a scikit-learn Pipeline and ColumnTransformer for imputation, one-hot encoding, and numerical scaling, followed by class-weighted logistic regression. The team used an 80/20 stratified train/test split.

04

Compare errors as well as scores

Compared accuracy, precision, recall, F1-score, ROC-AUC, and confusion matrices across three models. With only 19.05% of flights delayed, F1 and recall helped expose the trade-off between missed delays and false alerts.

04 / OUTCOMES

What the work produced

Team model comparison · saved notebook outputs
ModelF1-scoreROC-AUCRecall
Logistic regression0.38730.67270.6222
Decision tree0.49680.76950.6306
Random forest0.45900.74830.6269

Recorded on an 18,676-flight test set from an 80/20 stratified split. The decision tree had the highest F1-score and ROC-AUC in this run. These are saved evaluation results; the models have not been rerun for this portfolio.

  • Prepared 93,380 completed-flight records from 101,000 source rows.
  • Developed an interpretable logistic regression baseline and evaluated delayed-flight detection.
  • Helped establish a team comparison that made false alerts, missed delays, and class imbalance visible.

The study used historical data and scheduled inputs. Live weather, congestion, and aircraft rotation were outside the model, and future use would require validation on later flights.

UP NEXT / PROJECT 01

IdeaJoust

View case study

Want to talk about this project?

Get in touch