Advanced Data Analytics and Predictive Modelling Training Course
| Course code | SD-DA-002 |
|---|---|
| Duration | 5 days |
| Level | Intermediate to Advanced |
| Category | Data Analytics |
| Delivery | Classroom or live online |
| Language | English |
| Certificate | Certificate of completion |
Course overview
Organisations collect transactional, operational, customer and digital data, yet many analysis teams still spend too much time producing retrospective dashboards and too little time estimating what will happen next. This course addresses the practical gap between descriptive reporting and decision-ready predictive analysis: selecting a business question, preparing defensible data, building models, testing their reliability and communicating recommendations that operational teams can act on. It is designed for professionals who need to move beyond spreadsheet analysis or basic Python notebooks without treating machine learning as a black box.
Participants work through an end-to-end predictive modelling workflow using Python, Jupyter Notebook, pandas and scikit-learn. They profile and clean data, engineer meaningful features, select suitable regression and classification approaches, handle imbalanced classes, validate models with cross-validation and assess performance using metrics such as RMSE, precision, recall, F1 score, ROC-AUC and lift. The course also covers model interpretation, bias checks, data leakage controls, reproducible analysis and the practical choices involved in deploying model outputs into business processes.
Teaching combines instructor-led technical demonstrations with guided notebook labs and a realistic business case. Participants build a predictive model from a raw dataset, document assumptions and validation decisions, and present a recommendation for how the model should be used, monitored and governed. They leave with an annotated Jupyter Notebook, a model evaluation report, a feature and data-quality log, and an action plan for applying the workflow to a live organisational use case.
The course is best suited to analysts, data professionals and technically minded managers who already work with data and need a rigorous, business-focused route into applied predictive modelling.
Course objectives
By the end of this course, participants will be able to:
- Frame a predictive analytics problem as a measurable business objective, target variable and decision rule
- Profile, cleanse and join analytical datasets using pandas DataFrames
- Engineer, encode and scale features while controlling for missing data and data leakage
- Build and tune regression and classification models using scikit-learn pipelines
- Evaluate model performance with cross-validation, confusion matrices, ROC-AUC, lift and error metrics
- Diagnose overfitting, class imbalance, multicollinearity and unstable model behaviour
- Interpret model outputs using feature importance, partial dependence and scenario-based explanations
- Produce a reproducible predictive modelling notebook and model evaluation report for stakeholders
Benefits of attending
For you
- Build a portfolio-ready predictive modelling notebook that demonstrates applied Python and scikit-learn capability
- Gain confidence selecting evaluation metrics that match business costs rather than relying on model accuracy alone
- Strengthen credibility when challenging unsupported model claims, leakage risks or weak validation practices
- Develop the ability to explain model predictions and limitations to non-technical decision-makers
- Prepare to lead or contribute to forecasting, churn, propensity, risk-scoring and demand-planning initiatives
For your organisation
- Improve prioritisation decisions by converting historical data into tested probability scores and forecasts
- Reduce wasted analytical effort through repeatable data preparation, validation and documentation practices
- Lower model-risk exposure by identifying leakage, bias, overfitting and unsuitable performance measures before use
- Create clearer handover artefacts through reproducible notebooks, feature logs and model evaluation reports
- Enable managers to judge whether predictive models are reliable enough for operational workflows and monitoring
Target competencies
Who should attend
- Data Analysts — who need to progress from descriptive reporting to validated predictive models
- Business Intelligence Analysts — who need to add forecasting and classification evidence to dashboard-led decision support
- Data Scientists — who require a structured, business-grounded approach to model validation and interpretation
- Analytics Managers — who oversee analytical delivery and must assess whether models are fit for operational use
- Digital Product Analysts — who need to predict churn, conversion, demand or customer behaviour from product data
- Risk and Operations Analysts — who need defensible scoring models for prioritisation, intervention or resource allocation
Requirements and prerequisites
Participants should be comfortable working with structured data and interpreting common business metrics. Prior experience writing basic Python is required, including variables, functions, lists, conditional logic and reading CSV files; familiarity with pandas DataFrames and Jupyter Notebook is strongly recommended. Participants should also understand descriptive statistics, correlation, distributions and the distinction between a target variable and input variables. Bring a laptop able to run a current Python environment or cloud notebook. Prior machine learning experience, advanced calculus, formal statistical proofs and production software engineering experience are not required.
Training methodology
The five days alternate short instructor-led explanations with progressively more demanding Jupyter Notebook labs. Participants inspect a realistic customer or operations dataset, formulate a prediction question, prepare features, build competing models and compare their results using appropriate validation methods. Instructor demonstrations show the rationale behind each Python and scikit-learn workflow before participants apply it in pairs or individually. Group review sessions focus on model assumptions, business costs and stakeholder communication. The final workshop requires each participant to document and present a deployable modelling recommendation for their chosen case.
Course outline
Day 1: Framing predictive analytics and preparing data
- Translating business decisions into prediction targets and success criteria
- Analytical unit of analysis and time-window design
- Data profiling with pandas DataFrames
- Missing-value patterns and treatment strategies
- Outlier detection using distributions and robust statistics
- Data joins, duplicate records and entity-resolution checks
- Training, validation and test-set design
Workshop: Participants profile a raw customer dataset and produce a data-quality log, target definition and initial modelling dataset.
Day 2: Feature engineering and baseline models
- Feature engineering from dates, categories and transactional history
- Categorical encoding and numerical feature scaling
- Feature selection using business logic and statistical evidence
- Data leakage detection in historical datasets
- scikit-learn Pipeline and ColumnTransformer construction
- Linear regression and logistic regression baselines
- Regularisation with Ridge and Lasso models
Workshop: Participants build a reusable preprocessing pipeline and a baseline regression or classification model with documented feature choices.
Day 3: Classification, forecasting and model tuning
- Decision trees and random forest modelling
- Gradient boosting model principles
- Hyperparameter tuning with GridSearchCV and RandomizedSearchCV
- K-fold and time-series cross-validation
- Class imbalance treatment with weighting and resampling
- Probability calibration and decision thresholds
- Regression forecasting metrics including MAE, RMSE and MAPE
Workshop: Participants train and tune competing models, then produce a comparison table showing validation results and selected parameters.
Day 4: Evaluation, interpretation and model risk
- Confusion matrices, precision, recall and F1 score
- ROC curves, precision-recall curves and ROC-AUC
- Lift charts and gains analysis for prioritisation decisions
- Residual analysis and error segmentation
- Feature importance and permutation importance
- Partial dependence plots and scenario-based explanations
- Bias, fairness and model governance checks
Workshop: Participants complete a model review pack containing performance charts, error analysis, interpretation findings and model-risk controls.
Day 5: Operationalising predictive insight
- Selecting a model against business cost and benefit criteria
- Model documentation using assumptions, limitations and lineage
- Reproducible notebooks and version-controlled analytical workflows
- Batch scoring and integrating predictions into business processes
- Data drift, performance drift and monitoring thresholds
- Communicating uncertainty to executive stakeholders
- Predictive analytics application roadmap and governance ownership
Workshop: Participants present an end-to-end predictive modelling recommendation and produce an application plan for a workplace use case.
Tools & standards covered
Python, Jupyter Notebook, pandas, scikit-learn
A typical training day
| 08:30 – 10:30 | First session |
| 10:30 – 10:45 | Refreshment break |
| 10:45 – 12:30 | Second session |
| 12:30 – 13:30 | Lunch and networking |
| 13:30 – 15:00 | Third session |
| 15:00 – 15:15 | Refreshment break |
| 15:15 – 16:30 | Workshop and daily review |
Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.
What the fee includes
- Instruction by a practitioner facilitator
- Full course workbook and materials
- Exercise files, templates and case studies
- Certificate of completion
- Refreshments and lunch (classroom deliveries)
- Post-course application plan
- Facilitator follow-up on request
- Group rates from five participants
How you can take this course
Classroom
Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.
Live online
The same facilitator and materials, delivered live for distributed teams and individuals.
In-house
Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.
Certification
Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.
Frequently asked questions
Upcoming sessions
-
21 – 25 Sep 2026Book
Cape Town · USD 4,200 -
21 – 25 Sep 2026Book
Nairobi · USD 3,000 -
21 – 25 Sep 2026Book
Kigali · USD 3,500 -
28 Sep – 02 Oct 2026Book
Live Online · USD 1,500 -
05 – 09 Oct 2026Book
Live Online · USD 1,500 -
12 – 16 Oct 2026Book
Nairobi · USD 3,000 -
19 – 23 Oct 2026Book
Dubai · USD 4,500 -
19 – 23 Oct 2026Book
Live Online · USD 1,500
49 more dates — ask us.
Group of 5+?
Request in-house delivery or group rates →Related courses in Data Analytics
dbt Analytics Engineering and Data Quality Testing Training Course
Analytics teams often inherit SQL transformations that run without ownership, documentation or reliable checks. A dashboard can look credibl…
Healthcare Data Analytics for Quality and Patient Outcomes Training Course
Healthcare quality teams often hold fragmented data across electronic health records, claims, incident systems, patient surveys and operatio…
SAP Analytics Cloud Planning and Data Analysis Training Course
Finance, sales and operational teams often work from disconnected spreadsheets, static reports and planning cycles that cannot explain how a…
Alteryx Data Preparation and Workflow Analytics Training Course
Operational data is often spread across spreadsheets, CRM exports, finance systems, databases and shared folders, leaving analysts to repeat…