CRISP-DM Data Science Project Lifecycle Training Course
| Course code | SD-DS-007 |
|---|---|
| Duration | 5 days |
| Level | Intermediate |
| Category | Data Science |
| Delivery | Classroom or live online |
| Language | English |
| Certificate | Certificate of completion |
Course overview
Data science projects often stall after an impressive prototype because the team has not agreed the business question, defined a usable target variable, documented data limitations, or set acceptance criteria for deployment. CRISP-DM provides a repeatable lifecycle for moving from a decision problem to a governed analytical solution. This course helps professionals use the method to prevent scope drift, connect modelling work to measurable business value, and make project decisions traceable to stakeholders, data owners, and delivery teams.
Participants work through all six CRISP-DM phases: Business Understanding, Data Understanding, Data Preparation, Modelling, Evaluation, and Deployment. They learn to frame analytical objectives, create success criteria, profile data quality, select and justify features, establish baselines, compare models, interpret evaluation metrics, identify operational risks, and plan monitoring. The course also addresses the iteration points that distinguish real projects from linear textbook workflows, including revisiting business objectives when data evidence changes.
Instructor-led explanations are paired with a running case study, structured templates, data-analysis exercises in JupyterLab, and facilitated project reviews. Teams develop a CRISP-DM project pack containing a business objective statement, data inventory, preparation plan, experiment log, model evaluation record, deployment plan, and monitoring measures. Participants leave with a practical set of artefacts they can adapt for an active data science initiative and a completion certificate.
The course is designed for data professionals and project stakeholders who already understand the basic language of data, analytics, and machine learning but need a disciplined method for delivering data science work in a business setting.
Course objectives
By the end of this course, participants will be able to:
- Frame a business problem as a measurable CRISP-DM analytical objective with success criteria and constraints
- Produce a data understanding report covering source lineage, data quality, distributions, gaps, and risks
- Design a data preparation plan that documents cleaning, joining, feature engineering, and leakage controls
- Build and record a reproducible modelling experiment using baselines, train-test splits, and versioned assumptions
- Select evaluation metrics that connect model performance to business cost, benefit, and decision thresholds
- Conduct a CRISP-DM evaluation review that tests technical validity, business fitness, fairness, and operational readiness
- Create a deployment and monitoring plan covering ownership, refresh cycles, drift indicators, and retraining triggers
- Compile a CRISP-DM project pack that supports stakeholder approval, handover, and future auditability
Benefits of attending
For you
- Gain a repeatable way to lead or contribute to data science projects beyond isolated model-building tasks
- Build confidence in challenging vague business requests with measurable objectives, assumptions, and acceptance criteria
- Create portfolio-ready project artefacts that demonstrate structured data science delivery capability
- Improve credibility with business sponsors by explaining model results in terms of decisions, thresholds, costs, and benefits
- Prepare to take responsibility for analytics delivery roles that require lifecycle governance and cross-functional coordination
For your organisation
- Reduce investment in low-value modelling by requiring business objectives and measurable success criteria before experimentation
- Improve project predictability through consistent data assessment, preparation plans, experiment records, and stage-gate reviews
- Lower deployment risk by identifying data lineage, leakage, fairness, operational ownership, and monitoring requirements early
- Enable clearer investment decisions with evaluation evidence linked to business impact rather than accuracy metrics alone
- Create reusable CRISP-DM templates and a shared delivery vocabulary across analytics, technology, and business teams
Target competencies
Who should attend
- Data Scientists — who need a disciplined framework for converting exploratory work into deployable, business-approved solutions
- Data Analysts — who support analytical projects and need to define data, metrics, and decision requirements rigorously
- Machine Learning Engineers — who must translate model outputs into reliable deployment, monitoring, and retraining plans
- Analytics Managers — who govern portfolios of data science initiatives and need consistent stage gates and evidence
- Data Product Managers — who must align user needs, business value, data constraints, and release decisions
- Project Managers — who coordinate data science delivery and need realistic lifecycle artefacts, risks, and acceptance criteria
Requirements and prerequisites
Participants should be comfortable working with tabular data and understand basic analytical concepts such as variables, data types, descriptive statistics, training data, test data, and model performance metrics. Experience using spreadsheets, SQL, Python, or a business intelligence tool is helpful; the practical exercises use JupyterLab and Python notebooks, so attendees should be able to follow simple code cells and inspect outputs. Prior exposure to a machine learning project is beneficial but not essential. Advanced programming, calculus, cloud engineering, MLOps platforms, and prior knowledge of CRISP-DM are not required.
Training methodology
The five-day programme uses short instructor-led sessions to introduce each CRISP-DM phase, followed by guided work on a realistic predictive analytics case. Participants profile data in JupyterLab, complete lifecycle templates, review model evidence, and make documented go/no-go decisions in small groups. Facilitated discussions examine trade-offs such as accuracy versus business cost, incomplete data versus scope revision, and prototype performance versus deployment readiness. On the final day, each participant adapts the method to a current or proposed workplace initiative and receives feedback on their project pack.
Course outline
Day 1: Business Understanding and CRISP-DM initiation
- The six CRISP-DM phases and iterative project loops
- Business problem statements versus analytical problem statements
- Stakeholder mapping and decision-owner identification
- Analytical objectives, hypotheses, and target-variable definition
- Business success criteria and technical success criteria
- Project constraints, assumptions, dependencies, and risks
- CRISP-DM project charter and stage-gate structure
Workshop: Participants convert a sponsor brief into a CRISP-DM project charter with objectives, stakeholders, constraints, and measurable success criteria.
Day 2: Data Understanding and preparation design
- Data source inventory, ownership, and lineage mapping
- Data profiling with JupyterLab and pandas
- Distribution analysis, missingness, duplicates, and outliers
- Data quality dimensions and fitness-for-purpose assessment
- Exploratory analysis for target leakage and sampling bias
- Data dictionaries and feature-level documentation
- Preparation plans for cleaning, joins, transformations, and splits
Workshop: Participants profile a case-study dataset and produce a data understanding report with quality findings, risks, and a prioritised preparation plan.
Day 3: Modelling and experiment management
- Baseline models and the value of simple benchmarks
- Feature engineering hypotheses and transformation choices
- Training, validation, and test-set design
- Cross-validation and reproducible random-state controls
- Model selection criteria and hyperparameter search boundaries
- Experiment logs, assumptions, and result traceability
- Git-based version control for notebooks and project artefacts
Workshop: Participants build a baseline and candidate model in JupyterLab, then document the comparison in a structured experiment log.
Day 4: Evaluation, approval, and responsible use
- Classification and regression metric selection
- Confusion matrices, threshold setting, and error trade-offs
- Linking model performance to cost, benefit, and intervention capacity
- Business evaluation against original success criteria
- Bias, fairness, explainability, and data protection considerations
- Model limitations, failure modes, and decision boundaries
- Evaluation review meetings and go/no-go recommendations
Workshop: Teams conduct an evaluation review, calculate threshold trade-offs, and present a justified go/no-go recommendation to a sponsor panel.
Day 5: Deployment, monitoring, and workplace application
- Deployment patterns for batch scoring, APIs, dashboards, and decision support
- Production data contracts and input validation controls
- Model monitoring for performance, drift, and data-quality change
- Retraining triggers, review cycles, and model retirement criteria
- Operational ownership, support procedures, and escalation routes
- MLflow tracking concepts for model and experiment traceability
- CRISP-DM closure, lessons learned, and next-project planning
Workshop: Participants complete a deployment and monitoring plan, then assemble and peer-review their full CRISP-DM project pack for a workplace use case.
Tools & standards covered
CRISP-DM, JupyterLab, Git, MLflow
A typical training day
| 08:30 – 10:30 | First session |
| 10:30 – 10:45 | Refreshment break |
| 10:45 – 12:30 | Second session |
| 12:30 – 13:30 | Lunch and networking |
| 13:30 – 15:00 | Third session |
| 15:00 – 15:15 | Refreshment break |
| 15:15 – 16:30 | Workshop and daily review |
Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.
What the fee includes
- Instruction by a practitioner facilitator
- Full course workbook and materials
- Exercise files, templates and case studies
- Certificate of completion
- Refreshments and lunch (classroom deliveries)
- Post-course application plan
- Facilitator follow-up on request
- Group rates from five participants
How you can take this course
Classroom
Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.
Live online
The same facilitator and materials, delivered live for distributed teams and individuals.
In-house
Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.
Certification
Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.
Frequently asked questions
Upcoming sessions
-
28 Sep – 02 Oct 2026Book
Nairobi · USD 3,000 -
28 Sep – 02 Oct 2026Book
Kigali · USD 3,500 -
12 – 16 Oct 2026Book
Mombasa · USD 3,200 -
19 – 23 Oct 2026Book
Cape Town · USD 4,200 -
26 – 30 Oct 2026Book
Live Online · USD 1,500 -
26 – 30 Oct 2026Book
Mombasa · USD 3,200 -
09 – 13 Nov 2026Book
Live Online · USD 1,500 -
09 – 13 Nov 2026Book
Cape Town · USD 4,200
49 more dates — ask us.
Group of 5+?
Request in-house delivery or group rates →Related courses in Data Science
Public Sector Data Science and Policy Analytics Training Course
Public-sector teams hold large volumes of administrative, service, financial and operational data, yet many policy questions remain answered…
Retail Data Science and Demand Forecasting Training Course
Retailers hold transaction, promotion, product, store, inventory and digital-channel data, yet many planning teams still forecast with sprea…
Advanced Data Science and Machine Learning Training Course
Many data science teams can build a model that performs well in a notebook but struggle to demonstrate that it will make reliable, commercia…
Advanced Deep Learning for Data Science Training Course
Data science teams are increasingly asked to build models for images, text, time series and recommendation problems where tabular machine-le…