Dataiku Data Analytics and Machine Learning Workflow Training Course
| Course code | SD-DA-042 |
|---|---|
| Duration | 5 days |
| Level | Intermediate to Advanced |
| Category | Data Analytics |
| Delivery | Classroom or live online |
| Language | English |
| Certificate | Certificate of completion |
Course overview
Data teams often have capable analysts, data scientists and engineers working in separate tools, producing models and reports that are difficult to reproduce, govern or deploy. Dataiku provides a shared environment for preparing data, designing visual and code-based workflows, training machine learning models, and operationalising outputs. This course helps participants turn fragmented analysis activities into traceable Dataiku projects with clear datasets, reusable recipes, documented decisions and production-ready scenarios.
Participants work through the full Dataiku workflow: connecting and profiling data, building visual recipes, using SQL and Python where appropriate, creating joined and enriched datasets, and applying quality checks. They then build and compare machine learning models using Dataiku AutoML and custom code, interpret model performance, manage features, and publish scoring outputs. The course also covers Flow design, project variables, reusable components, Git integration, automation scenarios, monitoring and governance practices needed to move work beyond an individual desktop analysis.
Instructor-led demonstrations are followed by guided work in a Dataiku DSS project based on a realistic customer-retention and operational-data case. Participants progressively create a working Flow that ingests data, prepares an analytical dataset, trains and evaluates a model, and automates a scoring process. They leave with a documented Dataiku project blueprint, a reusable workflow pattern, model evaluation evidence and an implementation plan for applying Dataiku practices to a live business use case.
The course is suited to professionals who already analyse data or build models and need to deliver those outputs through a governed, collaborative Dataiku environment. It is equally valuable for managers establishing repeatable analytics delivery standards across analyst, engineering and data science teams.
Course objectives
By the end of this course, participants will be able to:
- Design a traceable Dataiku Flow using managed datasets, recipes, zones and clear dependency structure
- Build data preparation pipelines with visual recipes, SQL recipes, joins, window functions and data-quality checks
- Profile source data and define reusable validation rules for completeness, uniqueness, format and distribution anomalies
- Develop analytical features using Dataiku Prepare, Group, Window, Pivot and Python recipes
- Train, compare and interpret supervised machine learning models with Dataiku AutoML and model diagnostics
- Create reproducible scoring pipelines using saved models, scoring recipes, project variables and scenarios
- Apply Git-based project versioning, reusable components and documentation practices to a Dataiku project
- Produce a deployable Dataiku workflow blueprint including governance controls, ownership and monitoring actions
Benefits of attending
For you
- Build evidence of end-to-end Dataiku delivery capability rather than isolated dashboard or notebook work
- Gain practical confidence choosing visual, SQL and Python recipes for different transformation requirements
- Learn to defend model-selection decisions with Dataiku evaluation metrics, diagnostics and documented assumptions
- Create reusable Flow and automation patterns that support progression into analytics engineering or applied data science roles
- Leave with a project blueprint that demonstrates governed analytics workflow design to managers and stakeholders
For your organisation
- Reduce rework by standardising data preparation, validation and dependency management within Dataiku Flows
- Improve trust in analytical outputs through visible lineage, documented recipes and repeatable model evaluation
- Shorten the path from exploratory analysis to scheduled scoring and managed business outputs
- Lower operational risk by applying project permissions, version control, scenarios and ownership controls consistently
- Enable analysts, engineers and data scientists to collaborate on shared datasets and production workflows
Target competencies
Who should attend
- Data Analysts — who need to turn repeatable analysis into governed Dataiku workflows
- Data Scientists — who need to operationalise models alongside visual data preparation and business users
- Analytics Engineers — who build reliable transformation pipelines and shared analytical datasets
- BI Developers — who need curated, validated data outputs for reporting and self-service analysis
- Data Engineers — who support Dataiku projects, data connections and production automation
- Analytics Managers — who need consistent delivery, review and governance practices across data teams
Requirements and prerequisites
Participants should be comfortable working with tabular data, including filtering, joins, aggregations, calculated fields and basic data-quality checks. Practical experience using SQL is expected, as is familiarity with at least one analytics or reporting tool. Participants should understand the purpose of train/test data, common classification or regression use cases, and basic performance measures such as accuracy, precision, recall or RMSE. Prior Dataiku experience is not required. Python is helpful for the code-recipe sections but is not mandatory; no advanced programming, statistics degree or prior MLOps platform experience is assumed.
Training methodology
The programme alternates concise instructor demonstrations in Dataiku DSS with hands-on build sessions in a shared business case. Participants inspect source data, construct recipes, test data-quality rules, train models and configure automation directly in a Dataiku project. Instructor reviews focus on Flow readability, reproducibility and deployment choices rather than only getting a result. Small-group design discussions compare visual, SQL and Python implementation options. On the final day, each participant converts the completed workflow into an application plan identifying data owners, run schedules, controls and next production steps.
Course outline
Day 1: Dataiku projects, Flows and governed data access
- Dataiku DSS architecture, project structure and user roles
- Connections, managed datasets and external dataset concepts
- Flow canvas navigation, dataset lineage and dependency analysis
- Project zones, dataset naming conventions and Flow readability
- Data profiling with statistics, charts and column meaning detection
- Dataiku sampling, schema management and storage formats
- Project documentation, wiki pages and dataset descriptions
Workshop: Build a Dataiku project shell for a customer-retention case, connect source datasets, profile their quality and produce a documented Flow map.
Day 2: Data preparation and analytical dataset construction
- Prepare recipes for cleansing, parsing and standardising fields
- Join, Stack and Group recipes for combining operational datasets
- Window recipes for rankings, lags, rolling metrics and partitions
- Pivot and Unpivot recipes for reshaping analytical data
- SQL recipes and pushdown execution considerations
- Python recipes for custom transformations and reusable logic
- Data-quality rules, checks and schema drift detection
Workshop: Create a validated customer-level analytical dataset from transaction, service and demographic data, including a documented set of quality checks.
Day 3: Machine learning design, training and evaluation
- Defining prediction targets, prediction types and leakage controls
- Feature selection and feature handling in Dataiku AutoML
- Train-test splits, cross-validation and experiment design
- Classification algorithms available in Dataiku visual machine learning
- Regression workflows and error-based performance measures
- Model comparison using ROC curves, lift charts and confusion matrices
- Feature importance, partial dependence and model interpretation
Workshop: Train and compare churn-prediction models in Dataiku, then produce a model-selection note supported by performance and interpretability evidence.
Day 4: Automation, reuse and production workflow controls
- Saved models, scoring recipes and output dataset design
- Dataiku scenarios, triggers, steps and run conditions
- Project variables and scenario variables for parameterised workflows
- Metrics, checks, alerts and failure-handling patterns
- Reusable recipes, plugins and project templates
- Git integration and project versioning practices
- Project permissions, collaboration roles and governance responsibilities
Workshop: Configure a scheduled scoring scenario with quality gates and alerts, then version the project changes through a Git-connected workflow.
Day 5: Deployment planning, monitoring and business adoption
- Designing batch scoring outputs for downstream reporting and operations
- Deployment approaches with Dataiku Automation and Design nodes
- Model monitoring, data drift and performance review cycles
- Model documentation, approval evidence and audit trails
- Dataiku dashboards and webapps for communicating workflow outputs
- Operational handover, ownership and support runbooks
- Use-case prioritisation and Dataiku adoption roadmap planning
Workshop: Complete and present a Dataiku workflow blueprint covering deployment architecture, monitoring measures, accountable owners and a 90-day implementation plan.
Tools & standards covered
Dataiku DSS, SQL, Python, Git
A typical training day
| 08:30 – 10:30 | First session |
| 10:30 – 10:45 | Refreshment break |
| 10:45 – 12:30 | Second session |
| 12:30 – 13:30 | Lunch and networking |
| 13:30 – 15:00 | Third session |
| 15:00 – 15:15 | Refreshment break |
| 15:15 – 16:30 | Workshop and daily review |
Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.
What the fee includes
- Instruction by a practitioner facilitator
- Full course workbook and materials
- Exercise files, templates and case studies
- Certificate of completion
- Refreshments and lunch (classroom deliveries)
- Post-course application plan
- Facilitator follow-up on request
- Group rates from five participants
How you can take this course
Classroom
Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.
Live online
The same facilitator and materials, delivered live for distributed teams and individuals.
In-house
Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.
Certification
Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.
Frequently asked questions
Upcoming sessions
New dates are being scheduled. Ask us about the next session or an in-house delivery for your team.
Ask about datesGroup of 5+?
Request in-house delivery or group rates →Related courses in Data Analytics
Oil and Gas Data Analytics for Production Performance Training Course
Production teams often hold years of historian, well test, allocation, maintenance, drilling and laboratory data, yet struggle to turn it in…
Apache Spark Data Analytics with PySpark Training Course
Teams often have data spread across transaction systems, log files, APIs and cloud storage, yet struggle to turn large or messy datasets int…
KNIME Analytics Platform for Data Blending Training Course
Data teams often spend more time reconciling spreadsheets, database extracts, CRM exports and operational files than analysing them. Repeate…
Data Analytics Fundamentals for Business Decision Making Training Course
Business teams often have access to operational, customer, financial and project data but struggle to turn it into evidence that supports a …