Public Sector Data Science and Policy Analytics Training Course
| Course code | SD-DS-006 |
|---|---|
| Duration | 5 days |
| Level | Foundation to Intermediate |
| Category | Data Science |
| Delivery | Classroom or live online |
| Language | English |
| Certificate | Certificate of completion |
Course overview
Public-sector teams hold large volumes of administrative, service, financial and operational data, yet many policy questions remain answered through static reports, fragmented spreadsheets or assumptions that cannot be tested. Analysts and policy professionals need to turn data into evidence while accounting for data quality, privacy, fairness, transparency and the practical constraints of public decision-making. This course addresses the gap between technical analysis and policy use: how to frame a defensible question, prepare government data, select an appropriate analytical method and communicate findings that senior officials can act on.
Participants build a practical workflow for policy analytics using SQL, Python, Power BI and OpenRefine. They learn to profile and clean administrative data; join datasets; define policy-relevant measures; conduct exploratory analysis; build reproducible descriptive models; identify trends, geographic variation and service disparities; and distinguish correlation from causal claims. The course also covers data governance, disclosure control, algorithmic fairness, uncertainty and visual communication for ministerial briefings, performance reviews and service-improvement decisions.
Instruction combines expert demonstrations with guided lab work based on realistic public-sector scenarios, including service demand, programme uptake, case-processing performance and outcome monitoring. Participants work through a policy analytics case from problem statement to briefing-ready evidence. They leave with a documented analysis pack containing a data-quality assessment, reproducible analysis workflow, dashboard prototype, findings summary and a policy recommendation with stated assumptions, limitations and next-step questions.
The course is designed for professionals who need to use data more rigorously in public policy, operations, performance management or evaluation. It suits both aspiring analysts and experienced policy staff who want enough technical capability to commission, assess and explain data science work credibly.
Course objectives
By the end of this course, participants will be able to:
- Frame a policy question as a measurable analytical problem with defined population, outcome, decision owner and success criteria
- Profile administrative datasets to identify missing values, duplicates, invalid codes, outliers and join risks
- Clean and reshape public-sector data using repeatable workflows in Python, SQL and OpenRefine
- Construct policy measures, service-performance indicators and equity breakdowns from raw operational records
- Apply exploratory data analysis to identify trends, geographic variation, segmentation patterns and anomalous service outcomes
- Evaluate correlations, confounding risks and causal-claim limits before presenting findings as policy evidence
- Build an interactive Power BI dashboard with filters, measures and drill-down views for a public-service audience
- Produce a policy analytics briefing pack documenting methods, findings, uncertainty, governance considerations and recommended action
Benefits of attending
For you
- Gain a repeatable method for moving from a policy question to evidence, recommendation and documented limitations
- Build confidence using Python and SQL to inspect and analyse administrative data rather than relying solely on spreadsheets
- Learn to challenge weak causal claims, misleading comparisons and unsupported performance narratives
- Develop briefing and dashboard artefacts that demonstrate practical analytical capability to senior stakeholders
- Strengthen credibility for roles in policy analysis, service improvement, performance management and public-sector digital teams
For your organisation
- Improve the quality and consistency of evidence used in policy development, operational reviews and business cases
- Reduce decision risk by making data-quality issues, uncertainty and causal limitations visible before recommendations are approved
- Create more reusable analytical workflows for recurring service-demand, performance and equity reporting
- Enable managers to identify geographic, demographic and process disparities that may require targeted intervention
- Increase internal capability to specify, commission and quality-assure analytics work from specialist teams or suppliers
Target competencies
Who should attend
- Policy Analysts — who need to convert evidence into defensible options and recommendations
- Public Sector Data Analysts — who prepare administrative data and produce insight for service and policy teams
- Performance and Planning Managers — who monitor outcomes, demand, targets and operational delivery
- Programme and Service Managers — who need evidence to improve uptake, access, timeliness and service quality
- Monitoring and Evaluation Officers — who assess programme implementation and interpret outcome data carefully
- Digital, Transformation and PMO Professionals — who support data-enabled service redesign and benefits tracking
Requirements and prerequisites
This is a foundation-to-intermediate course. Participants should be comfortable working with spreadsheets, tables, percentages and basic charts, and should understand how to interpret simple business or policy metrics such as counts, rates and averages. Prior use of Excel or another reporting tool is helpful. No prior programming, SQL, statistics degree, data science role or Power BI experience is required; Python and SQL are introduced through guided exercises. Complete beginners should expect to work carefully through structured datasets and to practise basic code and queries rather than build advanced machine-learning models.
Training methodology
The course uses short instructor-led explanations followed by hands-on analysis of realistic public-sector datasets. Participants clean a service dataset in OpenRefine, query linked records in PostgreSQL, analyse patterns in Python and build a decision-focused view in Power BI. Case discussions examine how data quality, protected characteristics, confidentiality and causal uncertainty affect policy choices. Small groups critique findings and briefing language, then complete an end-of-course application plan that identifies a suitable workplace dataset, decision question, stakeholders, controls and first analytical steps.
Course outline
Day 1: Policy questions, public data and analytical foundations
- Policy analytics lifecycle from question to decision
- Translating policy objectives into measurable outcomes
- Units of analysis, populations and comparison groups
- Administrative data, survey data and operational data distinctions
- Data dictionaries, metadata and record-level provenance
- Public-sector data governance, privacy and disclosure risks
- Descriptive statistics and rate-based policy measures
Workshop: Participants turn a service-improvement scenario into an analytical problem statement, metric specification and data requirements checklist.
Day 2: Data preparation and quality assurance
- Data profiling for completeness, validity and consistency
- Missing-data patterns and treatment decisions
- Duplicate detection and entity-resolution risks
- Standardising categories, dates and geographic codes in OpenRefine
- Outlier checks and plausible-value rules
- Joining service, demographic and geographic datasets with SQL
- Documenting cleaning decisions in an auditable data-quality log
Workshop: Participants clean and join a simulated benefits-service dataset, producing a data-quality report and analysis-ready table.
Day 3: Exploratory analysis and policy insight
- Python notebooks and pandas data-analysis workflow
- Grouping, aggregation and cross-tabulation of service records
- Trend analysis using monthly and quarterly time series
- Segmentation by geography, service channel and user group
- Rates, denominators and the danger of misleading counts
- Distribution analysis and operational bottleneck detection
- Correlation analysis, confounding and non-causal interpretation
Workshop: Participants analyse variation in application processing times and produce an evidence table identifying priority service segments.
Day 4: Visualisation, dashboards and responsible communication
- Selecting charts for policy trends, comparisons and distributions
- Power BI data model relationships and calculated measures
- Dashboard filters, drill-through and audience-specific views
- Geographic analysis and area-level comparison cautions
- Confidence, uncertainty and small-number suppression
- Fairness checks across protected and underserved groups
- Writing findings, caveats and recommendations for senior officials
Workshop: Participants build a Power BI dashboard and draft a one-page briefing that explains a service disparity without overstating the evidence.
Day 5: From analysis to policy action
- Interpreting evidence for policy options and service interventions
- Logic models, assumptions and measurable implementation signals
- Prioritising actions using impact, feasibility and evidence strength
- Reproducible analysis folders, versioning and documentation
- Quality assurance checks before publishing findings
- Communicating analytical risk to non-technical decision-makers
- Workplace analytics roadmap and stakeholder engagement plan
Workshop: Participants complete and present a policy analytics pack containing a dashboard, findings summary, recommendation, limitations register and 90-day application plan.
Tools & standards covered
Python, PostgreSQL, Microsoft Power BI, OpenRefine
A typical training day
| 08:30 – 10:30 | First session |
| 10:30 – 10:45 | Refreshment break |
| 10:45 – 12:30 | Second session |
| 12:30 – 13:30 | Lunch and networking |
| 13:30 – 15:00 | Third session |
| 15:00 – 15:15 | Refreshment break |
| 15:15 – 16:30 | Workshop and daily review |
Live online deliveries follow the same structure in the East Africa Time zone, with shorter screen blocks and longer breaks.
What the fee includes
- Instruction by a practitioner facilitator
- Full course workbook and materials
- Exercise files, templates and case studies
- Certificate of completion
- Refreshments and lunch (classroom deliveries)
- Post-course application plan
- Facilitator follow-up on request
- Group rates from five participants
How you can take this course
Classroom
Scheduled sessions in Nairobi, Mombasa, Kigali, Dar es Salaam, Dubai and Cape Town.
Live online
The same facilitator and materials, delivered live for distributed teams and individuals.
In-house
Delivered privately for your team, at your offices or a venue of your choice, tailored to your context. Request a proposal.
Certification
Participants who complete the full five days receive the Skillset Development Certificate of Completion, stating the course title, course code, dates and delivery format — suitable for professional-development records and employer reimbursement.
Frequently asked questions
Upcoming sessions
-
28 Sep – 02 Oct 2026Book
Nairobi · USD 3,000 -
28 Sep – 02 Oct 2026Book
Mombasa · USD 3,200 -
05 – 09 Oct 2026Book
Live Online · USD 1,500 -
05 – 09 Oct 2026Book
Dubai · USD 4,500 -
12 – 16 Oct 2026Book
Live Online · USD 1,500 -
12 – 16 Oct 2026Book
Kigali · USD 3,500 -
19 – 23 Oct 2026Book
Dar es Salaam · USD 3,500 -
26 – 30 Oct 2026Book
Nairobi · USD 3,000
49 more dates — ask us.
Group of 5+?
Request in-house delivery or group rates →Related courses in Data Science
Data Science Foundations and Exploratory Analysis Training Course
Teams increasingly hold customer, operational, financial and digital-service data, yet many analysts and subject-matter professionals strugg…
Apache Spark Data Science for Large Scale Analytics Training Course
Data science teams often prove a model or analytical method on a sampled dataset, then struggle to run the same work reliably across billion…
Advanced Deep Learning for Data Science Training Course
Data science teams are increasingly asked to build models for images, text, time series and recommendation problems where tabular machine-le…
Advanced Data Science and Machine Learning Training Course
Many data science teams can build a model that performs well in a notebook but struggle to demonstrate that it will make reliable, commercia…