Skip to main content

Learn More: Data Science Lab

Learn more about our skill-based learning focused on applied data science and AI.

Written by Kate Porter

Whether you're launching a career in data science, developing expertise in deep learning and AI, or preparing for your next opportunity, WQU Labs provide a flexible way to build practical, in-demand technical skills through hands-on projects.

These self-paced, project-based programs are designed as applied credentials, giving learners the opportunity to develop and demonstrate their skills through authentic project work using real-world datasets. Working in cloud-based virtual machines, learners tackle the kinds of challenges faced by data scientists and AI practitioners across industries - entirely free of cost.

Self-Paced (Est. 10-16 weeks)

10-15 Hours Per Week

Data Science Lab:

The Data Science Lab takes you through eight real-world projects across two units, from foundational data wrangling and modeling to applied machine learning and deployment. You'll work with messy, real datasets to build predictive models, forecast trends, design experiments, and ship production-ready code, gaining the practical, end-to-end experience employers look for in a data scientist.

  • Prerequisites:

    Beginner-level Python skills

    Familiarity with basic statistics

    Familiarity with basic linear algebra

Application Requirements

  • Above Suggested Prerequisites

  • Passing score on Admissions Assessment (66% or higher)

    • Before you can start the Data Science Lab, you need to take a short Admissions Assessment. We want to make sure you have a solid foundation on which you can build the skills we teach in our program.

      We ask that you not use supplementary materials, so close your textbooks and any other browser windows. We recommend that you have a pencil and paper ready in case you want to write out a problem.

      To prepare for the Assessment, we recommend the following resources:

    • If you fail the Admissions Assessment

      • you have a second (and final) attempt after a 7-day waiting period.

      • Applicants who do not pass the test on their 2nd attempt are able to reapply to the Lab following a waiting period of 6 months from the date of their 2nd attempt.

      • Important Warning: Creating multiple accounts to attempt the Admissions Assessment is a violation of the University’s Academic Integrity policy. Any user identified for doing so will be immediately terminated and will not have the opportunity to be considered for enrollment in the Lab program.

Next Steps After Passing the Admissions Assessment

  1. Complete your Student Profile;

  2. Sign the User Agreement;

  3. Take the mandatory Orientation Course;

  4. Register for Project 1.

Once you have completed your profile and signed the User Agreement, you will be automatically enrolled in a mandatory Orientation Course, which can be accessed via My Courses in the top navigation of the WQU Learning Platform.

From walking you through the Lab’s structure to helping you navigate our Learning Platform and virtual machines, the Orientation Course takes about 1 hour to complete and covers everything you need to know to set you up for success.

Upon successful completion of the Orientation Course, course registration will NOT occur automatically. Instead, you are responsible to register for each of your Projects by navigating to My Courses and "Register".

The Data Science Lab is organized into two sequential units of four projects each. Every project is grounded in a real-world dataset from around the world and guides you through the full analytical workflow: from raw data to deployed, interpretable results.

Unit 1 covers the foundational data science workflow, taking you from wrangling and exploring data to building your first regression, classification, and time series models.
​

Unit 2 scales up to applied machine learning, tackling ensemble methods, unsupervised learning, experimentation, and real deployment workflows.

You’ll gain practical experience in:

  • Cleaning and analyzing the kind of messy data you actually encounter on the job, not tidy classroom examples

  • Building predictive models that forecast housing prices, detect bankruptcy risk, and classify earthquake damage

  • Working with time-based data to spot trends and make forecasts

  • Designing experiments to test whether interventions actually work, and communicating results to non-technical audiences

  • Delivering a complete, deployable data science project from raw data all the way to production-ready code

Credentials Earned Upon Successful Completion

  • Verified Digital Badge and Certificate

Project Descriptions

The Data Science Lab comprises two units, each with four end-to-end projects.

Each successful project completion unlocks the registration for the next.

UNIT 1

Hands-on Data Science in the Mexican Real Estate Market

Project 1 • self-paced

This project introduces the foundational data science workflow using real estate data from Mexico City. Students load, clean, and visualize property listings, then quantify relationships among variables using correlation analysis. The central question — whether price is driven more by size or location — runs through all four notebooks and prepares students for the modeling work in Project 2.

Predicting Housing Prices in Buenos Aires with Linear Regression

Project 2 • self-paced

Project 2 introduces supervised machine learning by predicting housing prices in Buenos Aires. Students prepare data for modeling, build their first linear regression model, and directly confront overfitting — addressing it with Ridge and Lasso regularization. The focus is on understanding generalization, not just fitting numbers.

Time Series Modeling of Air Quality in Nairobi

Project 3 • self-paced

This project moves from independent and identically distributed data to time-indexed data, using Nairobi air quality measurements stored in MongoDB. Students learn to query, prepare, and forecast time series, building from lagged regression to AR and ARMA models. The key conceptual shift: in time series, the past is a feature.

Predicting Earthquake Damage in Nepal with Classification Models

Project 4 • self-paced

Project 4 introduces classification using earthquake-damage data from Nepal stored in SQLite. Students work through the full pipeline — querying relational data, training logistic regression and decision tree models, and evaluating performance with metrics suited for discrete outcomes. It closes Unit 1 with a real humanitarian decision-making context.

UNIT 2

Random Forests and Gradient Boosting for Bankruptcy Prediction in Poland and Taiwan

Project 5 • self-paced

This project opens Unit 2 by scaling up to ensemble methods. Using corporate bankruptcy records from Poland and Taiwan in JSON format, students build Random Forest and Gradient Boosting classifiers to handle severe class imbalance. The emphasis is on moving from a working model to a production-ready prediction pipeline.

Unsupervised Learning for Consumer Finance Segmentation

Project 6 • self-paced

This project introduces unsupervised learning by segmenting consumer finance data. Without a target variable, students explore how clustering algorithms reveal latent structure in financial data. The key idea: clustering results are analytical constructs, not ground truth, and interpretability matters as much as the algorithm.

Experimentation and Decision-Making with WQU Applicant Data

Project 7 • self-paced

Project 7 project shifts from prediction to experimentation. Using applicant data from the WQU Data Science Lab admissions process stored in MongoDB, students build ETL pipelines, test hypotheses with chi-square statistics, and communicate results through interactive dashboards. The question is no longer "what will happen", it's "did this intervention work?”

Financial Volatility Modelling and End-to-End Deployment

Project 8 • self-paced

Project Overview: This culminating project functions as the capstone. Students model financial volatility using GARCH on real market data while applying test-driven development and proper software engineering practices. The goal is not just a working model; rather, it's a reproducible, packaged system that could plausibly be deployed.

Learning Outcomes

  • Explore, Clean, and Visualize Real-World Data

    Explore, clean, and visualize real-world datasets to extract patterns and formulate data-driven hypotheses.

  • Train and Regularize Regression Models

    Build, train, and evaluate supervised regression models, including regularization strategies for controlling model complexity.

  • Analyze and Forecast Time Series Data

    Analyze time-indexed data and apply forecasting techniques using temporal modeling frameworks.

  • Classify Outcomes and Interpret Results

    Implement classification models, evaluate performance on discrete outcomes, and interpret results in applied contexts.

  • Apply Ensemble Methods to Imbalanced Problems

    Apply ensemble methods — Random Forests and Gradient Boosting — to high-stakes classification problems with class imbalance.

  • Segment Customers with Clustering Algorithms

    Perform unsupervised learning and customer segmentation using clustering algorithms, with emphasis on interpretability and feature selection.

  • Design Experiments and Build Interactive Dashboards

    Design and evaluate data-driven experiments using hypothesis testing, ETL pipelines, and interactive dashboards.

  • Deploy Reproducible End-to-End Data Science Systems

    Build reproducible, deployable end-to-end data science systems that integrate modeling, testing, and software engineering best practices.

Did this answer your question?