Machine Learning Fundamentals · Course lab · about 300 minutes · 6 tasks · marked out of 100, pass at 60
The dropout model, trained, evaluated honestly and deployed
The situation
An academy wants to know which students are likely to drop out. You will build the dataset (plausibly, with one leaky feature planted on purpose), train two models, evaluate them the honest way — held-out data, a baseline beaten or not, precision and recall argued — watch the leak inflate the score, then save the model, write the script that uses it, the monitoring plan, and the ethics check. This is the interview answer to "walk me through an ML project".
What you'll be able to show
- Tell rules from ML, and spot the task that is neither
- Build a dataset with documented handling of missing values and a planted leak
- Evaluate on held-out data against a baseline, with precision/recall chosen for the business
- Save, use, monitor and ethically check a model
What you need
- Python with pandas and scikit-learn (or a notebook service)
- A spreadsheet for the 60 invented rows
- joblib for saving the model
Tasks
-
1Rule, ML, or trapSort the five tasks from lesson 1 into rule vs ML with one line each, and name the trap (no examples exist or the pattern is not learnable).A correct result: Five verdicts and the trap named with a reason.
-
2The dataset, with a planted leakBuild 60 plausible student rows: attendance genuinely correlating with completed/dropped labels, city one-hot encoded, one column with missing values handled and documented, and one deliberately LEAKY feature. Write the sentence explaining why the leak must be removed before training.A correct result: A 60-row file, the missing-value decision written down, and the leak identified with its sentence.
-
3Train two, read the treeTrain a decision tree and a logistic regression on the clean features. Predict three invented new students. Plot the tree and read its rules aloud — do they make real-world sense? Rank features by logistic weights. Write one paragraph on which model you would hand the principal and why.A correct result: Two trained models, three predictions, the tree's rules assessed, the paragraph.
-
4Evaluate honestlyProper train/test split; both models scored on held-out data only; the majority-class baseline computed FIRST; precision and recall for the dropout class with one sentence choosing your trade-off and its business cost; the unlimited vs depth-3 tree comparison showing overfitting numerically.A correct result: A results table with baseline, both models, precision/recall, and the overfitting comparison.
-
5Let the leak in, then autopsyRe-add the leaky feature and re-run the evaluation. Watch the score soar. Write the one-line autopsy.A correct result: The inflated score next to the honest one, and the autopsy line.
-
6Ship, monitor, checkjoblib-save the chosen model. Write a 10-line script that loads it and prints the risk for a new student passed as input. Write the monitoring plan in three lines: what fresh data arrives when, what score triggers retraining, who is told. Then the ethics pass: one group your invented data might mistreat and the check you would run.A correct result: The saved model, the script working on one input, the three-line plan, and the ethics check.
What to hand in
The dataset file, the notebook or scripts, the results table, the autopsy line, the saved model with its loader script, the monitoring plan and the ethics check.
How it is marked
| Criterion | Points |
|---|---|
| Rule/ML/trap sorted with reasons | 10 |
| Dataset built with documented missing-value handling and the leak identified | 15 |
| Two models trained, tree rules assessed, model choice argued | 15 |
| Honest evaluation: held-out, baseline first, precision/recall argued, overfitting shown | 30 |
| Leak demonstrated and autopsied | 10 |
| Model saved, loader works, monitoring plan and ethics check written | 20 |
| Total · pass at 60 | 100 |
صورتحال
ایک academy جاننا چاہتی ہے کہ کون سے students کے چھوڑ جانے کا امکان ہے۔ آپ dataset بنائیں گے (قابلِ یقین، ایک leaky feature جان بوجھ کر رکھا ہوا)، دو models train کریں گے، ایماندار طریقے سے پرکھیں گے — held-out data، baseline کو پیچھے چھوڑا یا نہیں، precision اور recall پر دلیل — leak کو score بڑھاتے دیکھیں گے، پھر model save کریں گے، اسے استعمال کرنے والی script، monitoring plan اور ethics check لکھیں گے۔ یہی "walk me through an ML project" کا interview جواب ہے۔
آپ کیا دکھا سکیں گے
- rules کو ML سے الگ کرنا، اور وہ کام پہچاننا جو دونوں میں سے کوئی نہیں
- غائب values کی درج شدہ handling اور رکھے ہوئے leak کے ساتھ dataset بنانا
- baseline کے مقابل held-out data پر پرکھنا، کاروبار کے لیے چنے ہوئے precision/recall کے ساتھ
- model کو save، استعمال، monitor اور اخلاقی طور پر چیک کرنا
آپ کو کیا چاہیے
- pandas اور scikit-learn کے ساتھ Python (یا کوئی notebook service)
- 60 فرضی قطاروں کے لیے ایک spreadsheet
- model save کرنے کے لیے joblib
کام
-
1rule، ML، یا جالسبق 1 کے پانچ کاموں کو ایک ایک line کے ساتھ rule بمقابلہ ML میں بانٹیں، اور جال کا نام لیں (مثالیں موجود نہیں یا pattern سیکھا نہیں جا سکتا)۔درست نتیجہ: پانچ فیصلے اور وجہ کے ساتھ نامزد جال۔
-
2dataset، رکھے ہوئے leak کے ساتھ60 قابلِ یقین student قطاریں بنائیں: attendance کا completed/dropped labels سے حقیقی تعلق، city one-hot encoded، ایک column جس میں غائب values سنبھالی اور درج ہوں، اور ایک جان بوجھ کر LEAKY feature۔ وہ جملہ لکھیں جو بتائے کہ training سے پہلے leak کیوں ہٹانا ضروری ہے۔درست نتیجہ: 60 قطاروں کی file، غائب values کا فیصلہ لکھا ہوا، اور جملے کے ساتھ پہچانا ہوا leak۔
-
3دو train کریں، tree پڑھیںصاف features پر decision tree اور logistic regression train کریں۔ تین فرضی نئے students کی پیش گوئی کریں۔ tree بنا کر اس کے اصول بلند آواز میں پڑھیں — کیا وہ حقیقی دنیا میں معنی رکھتے ہیں؟ logistic weights سے features کی درجہ بندی کریں۔ ایک paragraph لکھیں کہ کون سا model آپ principal کو دیں گے اور کیوں۔درست نتیجہ: دو train شدہ models، تین پیش گوئیاں، tree کے اصول پرکھے ہوئے، paragraph۔
-
4ایمانداری سے پرکھیںصحیح train/test split؛ دونوں models صرف held-out data پر score؛ majority-class baseline پہلے نکالی ہوئی؛ dropout class کے لیے precision اور recall مع ایک جملہ جو آپ کا trade-off اور اس کی کاروباری قیمت چنے؛ unlimited بمقابلہ depth-3 tree کا موازنہ جو overfitting نمبروں میں دکھائے۔درست نتیجہ: baseline، دونوں models، precision/recall اور overfitting کے موازنے والی results table۔
-
5leak اندر آنے دیں، پھر postmortemleaky feature دوبارہ ڈالیں اور evaluation دوبارہ چلائیں۔ score کو اڑتے دیکھیں۔ ایک line کا postmortem لکھیں۔درست نتیجہ: ایماندار score کے ساتھ بڑھا ہوا score، اور postmortem کی line۔
-
6بھیجیں، monitor کریں، چیک کریںچنا ہوا model joblib سے save کریں۔ 10 lines کی script لکھیں جو اسے load کر کے input کے طور پر دیے گئے نئے student کا risk print کرے۔ monitoring plan تین lines میں لکھیں: کون سا تازہ data کب آتا ہے، کون سا score دوبارہ training شروع کرے، کسے بتایا جائے۔ پھر ethics کا جائزہ: ایک گروہ جس کے ساتھ آپ کا فرضی data ناانصافی کر سکتا ہے اور وہ check جو آپ چلائیں گے۔درست نتیجہ: save شدہ model، ایک input پر کام کرتی script، تین lines کا plan، اور ethics check۔
کیا جمع کروانا ہے
dataset کی file، notebook یا scripts، results table، postmortem کی line، loader script کے ساتھ save شدہ model، monitoring plan اور ethics check۔
نمبر کیسے ملیں گے
| معیار | نمبر |
|---|---|
| rule/ML/جال وجوہات کے ساتھ بانٹے گئے | 10 |
| dataset درج شدہ غائب values کی handling کے ساتھ بنا اور leak پہچانا گیا | 15 |
| دو models train ہوئے، tree کے اصول پرکھے گئے، model کا انتخاب دلیل سے | 15 |
| ایماندار evaluation: held-out، baseline پہلے، precision/recall پر دلیل، overfitting دکھایا | 30 |
| leak دکھایا اور postmortem ہوا | 10 |
| model save، loader کام کرتا ہے، monitoring plan اور ethics check لکھے گئے | 20 |
| کل · پاس 60 | 100 |
Hand in your lab
Create a free BvLogic ID to hand in your lab, get it marked, and have it on your certificate.
Create your BvLogic IDAlready have one? Sign in and this course will be added to it.