Episode 12: Phase 1: EDA & Data Understanding
Act 1 — The Setup
Every project in this series so far has started the same way: math first, then code.
Linear Regression. Logistic Regression. Neural Networks. Decision Trees. Random Forests. All NumPy. All from scratch. All about understanding why before touching a library.
Tabular ML is different.
This isn't a from-scratch implementation. There's no single algorithm to derive. Instead, it's the closest thing to what ML engineers actually do day-to-day — take a messy real-world dataset, understand it deeply, build a clean pipeline, and squeeze the best signal out of structured data.
The dataset: Kaggle's Telco Customer Churn. 7,043 customers. 19 features. One question — who's about to leave?
Before I built anything, I let the data talk.
Act 2 — What the Data Said
The first thing I checked was the target distribution. ~73.5% of customers didn't churn. ~26.5% did.
That number matters more than it looks. On an imbalanced dataset, a model that predicts "No churn" for everyone gets 73.5% accuracy and learns absolutely nothing. Accuracy is a liar here. ROC-AUC and Recall are the metrics that actually matter.
Then came the first real gotcha.
TotalCharges — a column that should be numeric — was loading as an object dtype. Not because of bad data. Because 11 rows had whitespace strings instead of nulls. Brand new customers. Never billed. The fix is one line:
python
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
But the lesson is bigger than the fix: always audit your dtypes before trusting your data.
The insight that hit hardest came from churn rates by contract type. Month-to-month customers churn at ~42%. Two-year contract customers? ~3%. Same product. Same company. The contract structure alone creates a 39-point gap in retention.
A few more patterns that stood out:
Customers with no OnlineSecurity and no TechSupport churn significantly more — they feel unsupported, not just unsubscribed
Fiber optic internet has higher churn than DSL, despite being the premium tier — likely a pricing sensitivity signal
Electronic check users churn more than any other payment method — a proxy for lower engagement and trust
Churned customers have dramatically lower tenure — they leave early or not at all
The three numerical features told a clean story too. Tenure is bimodal — lots of brand new customers and lots of very loyal ones, with a gap in the middle. Monthly charges for churned customers skew higher. Total charges skew lower — because they didn't stay long enough for them to accumulate.
Act 3 — Why This Matters
EDA is the phase most people rush. You download the dataset, run a correlation matrix, and jump to modeling.
But EDA is where your feature engineering instincts come from. It's where you decide what to encode, what to scale, what to flag. It's where you catch the TotalCharges bug before it silently corrupts your pipeline.
Every insight from today becomes a decision in Phase 2.
The class imbalance → SMOTE and class_weight='balanced' The contract type signal → a feature that will carry weight in every model The dtype bug → a preprocessing step that goes into the pipeline before anything else
The data already knows what's important. Phase 1 is just learning to listen.
Next up: building the sklearn Pipeline with ColumnTransformer — the engineering pattern that keeps preprocessing clean and leakage-free.
Six projects in. One final chapter to go.

