# Episode 12: 
Phase 1: EDA & Data Understanding

### Act 1 — The Setup

Every project in this series so far has started the same way: math first, then code.

Linear Regression. Logistic Regression. Neural Networks. Decision Trees. Random Forests. All NumPy. All from scratch. All about understanding *why* before touching a library.

Tabular ML is different.

This isn't a from-scratch implementation. There's no single algorithm to derive. Instead, it's the closest thing to what ML engineers actually do day-to-day — take a messy real-world dataset, understand it deeply, build a clean pipeline, and squeeze the best signal out of structured data.

The dataset: Kaggle's Telco Customer Churn. 7,043 customers. 19 features. One question — who's about to leave?

Before I built anything, I let the data talk.

* * *

### Act 2 — What the Data Said

The first thing I checked was the target distribution. ~73.5% of customers didn't churn. ~26.5% did.

That number matters more than it looks. On an imbalanced dataset, a model that predicts "No churn" for everyone gets 73.5% accuracy and learns absolutely nothing. Accuracy is a liar here. ROC-AUC and Recall are the metrics that actually matter.

Then came the first real gotcha.

`TotalCharges` — a column that should be numeric — was loading as an `object` dtype. Not because of bad data. Because 11 rows had whitespace strings instead of nulls. Brand new customers. Never billed. The fix is one line:

python

```python
df['TotalCharges'] = pd.to_numeric(df['TotalCharges'], errors='coerce')
```

But the lesson is bigger than the fix: *always audit your dtypes before trusting your data.*

The insight that hit hardest came from churn rates by contract type. Month-to-month customers churn at ~42%. Two-year contract customers? ~3%. Same product. Same company. The contract structure alone creates a 39-point gap in retention.

A few more patterns that stood out:

*   Customers with no OnlineSecurity and no TechSupport churn significantly more — they feel unsupported, not just unsubscribed
    
*   Fiber optic internet has *higher* churn than DSL, despite being the premium tier — likely a pricing sensitivity signal
    
*   Electronic check users churn more than any other payment method — a proxy for lower engagement and trust
    
*   Churned customers have dramatically lower tenure — they leave early or not at all
    

The three numerical features told a clean story too. Tenure is bimodal — lots of brand new customers and lots of very loyal ones, with a gap in the middle. Monthly charges for churned customers skew higher. Total charges skew lower — because they didn't stay long enough for them to accumulate.

* * *

### Act 3 — Why This Matters

EDA is the phase most people rush. You download the dataset, run a correlation matrix, and jump to modeling.

But EDA is where your feature engineering instincts come from. It's where you decide what to encode, what to scale, what to flag. It's where you catch the `TotalCharges` bug before it silently corrupts your pipeline.

Every insight from today becomes a decision in Phase 2.

The class imbalance → SMOTE and `class_weight='balanced'` The contract type signal → a feature that will carry weight in every model The dtype bug → a preprocessing step that goes into the pipeline before anything else

The data already knows what's important. Phase 1 is just learning to listen.

* * *

*Next up: building the sklearn Pipeline with ColumnTransformer — the engineering pattern that keeps preprocessing clean and leakage-free.*

*Six projects in. One final chapter to go.*
