Using CatBoost for mental health assessment: tradeoffs and lessons
Building an ML model for mental health assessment is technically interesting and ethically complicated. Here's what I learned doing both.
Why CatBoost
The training data had a lot of categorical features — symptom types, frequency descriptors, demographic categories. CatBoost handles categorical features natively without one-hot encoding, which kept the preprocessing pipeline simple and reduced the risk of introducing encoding artifacts.
It also trains faster than XGBoost on this type of data and requires less hyperparameter tuning out of the box. For a solo project with limited compute, that matters.
The imbalanced data problem
Mental health datasets are almost always imbalanced. Severe cases are rare, which means a naive model learns to predict "mild" for everything and still gets high accuracy. That's useless — and potentially harmful.
I used SMOTE (Synthetic Minority Oversampling Technique) to generate synthetic samples for underrepresented classes. Combined with class-weighted loss, this pushed the model to actually learn the minority classes.
Knowing when not to trust your model
The validation metrics looked reasonable. But I'm cautious about what that means in practice.
The training data was limited and came from a specific population. The model hasn't been reviewed by mental health professionals. It hasn't been tested on the actual user population it will serve.
So the assessment is framed as a starting point, not a diagnosis. The UI makes this explicit. The model output is one input into a conversation with a real therapist — not a verdict.
What I'd do differently
Get clinical review earlier. I built the model first and thought about clinical validity second. It should be the other way around — define what "good" looks like with a professional before writing any code.