ArticleFrontiers in medicine2026
Identifying predictors of student depression through validated machine learning pipelines.
Article in Frontiers in medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Depression among student populations has become a growing public health concern, with prevalence rates ranging from 10 to 30 percent across studies. Machine learning methods offer the potential not only to predict depression risk but also to identify which factors most strongly predict depression through feature importance analysis. However, the validity of such rankings depends critically on the quality of the underlying model's learning dynamics; which is a consideration often overlooked when aggregate performance metrics appear favorable. Methods: This study evaluated six baseline classification algorithms (Logistic Regression, Decision Tree, k-Nearest Neighbors, Support Vector Machine, Naive Bayes, and Random Forest) and two RUSBoost pipeline configurations on a student depression dataset comprising 27,901 records. Pipeline A employed fixed hyperparameters with constant model complexity during learning curve construction, while Pipeline B implemented systematic hyperparameter optimization through grid search and scaled model complexity proportionally with training data availability. Learning curves were generated by training models on progressively larger subsets of training data (10% to 100%) to assess whether each algorithm exhibited healthy learning dynamics characterized by monotonically increasing validation accuracy and convergent training-validation gaps. Results: Despite achieving the highest F1-score (87.5%), Logistic Regression and other baseline algorithms exhibited pathological learning dynamics including flat curves indicating no benefit from additional training data, severe overfitting with training-validation gaps exceeding 15 percentage points, and erratic non-monotonic behavior. Pipeline A's RUSBoost implementation showed oscillatory validation accuracy and failed to converge to a stable asymptote. Only Pipeline B demonstrated textbook healthy learning dynamics: training accuracy decreased monotonically from 88.8% to 85.4% while validation accuracy increased monotonically from 82.2% to 83.1%, with progressive gap convergence. Feature importance analysis from the validated Pipeline B model identified history of suicidal thoughts as the dominant predictor (normalized importance: 1.0), followed by academic pressure (0.57), financial stress (0.31), age (0.18), work/study hours (0.13), dietary habits (0.11), and study satisfaction (0.07). Conclusions: This study demonstrates that aggregate performance metrics are insufficient indicators of model reliability for scientific inference. Learning curve diagnostics must precede interpretation of feature importance rankings to ensure conclusions rest on demonstrably healthy learning processes rather than artifacts of pathological training dynamics. The validated model's identification of suicidal ideation history, academic pressure, and financial stress as leading predictors suggests that targeted screening for suicidal thoughts, academic workload management programs, and financial support initiatives may prove most effective for reducing depression prevalence among student populations.
Indexed as
Identifiers
What Socratic holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.