Evidence map›Paper›PMID 42210272›Full record

ArticleBMC medicine2026

Machine learning and natural language processing for the identification of potential mental disorders among school-age children: a prospective birth cohort study.

Shanquan Chen, Ting Dang, Mengjie Qian, Huizhi Liang, Diribsa Tsegaye Bedada, Quinette Abegail Louw, Anna Moore, Rudolf N Cardinal, Tamsin J Ford, Fan Jiang

Abstract read
In one paragraph

Article in BMC medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Shanquan ChenSchool of Public Health, Li Ka Shing Faculty of Medicine, The University of Hong Kong, Hong Kong, China. shanquan.chen@hku.hk.
Ting DangSchool of Computing and Information Systems, University of Melbourne, Melbourne, VIC, 3010, Australia.
Mengjie QianDepartment of Engineering, University of Cambridge, Cambridge, CB2 1PZ, UK.
Huizhi LiangSchool of Computing, Newcastle University, Newcastle upon Tyne, UK.
Diribsa Tsegaye BedadaDepartment of Health and Rehabilitation Sciences, Faculty of Medicine and Health Sciences, Stellenbosch University, Cape Town, 7505, South Africa.
Quinette Abegail LouwDepartment of Health and Rehabilitation Sciences, Faculty of Medicine and Health Sciences, Stellenbosch University, Cape Town, 7505, South Africa.
Anna MooreDepartment of Psychiatry, University of Cambridge, Cambridge, CB2 0SZ, UK.
Rudolf N CardinalDepartment of Psychiatry, University of Cambridge, Cambridge, CB2 0SZ, UK.
Tamsin J FordDepartment of Psychiatry, University of Cambridge, Cambridge, CB2 0SZ, UK.
Fan JiangSchool of public health, Shandong ENT Hospital, Shandong University, Jinan, Shandong, China, 250012. Jiang.fan@sdu.edu.cn.

Funding

Medical Research Council MR/Z504816/1National Natural Science Foundation of China 72204143Natural Science Foundation of Shandong Province of China ZR2022QG081NIHR Cambridge Biomedical Research Centre NIHR203312
6 · The paper itself

Abstract

backgroundEarly identification of childhood mental health disorders is a critical public health objective. Existing screening approaches, largely dependent on observer reports, are resource-intensive and may overlook subtle internalized symptoms. The analysis of children's linguistic expression presents a scalable and potentially more objective alternative. This study evaluates whether combining natural language processing (NLP) of children's essays with conventional risk factors improves the detection of mental health difficulties in school-age populations, relative to models based on a single data source.

methodsWe conducted a prospective analysis using data from the UK-based National Child Development Study (NCDS), a national birth cohort initiated in 1958. Data from birth, age 7, and age 11 assessments were analyzed. The final sample included 8,981 children (4,428 [49.3%] female) who completed a creative writing essay at age 11 describing their imagined life at age 25. Predictors comprised traditional risk factors (perinatal, socioeconomic, and parental engagement variables) and linguistic features computationally extracted from the essays. The primary outcome was potential mental health disorder at age 11, defined as scoring above the 95th or 90th percentile on the teacher-completed Bristol Social Adjustment Guide (BSAG). The mother-completed Rutter A Scale was used for sensitivity analysis. Machine learning models incorporating various predictor combinations were developed, and their predictive performance was evaluated using area under the receiver operating characteristic (AUROC) values.

resultsUsing BSAG 95th percentile threshold, models combining top five selected variables with essay features achieved significantly higher predictive capability (AUROC:0.77, 95%CI:0.71-0.83) compared to models using all variables (AUROC:0.70, 95%CI:0.63-0.76) or essay features alone (AUROC:0.67, 95%CI:0.60-0.74). At 90th percentile threshold, this integrated approach showed similar improvement (AUROC:0.81, 95%CI:0.78-0.85). Key predictors included gestational length, maternal parity, parental age, residential characteristics, parental engagement metrics, and children's body mass index. Sensitivity analyses using Rutter A Scale confirmed these findings.

conclusionsIn this prospective birth cohort study, integrating NLP analysis of children's essays with a small set of key risk factors substantially improved the identification of potential mental health disorders. This integrated approach represents a potential paradigm for developing scalable, objective screening tools, but requires validation in contemporary, diverse pediatric populations before clinical consideration.

Indexed as

Machine LearningMental DisordersNatural Language ProcessingBirth CohortChildFemaleHumansMalePredictive Learning ModelsProspective StudiesRisk FactorsUnited KingdomChildrenMachine learningMental health screeningNatural language processing

Identifiers

PMID42210272
PMCPMC13439860

What Socratic holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.