Evidence mapPaperPMID 41360938Full record

ArticleNPP - digital psychiatry and neuroscience2025

Mindbench.ai: an actionable platform to evaluate the profile and performance of large language models in a mental healthcare context.

Bridget Dwyer, Matthew Flathers, Akane Sano, Allison Dempsey, Andrea Cipriani, Asim H Gazi, Bryce Hill, Carla Gorban, Carolyn I Rodriguez, Charles Stromeyer and 23 more

Erratum issuedAbstract read
In one paragraph

Article in NPP - digital psychiatry and neuroscience, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 8 papers.

0numbers the graph read from it
0cells of the map it votes in
8citing papers in PubMed
field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

8 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Review
  5. Article
  6. Article
  7. Review
  8. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

33 authors.

Bridget Dwyer *Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA.
Matthew Flathers *Division of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA.ORCID http://orcid.org/0009-0004-3351-0688
Akane SanoDepartment of Electrical and Computer Engineering, Rice University, Houston, TX, USA.
Allison DempseyDepartment of Psychiatry and Family Medicine, School of Medicine, Centers for American Indian and Alaska Native Health, Colorado School of Public Health, Telemedicine Helen and Arthur E. Johnson Depression Center, University of Colorado Anschutz Medical Campus, Aurora, CO, USA.
Andrea CiprianiDepartment of Psychiatry, University of Oxford, UK; Oxford Precision Psychiatry Lab, National Institute for Health and Care Research (NIHR) Oxford Health Biomedical Research Centre, Oxford, UK.
Asim H GaziSchool of Engineering and Applied Sciences, Harvard University, Cambridge, MA, USA.
Bryce HillDivision of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA.
Carla GorbanThe Brain and Mind Centre, The University of Sydney, Camperdown, New South Wales, Australia.
Carolyn I RodriguezDepartment of Psychiatry and Behavioral Sciences, Stanford University, Stanford, CA, USA.
Charles StromeyerPeer Advisory Advocacy and Research Council, Massachusetts Mental Health Center Public Psychiatry Division of the Beth Israel Deaconess Medical Center, Boston, MA, USA.ORCID http://orcid.org/0009-0006-5957-9258
Darlene KingDepartment of Psychiatry, The University of Texas Southwestern Medical Center, Dallas, TX, USA.
Eden RozenblitDivision of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA.
Gillian StrudwickCentre for Addiction and Mental Health, Toronto, Ontario, Canada.
Jake LinardonSchool of Psychology, Deakin University, Geelong, VIC, Australia.
Jiaee CheongDivision of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA.
Joseph FirthDivision of Psychology and Mental Health, University of Manchester, Manchester Academic Health Science Centre, Manchester, UK.
Julian HerpertzDepartment of Psychiatry and Neuroscience, Campus Benjamin Franklin, Charité-Universitätsmedizin Berlin, Berlin, Germany.
Julian SchwarzDepartment of Psychiatry and Psychotherapy, Center for Mental Health, Immanuel Hospital Rüdersdorf, Brandenburg Medical School Theodor Fontane, Rüdersdorf, Germany.
Khai TruongDivision of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA.
Margaret EmersonCollege of Nursing, University of Nebraska Medical Center, Omaha, USA.
Martin P PaulusLaureate Institute for Brain Research, Tulsa, OK, USA.ORCID http://orcid.org/0000-0002-0825-3606
Michelle PatriquinBaylor College of Medicine, One Baylor Plaza, Houston, TX, USA.
Yining HuaDepartment of Epidemiology, Harvard T.H. Chan School of Public Health, Boston, MA, USA.
Soumya ChoudharyDepartment of Psychiatry, National Institute of Mental Health and Neurosciences, Bangalore, India.
Steven SiddalsDivision of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA.ORCID http://orcid.org/0009-0007-0514-8978
Laura Ospina PinillosInstitute for Life Course Health Research, Department of Global Health, Faculty of Medicine and Health Sciences, Stellenbosch University, Stellenbosch, South Africa.
Jason BantjesDepartment of Psychiatry and Mental Health, Faculty of Medicine, Pontificia Universidad Javeriana, Bogota, Colombia.
Stephen M SchuellerDepartment of Psychological Science, University of California, Irvine, CA, USA.
Xuhai XuDepartment of Biomedical Informatics, Columbia University, NY, USA.
Ken DuckworthNational Alliance on Mental Illness (NAMI), Arlington, VA, USA.
Daniel H GillisonNational Alliance on Mental Illness (NAMI), Arlington, VA, USA.
Michael WoodNational Alliance on Mental Illness (NAMI), Arlington, VA, USA.
John TorousDivision of Digital Psychiatry, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA. jtorous@bidmc.harvard.edu.ORCID http://orcid.org/0000-0002-5362-7937

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Individuals are increasingly utilizing large language model (LLM)-based tools for mental health guidance and crisis support in place of human experts. While AI technology has great potential to improve health outcomes, insufficient empirical evidence exists to suggest that AI technology can be deployed as a clinical replacement; thus, there is an urgent need to assess and regulate such tools. Regulatory efforts have been made and multiple evaluation frameworks have been proposed, however,field-wide assessment metrics have yet to be formally integrated. In this paper, we introduce a comprehensive online platform that aggregates evaluation approaches and serves as a dynamic online resource to simplify LLM and LLM-based tool assessment: MindBench.ai. At its core, MindBench.ai is designed to provide easily accessible/interpretable information for diverse stakeholders (patients, clinicians, developers, regulators, etc.). To create MindBench.ai, we built off our work developing MINDapps.org to support informed decision-making around smartphone app use for mental health, and expanded the technical MINDapps.org framework to encompass novel large language model (LLM) functionalities through benchmarking approaches. The MindBench.ai platform is designed as a partnership with the National Alliance on Mental Illness (NAMI) to provide assessment tools that systematically evaluate LLMs and LLM-based tools with objective and transparent criteria from a healthcare standpoint, assessing both profile (i.e. technical features, privacy protections, and conversational style) and performance characteristics (i.e. clinical reasoning skills). With infrastructure designed to scale through community and expert contributions, along with adapting to technological advances, this platform establishes a critical foundation for the dynamic, empirical evaluation of LLM-based mental health tools-transforming assessment into a living, continuously evolving resource rather than a static snapshot.

Identifiers

PMID41360938
PMCPMC12624894

What Socratic holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the Socratic graph.