← Evidence map

ArticleResearch square2026

Auditing frontier general-purpose large language models in biomedical tasks: reasoning gains, extraction limits, and benchmark reliability.

Yu Hou et al.PubMed ↗Full text ↗Publisher ↗

No numbers read from the abstract.

Not cited yet

Full record →Abstract, authors, funding and every citing paper · PMID 41756442