Ciência_Iscte Publicações Descrição Detalhada da Publicação Exportar

Exportar Publicação

A publicação pode ser exportada nos seguintes formatos: referência da APA (American Psychological Association), referência do IEEE (Institute of Electrical and Electronics Engineers), BibTeX e RIS.

Exportar Referência (APA)

Santos, J. M., Shah, S., Gupta, A., Mann, A., Vaz, A., Caldwell, B. E....Rousmaniere, T. (N/A). Evaluating the clinical safety of large language models in response to high-risk mental health disclosures. Practice Innovations. N/A

Exportar Referência (IEEE)

J. M. Santos et al.,  "Evaluating the clinical safety of large language models in response to high-risk mental health disclosures", in Practice Innovations, vol. N/A, N/A

Exportar BibTeX

@article{santosN/A_1784140111967,
	author = "Santos, J. M. and Shah, S. and Gupta, A. and Mann, A. and Vaz, A. and Caldwell, B. E. and Scholz, R. and Awad, P. and Allemandi, R. and Faust, D. and Banka, H. and Rousmaniere, T.",
	title = "Evaluating the clinical safety of large language models in response to high-risk mental health disclosures",
	journal = "Practice Innovations",
	year = "N/A",
	volume = "N/A",
	number = "",
	doi = "10.1037/pri0000316",
	url = "https://www.apa.org/pubs/journals/pri"
}

Exportar RIS

TY  - JOUR
TI  - Evaluating the clinical safety of large language models in response to high-risk mental health disclosures
T2  - Practice Innovations
VL  - N/A
AU  - Santos, J. M.
AU  - Shah, S.
AU  - Gupta, A.
AU  - Mann, A.
AU  - Vaz, A.
AU  - Caldwell, B. E.
AU  - Scholz, R.
AU  - Awad, P.
AU  - Allemandi, R.
AU  - Faust, D.
AU  - Banka, H.
AU  - Rousmaniere, T.
PY  - N/A
SN  - 2377-889X
DO  - 10.1037/pri0000316
UR  - https://www.apa.org/pubs/journals/pri
AB  - As large language models increasingly mediate emotionally sensitive conversations, especially in mental health contexts, their ability to recognize and respond to high-risk situations becomes a matter of public safety. This study evaluates the responses of six popular large language models—Claude, Gemini, DeepSeek, ChatGPT, Grok 3, and LLAMA—to user prompts simulating crisis-level mental health disclosures. Drawing on a coding framework developed by licensed clinicians, five safety-oriented behaviors were assessed: explicit risk acknowledgment, empathy, encouragement to seek help, provision of specific resources, and invitation to continue the conversation. Claude outperformed all others in a global assessment, while Grok 3, ChatGPT, and LLAMA underperformed across multiple domains. Notably, most models exhibited empathy, but few consistently provided practical support or kept the conversation open. These findings suggest that while large language models show potential for emotionally attuned communication, none currently meet satisfactory clinical standards for crisis response. Ongoing development and targeted fine-tuning are essential to ensure ethical deployment of AI in mental health settings.
ER  -