Identifying Chinese Microblog Users With High Suicide Probability Using Internet-Based Profile and Linguistic Features: Classification Model

Guan, L; Hao, B; Cheng, Q; Yip, PSF; Zhu, T

File Download

There are no files associated with this item.

Links for fulltext

(May Require Subscription)

Publisher Website: 10.2196/mental.4227
Scopus: eid_2-s2.0-84941367910
WOS: WOS:000414978100005

Supplementary

Citations:
- Scopus: 0
- Web of Science: 0
Appears in Collections:
- Hong Kong Jockey Club Centre for Suicide Research and Prevention: Journal/Magazine Articles

Article: Identifying Chinese Microblog Users With High Suicide Probability Using Internet-Based Profile and Linguistic Features: Classification Model

Title	Identifying Chinese Microblog Users With High Suicide Probability Using Internet-Based Profile and Linguistic Features: Classification Model
Authors	Guan, L Hao, B Cheng, Q Yip, PSF Zhu, T
Issue Date	2015
Citation	JMIR Mental Health, 2015, v. 2, p. e17 How to Cite? DOI: http://dx.doi.org/10.2196/mental.4227
Abstract	Background: Traditional offline assessment of suicide probability is time consuming and difficult in convincing at-risk individuals to participate. Identifying individuals with high suicide probability through online social media has an advantage in its efficiency and potential to reach out to hidden individuals, yet little research has been focused on this specific field. Objective: The objective of this study was to apply two classification models, Simple Logistic Regression (SLR) and Random Forest (RF), to examine the feasibility and effectiveness of identifying high suicide possibility microblog users in China through profile and linguistic features extracted from Internet-based data. Methods: There were nine hundred and nine Chinese microblog users that completed an Internet survey, and those scoring one SD above the mean of the total Suicide Probability Scale (SPS) score, as well as one SD above the mean in each of the four subscale scores in the participant sample were labeled as high-risk individuals, respectively. Profile and linguistic features were fed into two machine learning algorithms (SLR and RF) to train the model that aims to identify high-risk individuals in general suicide probability and in its four dimensions. Models were trained and then tested by 5-fold cross validation; in which both training set and test set were generated under the stratified random sampling rule from the whole sample. There were three classic performance metrics (Precision, Recall, F1 measure) and a specifically defined metric “Screening Efficiency” that were adopted to evaluate model effectiveness. Results: Classification performance was generally matched between SLR and RF. Given the best performance of the classification models, we were able to retrieve over 70% of the labeled high-risk individuals in overall suicide probability as well as in the four dimensions. Screening Efficiency of most models varied from 1/4 to 1/2. Precision of the models was generally below 30%. Conclusions: Individuals in China with high suicide probability are recognizable by profile and text-based information from microblogs. Although there is still much space to improve the performance of classification models in the future, this study may shed light on preliminary screening of risky individuals via machine learning algorithms, which can work side-by-side with expert scrutiny to increase efficiency in large-scale-surveillance of suicide probability from online social media.
Persistent Identifier	http://hdl.handle.net/10722/219209
ISI Accession Number ID	WOS:000414978100005

DC Field	Value	Language
dc.contributor.author	Guan, L	-
dc.contributor.author	Hao, B	-
dc.contributor.author	Cheng, Q	-
dc.contributor.author	Yip, PSF	-
dc.contributor.author	Zhu, T	-
dc.date.accessioned	2015-09-18T07:17:34Z	-
dc.date.available	2015-09-18T07:17:34Z	-
dc.date.issued	2015	-
dc.identifier.citation	JMIR Mental Health, 2015, v. 2, p. e17	-
dc.identifier.uri	http://hdl.handle.net/10722/219209	-
dc.description.abstract	Background: Traditional offline assessment of suicide probability is time consuming and difficult in convincing at-risk individuals to participate. Identifying individuals with high suicide probability through online social media has an advantage in its efficiency and potential to reach out to hidden individuals, yet little research has been focused on this specific field. Objective: The objective of this study was to apply two classification models, Simple Logistic Regression (SLR) and Random Forest (RF), to examine the feasibility and effectiveness of identifying high suicide possibility microblog users in China through profile and linguistic features extracted from Internet-based data. Methods: There were nine hundred and nine Chinese microblog users that completed an Internet survey, and those scoring one SD above the mean of the total Suicide Probability Scale (SPS) score, as well as one SD above the mean in each of the four subscale scores in the participant sample were labeled as high-risk individuals, respectively. Profile and linguistic features were fed into two machine learning algorithms (SLR and RF) to train the model that aims to identify high-risk individuals in general suicide probability and in its four dimensions. Models were trained and then tested by 5-fold cross validation; in which both training set and test set were generated under the stratified random sampling rule from the whole sample. There were three classic performance metrics (Precision, Recall, F1 measure) and a specifically defined metric “Screening Efficiency” that were adopted to evaluate model effectiveness. Results: Classification performance was generally matched between SLR and RF. Given the best performance of the classification models, we were able to retrieve over 70% of the labeled high-risk individuals in overall suicide probability as well as in the four dimensions. Screening Efficiency of most models varied from 1/4 to 1/2. Precision of the models was generally below 30%. Conclusions: Individuals in China with high suicide probability are recognizable by profile and text-based information from microblogs. Although there is still much space to improve the performance of classification models in the future, this study may shed light on preliminary screening of risky individuals via machine learning algorithms, which can work side-by-side with expert scrutiny to increase efficiency in large-scale-surveillance of suicide probability from online social media.	-
dc.language	eng	-
dc.relation.ispartof	JMIR Mental Health	-
dc.title	Identifying Chinese Microblog Users With High Suicide Probability Using Internet-Based Profile and Linguistic Features: Classification Model	-
dc.type	Article	-
dc.identifier.email	Cheng, Q: chengqj@connect.hku.hk	-
dc.identifier.email	Yip, PSF: sfpyip@hku.hk	-
dc.identifier.authority	Cheng, Q=rp02018	-
dc.identifier.authority	Yip, PSF=rp00596	-
dc.identifier.doi	10.2196/mental.4227	-
dc.identifier.scopus	eid_2-s2.0-84941367910	-
dc.identifier.hkuros	253568	-
dc.identifier.volume	2	-
dc.identifier.spage	e17	-
dc.identifier.epage	e17	-
dc.identifier.eissn	2368-7959	-
dc.identifier.isi	WOS:000414978100005	-
dc.identifier.issnl	2368-7959	-

File Download

Links for fulltext

(May Require Subscription)

Supplementary

Article: Identifying Chinese Microblog Users With High Suicide Probability Using Internet-Based Profile and Linguistic Features: Classification Model

Export via OAI-PMH Interface in XML Formats

OR

Export to Other Non-XML Formats