Detecting fake news over online social media via domain reputations and content understanding

Kuai Xu; Feng Wang; Haiyan Wang; Bo Yang

doi:10.26599/TST.2018.9010139

Detecting fake news over online social media via domain reputations and content understanding

Kuai Xu, Feng Wang, Haiyan Wang, Bo Yang

Mathematical and Natural Sciences, School of (SMNS)

Research output: Contribution to journal › Article › peer-review

55 Scopus citations

Abstract

Fake news has recently leveraged the power and scale of online social media to effectively spread misinformation which not only erodes the trust of people on traditional presses and journalisms, but also manipulates the opinions and sentiments of the public. Detecting fake news is a daunting challenge due to subtle difference between real and fake news. As a first step of fighting with fake news, this paper characterizes hundreds of popular fake and real news measured by shares, reactions, and comments on Facebookfrom two perspectives: domain reputations and content understanding. Our domain reputation analysis reveals that the Web sites of the fake and real news publishers exhibit diverse registration behaviors, registration timing, domain rankings, and domain popularity. In addition, fake news tends to disappear from the Web after a certain amount of time. The content characterizations on the fake and real news corpus suggest that simply applying term frequency-inverse document frequency (tf-idf) and Latent Dirichlet Allocation (LDA) topic modeling is inefficient in detecting fake news, while exploring document similarity with the term and word vectors is a very promising direction for predicting fake and real news. To the best of our knowledge, this is the first effort to systematically study domain reputations and content characteristics of fake and real news, which will provide key insights for effectively detecting fake news on social media.

Original language	English (US)
Article number	8768083
Pages (from-to)	20-27
Number of pages	8
Journal	Tsinghua Science and Technology
Volume	25
Issue number	1
DOIs	https://doi.org/10.26599/TST.2018.9010139
State	Published - Feb 2020

Keywords

content modeling
domain reputations
fake news detection
social media

ASJC Scopus subject areas

General

Access to Document

10.26599/TST.2018.9010139

Cite this

@article{c6b5aa43eaa140099d824ba2b6cc2865,

title = "Detecting fake news over online social media via domain reputations and content understanding",

abstract = "Fake news has recently leveraged the power and scale of online social media to effectively spread misinformation which not only erodes the trust of people on traditional presses and journalisms, but also manipulates the opinions and sentiments of the public. Detecting fake news is a daunting challenge due to subtle difference between real and fake news. As a first step of fighting with fake news, this paper characterizes hundreds of popular fake and real news measured by shares, reactions, and comments on Facebookfrom two perspectives: domain reputations and content understanding. Our domain reputation analysis reveals that the Web sites of the fake and real news publishers exhibit diverse registration behaviors, registration timing, domain rankings, and domain popularity. In addition, fake news tends to disappear from the Web after a certain amount of time. The content characterizations on the fake and real news corpus suggest that simply applying term frequency-inverse document frequency (tf-idf) and Latent Dirichlet Allocation (LDA) topic modeling is inefficient in detecting fake news, while exploring document similarity with the term and word vectors is a very promising direction for predicting fake and real news. To the best of our knowledge, this is the first effort to systematically study domain reputations and content characteristics of fake and real news, which will provide key insights for effectively detecting fake news on social media.",

keywords = "content modeling, domain reputations, fake news detection, social media",

author = "Kuai Xu and Feng Wang and Haiyan Wang and Bo Yang",

note = "Funding Information: This work was supported in part by National Science Foundation (NSF) Algorithms for Threat Detection (ATD) Program (No. DMS #1737861) and NSF Computer and Network Systems (CNS) Program (No. CNS #1816995). Publisher Copyright: {\textcopyright} 1996-2012 Tsinghua University Press.",

year = "2020",

month = feb,

doi = "10.26599/TST.2018.9010139",

language = "English (US)",

volume = "25",

pages = "20--27",

journal = "Tsinghua Science and Technology",

issn = "1007-0214",

publisher = "Tsing Hua University",

number = "1",

}

TY - JOUR

T1 - Detecting fake news over online social media via domain reputations and content understanding

AU - Xu, Kuai

AU - Wang, Feng

AU - Wang, Haiyan

AU - Yang, Bo

N1 - Funding Information: This work was supported in part by National Science Foundation (NSF) Algorithms for Threat Detection (ATD) Program (No. DMS #1737861) and NSF Computer and Network Systems (CNS) Program (No. CNS #1816995). Publisher Copyright: © 1996-2012 Tsinghua University Press.

PY - 2020/2

Y1 - 2020/2

N2 - Fake news has recently leveraged the power and scale of online social media to effectively spread misinformation which not only erodes the trust of people on traditional presses and journalisms, but also manipulates the opinions and sentiments of the public. Detecting fake news is a daunting challenge due to subtle difference between real and fake news. As a first step of fighting with fake news, this paper characterizes hundreds of popular fake and real news measured by shares, reactions, and comments on Facebookfrom two perspectives: domain reputations and content understanding. Our domain reputation analysis reveals that the Web sites of the fake and real news publishers exhibit diverse registration behaviors, registration timing, domain rankings, and domain popularity. In addition, fake news tends to disappear from the Web after a certain amount of time. The content characterizations on the fake and real news corpus suggest that simply applying term frequency-inverse document frequency (tf-idf) and Latent Dirichlet Allocation (LDA) topic modeling is inefficient in detecting fake news, while exploring document similarity with the term and word vectors is a very promising direction for predicting fake and real news. To the best of our knowledge, this is the first effort to systematically study domain reputations and content characteristics of fake and real news, which will provide key insights for effectively detecting fake news on social media.

AB - Fake news has recently leveraged the power and scale of online social media to effectively spread misinformation which not only erodes the trust of people on traditional presses and journalisms, but also manipulates the opinions and sentiments of the public. Detecting fake news is a daunting challenge due to subtle difference between real and fake news. As a first step of fighting with fake news, this paper characterizes hundreds of popular fake and real news measured by shares, reactions, and comments on Facebookfrom two perspectives: domain reputations and content understanding. Our domain reputation analysis reveals that the Web sites of the fake and real news publishers exhibit diverse registration behaviors, registration timing, domain rankings, and domain popularity. In addition, fake news tends to disappear from the Web after a certain amount of time. The content characterizations on the fake and real news corpus suggest that simply applying term frequency-inverse document frequency (tf-idf) and Latent Dirichlet Allocation (LDA) topic modeling is inefficient in detecting fake news, while exploring document similarity with the term and word vectors is a very promising direction for predicting fake and real news. To the best of our knowledge, this is the first effort to systematically study domain reputations and content characteristics of fake and real news, which will provide key insights for effectively detecting fake news on social media.

KW - content modeling

KW - domain reputations

KW - fake news detection

KW - social media

UR - http://www.scopus.com/inward/record.url?scp=85069524460&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85069524460&partnerID=8YFLogxK

U2 - 10.26599/TST.2018.9010139

DO - 10.26599/TST.2018.9010139

M3 - Article

AN - SCOPUS:85069524460

SN - 1007-0214

VL - 25

SP - 20

EP - 27

JO - Tsinghua Science and Technology

JF - Tsinghua Science and Technology

IS - 1

M1 - 8768083

ER -

Detecting fake news over online social media via domain reputations and content understanding

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this