emoji2vec: Learning Emoji Representations from their Description

Eisner, Ben; Rocktäschel, Tim; Augenstein, Isabelle; Bošnjak, Matko; Riedel, Sebastian

Computer Science > Computation and Language

arXiv:1609.08359 (cs)

[Submitted on 27 Sep 2016 (v1), last revised 20 Nov 2016 (this version, v2)]

Title:emoji2vec: Learning Emoji Representations from their Description

Authors:Ben Eisner, Tim Rocktäschel, Isabelle Augenstein, Matko Bošnjak, Sebastian Riedel

View PDF

Abstract:Many current natural language processing applications for social media rely on representation learning and utilize pre-trained word embeddings. There currently exist several publicly-available, pre-trained sets of word embeddings, but they contain few or no emoji representations even as emoji usage in social media has increased. In this paper we release emoji2vec, pre-trained embeddings for all Unicode emoji which are learned from their description in the Unicode emoji standard. The resulting emoji embeddings can be readily used in downstream social natural language processing applications alongside word2vec. We demonstrate, for the downstream task of sentiment analysis, that emoji embeddings learned from short descriptions outperforms a skip-gram model trained on a large collection of tweets, while avoiding the need for contexts in which emoji need to appear frequently in order to estimate a representation.

Comments:	7 pages, 4 figures, 1 table, In Proceedings of the 4th International Workshop on Natural Language Processing for Social Media at EMNLP 2016 (SocialNLP at EMNLP 2016)
Subjects:	Computation and Language (cs.CL)
MSC classes:	68T50
ACM classes:	I.2.7
Cite as:	arXiv:1609.08359 [cs.CL]
	(or arXiv:1609.08359v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.1609.08359

Submission history

From: Isabelle Augenstein [view email]
[v1] Tue, 27 Sep 2016 11:32:25 UTC (2,655 KB)
[v2] Sun, 20 Nov 2016 22:43:46 UTC (2,655 KB)

Computer Science > Computation and Language

Title:emoji2vec: Learning Emoji Representations from their Description

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:emoji2vec: Learning Emoji Representations from their Description

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators