An analysis of hierarchical text classification using word embeddings

An analysis of hierarchical text classification using word embeddings

Stein, Roger Alan

URI: http://www.repositorio.jesuita.org.br/handle/UNISINOS/7624

Date: 2018-03-28

xmlui.dri2xhtml.METS-1.0.item-contributorAdvisor: Maillard, Patrícia Augustin Jaques

Abstract:

Efficient distributed numerical word representation models (word embeddings) combined with modern machine learning algorithms have recently yielded considerable improvement on automatic document classification tasks. However, the effectiveness of such techniques has not been assessed for the hierarchical text classification (HTC) yet. This study investigates application of those models and algorithms on this specific problem by means of experimentation and analysis. Classification models were trained with prominent machine learning algorithm implementations—fastText, XGBoost, and Keras’ CNN—and noticeable word embeddings generation methods—GloVe, word2vec, and fastText—with publicly available data and evaluated them with measures specifically appropriate for the hierarchical context. FastText achieved an LCAF1 of 0.871 on a single-labeled version of the RCV1 dataset. The results analysis indicates that using word embeddings is a very promising approach for HTC.

Show full item record

Files in this item

Name: Roger Alan Stein_.pdf

Size: 465.0Kb

Format: PDF

Description: analysis_hierarchical

View/Open

This item appears in the following Collection(s)

PPG Computação Aplicada [366]
PPG Computação Aplicada

Search

Browse

All of RDBU
- Communities & Collections
This Collection

My Account

Statistics

View Usage Statistics

An analysis of hierarchical text classification using word embeddings

An analysis of hierarchical text classification using word embeddings

Abstract:

Files in this item

This item appears in the following Collection(s)

Search

Browse

All of RDBU

This Collection

My Account

Statistics