An analysis of hierarchical text classification using word embeddings

Carregando...
Imagem de Miniatura

Data de defesa

Autor(a)



Orientador(a)


Coorientador(a)


Título do periódico

ISSN

Título do volume

Nome da instituição

Universidade do Vale do Rio dos Sinos

Departamento

Escola Politécnica

Programa

Programa de Pós-Graduação em Computação Aplicada

Agência de fomento

CAPES - Coordenação de Aperfeiçoamento de Pessoal de Nível Superior
Resumo

Efficient distributed numerical word representation models (word embeddings) combined with modern machine learning algorithms have recently yielded considerable improvement on automatic document classification tasks. However, the effectiveness of such techniques has not been assessed for the hierarchical text classification (HTC) yet. This study investigates application of those models and algorithms on this specific problem by means of experimentation and analysis. Classification models were trained with prominent machine learning algorithm implementations—fastText, XGBoost, and Keras’ CNN—and noticeable word embeddings generation methods—GloVe, word2vec, and fastText—with publicly available data and evaluated them with measures specifically appropriate for the hierarchical context. FastText achieved an LCAF1 of 0.871 on a single-labeled version of the RCV1 dataset. The results analysis indicates that using word embeddings is a very promising approach for HTC.