The paper describes an Italian language text categorizer by Lemmatization and support vector machines. The categorizer is composed of six modules. The first module performs the tokenization, removing the punctuation signs; the second and third ones carry out stopping and lemmatization, respectively; the fourth module implements the bag-of-words approach; the fifth one performs feature dimensionality reduction eliminating poor discriminant features; finally, the last module does the classification. The Italian text categorizer has been validated on a database composed of more than 1100 articles, extracted from online edition of three Italian language newspapers, belonging to eight different categories. The work is highly novel, since to the best our knowledge, there are no works in literature on Italian text categorization.

Italian Text Categorization with Lemmatization and Support Vector Machines

Camastra F.
;
2020-01-01

Abstract

The paper describes an Italian language text categorizer by Lemmatization and support vector machines. The categorizer is composed of six modules. The first module performs the tokenization, removing the punctuation signs; the second and third ones carry out stopping and lemmatization, respectively; the fourth module implements the bag-of-words approach; the fifth one performs feature dimensionality reduction eliminating poor discriminant features; finally, the last module does the classification. The Italian text categorizer has been validated on a database composed of more than 1100 articles, extracted from online edition of three Italian language newspapers, belonging to eight different categories. The work is highly novel, since to the best our knowledge, there are no works in literature on Italian text categorization.
2020
978-981-13-8949-8
978-981-13-8950-4
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11367/82199
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 13
  • ???jsp.display-item.citation.isi??? ND
social impact