| Nome: | Descrição: | Tamanho: | Formato: | |
|---|---|---|---|---|
| 2.93 MB | Adobe PDF |
Autores
Orientador(es)
Resumo(s)
The continuous increase in the production and availability of textual content presents both challenges and opportunities for the field of Authorship Identification. The main objective of this dissertation is to develop a strategy that utilizes Natural Language Processing (NLP) methodologies and Machine Learning algorithms to identify the authorship of texts.
Authorship Identification consists of determining the identity of an author based on a specific text, relying on their unique linguistic and stylistic patterns. This plays a crucial role, not only in understanding writing styles but also in combating practices such as plagiarism. The study, framed as a supervised classification problem, uses data from the Twitter platform (currently known as X) to explore Authorship Identification strategies and build models capable of predicting the authorship of texts based on labeled data.
The study proposes an approach that utilizes NLP methodologies to transform texts into structured data, enabling the application of Machine Learning techniques to create models capable of identifying writing patterns and associating them with specific authors. To evaluate the models' effectiveness, traditional Machine Learning algorithms, such as Naive Bayes (NB) and Random Forest (RF), as well as advanced Deep Learning (DL) models, including Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN), will be explored.
This dissertation adopts the KDD (Knowledge Discovery in Databases) methodology, starting with the definition of the problem and objectives related to Authorship Identification. During the data preparation and exploration stages, NLP techniques, such as tokenization and stopwords removal, are applied to structure the data for the Machine Learning models. The modeling phase consists of the implementation and optimization of Machine Learning and DL algorithms, with the development of a practical case of Authorship Identification on the Twitter platform. Finally, the model evaluation phase will be carried out based on performance metrics, where the Passive-Aggressive Classifier (PAC) model was identified as the most effective, standing out for its precision in distinguishing writing styles and contributing significantly to the improvement of the Authorship Identification process.
Descrição
Palavras-chave
Authorship identification Natural Language Processing (NLP) Machine Learning
