| Nome: | Descrição: | Tamanho: | Formato: | |
|---|---|---|---|---|
| 88.51 KB | Adobe PDF |
Orientador(es)
Resumo(s)
A corpus is a set of texts related to some topic, theme or subject This means that the texts in a corpus, although produced by distinct authors, have meanings related to that subject, topic or theme. This work addresses the problem of descriptive statistics of variables whose observed values are texts, aiming at descriptive statistics of corpus variables (CVs) and its interpretation. Specifically, the aim is to build text mining methods to compare and relate CVs, assuming texts as its values, using intersection graphs and hypergraphs as the basic mathematical representation of CVs and its relations. The data sets involved have the usual tabular structure, crossing the values assumed by p CVs, observed on n individuals, objects , cases or documents. The values assumed by these CVs are texts of any length, formed by sets of words in a specific language, expressing opinions or other content created by its authors. If T is the set of all texts (observed values) in such data set, let t ε T; t is a subset of V, formed by words of that specific language, obtained by tokenization of all texts in T. It is shown that hypergraphs and intersection graphs occur in this context as natural mathematical representation tools for texts and its relations. This mathematical representation is a powerful support for statistical mining of texts and its relations, namely text mining in the context of maintenance management and other engineering disciplines.
Descrição
Palavras-chave
Maintenance Dynamic Ship Risk FMECA Threat
