A Taxonomy of Data Quality Challenges in Empirical Software Engineering

Bosu, Michael Franklin; MacDonell, Stephen G.

doi:10.1109/aswec.2013.21

Cited by 32 publications

(29 citation statements)

References 63 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“… Amount of data: the amount of data available for model building contributes to relevance in terms of goal attainment [7]; small and imbalanced data sets build inaccurate models.…”

Section: Journal Of Computersmentioning

confidence: 99%

“… Inconsistency: refers to a lack of harmony between different parts or elements; instances that are self-contradictory, or lacking in agreement when it is expected [7]. This problem is also known as mislabeled data or class noise.…”

Section: Data Quality Diagnosismentioning

confidence: 99%

“…observations are not fit for model-building (accuracy); second are data set characteristics that lead to concerns about the suitability of applying one model to another data set (relevance); and third is a set of factors that limit data accessibility and trust (provenance) [7]. Fig.…”

Section: Journal Of Computersmentioning

confidence: 99%

“…Therefore, in this paper we proposed a conceptual framework for data quality in knowledge discovery tasks based on ESE taxonomy [7]. This framework is a result of filtering elements of CRISP-DM, SEMMA and Data Science, and checking their suitability to the nature in data mining and machine learning projects.…”

Section: Introductionmentioning

confidence: 99%

See 3 more Smart Citations

A Conceptual Framework for Data Quality in Knowledge Discovery Tasks (FDQ-KDT): A Proposal

Corrales¹,

Ledezma²,

Corrales³

2015

JCP

View full text Add to dashboard Cite

Large Volume of Data is growing because the organizations are continuously capturing the collective amount of data for better decision-making process. The most fundamental challenge is to explore the large volumes of data and extract useful knowledge for future actions through data mining and data science methodologies. Nevertheless these not tackle the issues in data quality clearly, leaving out relevant activities. We proposed a conceptual framework for data quality in knowledge discovery tasks based on CRISP-DM, SEMMA and Data Science, considering the issues of ESE Taxonomy.

show abstract

“… Amount of data: the amount of data available for model building contributes to relevance in terms of goal attainment [7]; small and imbalanced data sets build inaccurate models.…”

Section: Journal Of Computersmentioning

confidence: 99%

Section: Data Quality Diagnosismentioning

confidence: 99%

Section: Journal Of Computersmentioning

confidence: 99%

Section: Introductionmentioning

confidence: 99%

See 2 more Smart Citations

A Conceptual Framework for Data Quality in Knowledge Discovery Tasks (FDQ-KDT): A Proposal

Corrales¹,

Ledezma²,

Corrales³

2015

JCP

View full text Add to dashboard Cite

show abstract

“…In this paper we present a systematic review for data quality issues in knowledge discovery tasks as: heterogeneity, outliers, noise, inconsistency, incompleteness, amount of data, redundancy and timeliness which are defined in [7][8] and a case study in agricultural diseases: the coffee rust. This paper is organized as follows.…”

Section: Introductionmentioning

confidence: 99%

A systematic review of data quality issues in knowledge discovery tasks

Corrales

Ledezma

Corrales

2016

Rev. ing. univ. Medellín

View full text Add to dashboard Cite

Large volume of data is growing because organizations are continuously capturing the collective amount of data for a better decision-making process. The most fundamental challenge is to explore the large volumes of data and extract useful knowledge for future actions through knowledge discovery tasks, nevertheless many data has poor quality. We presented a systematic review of the data quality issues in knowledge discovery tasks and a case study applied to agricultural disease named coffee rust. ResumenHay un gran crecimiento en el volumen de datos porque las organizaciones capturan permanentemente la cantidad colectiva de datos para lograr un mejor proceso de toma de decisiones. El desafío mas fundamental es la exploración de los grandes volúmenes de datos y la extracción de conocimiento útil para futuras acciones por medio de tareas para el descubrimiento del conocimiento; sin embargo, muchos datos presentan mala calidad. Presentamos una revisión sistemática de los asuntos de calidad de datos en las áreas del descubrimiento de conocimiento y un estudio de caso aplicado a la enfermedad agrícola conocida como la roya del café.

show abstract