COMMON MISTAKES STUDENTS MAKE ABOUT DATA ANALYSIS

Data collected from the field must be analysed, presented, and interpreted to make it meaningful to the audiences of the research work. Raw data are not yet findings. Analysis transforms raw data into evidence that can be used to answer research questions.

The process of reducing research data to manageable summaries is what is called data analysis. In most social science research Projects, dissertations, and theses, data analysis is done in chapter four.

Data analysis is purpose-driven and dependent on the type of data collected; therefore, do not analyse data simply because you have data. Analyse data because you have a research question to answer.

The sequence of analysis should be:

Research question → Type of data/scale of measurement → Analytical technique / statistical tool → Analysis → presentation of findings → Interpretation of findings

and not:

Data → SPSS → Click buttons → Find a significant result → Develop a research question.

This distinction makes a major difference in the quality of students’ research.

The following are some mistakes that students make while analysing data.

Mistake 1: Thinking data analysis means using SPSS

Many students imagine that data analysis means using SPSS.  Data analysis is the intellectual process of deciding what needs to be analysed, why, how, and what the results mean. What one should remember is that SPSS is a tool; it’s a software to analyse quantitative data. Other tools include:

  • Excel
  • R
  • Stata
  • Python
  • NVivo

If you use it wrongly, it will give you wrong results.

Having these tools does not mean having the best results. One must be aware of the type of data they have collected and the appropriate statistical tool to apply.


Mistake 2: Choosing statistical tests simply because they are available

A statistical test is a procedure for deciding whether an assertion (e. g. a hypothesis) about a quantitative feature of a population is true or false. We test a hypothesis of this sort by drawing a random sample from the population in question and calculating an appropriate statistic on its items (Ille & Milic, 2008). The intent is to determine whether there is enough evidence to “reject” such an assertion (inference) or hypothesis about the process. There are two types of statistical tests: parametric and non-parametric tests.

  1. Parametric tests – these are tests that make certain assumptions about the population distribution. These tests are called parametric because the assumptions are about population parameters
  2. Non-parametric tests – These are sometimes known as assumption-free tests because they make fewer assumptions. Most of these tests work on the principle of ranking the data. Analysis is then carried on ranks and not on actual data

Students sometimes say:

“I used correlation because SPSS provides correlation.”

The correct reasoning is the reverse:

The type of data collected /scales of measurement determine the appropriate analytical technique; the software merely helps perform the analysis.


Mistake 3: Confusing percentages with interpretation

Students always confuse presentation of data and interpretation of data. Presentation of data is the arrangement of data using tables, figures, and numbers to make it clear, while interpretation means searching for meaning and implications of research results in order to make inferences, draw conclusions, and relate to the theory; i.e., based on the research problem, what is the implication of the finding presented?

  • Suppose 65.0% of respondents are dissatisfied with a service.

Writing:

“65.0% of respondents were dissatisfied” is a finding presented in numbers.

The appropriate interpretation based on the research problem would be:

“This may suggest that the government or the organization needs to examine factors contributing to customer dissatisfaction.”


Mistake 4: Believing every dataset requires sophisticated statistics

Statistics is a body of mathematical techniques or processes for gathering, organizing, analysing, and interpreting numerical data (Best and Khan, 2009, p. 354).  This data is based on a sample. There are two types of statistics:

  1. Descriptive statistics – this is a way of summarizing data by letting one number stand for a group of numbers. It describes a sample
  2. Inferential statistics – this is the type of statistics that allows researchers to make inferences and predictions about a population based on a sample of data taken from the population in question. It is based on probability (to infer). This inference is made possible by testing hypotheses.

Not every research question requires regression, ANOVA, or structural equation modelling. Secondly, inferential statistics are mainly useful when one is testing hypotheses.

Sometimes:

  • Frequencies
  • Percentages
  • Means
  • Tables
  • Graphs
  • Themes for qualitative research

may be sufficient.

The principle is:

Use the simplest appropriate analytical technique that adequately answers the research question.


Mistake 5: Confusing correlation with causation

 Correlation establishes a relationship between variables. Variable X is positively or negatively related to variable Y, or there is no relationship between variable X and variable Y. Causation means that variable X causes variable Y. If you notice a correlation between two variables, it’s tempting to think one causes the other. But that’s not the case unless established through determination of causation. Many times, the correlation is purely a coincidence. To find out whether two factors are related, you should look at the context. Are there any other factors that could cause the correlation? Don’t assume a connection without conducting more research.


In conclusion, researchers analyse data to:

  1. Answer research questions
  2. Achieve research objectives
  3. Test hypotheses
  4. Identify patterns and trends
  5. Establish relationships
  6. Compare groups
  7. Generate evidence-based conclusions
  8. Support decision-making

Data analysis and interpretation are different: analysis establishes what the data shows, while interpretation explains what the findings mean.

Leave a Reply

Your email address will not be published. Required fields are marked *