Data is the information collected by researchers during a study. In psychology, data can take different forms depending on what is being measured and how it is collected. Researchers also need to consider the level of measurement used as this affects how the data can be analysed.

This page covers the key concepts needed to understand and use data in psychology:


Raw data


Raw data is the original information collected by researchers before it has been processed or summarised. It could be scores from a psychological test, reaction times, questionnaire responses, the number of participants who showed a particular behaviour – stuff like that.

Raw data recording tables

Researchers often use a raw data recording table to organise the data they collect. A well-designed table makes it easy to record and interpret the results.

For example, if a researcher measures the reaction time of five participants, they could use a table like this:

Participant Reaction time (ms)
1 420
2 385
3 451
4 397
5 412
6 1200

A raw data table should:

  • Give each column a clear heading
  • Include the units of measurement where appropriate, such as seconds (s) or milliseconds (ms)
  • Only contain the raw data – so no calculations or summaries of the results.

Outliers

An outlier is a data value that is unusually high or low compared with the other values.

For example, in the reaction times data set above:

420, 385, 451, 397, 412, 1200

The value 1200 is much higher than the other reaction times, so it could be an outlier.

Outliers can affect calculations (e.g. the mean) because an unusually high or low value can pull the mean towards it. As such, researchers may check for and remove outliers before analysing their data, as this may produce more accurate results. However, an outlier should not automatically be removed. Researchers may then consider why the value is unusual and whether there is a valid reason for excluding it.

Standard form and decimal form

Very large or very small numbers can be written in standard form to make them easier to read and work with. Numbers can also be written in decimal form, which is the ordinary way of writing a number.

Decimal form Standard form
3,000,000,000 3 x 109
0.000003 3 x 10−6

For example, researchers estimate the human brain has around 86 billion (86,000,000,000) neurons and 100 trillion (100,000,000,000,000) synapses. Instead of writing these massive numbers out as decimal forms, you can just say 8.6 x 10¹⁰ and 1 x 10¹⁴.

Significant figures

Significant figures are the digits in a number that show how precise the measurement is.

For example:

  • 4.673 to 3 significant figures = 4.67
  • 15.28 to 3 significant figures = 15.3
  • 0.004826 to 2 significant figures = 0.0048

When rounding to significant figures, look at the digit immediately after the final significant figure:

  • If it is 5 or more, round up
    • E.g. 4.678 to 3 significant figures = 4.68
  • If it is 4 or less, leave the final digit unchanged
    • E.g. 4.673 to 3 significant figures = 4.67

Researchers may use significant figures when reporting measurements or calculated results so that the results are not presented with unnecessary precision.

Making estimations from data

Researchers can use the data they collect to make estimations about what might happen in similar situations or in the wider population.

For example, imagine a researcher finds that 18 out of 20 participants followed an instruction in an experiment. This is 90% of the sample. The researcher could use this result to estimate that around 90% of similar participants might follow the instruction under the same conditions.

However, an estimation is not guaranteed to be accurate. The sample may not perfectly represent the wider population, so researchers should treat estimates as approximate rather than exact.


Types of data


Quantitative vs. qualitative

Qualitative vs. quantitative is about whether the data can be expressed and compared in numbers.

Quantitative data

Quantitative data is data that can be measured and expressed as numbers. It is usually collected by measuring or counting something. This data can be analysed using statistical calculations such as calculating the mean, median, or range.

Examples of quantitative data include:

  • A participant’s reaction time in milliseconds
    • E.g. 397
  • The number of words a participant recalls on a memory test
    • E.g. 19
  • A participant’s age in years
    • E.g. 35
Strengths of quantitative data:

  • It is objective because it is based on numerical measurements rather than researchers’ interpretations.
  • It can be easily compared between participants or groups.
  • It can be analysed using statistical tests, allowing researchers to identify patterns, differences and relationships in the data.
Weaknesses of quantitative data:

  • It can lack detail and depth, as numbers may not explain why participants behaved or responded in a particular way.
  • Some psychological experiences, such as thoughts and feelings, can be difficult to represent accurately using numbers.
  • The way researchers measure something can affect the results. For example, a questionnaire score may not fully capture someone’s actual level of anxiety.

Qualitative data

Qualitative data is data that describes experiences, thoughts, feelings, or behaviour using words rather than numbers. It is usually collected through methods such as interviews and observations.

Qualitative data is often more detailed than quantitative data and can give researchers a better understanding of why people think, feel or behave in particular ways.

Examples of qualitative data include:

  • A participant’s description of how they felt during an experiment
  • A researcher’s written description of a participant’s behaviour
  • Responses to an open question such as “how did you feel when completing the task?”
Strengths of qualitative data:

  • It provides detailed and in-depth information about people’s experiences, thoughts, and feelings.
  • It can provide a better understanding of why people behave in particular ways.
  • It allows participants to give answers in their own words, rather than being restricted to fixed responses.
Weaknesses of qualitative data:

  • It can be difficult to analyse because researchers have to identify and interpret themes or patterns in people’s responses.
  • The analysis can be more subjective, as researchers may interpret participants’ responses differently.
  • It can be time-consuming to collect and analyse – particularly when participants provide long or detailed responses.

Qualitative data can sometimes be converted into quantitative data.

For example, researchers might ask participants:

“How did you feel during the experiment?”

By placing responses into categories and counting how often each category occurs, participants’ qualitative answers can be converted into quantitative numbers. For example, researchers could group the answers into categories such as ‘happy’, ‘anxious’, and ‘neutral’, and then count how many participants gave an answer in each category like this:

Response category Number of participants
Happy 8
Anxious 7
Neutral 5

The original answers are qualitative data because they describe participants’ feelings in words. Once the researchers have counted the responses in each category, the results are quantitative data because they are expressed as numbers.

This can make the data easier to compare and analyse statistically, although some of the detail from the original qualitative responses may be lost.


Primary vs. secondary

Primary vs. secondary data is about whether the researchers generate the data themselves or use pre-existing data collected by someone else.

Primary data

Primary data is data that is collected by the researchers themselves for their own study – they decide what data they want and go and collect it from participants themselves.

For example, a psychologist could carry out an experiment to investigate whether sleep affects memory. They might recruit participants, give them a memory test and record their scores. These scores would be primary data because the researcher collected them specifically for the study.

Other examples of methods that can produce primary data include:

  • Experiments
  • Questionnaires
  • Interviews
  • Observations
Strengths of primary data:

  • The researcher can collect data that is directly relevant to their research question.
  • The researcher has control over how the data is collected, such as which participants are used and which measures are taken.
  • The data is up to date because it has been collected specifically for the current study.
Weaknesses of primary data:

  • Collecting primary data can be time-consuming, particularly when recruiting participants and collecting large amounts of data.
  • It can be expensive to collect, especially for large-scale studies.
  • The researcher may need to recruit a large or representative sample, which can be difficult and may affect how well the findings can be generalised.

Secondary data

Secondary data is data that was collected by someone else, usually for a different purpose, but is then used by a researcher in their own study.

For example, a psychologist investigating changes in depression rates could use statistics collected by a government organisation or health service rather than collecting the data themselves.

Other examples of sources of secondary data include:

  • Previous psychological research
  • Government statistics
  • Medical or health records
  • Books, journals, online databases, etc.
Strengths of secondary data:

  • It is often quick and inexpensive to obtain because the researcher does not have to collect the data themselves.
  • Researchers can access large amounts of data, particularly when using government statistics or large databases.
  • It allows researchers to investigate past trends or events where collecting the original data would no longer be possible.
Weaknesses of secondary data:

  • The data may not have been collected for the researcher’s specific research question so it may not contain exactly the information they need.
  • The data may be out of date – particularly if the researcher is investigating a changing behaviour or population.
  • The researcher has less control over how the data was collected, so they may not know how reliable or representative it is.

Levels of data


Levels of data describe the different ways that data can be measured and organised. In psychology, there are three main levels of data:

Nominal Ordinal Interval
Categories with no meaningful order or ranking. Data that can ranked in order but the differences between values are not necessarily equal. Numerical data where the differences between values are equal and meaningful.
E.g. participants’ eye colours or occupations E.g. rankings of favourite to least favourite E.g. times in seconds, distance in metres, etc.

Nominal data

Nominal data is data that consists of categories or labels with no meaningful order or ranking. The categories are simply different from each other.

For example, a researcher could record participants’ eye colours:

  • Blue
  • Brown
  • Green

There is no meaningful way to rank these categories as higher or lower than each other, so this is nominal data.

Another example would be from Milgram (1963): Participants either obeyed or disobeyed the experimenter’s instructions. These are nominal categories because there is no meaningful way to rank ‘obeyed’ and ‘disobeyed’ as higher or lower than each other.

Nominal data is usually recorded as frequencies or tallies, i.e. the number of participants or observations in each category.

Strengths of nominal data:

  • It can be used to record categorical information that cannot easily be measured numerically.
  • It is relatively simple to collect and organise.
  • It can be analysed using non-parametric tests, such as the Sign Test and Chi-squared, when the data meets the conditions for these tests.
Weaknesses of nominal data:

  • It provides less information than ordinal or interval data because the categories cannot be ranked.
  • It is not possible to calculate a meaningful mean average from nominal data.

Ordinal data

Ordinal data is data that can be placed in a meaningful order or ranked but the differences between the values are not necessarily equal.

For example, in Lee et al (1997), children were asked to rate how naughty or good a character was using a 7-point numerical rating scale, ranging from very very good to very very naughty:

  • 1 = Very very good
  • 2 = Very good
  • 3 = Good
  • 4 = Neither good nor naughty
  • 5 = Naughty
  • 6 = Very naughty
  • 7 = Very very naughty

This is ordinal data because the scores can be ranked in order. A score of 7 represents more naughtiness than a score of 6, for example. However, we cannot assume that the difference between each point on the scale is exactly equal.

Another example of ordinal data would be asking participants to rank different types of therapy from most to least preferred. If a participant ranked their options like this:

  • 1 = CBT
  • 2 = Counselling
  • 3 = Medication

It doesn’t necessarily mean they prefer CBT over counselling by the same amount they prefer counselling over medication.

Strengths of ordinal data:

  • It provides more information than nominal data because the categories can be ranked.
  • It allows researchers to identify patterns and differences in people’s responses.
  • It can be analysed using statistical tests designed for ordinal data such as Spearman’s Rho or the Mann-Whitney U test.
Weaknesses of ordinal data:

  • The differences between scores are not necessarily equal so calculations such as the mean may not be appropriate.
  • Less precise than interval data – a ranking doesn’t show how large the difference is between two scores.

Interval data

Interval data is precise numerical data where the differences between values are equal and meaningful. For example, the gap between 10 and 20 is the same as the gap between 20 and 30.

For example, in a psychological study, reaction time measured in milliseconds can be treated as interval data. A reaction time of 300 ms is 100 ms longer than 200 ms and the difference between 400 ms and 500 ms is also 100 ms.

Another example would be Grant et al (1998): participants completed a memory test and their scores could be treated as interval data because the difference between scores is meaningful. For example, a score of 80 represents 10 more points than a score of 70 in the same way that 70 is 10 points higher than 60.

Interval data can be analysed using calculations such as the mean and standard deviation, and can be used with parametric statistical tests when the other conditions for those tests are met.

Strengths of interval data:

  • It provides more precise information than nominal or ordinal data.
  • The equal intervals between values allow researchers to calculate statistics such as the mean and standard deviation.
  • It can be analysed using parametric tests, such as the t tests and Pearson’s r, when the data is normally distributed.
Weaknesses of interval data:

  • Psychological concepts such as intelligence or anxiety can be difficult to measure as true interval data.