Data is the information collected by researchers during a study. In psychology, data can take different forms depending on what is being measured and how it is collected. Researchers also need to consider the level of measurement used as this affects how the data can be analysed.
This page covers the key concepts needed to understand and use data in psychology:
- Raw data – the original data collected by researchers, including how it can be recorded, presented, and interpreted.
- Types of data – quantitative vs. qualitative data and primary vs. secondary data.
- Levels of data – nominal, ordinal, and interval data formats.
Raw data
Raw data is the original information collected by researchers before it has been processed or summarised. It could be scores from a psychological test, reaction times, questionnaire responses, the number of participants who showed a particular behaviour – stuff like that.
Raw data recording tables
Researchers often use a raw data recording table to organise the data they collect. A well-designed table makes it easy to record and interpret the results.
For example, if a researcher measures the reaction time of five participants, they could use a table like this:
| Participant | Reaction time (ms) |
| 1 | 420 |
| 2 | 385 |
| 3 | 451 |
| 4 | 397 |
| 5 | 412 |
| 6 | 1200 |
A raw data table should:
- Give each column a clear heading
- Include the units of measurement where appropriate, such as seconds (s) or milliseconds (ms)
- Only contain the raw data – so no calculations or summaries of the results.
Outliers
An outlier is a data value that is unusually high or low compared with the other values.
For example, in the reaction times data set above:
420, 385, 451, 397, 412, 1200
The value 1200 is much higher than the other reaction times, so it could be an outlier.
Outliers can affect calculations (e.g. the mean) because an unusually high or low value can pull the mean towards it. As such, researchers may check for and remove outliers before analysing their data, as this may produce more accurate results. However, an outlier should not automatically be removed. Researchers may then consider why the value is unusual and whether there is a valid reason for excluding it.
Standard form and decimal form
Very large or very small numbers can be written in standard form to make them easier to read and work with. Numbers can also be written in decimal form, which is the ordinary way of writing a number.
| Decimal form | Standard form |
| 3,000,000,000 | 3 x 109 |
| 0.000003 | 3 x 10−6 |
For example, researchers estimate the human brain has around 86 billion (86,000,000,000) neurons and 100 trillion (100,000,000,000,000) synapses. Instead of writing these massive numbers out as decimal forms, you can just say 8.6 x 10¹⁰ and 1 x 10¹⁴.
Significant figures
Significant figures are the digits in a number that show how precise the measurement is.
For example:
- 4.673 to 3 significant figures = 4.67
- 15.28 to 3 significant figures = 15.3
- 0.004826 to 2 significant figures = 0.0048
When rounding to significant figures, look at the digit immediately after the final significant figure:
- If it is 5 or more, round up
- E.g. 4.678 to 3 significant figures = 4.68
- If it is 4 or less, leave the final digit unchanged
- E.g. 4.673 to 3 significant figures = 4.67
Researchers may use significant figures when reporting measurements or calculated results so that the results are not presented with unnecessary precision.
Making estimations from data
Researchers can use the data they collect to make estimations about what might happen in similar situations or in the wider population.
For example, imagine a researcher finds that 18 out of 20 participants followed an instruction in an experiment. This is 90% of the sample. The researcher could use this result to estimate that around 90% of similar participants might follow the instruction under the same conditions.
However, an estimation is not guaranteed to be accurate. The sample may not perfectly represent the wider population, so researchers should treat estimates as approximate rather than exact.
Types of data
Quantitative vs. qualitative
Qualitative vs. quantitative is about whether the data can be expressed and compared in numbers.
Quantitative data
Quantitative data is data that can be measured and expressed as numbers. It is usually collected by measuring or counting something. This data can be analysed using statistical calculations such as calculating the mean, median, or range.
Examples of quantitative data include:
- A participant’s reaction time in milliseconds
- E.g. 397
- The number of words a participant recalls on a memory test
- E.g. 19
- A participant’s age in years
- E.g. 35
Strengths of quantitative data:
|
Weaknesses of quantitative data:
|
Qualitative data
Qualitative data is data that describes experiences, thoughts, feelings, or behaviour using words rather than numbers. It is usually collected through methods such as interviews and observations.
Qualitative data is often more detailed than quantitative data and can give researchers a better understanding of why people think, feel or behave in particular ways.
Examples of qualitative data include:
- A participant’s description of how they felt during an experiment
- A researcher’s written description of a participant’s behaviour
- Responses to an open question such as “how did you feel when completing the task?”
Strengths of qualitative data:
|
Weaknesses of qualitative data:
|
Qualitative data can sometimes be converted into quantitative data.
For example, researchers might ask participants:
“How did you feel during the experiment?”
By placing responses into categories and counting how often each category occurs, participants’ qualitative answers can be converted into quantitative numbers. For example, researchers could group the answers into categories such as ‘happy’, ‘anxious’, and ‘neutral’, and then count how many participants gave an answer in each category like this:
| Response category | Number of participants |
| Happy | 8 |
| Anxious | 7 |
| Neutral | 5 |
The original answers are qualitative data because they describe participants’ feelings in words. Once the researchers have counted the responses in each category, the results are quantitative data because they are expressed as numbers.
This can make the data easier to compare and analyse statistically, although some of the detail from the original qualitative responses may be lost.
Primary vs. secondary
Primary vs. secondary data is about whether the researchers generate the data themselves or use pre-existing data collected by someone else.
Primary data
Primary data is data that is collected by the researchers themselves for their own study – they decide what data they want and go and collect it from participants themselves.
For example, a psychologist could carry out an experiment to investigate whether sleep affects memory. They might recruit participants, give them a memory test and record their scores. These scores would be primary data because the researcher collected them specifically for the study.
Other examples of methods that can produce primary data include:
- Experiments
- Questionnaires
- Interviews
- Observations
Strengths of primary data:
|
Weaknesses of primary data:
|
Secondary data
Secondary data is data that was collected by someone else, usually for a different purpose, but is then used by a researcher in their own study.
For example, a psychologist investigating changes in depression rates could use statistics collected by a government organisation or health service rather than collecting the data themselves.
Other examples of sources of secondary data include:
- Previous psychological research
- Government statistics
- Medical or health records
- Books, journals, online databases, etc.
Strengths of secondary data:
|
Weaknesses of secondary data:
|
Levels of data
Levels of data describe the different ways that data can be measured and organised. In psychology, there are three main levels of data:
| Nominal | Ordinal | Interval |
|---|---|---|
| Categories with no meaningful order or ranking. | Data that can ranked in order but the differences between values are not necessarily equal. | Numerical data where the differences between values are equal and meaningful. |
| E.g. participants’ eye colours or occupations | E.g. rankings of favourite to least favourite | E.g. times in seconds, distance in metres, etc. |
Nominal data
Nominal data is data that consists of categories or labels with no meaningful order or ranking. The categories are simply different from each other.
For example, a researcher could record participants’ eye colours:
- Blue
- Brown
- Green
There is no meaningful way to rank these categories as higher or lower than each other, so this is nominal data.
Another example would be from Milgram (1963): Participants either obeyed or disobeyed the experimenter’s instructions. These are nominal categories because there is no meaningful way to rank ‘obeyed’ and ‘disobeyed’ as higher or lower than each other.
Nominal data is usually recorded as frequencies or tallies, i.e. the number of participants or observations in each category.
Strengths of nominal data:
|
Weaknesses of nominal data:
|
Ordinal data
Ordinal data is data that can be placed in a meaningful order or ranked but the differences between the values are not necessarily equal.
For example, in Lee et al (1997), children were asked to rate how naughty or good a character was using a 7-point numerical rating scale, ranging from very very good to very very naughty:
- 1 = Very very good
- 2 = Very good
- 3 = Good
- 4 = Neither good nor naughty
- 5 = Naughty
- 6 = Very naughty
- 7 = Very very naughty
This is ordinal data because the scores can be ranked in order. A score of 7 represents more naughtiness than a score of 6, for example. However, we cannot assume that the difference between each point on the scale is exactly equal.
Another example of ordinal data would be asking participants to rank different types of therapy from most to least preferred. If a participant ranked their options like this:
- 1 = CBT
- 2 = Counselling
- 3 = Medication
It doesn’t necessarily mean they prefer CBT over counselling by the same amount they prefer counselling over medication.
Strengths of ordinal data:
|
Weaknesses of ordinal data:
|
Interval data
Interval data is precise numerical data where the differences between values are equal and meaningful. For example, the gap between 10 and 20 is the same as the gap between 20 and 30.
For example, in a psychological study, reaction time measured in milliseconds can be treated as interval data. A reaction time of 300 ms is 100 ms longer than 200 ms and the difference between 400 ms and 500 ms is also 100 ms.
Another example would be Grant et al (1998): participants completed a memory test and their scores could be treated as interval data because the difference between scores is meaningful. For example, a score of 80 represents 10 more points than a score of 70 in the same way that 70 is 10 points higher than 60.
Interval data can be analysed using calculations such as the mean and standard deviation, and can be used with parametric statistical tests when the other conditions for those tests are met.
Strengths of interval data:
|
Weaknesses of interval data:
|