You show a questionnaire is valid by gathering evidence that it measures what you claim: use a published instrument where one exists, have experts judge the items, and pilot it on people like your participants. You show it is reliable by reporting consistency from your own data, usually Cronbach’s alpha for each scale, and by explaining any adaptation you made.
Markers look for this in the methodology chapter, and they notice when it is missing. This guide explains the difference between validity and reliability, lists the forms of evidence you can realistically collect in an undergraduate or master’s project, and gives a worked example with the arithmetic done for you.
What is the difference between validity and reliability?
Validity asks whether your instrument measures the thing you say it measures. Reliability asks whether it measures that thing consistently. A bathroom scale that always reads two kilograms too heavy is reliable but not valid. You need both, and you need to show evidence for each rather than simply assert it. If you are still turning concepts into measurable variables, start with our guide to the operationalisation table.
Which evidence can you collect, and what should you report?
| Type of evidence | What it shows | How to do it in a student project | What to report |
|---|---|---|---|
| Face validity | The items look sensible to the people answering them | Pilot with a few people from your target group and ask what each item meant to them | Who piloted, what was unclear, what you changed |
| Content validity | The items cover the construct and are relevant to it | Ask a small panel of experts to rate each item’s relevance and compute a content validity index | Number and background of experts, rating scale, item-level and scale-level indices, which method you used for the scale-level index |
| Construct validity | The items behave as the theory predicts, for example loading on the expected factors | Usually cited from the original validation of a published scale; factor analysis needs a large sample | The source of the validation evidence, or your own analysis if your sample allows it |
| Internal consistency | The items in a scale hang together | Cronbach’s alpha for each subscale in your own data | Alpha per scale with the number of items, and a comment on any low value |
| Test-retest reliability | Scores are stable over time | Same people complete the questionnaire twice, a short interval apart, and you correlate the scores | Interval, sample size, and the coefficient used |
| Inter-rater agreement | Two coders apply a coding scheme the same way | Double-code a sample of qualitative data and calculate agreement | Proportion coded twice and the agreement statistic |
Step 1: Use a validated instrument if one exists
The most defensible way to establish validity is to adopt a published instrument and cite the evidence from its original validation. You still have to report reliability from your own data, because reliability belongs to the scores in your sample, not to the scale itself. Check that you may use the scale, as explained in our guide to which psychology scales you can use in a dissertation.
Step 2: Be honest about any adaptation
If you change wording, drop items, shorten a scale, translate it or use it with a new population, you no longer have exactly the validated version. Say what you changed and why, run a pilot, and check reliability again. A translated scale needs a proper forward and back translation, so ask your supervisor whether that is realistic before you commit.
Step 3: Get content validity from experts
- Choose your panel. Invite people with relevant expertise, for example academics who teach the topic or practitioners in the field. Keep a short record of why each is qualified.
- Give them a rating task. For each item, ask experts to rate relevance to the construct on a four-point scale, from not relevant to highly relevant, and invite comments on wording and missing content.
- Calculate the item-level index (I-CVI). This is the proportion of experts who rate the item as relevant, commonly 3 or 4 on a four-point scale.
- Calculate the scale-level index (S-CVI). Polit and Beck (2006) point out that there are two methods: one requires universal agreement among experts, and the other averages the item-level indices. The two can give different values, so state which you used.
- Revise or drop weak items, and report what you changed.
Lynn (1986) is the classic source for quantifying content validity in nursing research, and Polit and Beck (2006) is the standard critique, so cite them if you use the index.

Step 4: Pilot the questionnaire
Pilot with a small group from your target population, not with friends who know your topic. Watch for items they reread, skip or interpret differently from your intention. Time the completion, because a long questionnaire reduces your response rate. Record the changes in a short table, which also serves as evidence of face validity. Our questionnaire and interview templates include a layout you can adapt.
Step 5: Calculate reliability from your own data
For multi-item scales, report Cronbach’s alpha separately for each subscale. Tavakol and Dennick (2011) note that reported acceptable values range from 0.70 to 0.95, that a maximum of 0.90 has been recommended, and that alpha is affected by test length and dimensionality. A very high alpha can suggest redundant items, and a low alpha for a scale with only two or three items is common, so interpret it rather than just comparing it with a cut-off. Our guide to acceptable Cronbach’s alpha values covers this in detail.
A worked example (illustrative)
This example uses invented numbers to show the arithmetic. Suppose you wrote five items to measure study confidence and asked five experts to rate each item’s relevance on a four-point scale, counting 3 or 4 as relevant. Four experts rate item 1 as relevant, five rate item 2, five rate item 3, three rate item 4 and five rate item 5.
| Item | Experts rating relevant | I-CVI (relevant divided by 5) |
|---|---|---|
| 1 | 4 | 0.80 |
| 2 | 5 | 1.00 |
| 3 | 5 | 1.00 |
| 4 | 3 | 0.60 |
| 5 | 5 | 1.00 |
Averaging the item indices gives (0.80 + 1.00 + 1.00 + 0.60 + 1.00) divided by 5, which is 4.40 divided by 5, so S-CVI by averaging is 0.88. Universal agreement requires all five experts to rate an item relevant, which happens for three of the five items, so the universal agreement index is 3 divided by 5, which is 0.60. The two values differ, which is exactly why Polit and Beck advise you to say which method you used. In this example item 4 is the weak item: you would read the experts’ comments, reword it or drop it, and report the change.
What should you do if your dissertation is qualitative?
Qualitative studies do not use alpha or content validity indices. Instead, you demonstrate trustworthiness, commonly through credibility, transferability, dependability and confirmability, following Lincoln and Guba. In practice that means keeping an audit trail of your coding decisions, checking your interpretations with participants or a second coder where appropriate, describing the context in enough detail for readers to judge transferability, and being explicit about your own position. Our guide to thematic analysis shows where these checks fit into the six phases.

How do you write it up in the methodology?
An illustrative paragraph: “The questionnaire used the published scale of the original authors, adapted by changing the context in three items. Five experts rated each item’s relevance on a four-point scale; item-level content validity indices ranged from 0.60 to 1.00, and the scale-level index, calculated by averaging item indices, was 0.88. One item with a low index was reworded following expert comments. A pilot with eight students led to two further changes to wording. In the main sample, Cronbach’s alpha was calculated for each subscale and is reported in Table 3.1.” Replace every detail with what you actually did, and include the expert instructions and pilot log in an appendix. Plan your sample for the main study with our guide to sample size for an undergraduate dissertation.
What if your alpha comes out low?
A low alpha is a finding, not a disaster, provided you diagnose it. Work through these checks in order:
- Check the data entry. Reverse-scored items that were not reversed are the most common cause of a surprisingly low value.
- Check the number of items. Short scales tend to produce lower alpha, so a two- or three-item subscale needs a cautious interpretation.
- Check dimensionality. If your items actually measure two ideas, a single alpha for both will mislead. Report each subscale separately.
- Look at item-total statistics. If dropping one item raises alpha noticeably, consider whether that item is worded badly, and report what you decided.
- Say so in the limitations. Explain what the low value means for the confidence readers can place in that scale’s results.
Whatever you decide, report the original value and the revised value, and do not quietly delete items until the number looks acceptable.
What does a good pilot checklist look like?
- Participants: a handful of people who resemble your target group but will not be in the main sample.
- Timing: record how long the questionnaire takes and compare it with what you promised in the information sheet.
- Wording: ask what each item meant, and note any item that people interpret differently.
- Layout: check the questionnaire on a phone as well as a laptop, since many respondents will use a phone.
- Technology: test the survey link, the consent page and the data export before you send the real invitation.
- Log: keep a table of problems found, changes made and the date of each change.
Treat the pilot as part of your evidence, not as an admin step. A short, honest account of what the pilot changed is more convincing to a marker than a claim that everything worked first time.
What mistakes do markers see?
- Asserting validity without evidence. “The questionnaire was valid” is not a finding.
- Quoting alpha from the original paper only. Report it for your own sample.
- Treating one cut-off as a law. Interpret alpha in light of number of items and dimensionality.
- Changing a validated scale silently. Report every change and re-check reliability.
- Using friends as the pilot group. Use people similar to your participants.
- Applying quantitative tests to qualitative work. Use trustworthiness criteria instead.
Tesify helps you structure the methodology around your instrument, so the validity evidence, pilot and reliability sections stay consistent while you write every word yourself. Over 9,000 students have used it across more than 15,000 chapters. Plan your methodology chapter with Tesify.
Validity and reliability FAQ
What is the difference between validity and reliability in a dissertation?
Validity is whether the questionnaire measures what you claim it measures. Reliability is whether it measures it consistently. You need evidence for both.
How do I show content validity?
Ask a small panel of experts to rate each item’s relevance, calculate item-level and scale-level indices, and report which scale-level method you used.
Do I need to pilot my questionnaire?
Yes, if you wrote or adapted it. Pilot with people like your participants and report the changes you made.
What Cronbach’s alpha is acceptable?
Reported acceptable values range from 0.70 to 0.95, with a maximum of 0.90 recommended in one widely cited paper, but you should interpret alpha alongside the number of items and the dimensionality of the scale.
Can I use a validated scale without further checks?
You can cite the original validation evidence, but you still need to report reliability from your own data and explain any adaptation.
What is a content validity index?
It is the proportion of experts who rate an item as relevant, calculated for each item and then combined into a scale-level figure.
How many experts do I need?
There is no single number, so follow the source you cite and explain your choice. Make sure the panel has relevant expertise and report how they were selected.
What is test-retest reliability?
It is the stability of scores when the same people complete the questionnaire twice with a short gap, summarised by a correlation or agreement statistic.
Does a qualitative dissertation need validity and reliability?
It needs trustworthiness instead, such as credibility, transferability, dependability and confirmability, shown through an audit trail and transparent coding.
Where does the validity and reliability section go?
In the methodology chapter, usually after you describe the instrument and before the data analysis section.
