To test validity and reliability of a questionnaire you need two different sets of evidence. Reliability asks whether the instrument gives stable, consistent answers across items, occasions and raters; validity asks whether it actually measures the construct it claims to measure. Both are established from a pilot administration you analyse before you collect the main data, never from the full study afterwards.
Most of the difficulty is not the maths. It is choosing the right test for the design you have, running it in the right order, and knowing what a disappointing number means when you get one. This guide walks through that whole process, including the SPSS click-paths, the thresholds worth using, and the mistakes that get questioned by supervisors and reviewers.
Budget a couple of working days end to end. The pilot itself takes longer, but the analysis and revision cycle is where most of the effort sits.
Table of Contents
- 1What You Need
- 2Step-by-Step: How to Test Validity and Reliability of a Questionnaire
- 3Step 1: How to Test Validity and Reliability of a Questionnaire by Defining Its Purpose
- 4Step 2: Review Item Wording, Content and Structure
- 5Step 3: Pilot the Questionnaire and Examine Response Quality
- 6Step 4: Measure Internal Consistency
- 7Step 5: Check Construct, Convergent and Discriminant Validity
- 8Step 6: Assess Criterion and Test-Retest Evidence
- 9Step 7: Report Results and Make a Decision
- 10Common Mistakes
- 11Frequently Asked Questions
- 12What sample size is needed to test a questionnaire’s reliability and validity?
- 13Is Cronbach’s alpha above 0.70 always acceptable?
- 14How do I choose the right validity test for my questionnaire?
- 15What should I do if Cronbach’s alpha or McDonald’s omega is low?
- 16Do factor analysis and Cronbach’s alpha test the same thing?
- 17Can a questionnaire be reliable without being valid?
- 18Conclusion
What You Need

Gather these before you open any software. Missing one of them is what turns the analysis into guesswork.
- The questionnaire draft with every item numbered and its response scale written out. Include the instructions given to respondents.
- A construct definition sheet: for each variable, the definition, the dimensions, and the theoretical source. If you cannot write one sentence per construct, you are not ready to test anything.
- Target respondent profile and a sampling plan for the pilot, including how you will reach people who match the main study.
- Pilot size target, calculated from item count rather than guesswork (covered in Step 3).
- Prior research: validated questionnaires on the same topic, so you know what a good result looks like for this population.
- Statistical software. SPSS and JASP handle reliability analysis and exploratory factor analysis. Jamovi does both from a spreadsheet-style interface. R covers everything through the
psychandlavaanpackages. AMOS or its alternatives handle confirmatory factor analysis, and SmartPLS covers partial least squares structural equation modelling. - A reporting sheet — one row per scale or construct, with columns for alpha or omega, item count, validity evidence and any action taken.
Also decide now whether your design is exploratory, confirmatory, or model-based, because that choice determines which tests you run in Steps 4 and 5.
Step-by-Step: How to Test Validity and Reliability of a Questionnaire
The process runs in a fixed order: define what you are measuring, check the instrument reads the way you think it does, pilot it, measure internal consistency, gather validity evidence, add stability and criterion evidence where they fit, then report. Skipping ahead produces numbers that look impressive and mean very little.
No single coefficient makes a questionnaire valid and reliable. Alpha of 0.85 is evidence of consistency, nothing more. Validity is an accumulation of argument: content, structure, relationships with other measures, and agreement with external outcomes.
Step 1: How to Test Validity and Reliability of a Questionnaire by Defining Its Purpose
Start with the interpretations you intend to make. Write down each construct, its dimensions, the population it applies to, and the minimum measurement quality your institution or journal will accept.
Then map your design to the evidence you need. Students routinely run the wrong battery because they never made this decision explicitly.
| Your study design | Reliability evidence | Validity evidence |
|---|---|---|
| New or adapted questionnaire, thesis project | Internal consistency; test-retest if the construct is stable | Expert panel and content validity index; exploratory factor analysis |
| Confirmatory study of an existing model | Cronbach’s alpha and composite reliability per construct | Confirmatory factor analysis with fit indices; convergent and discriminant validity |
| PLS-SEM or path model | Cronbach’s alpha, composite reliability | Outer loadings, AVE, Fornell-Larcker and HTMT |
| Open-ended items or qualitative instrument | Inter-rater agreement where coding is involved | Content validity, face validity, member checking, expert review |
| Objective or performance measure rated by two people | Inter-rater reliability (Cohen’s kappa or weighted kappa) | Criterion validity against an external outcome |
Set the acceptance rules before you see the data. Writing down “alpha of at least 0.70 per scale, loadings above 0.40, no cross-loading above 0.32” keeps you from quietly picking the friendlier threshold afterwards.
Step 2: Review Item Wording, Content and Structure
Review the instrument before anyone fills it in. Reliability statistics tell you whether items hang together; they cannot tell you that an item asked the wrong question.
Run an expert panel of three to seven people who know the construct. Give them the item definitions and ask them to judge each item for relevance and clarity on a four-point scale. The proportion of experts rating an item 3 or 4 gives you the item content validity index, or I-CVI; the average across items for a scale is the scale level CVI. Common practice treats I-CVI of 0.78 or above as acceptable and a scale-level CVI of 0.90 or above as good, based on the Lynn (1986) calculation for panels of this size. Revise or drop anything that fails and ask the panel to re-judge.
Then check the mechanics:
- Construct alignment. Every item maps to one dimension and one only. Build a two-column table of item and dimension and look for orphans.
- Scale consistency. Mixing a 5-point Likert block with a 7-point block in the same scale changes the statistics. Keep one response format per scale.
- Reverse-coded items. Fix reverse-coded items in the data preparation step, before any reliability run, or alpha comes out misleadingly low.
- Question order. Leading questions, double-barrelled items such as “how satisfied and how fast are the service”, and absolute terms like “always” all distort responses.
- Missing-value rules. Decide whether “not applicable” is missing and how it will be handled, so the cleaning step is not improvised later.
Face validity is the quickest check in this whole guide: hand the draft to five people outside the study and ask them what the questionnaire is measuring. If they describe different constructs, no statistic will rescue it.
Step 3: Pilot the Questionnaire and Examine Response Quality
Pilot with 30 to 50 respondents drawn from the target population, or about 5 to 10 respondents per item, whichever is larger. A pilot of 15 people produces an alpha that swings wildly between runs, and the number tells you nothing stable.
Ask pilot participants to think aloud while they answer. That single technique surfaces more wording problems than any amount of silent reviewing.
Before running statistics, clean and inspect:
- Completion time. Compare the shortest time against the median. A batch of responses far faster than the rest is usually straightlining, not enthusiasm.
- Item distributions. An item where 90% of respondents pick the same box carries almost no variance and will drag alpha down.
- Missing patterns. Systematic missingness on one item is a clue that the wording or the response option was confusing.
- Attention and duplicate checks. Look for identical response strings across long item blocks and for repeated records.
Revise the wording based on what you found, then run the statistics on the cleaned pilot data. Note every change you made, because your methods chapter has to account for the final version of the instrument.
Step 4: Measure Internal Consistency
Internal consistency asks whether the items of a scale behave as though they measure one thing. You compute it separately for each scale, not for the whole questionnaire pooled.
Cronbach’s alpha remains the default. In SPSS the path is Analyze, then Scale, then Reliability Analysis; move the items into the Items box, set Model to Alpha, and under Statistics tick Item, Scale and Scale if item deleted. JASP has a Reliability module in the Descriptives ribbon, and Jamovi does the same under Reliability. In R, psych::alpha() returns the same table.
McDonald’s omega is the better choice when your items are congeneric but not tau-equivalent, which is common with Likert scales. It needs an exploratory factor analysis first so R can estimate the item loadings. Report it alongside alpha rather than replacing one with the other without comment.
| Coefficient | Value | Reading |
|---|---|---|
| Cronbach’s alpha | 0.90 and above | Items may be redundant; check for overlap before celebrating |
| Cronbach’s alpha | 0.70 to 0.89 | Acceptable for most basic research and thesis work |
| Cronbach’s alpha | 0.60 to 0.69 | Questionable; only defensible for short scales or exploratory work |
| Cronbach’s alpha | below 0.60 | Poor; the items do not behave as one scale |
| McDonald’s omega | 0.70 and above | Acceptable |
| Corrected item-total correlation | 0.30 and above | Item contributes to the scale; below 0.30 needs investigation |
Read the item statistics, not just the total. A corrected item-total correlation below 0.30 means the item does not correlate well with the rest of the scale, which usually points to unclear wording rather than to a genuinely different trait.
On the “Alpha if item deleted” column: treat it as a diagnostic prompt, not an instruction. If deleting item 7 raises alpha from 0.68 to 0.74, ask what item 7 is doing. Sometimes it is worded inconsistently, which is worth fixing. If it genuinely measures the dimension you care about, deleting it narrows your instrument to make a number look better, and reviewers notice that pattern. Never delete items automatically, and never delete an item that carries important content coverage.
Step 5: Check Construct, Convergent and Discriminant Validity
Construct validity asks whether the instrument behaves the way your theory says it should. The main tools are factor analysis for structure, and correlation patterns against related measures for convergent and discriminant validity.
Run the prerequisites first. In SPSS: Analyze, then Descriptive Statistics, then Explore; open Statistics and tick KMO and Bartlett’s test of sphericity. The Kaiser-Meyer-Olkin measure should reach at least 0.60, and 0.80 is comfortable. Bartlett’s test should come out significant, usually reported as p below 0.05. A low KMO means your items overlap too little to share common factors, and factor analysis will not give you a structure you can defend.
Then choose the right analysis:
- Exploratory factor analysis (EFA) asks what structure exists in the data. It suits a new instrument. In SPSS use Analyze, then Dimension Reduction, then Factor, with principal components or principal axis factoring and varimax or oblique rotation. Judge the solution by the scree plot, eigenvalues above 1, and the percentage of variance explained.
- Confirmatory factor analysis (CFA) tests a structure you have already specified. It belongs in a second, independent sample or a later study, not in the same data you used to discover the structure.
| Validity check | Acceptable | Notes |
|---|---|---|
| Factor loading (standardised) | 0.40 and above, 0.70 preferable | Below 0.40 the item is a poor measure of its factor |
| Cross-loading | below 0.32, or 0.45 at the widest | A high cross-loading signals items that belong to two factors |
| KMO sampling adequacy | 0.60 minimum, 0.80 preferred | Below 0.60, do not run factor analysis as specified |
| Bartlett’s test of sphericity | p below 0.05 | Confirms the correlation matrix is factorable |
| Variance explained | 50 to 60% or more across the retained factors | Watch for a first factor that swallows everything |
| Convergent validity | loadings above 0.70 and AVE above 0.50 | Related measures and indicators agree |
| Discriminant validity | HTMT below 0.85, Fornell-Larcker satisfied | Distinct constructs do not overlap more than they should |
Convergent validity means the measure behaves like other measures of the same thing. Check it by correlating your scale against an established instrument: high positive correlation with a similar construct supports it, while a high correlation with something unrelated is a warning sign about your definitions.
Discriminant validity means separate constructs stay separate. Look at the correlation matrix between your scales: two constructs correlated at 0.90 are probably not two constructs. In PLS-SEM, compute AVE per construct, compare the square root of AVE against inter-construct correlations for the Fornell-Larcker criterion, and check the HTMT ratio, which should stay below 0.85 and below 0.90 under stricter expectations. Loadings should exceed 0.70 with composite reliability above 0.70.
Step 6: Assess Criterion and Test-Retest Evidence
Criterion validity tests your scores against something external. Concurrent validity compares them with a gold-standard measure collected at the same time, such as a validated instrument, an administrative record or a clinical assessment. Predictive validity checks whether the scores predict a later outcome. Known-groups validity checks whether the instrument separates groups that should differ on the construct, such as novices from experts.
Test-retest reliability estimates stability. Administer the same questionnaire to the same participants twice, then compute the intraclass correlation coefficient rather than a plain Pearson correlation, since ICC handles both the size of the differences and the agreement between the two runs. Choose the interval to match the construct: one to two weeks for a stable trait such as job satisfaction, a few days for attitudes that move quickly. Keep the conditions identical, because a changed setting, a different administration mode or a different respondent state will show up as instability.
Interpret ICC bands as a guide rather than law. Below 0.50 is poor, 0.50 to 0.75 is moderate, 0.75 to 0.90 is good, and 0.90 and above is excellent. If you also need agreement between two raters scoring the same responses, use Cohen’s kappa for two raters or weighted kappa for ordered ratings; above 0.80 is strong agreement, 0.60 to 0.80 is moderate.
Step 7: Report Results and Make a Decision
Report in a fixed sequence so a reader can check your reasoning: sample size and how respondents were selected, which items were revised after the pilot and why, reliability coefficients per scale with item counts, the validity evidence with its statistics, and the limitations.
A methods paragraph you can adapt, with the placeholders filled in:
A pilot study was conducted with [n] participants drawn from [population]. Participants completed the [n]-item questionnaire, and internal consistency was assessed using Cronbach’s alpha. Alpha values ranged from [lowest] to [highest] across the [number] subscales. Exploratory factor analysis with principal axis factoring and varimax rotation yielded a [number]-factor solution explaining [x]% of the variance; the Kaiser-Meyer-Olkin measure was [value] and Bartlett’s test of sphericity was significant, p < .05. Factor loadings ranged from [lowest] to [highest]. Convergent validity was supported by the correlation of [scale] with [comparison instrument], r = [value], p < .01. Test-retest reliability across a [interval] interval was [coefficient]. Based on this evidence the instrument was [judgement] for use in the main study, with the following limitations noted.
Judge the battery, not the number. An alpha of 0.74 with clean factor structure, expert-confirmed content coverage and sensible convergent evidence is a stronger instrument than one with alpha of 0.88 and cross-loadings everywhere.
Common Mistakes
Treating reliability as validity. A consistent instrument can measure the wrong thing. Fix: report reliability and validity separately, and back validity with content, structural and criterion evidence.
Applying one universal cutoff. A threshold of 0.70 suits exploratory thesis work, but exploratory studies and some applied research tolerate lower values, and an alpha of 0.95 usually signals overlapping items. Fix: state your rationale for the threshold, and justify it against published instruments in your area.
Testing a combined scale. If your questionnaire mixes three unrelated dimensions into one score, the low alpha is the correct result. Fix: run alpha per intended subscale, never across pooled items from different constructs.
Deleting items to raise alpha. Chasing 0.70 by removing items removes content and invites criticism. Fix: investigate low items for wording faults, keep a decision log, and only drop an item with a stated substantive reason.
Reporting a test the design cannot support. Confirmatory factor analysis on the same data that produced your exploratory solution, or test-retest separated by six months on an attitude measure, will not survive review. Fix: match each test to the design, as in the Step 1 table.
Claiming the instrument is fully valid. Validity evidence supports particular interpretations for particular populations. Fix: write “the evidence supports using this scale to measure X in population Y”, and name the limitations.
Two habits help more than anything else. Run the tests on pilot data, before the main collection, so you can still change the instrument. And keep every output, including the messy ones, because the exclusions you made are part of the argument.
Frequently Asked Questions
What sample size is needed to test a questionnaire’s reliability and validity?
For a pilot, aim for 30 to 50 respondents drawn from your target population, or roughly 5 to 10 respondents per item, whichever is larger. That range gives stable Cronbach’s alpha and enough cases for a first factor analysis, though a factor analysis of any real complexity needs several hundred. Five is the common floor per item for exploratory work, and for confirmatory factor analysis you need far more than a pilot: typically ten to twenty cases per parameter. Treat pilot numbers as a testing tool, not as evidence for a final claim.
Is Cronbach’s alpha above 0.70 always acceptable?
No. Alpha of 0.70 or above is a reasonable target for basic research and thesis work, but it is not a guarantee. An alpha above 0.90 often means items overlap heavily rather than that the scale is excellent. Some applied and exploratory research tolerates 0.60 to 0.69, and reviewers in your field may expect 0.80. Always state which threshold you used and why, and report the item count alongside the coefficient, since alpha rises mechanically as items are added.
How do I choose the right validity test for my questionnaire?
Match the test to the evidence your design can produce. New instruments need content validity from an expert panel and exploratory factor analysis. Studies testing an existing model need confirmatory factor analysis with fit indices. PLS-SEM work adds outer loadings, AVE, Fornell-Larcker and HTMT. Tests against outcomes or records call for concurrent, predictive or known-groups criterion validity. Avoid confirmatory analysis on the same data you used to explore the structure, since the fit will be optimistic.
What should I do if Cronbach’s alpha or McDonald’s omega is low?
Check the items before touching the model. Confirm reverse-coded items were recoded correctly, look for items with low corrected item-total correlations, and check whether one scale was accidentally pooled with items from another dimension. Rewording inconsistent items and running the analysis on a larger sample often fixes the number on its own. If the items genuinely measure different things, split the scale rather than deleting items to chase a target value, and record the reasoning.
Do factor analysis and Cronbach’s alpha test the same thing?
No. Alpha measures internal consistency: whether items of a scale hang together as one set. Factor analysis examines structure: whether the items group into factors and how many. They answer different questions and can disagree. A single scale can produce a high alpha and still emerge as two factors, because items may load together while measuring two things at once. Report both, run the KMO and Bartlett’s prerequisites first, and treat factor analysis as the stronger evidence about what your instrument actually contains.
Can a questionnaire be reliable without being valid?
Yes, and it is a common problem. A set of items can be asked consistently by everyone, week after week, while measuring the wrong construct entirely. Reliability is a necessary but not sufficient condition for validity: an instrument cannot be valid without being reasonably reliable, but reliability alone guarantees nothing about what is being measured. Validity is supported by accumulating evidence: expert content review, factor structure, correlation with related measures, and agreement with external outcomes.
Conclusion
Start by writing down what each construct means and which decisions the scores will support. Align every item to that definition, run an expert review, and pilot the questionnaire with 30 to 50 people from the target population before you collect anything you cannot change.
From the pilot, compute Cronbach’s alpha or McDonald’s omega for each scale separately, check the KMO and Bartlett’s prerequisites before factor analysis, and add content, convergent, discriminant or criterion evidence that fits your design. Report every coefficient, every item you removed and why, alongside the limitations.
The sequence is the answer to how to test validity and reliability of a questionnaire without overclaiming: define the construct, align the items, pilot it, then run only the checks your design can actually support. Getting that order right is what turns a pile of SPSS output into evidence you can defend.


