What Is Questionnaire Reliability?
In measurement theory, reliability refers to the degree to which a measurement instrument produces stable, consistent results across repeated administrations, different raters, or across items within the same construct. A reliable questionnaire behaves like a well-calibrated instrument: administer it today, administer it again next week under the same conditions, and you get substantially the same result.
For medical researchers, reliability is not an abstract psychometric concern it is a fundamental precondition for valid scientific inference. If your Patient Health Questionnaire score varies wildly from one administration to the next without any real change in the patient's condition, no statistical test will rescue your study from the measurement noise you have introduced.
There are four classical forms of reliability, each addressing a different source of measurement inconsistency:
- Internal consistency reliability Do all items within a scale measure the same underlying construct? This is what Cronbach's alpha assesses.
- Test-retest reliability Does the scale produce the same results when administered to the same individuals on two separate occasions? Measured by the intraclass correlation coefficient (ICC) or Pearson r over time.
- Inter-rater reliability Do two independent raters or observers applying the same scale produce the same scores? Measured by Cohen's kappa or ICC.
- Parallel forms reliability Do two different versions of the same scale produce equivalent results? Less commonly used in modern clinical research.
In the vast majority of medical thesis studies using Likert-scale questionnaires, internal consistency is the primary reliability concern, and Cronbach's alpha is the appropriate statistic. Understanding what it measures and what it does not is the first step. For the full analysis pipeline after establishing reliability, see our guide on how to analyze questionnaire data for a medical thesis.
Validity vs Reliability: A Critical Distinction
One of the most persistent conceptual errors in medical research is treating reliability and validity as interchangeable. They measure fundamentally different properties of a research instrument, and a questionnaire can possess one without the other.
Consistency of Measurement
Does the instrument produce the same results under the same conditions? Reliability is about precision the absence of random error.
- Measured by: Cronbach's alpha, ICC, kappa
- Question answered: "Is the scale consistent?"
- A necessary but not sufficient condition for validity
- Can be high even when measuring the wrong thing
Accuracy of Measurement
Does the instrument actually measure what it claims to measure? Validity is about accuracy the absence of systematic error.
- Measured by: content validity, construct validity, criterion validity
- Question answered: "Is the scale measuring the right thing?"
- Cannot exist without reliability (unreliable = automatically invalid)
- Cannot be established by Cronbach's alpha alone
The classic illustration: imagine using a thermometer to measure blood pressure. Your instrument will produce consistent readings every time you use it (high reliability), but it measures the wrong construct entirely (zero validity). Cronbach's alpha will tell you nothing about this problem it will cheerfully report high internal consistency while your study measures something entirely different from what you intend.
Types of Validity You Must Also Assess
- Content validity Do experts in the field agree that the items comprehensively cover the intended construct? Often reported as the Content Validity Index (CVI 0.80).
- Construct validity Does the scale structure match the theoretical model? Assessed via exploratory factor analysis (EFA) or confirmatory factor analysis (CFA).
- Criterion validity Does the scale correlate appropriately with an established gold-standard measure? Split into concurrent validity (measured at the same time) and predictive validity (measured at a later time).
What Is Internal Consistency?
Internal consistency is the degree to which all items in a scale or subscale correlate with one another, reflecting a common underlying construct. The underlying logic is straightforward: if your 10-item "patient satisfaction" scale genuinely measures a single latent construct called "satisfaction," then patients who score high on any one item should tend to score high on all the others. Items that consistently move in the opposite direction from the rest are either measuring something different or are reverse-coded but not recoded before analysis.
Internal consistency is appropriate to evaluate when:
- Your questionnaire contains multiple items intended to measure one construct (or one subscale at a time).
- Items use the same or comparable response format (e.g., 5-point Likert, 010 visual analog scale).
- The scale is being used to produce a composite score (sum or mean of items).
The Cronbach Alpha Formula Explained
Cronbach's alpha (+/-) was introduced by Lee Joseph Cronbach in 1951 as a generalization of earlier reliability coefficients. It is sometimes called the coefficient alpha or the reliability coefficient. The formula is:
The formula captures internal consistency by comparing the sum of individual item variances to the total composite variance. If items are highly correlated with one another, much of the composite variance will be shared (covariance) rather than individual, driving sigma_t2 much higher than 2. The ratio 2/2 will therefore be small, making the term (1 small fraction) close to 1, and alpha will be high.
Conversely, if items are poorly correlated each measuring something different individual variances dominate, 2 approaches 2, the ratio approaches 1, and (1 ratio) approaches 0, resulting in a low alpha.
Key Mathematical Properties
- Alpha ranges theoretically from 0 to 1, though negative values are possible (and always indicate a data problem see Common Mistakes below).
- Alpha is sensitive to the number of items: adding more items, even mediocre ones, will generally increase alpha. This is known as the Spearman-Brown effect and means alpha alone should never drive decisions about scale length.
- Alpha is a lower bound on true reliability the actual reliability is always at least as high as alpha, and often higher when items are tau-equivalent rather than congeneric.
- Alpha assumes that all items contribute equally to the underlying construct (the tau-equivalence assumption). When this assumption is violated as it often is in clinical scales McDonald's omega () is a more accurate reliability estimate.
How to Interpret Cronbach Alpha Values
The most widely cited interpretation framework derives from George and Mallery (2003) and Nunnally (1978), though their original contexts were not specifically medical research. The thresholds below represent the consensus range across medical and health sciences literature:
| Alpha Value | Interpretation | Action |
|---|---|---|
| < 0.60 | Poor Unacceptable | Do not use scale; major revision or item replacement required before data collection |
| 0.60 0.69 | Questionable Marginal | Acceptable only for exploratory studies or newly developed scales; requires improvement |
| 0.70 0.79 | Acceptable Adequate | Satisfactory for most research purposes; minimum standard for most peer-reviewed journals |
| 0.80 0.89 | Good Strong | Ideal range for established clinical scales; suitable for group comparisons and inferential analysis |
| 0.90 0.95 | Excellent | Excellent reliability; acceptable for high-stakes clinical measurement tools |
| > 0.95 | Caution Potential redundancy | Items may be too similar; scale may be unnecessarily long; consider item reduction |
What Is an Acceptable Cronbach Alpha in Medical Research?
The single most common mistake in medical thesis reliability sections is treating alpha thresholds as universal laws rather than context-dependent guidelines. The "correct" acceptable value depends on several factors:
Research Stage and Purpose
- Exploratory or pilot studies developing a new instrument for the first time: alpha 0.60 is acceptable because the scale is still being refined. The goal at this stage is identifying poorly performing items, not demonstrating final reliability.
- Confirmatory studies using a well-established validated instrument (PHQ-9, SF-36, VAS, etc.): alpha 0.80 is expected. Lower values suggest translation or cultural adaptation problems.
- High-stakes clinical decision-making (e.g., a scale used to decide on medication dosing or surgical eligibility): alpha 0.90 should be the minimum standard.
Number of Items in the Scale
A 3-item subscale with alpha = 0.72 demonstrates better item-level performance than a 20-item scale with alpha = 0.82, because alpha rises mechanically with item count. The average inter-item correlation (AIIC) is often a more informative measure for short scales:
Construct Dimensionality
Alpha assumes you are measuring a single underlying construct (unidimensionality). Before interpreting alpha, run an exploratory factor analysis or at least examine a scree plot to confirm your scale is not measuring multiple separate dimensions. If it is multidimensional, divide it into subscales and calculate alpha for each subscale separately.
Cronbach Alpha in Medical Thesis Research: Worked Examples
The following three examples demonstrate how Cronbach's alpha is calculated, reported, and acted on in typical medical thesis contexts.
Context: A postgraduate medical student is developing a 12-item "Patient Satisfaction with Outpatient Services" questionnaire using a 5-point Likert scale (1 = Strongly Disagree to 5 = Strongly Agree). A pilot test is conducted with 40 patients.
Result: Cronbach's alpha = 0.84. The Item-Total Statistics table reveals that Item 7 ("I was given a parking pass upon arrival") has a corrected item-total correlation of 0.11 and an alpha-if-deleted value of 0.87. The student removes Item 7 (which clearly measures a logistics variable, not satisfaction with care), reporting a final alpha of 0.87 for the 11-item scale.
Context: An Egyptian psychiatry resident is validating the Arabic translation of the PHQ-9 in a sample of 150 Type 2 diabetes outpatients. The PHQ-9 measures a single construct (depression symptom severity) across 9 items.
Result: Cronbach's alpha = 0.78. The average inter-item correlation is 0.31. All corrected item-total correlations are above 0.30, and no single item deletion improves alpha by more than 0.02.
Context: A surgery resident uses the 36-item Short Form Health Survey (SF-36) which measures 8 distinct health domains. He calculates a single Cronbach's alpha for all 36 items and obtains +/- = 0.61, which he reports as "questionable reliability."
Result: The supervisor correctly identifies the error. The SF-36 is explicitly multidimensional combining all 36 items violates the unidimensionality assumption. The resident recalculates alpha separately for each of the 8 subscales and obtains values ranging from 0.74 (Social Functioning, 2 items) to 0.92 (Physical Functioning, 10 items).
How to Calculate Cronbach Alpha in SPSS: Step-by-Step
SPSS (Statistical Package for the Social Sciences) remains the most widely used statistical software in medical research and is the standard platform for Cronbach's alpha calculation in clinical thesis work. The following steps apply to SPSS versions 20 through 29 (the menu path is identical).
Prepare Your Dataset
Ensure all questionnaire items are entered as separate numeric columns (e.g., Q1, Q2, Q3 Q12). Reverse-code any negatively worded items before analysis: for a 5-point scale, replace the original value (x) with the recoded value using the formula: 6 x. In SPSS: Transform Recode Into Different Variables.
Open the Reliability Analysis Menu
Navigate to: Analyze Scale Reliability Analysis. The Reliability Analysis dialog box will open.
Move Items Into the Analysis Box
Select all questionnaire items (Q1 through Q12, or whichever items belong to this subscale) from the left variable list and move them into the Items box using the arrow button. Confirm the Model dropdown is set to Alpha (this is the default).
Request Diagnostic Statistics
Click the Statistics button. Under "Descriptives for", check: Item, Scale, and Scale if item deleted. Under "Inter-Item", check Correlations. Click Continue.
Run the Analysis and Read the Output
Click OK. SPSS will generate three key output tables: (1) Reliability Statistics showing the overall Cronbach's alpha and number of items; (2) Item Statistics showing mean and SD per item; (3) Item-Total Statistics the most important table, showing corrected item-total correlations and alpha if each item were deleted.
Interpret the Item-Total Statistics Table
Review the Corrected Item-Total Correlation column. Items with correlations below 0.30 are weakly related to the overall scale and are candidates for review or removal. Review the Cronbach's Alpha if Item Deleted column if removing a specific item would substantially increase alpha (by 0.05 or more), consider removing that item after applying clinical and theoretical judgment.
Reading the Item-Total Statistics Table: A Worked Example
Below is a representative SPSS Item-Total Statistics output for a 7-item "Physician Communication Skills" scale:
| Item | Scale Mean if Item Deleted | Scale Variance if Item Deleted | Corrected Item-Total Correlation | Alpha if Item Deleted |
|---|---|---|---|---|
| Q1 Explained diagnosis clearly | 21.4 | 18.3 | 0.68 | 0.81 |
| Q2 Used understandable language | 21.1 | 17.9 | 0.71 | 0.80 |
| Q3 Answered all questions | 21.6 | 18.7 | 0.65 | 0.82 |
| Q4 Showed empathy and respect | 21.3 | 18.1 | 0.70 | 0.81 |
| Q5 Waiting time was acceptable | 22.8 | 24.6 | 0.14 | 0.89 |
| Q6 Treatment plan was explained | 21.4 | 18.4 | 0.62 | 0.82 |
| Q7 Follow-up instructions were clear | 21.5 | 18.0 | 0.67 | 0.81 |
The interpretation here is straightforward: Q5 ("Waiting time was acceptable") has a corrected item-total correlation of only 0.14 far below the 0.30 threshold and removing it would increase alpha from 0.83 to 0.89. This item clearly measures a logistical/administrative construct (waiting time management) rather than communication skills. After clinical review, the researcher removes Q5 and reports a final alpha of 0.89 for the 6-item scale, documenting the reason for removal in the methods section.
When Alpha Is Too Low or Too High
If Alpha Is Too Low (below 0.70)
A low Cronbach's alpha before data collection is actionable after data collection, your options are more limited. Pre-collection strategies:
- Conduct a comprehensive pilot study (n = 3050) and review all items with poor item-total correlations (< 0.30).
- Have domain experts review item wording for ambiguity, double-barreled phrasing, or cultural irrelevance.
- Verify that all items are measuring the same construct low alpha may correctly identify that your "scale" is actually two or three different instruments merged together.
- Ensure reverse-coded items have been properly recoded before running the analysis.
- Consider whether the sample size for the pilot was adequate (n < 30 can produce unstable alpha estimates).
If Alpha Is Too High (above 0.95)
Paradoxically, an extremely high alpha is also a warning sign. Values above 0.95 typically indicate item redundancy your questionnaire contains items that are phrased so similarly that they are essentially asking the same question multiple times. This inflates the scale length without adding measurement information, increases respondent burden, and can reduce the discriminative power of the scale.
Common Mistakes Researchers Make with Cronbach Alpha
Mistake 1: Calculating One Alpha for a Multidimensional Questionnaire
Reporting a single Cronbach's alpha for a questionnaire with multiple subscales (e.g., SF-36, WHOQOL, or a 3-domain patient experience tool) violates the unidimensionality assumption and produces a meaningless coefficient often artificially low that misrepresents the actual reliability of each subscale.
Mistake 2: Failing to Reverse-Score Negatively Worded Items
Items worded in the opposite direction (e.g., "I was NOT satisfied with the care I received" on a satisfaction scale) must be reverse-coded before analysis. If left uncoded, they produce negative item-total correlations that can drive alpha to near-zero or even negative values not because the scale is bad, but because the coding is wrong.
Mistake 3: Deleting Items Purely to Maximize Alpha
Some researchers systematically delete any item whose removal would increase alpha, without applying theoretical or clinical judgment. This approach can strip the scale of clinically important content and produce a scale that is statistically elegant but clinically impoverished. It also raises serious concerns about data manipulation in peer review.
Mistake 4: Using Alpha to Claim Validity
"The questionnaire demonstrated high validity (+/- = 0.87)" is a statement that appears in published medical papers and theses and it is categorically incorrect. Cronbach's alpha measures internal consistency reliability, not validity. Citing alpha as evidence of validity conflates two separate measurement properties and will draw immediate criticism from any knowledgeable reviewer or examiner.
Mistake 5: Running Cronbach Alpha on a Sample Too Small
Cronbach's alpha is a sample-based estimate. With n < 30, the estimate is highly unstable the confidence interval around alpha can span 0.30 or more, making the reported value essentially uninformative. Many researchers compute alpha from a pilot of 1015 participants and report it as definitive. For cross-sectional studies, the minimum sample size for a reliable pilot should be formally calculated using the StatClinic sample size calculator.
Mistake 6: Not Reporting Which Version of the Scale Was Used
Researchers often adopt existing questionnaires and report their own alpha value without specifying whether they used the original validated version, a translated version, or a modified version. Alpha values are not transferable between populations or translations a questionnaire validated in English with +/- = 0.88 may perform very differently in a new language or cultural context.
How to Report Cronbach Alpha in a Medical Thesis or Paper
Clear, transparent reporting of reliability analysis is essential for academic credibility. The following conventions are standard across medical and health sciences journals:
Example Reporting Language (Methods Section)
Example Reporting Language (Results Section)
Frequently Asked Questions
Need Help with Questionnaire Analysis?
Not sure which statistical test to use for your thesis data? StatClinic's AI Statistical Assistant guides you to the right test in under 2 minutes free, with no registration required.
Try StatClinic AI Statistical Assistant