Tag: scale validation

  • What Is an Acceptable Cronbach’s Alpha for a Dissertation? (2026)

    What Is an Acceptable Cronbach’s Alpha for a Dissertation? (2026)

    There is no single acceptable value. The convention is that alpha of .70 or above is adequate, but that figure comes from Nunnally’s advice for the early stages of research; he recommended .80 for basic research and .90 as a minimum where important decisions rest on individual scores. For an undergraduate dissertation using an established scale, .70 to .95 is the defensible range — and you must interpret the number, not just report it.

    That is the honest answer, and it is more useful than a threshold, because markers rarely penalise a modest alpha that has been discussed intelligently and frequently penalise a high one that has been pasted in without comment. Here is what the coefficient is doing, what it cannot do, and how to write it up.

    What does Cronbach’s alpha actually measure?

    It estimates internal consistency: the extent to which the items in one scale correlate with each other, and so appear to be tapping the same underlying construct. If your six questions about workplace stress genuinely hang together, people who score high on one will tend to score high on the others, and alpha will be high.

    Two things follow that students routinely get wrong. Alpha is a property of the scores from your sample, not a permanent property of the instrument — which is why you report your own alpha even for a well-established scale rather than quoting the original authors’. And alpha is about consistency, not accuracy: a scale can be highly consistent and consistently measure the wrong thing.

    Where did the .70 rule come from, and what did Nunnally really say?

    Almost every dissertation that justifies a threshold cites Jum Nunnally’s Psychometric Theory (2nd edition, 1978). Very few quote him. He wrote, at pages 245–246:

    “what a satisfactory level of reliability is depends on how a measure is being used. In the early stages of research … one saves time and energy by working with instruments that have only modest reliability, for which purpose reliabilities of .70 or higher will suffice. … In contrast to the standards in basic research, in many applied settings a reliability of .80 is not nearly high enough. In basic research, the concern is with the size of correlations and with the differences in means for different experimental treatments, for which purposes a reliability of .80 for the different measures is adequate. In many applied problems, a great deal hinges on the exact score made by a person on a test. … In those applied settings where important decisions are made with respect to specific test scores, a reliability of .90 is the minimum that should be tolerated, and a reliability of .95 should be considered the desirable standard.”

    Read that in full and the familiar rule inverts. Nunnally offers .70 as a concession for exploratory work, treats .80 as the ordinary standard for basic research, and reserves his real severity for applied decisions. The point he was making is that .70 is not usually sufficient and that we should be working to a considerably higher standard most of the time.

    This misreading is not a private observation. Lance, Butts and Michels traced four widely repeated cutoff criteria back to their original sources in Organizational Research Methods (2006, volume 9, issue 2, pages 202–220) and found that the sources did not say what they are routinely cited as saying. Citing “Nunnally (1978)” for a flat .70 threshold is citing a source against its own argument — and it is the kind of thing a well-read marker enjoys pointing out.

    The practical move for your dissertation is not to panic but to be precise: state the value you obtained, say what standard you are judging it against and why that standard fits your purpose, and cite honestly.

    Does a higher alpha always mean a better scale?

    No, and this is the second thing markers look for. Alpha is a function of the average correlation between items and the number of items. Add more items saying roughly the same thing and alpha rises even if the average inter-item correlation stays modest — a point Cortina made directly in Journal of Applied Psychology (1993, volume 78, issue 1, pages 98–104).

    So a twenty-item scale reporting alpha of .92 may be less impressive than a five-item scale reporting .78. And a very high alpha, above roughly .95, is usually a warning rather than a triumph: it suggests item redundancy, that you have asked the same question five times in slightly different words. Redundant items lengthen your questionnaire, increase drop-out, and add nothing.

    Alpha Conventional description What to actually think
    Below .60 Unacceptable Do not compute a total score from these items; investigate why
    .60–.69 Questionable Reportable with discussion; treat findings from the scale cautiously
    .70–.79 Acceptable Fine for exploratory undergraduate work; Nunnally’s floor, not his standard
    .80–.89 Good The ordinary target for a basic-research design
    .90–.94 Excellent Required where decisions about individuals rest on the score
    .95 and above “Better still” Check for redundant items before celebrating

    Descriptive labels like these circulate widely and vary between textbooks. Taber’s review of how alpha is used and described in science education research (Research in Science Education, 2018, volume 48, issue 6, pages 1273–1296) documents just how inconsistently the same numbers get labelled across published studies. Use the table as orientation, and let your discussion, not the adjective, carry the argument.

    A printed Likert-scale questionnaire being completed by a participant
    Alpha describes how your respondents answered these items — not a fixed property of the questionnaire itself.

    Does a good alpha prove my scale measures one thing?

    No. This is the most consequential misunderstanding of the coefficient. Alpha is not a test of unidimensionality, and a multidimensional set of items can produce a perfectly respectable alpha. If your scale has established subscales, compute alpha separately for each subscale as well as for the total, and say so. Demonstrating that items form a single dimension requires factor analysis, not a reliability coefficient.

    Should I delete items to raise my alpha?

    Sometimes, carefully, and always transparently. Your reliability output gives you two diagnostics: the corrected item-total correlation, which is the correlation between each item and the scale score computed without that item, and the alpha-if-item-deleted value. An item with a very low corrected item-total correlation, whose removal raises alpha above the full-scale value, is a genuine candidate for deletion.

    Three cautions. Check first that you reverse-scored every negatively worded item before running the analysis — a forgotten reverse-score is the most common cause of a mysteriously terrible alpha and of one item that looks catastrophic on its own. Do not strip a validated scale down to whatever maximises alpha, because you then no longer have the instrument you cited and cannot claim its published validity evidence. And report every deletion, with the reason and both alpha values, rather than quietly presenting the improved figure.

    What should I do if my alpha is low?

    Diagnose before you despair. Reverse-scoring errors come first. Then check whether one item was ambiguously worded or interpreted differently by your participants, whether the scale was written for a different population from yours, and whether your sample is simply small — alpha estimated from thirty responses is unstable, and your justification for that number should already be in your methods, as we set out in our guide to sample size for an undergraduate dissertation.

    If it stays low, report it and discuss it. Low reliability attenuates correlations, pulling them towards zero, which means it makes you less likely to find a significant relationship rather than more — so a low alpha alongside a non-significant result is a limitation worth stating explicitly, because it is a plausible reason the effect did not appear. A limitations section that identifies that mechanism reads as competence. A silently reported .54 reads as something else.

    A published scale's item list being checked against a dissertation questionnaire
    If you shortened or reworded a validated scale, its published reliability evidence no longer transfers — you are reporting on a new instrument.

    Should I use McDonald’s omega instead?

    There is a real methodological argument that you should. Alpha rests on assumptions — notably that all items relate equally strongly to the underlying construct — that real scales frequently violate, and McDonald’s omega relaxes them. Hayes and Coutts made the case directly in a paper titled “Use Omega Rather than Cronbach’s Alpha for Estimating Reliability. But…” (Communication Methods and Measures, 2020, volume 14, issue 1, pages 1–24), and the trailing “But…” is doing real work: they qualify the recommendation rather than issuing it flatly.

    For an undergraduate dissertation the sensible position is that alpha remains the expected convention and is what most UK departments teach. Report omega alongside it if your software gives it to you easily and you can explain what it is; do not substitute a coefficient you cannot define. A marker asking “why omega?” and receiving a confident answer is a good moment. Receiving silence is not.

    How do I report alpha in my dissertation?

    In the methods chapter, name the scale, its source, the number of items and the response format. In the results, give alpha for your own sample to two decimal places, with a leading zero omitted in APA style: α = .84. Report each subscale separately where subscales exist. If you deleted items, say which and why, and give alpha before and after.

    A worked sentence you can adapt: “Internal consistency for the six-item scale was good in the present sample (α = .84), comparable to the .87 reported by the scale’s authors. One item was retained despite a low corrected item-total correlation (.19) because removing it would have departed from the validated instrument; alpha excluding that item would have been .88.”

    That sentence does everything a marker wants: it reports, it compares, it makes a decision, and it justifies the decision. The wider architecture it sits inside — design, sampling, instruments, analysis — is set out in our guide to writing a methodology chapter, and the reliability run itself is a couple of clicks in whichever package your course uses, compared in SPSS vs R vs jamovi.

    Reliability is only half the picture, too. Once you know your scale is consistent, the analysis question is which test its scores belong in — see choosing the right statistical test. If your project turned out to be qualitative instead, reliability coefficients do not transfer at all; the equivalent quality debate is covered in our guide to doing a thematic analysis.

    When the analysis is settled and the writing is the bottleneck, Tesify can structure and draft your dissertation around your own results — every word still written by you, with the structure and bibliography handled.

    Frequently asked questions

    Is 0.7 a good Cronbach’s alpha?

    It is conventionally described as acceptable, and it is adequate for exploratory undergraduate work. But .70 was Nunnally’s figure for the early stages of research, not his general standard — he treated .80 as adequate for basic research. Report .70 with a sentence of discussion rather than presenting it as a pass mark.

    What if my Cronbach’s alpha is 0.6?

    Report it, investigate it and discuss it. Check reverse-scoring first, then item wording and sample size. You can still use the scale if you are explicit about the limitation and cautious in your conclusions, and the attenuation argument gives you something intelligent to say about any non-significant results.

    Can Cronbach’s alpha be too high?

    Yes. Above about .95 the usual cause is redundant items asking the same question repeatedly. That is a design weakness rather than a strength, and it is worth a line in your discussion.

    Do I report alpha for the whole scale or each subscale?

    Both, when the instrument has established subscales. A respectable total-scale alpha can hide a weak subscale, and reporting only the total looks like concealment even when it is not.

    Do I need to calculate alpha if I used a published validated scale?

    Yes. Alpha describes your data, not the instrument in the abstract, and reliability varies by sample and population. Report your own value and compare it with the published one — a sentence that does both is stronger than either alone.

    Can I use Cronbach’s alpha with a two-item scale?

    It is not informative with two items. Report the correlation between the two items instead, and say why. Alpha’s dependence on item count makes it a poor summary for very short scales.

    Does alpha apply to a scale I wrote myself?

    You can compute it, but a good alpha on a self-written scale demonstrates only internal consistency, not that the scale measures what you claim. Expect a marker to ask about validity, and prefer an established instrument wherever one exists.

    Is Cronbach’s alpha the same as validity?

    No. Reliability is consistency; validity is whether the instrument measures the intended construct. A scale can be reliably wrong. Your methods chapter should address both, and they need different evidence.

    Do I need alpha for a single-item measure?

    No — internal consistency has no meaning for one item. Single-item measures are acceptable for some constructs, such as a straightforward demographic or a global rating, but you should say why one item was sufficient.

    How do I write alpha in APA style?

    Use the Greek letter with no leading zero, italicised, to two decimal places: α = .84. Put it in the results text or in a table of scale statistics, and keep the format consistent throughout the dissertation.