Skip to main content

How to Run a Factor Analysis (EFA and CFA) — A Scale Validity Guide

Factor analysis is the validity analysis that reveals which latent constructs the items of a scale actually measure, and it comes in two forms: exploratory factor analysis (EFA), which discovers the dimensional structure from the data, and confirmatory factor analysis (CFA), which tests whether a pre-specified structure fits the data. If you are developing or adapting a scale, you run EFA first and then CFA on a separate sample; if the scale's dimensional structure is already established in the literature, CFA alone is enough. Before running EFA, check that the data are suitable: KMO should be at least 0.60 (0.80 or above is good) and Bartlett's test of sphericity should be significant (p < 0.05); expect 5-10 respondents per item and typically 200+ in total. Items should load at 0.40 or above, and cross-loading items — those loading on two factors with less than 0.10 difference — need a documented judgement call.

This guide walks through the steps of EFA and CFA, the decision criteria at each stage, and the points reviewers most often flag. If you are running a scale development or adaptation study, send us your data: we handle everything from suitability checks to factor retention, rotation choice, item removal decisions and CFA fit indices, and deliver tables and a path diagram you can paste straight into your thesis or manuscript. The initial review is free.

Who is this guide for?

  • Master's and doctoral students developing a scale or adapting one to a new language

  • Researchers who need to establish which dimensions their questionnaire items fall into

  • Authors whose reviewers noted that construct validity wasn't shown or CFA fit indices weren't reported

  • Anyone unsure whether their study calls for EFA or CFA

  • Researchers dealing with scattered loadings and cross-loading items

Exploratory factor analysis (EFA), step by step

EFA derives the dimensional structure from the data. The order of the steps — and the justification recorded at each one — is what makes the result defensible:

  1. 01

    Check that the data are suitable

    The KMO measure of sampling adequacy should be at least 0.60 (0.80 and above is considered good), and Bartlett's test of sphericity should be significant (p < 0.05), showing there is enough inter-item correlation to factor. Proceeding without both is not meaningful.

  2. 02

    Choose the extraction method

    If your aim is to model a latent construct, use principal axis factoring or maximum likelihood; if it is to reduce the data to fewer components, principal component analysis (PCA) applies. For scale validity in the social sciences, maximum likelihood and principal axis factoring are the theoretically appropriate choices; PCA is not technically factor analysis, and if you use it anyway — as is common in the literature — that choice should be stated explicitly.

  3. 03

    Decide how many factors to retain

    Don't rely on the 'eigenvalue greater than 1' rule alone; it systematically over-extracts. Read the scree plot, parallel analysis and — above all — the theoretical interpretability of the dimensions together. Parallel analysis is currently regarded as the most reliable criterion.

  4. 04

    Choose a rotation method

    Use orthogonal rotation (varimax) if you expect the factors to be uncorrelated, and oblique rotation (oblimin, promax) if you expect them to correlate. In the social and health sciences, dimensions usually do correlate in practice, so inspect the inter-factor correlations and decide rather than defaulting to varimax.

  5. 05

    Make the item-removal decisions

    Items loading below 0.40, and cross-loading items whose two loadings differ by less than 0.10, are removal candidates. But removal is not a purely numerical decision: assess whether the item is needed for the scale's content coverage, and report the reasoning. Re-run the analysis after each removal.

  6. 06

    Evaluate variance explained and name the factors

    For multi-factor scales, total variance explained above 50% is the commonly accepted benchmark; lower proportions can be acceptable in unidimensional structures. Finally, name each factor theoretically, based on the shared meaning of the items it contains.

  7. 07

    Add reliability

    Compute Cronbach's alpha (or McDonald's omega) for each dimension that emerges. A scale study is incomplete unless validity and reliability are reported together — see our scale reliability analysis guide for the details.

EFA or CFA? What separates them

The two are not alternatives — they are two stages of the same scale validity process:

CriterionExploratory factor analysis (EFA)Confirmatory factor analysis (CFA)
PurposeDiscover the dimensional structure from the dataTest a pre-specified structure
WhenNew scale development, item pool of unknown structureAdaptation studies, structures established in the literature
Theoretical inputItems are not assigned to factors in advanceItem-to-factor assignment is specified up front
Key outputsKMO, Bartlett, factor loadings, variance explainedFit indices (χ²/df, RMSEA, CFI, TLI, SRMR), standardised loadings
Sample5-10 respondents per item, typically 200+Typically 200+; more for complex models
SequenceComes first in scale developmentAfter EFA, preferably on a separate sample

EFA decision criteria at a glance

All of these values are expected in the write-up; borderline cases require an explicit justification:

CriterionAcceptable valueNote
KMO sampling adequacy≥ 0.60 (0.80+ good)Below 0.50: factor analysis is not appropriate
Bartlett's test of sphericityp < 0.05Non-significant means inter-item correlation is insufficient
Sample size5-10 per item; generally ≥ 200Smaller samples are defensible when loadings are high and factors clean
Factor loading≥ 0.40 (0.30 as a floor)Higher loadings mean the item represents the factor strongly
Cross-loadingGap between two loadings ≥ 0.10A small gap leaves the item's dimension ambiguous
Total variance explained> 50% in multi-factor structuresLower is acceptable for unidimensional scales
Factor retention criterionParallel analysis + scree plot + theory'Eigenvalue > 1' alone over-extracts

CFA fit indices and thresholds

No single index decides model fit in CFA; they are reported and read together:

IndexGood fitAcceptable fit
χ² / df (chi-square / degrees of freedom)≤ 3≤ 5
RMSEA≤ 0.05≤ 0.08
SRMR≤ 0.05≤ 0.08
CFI≥ 0.95≥ 0.90
TLI (NNFI)≥ 0.95≥ 0.90
Standardised factor loadings≥ 0.70≥ 0.50 (significant, p < 0.05)

Common mistakes

  • Running EFA and CFA on the same sample — confirming a structure on the very data it was derived from is not confirmation; use a separate sample, or split the sample in half if that isn't possible

  • Retaining factors on the 'eigenvalue > 1' rule alone, which systematically produces too many factors

  • Reaching for varimax by reflex; oblique rotation (oblimin/promax) fits better when the dimensions are theoretically expected to correlate

  • Reporting principal component analysis (PCA) as factor analysis; they are different methods and the one used must be stated

  • Adding theoretically ungrounded error covariances chased from modification indices to push fit indices over the line — the intervention reviewers question most

  • Deleting an item purely for a low loading, without ever weighing its contribution to content coverage

  • Reporting a single Cronbach's alpha for the whole scale before the dimensions are settled; alpha is computed per dimension

  • Running CFA on a very small sample (say 60-80 respondents) and presenting the resulting fit indices as conclusive

If you are adapting an existing scale

In adaptation studies the statistics are only the final link in the chain: linguistic equivalence is established first through translation and back-translation, content validity is assessed via expert review, and only then is construct validity tested with factor analysis. When the original scale's dimensional structure is known, CFA tests directly whether that structure holds in the new sample; if it doesn't fit, the structure may genuinely differ in this population, so you return to EFA and report the emerging structure with its justification. That is not a failed study — it is an ordinary and publishable finding in adaptation research.

If you plan to compare groups (does the scale measure the same thing for women and men, or across two age groups?), you need measurement invariance testing; mean comparisons can mislead when equivalence across groups hasn't been demonstrated. Convergent and discriminant validity may also call for additional indices such as AVE and CR. During the free initial review we work out which of these your study genuinely needs — without padding it with analyses that add nothing.

Frequently asked questions

Should I run EFA or CFA?

If you are developing the scale yourself, or its dimensional structure is uncertain, start with exploratory factor analysis (EFA); once the structure emerges, test it with confirmatory factor analysis (CFA), preferably on a separate sample. If you are using or adapting a scale whose structure is established in the literature, CFA alone is sufficient — though returning to EFA when CFA fit indices don't hold is a legitimate and common route.

How large a sample does factor analysis need?

The common rule is 5-10 respondents per item, so a 20-item scale calls for roughly 100-200 people. In practice 200 and above is considered safe, and 300+ good. Sample size alone doesn't settle it: with high loadings (0.60+) and cleanly separated dimensions, smaller samples can be defended, while low, scattered loadings won't be rescued even by a large sample.

What KMO value do I need, and what if it's too low?

KMO should be at least 0.60, with 0.80 and above considered good; below 0.50 the data aren't suitable for factor analysis. When it comes out low, the first place to look is the item-level anti-image correlations: identifying and removing items that correlate weakly with the rest usually raises KMO. If the problem persists, revisit the sample size or the theoretical coherence of the item pool.

What factor loading is high enough, and what is a cross-loading?

A loading of 0.40 or above is the widely used criterion (0.30 can serve as a floor). A cross-loading item loads on more than one factor at similar magnitudes; when the gap between the two loadings is under 0.10, the item's dimension is ambiguous and it becomes a removal candidate. Weigh the item's contribution to the scale's content coverage before removing it, and state the reasoning in the write-up.

Do you use AMOS, LISREL or SmartPLS?

No — we run the analyses on our own Python-based stack. The results are identical to what you'd get from those programs: the same fit indices (χ²/df, RMSEA, CFI, TLI, SRMR), standardised factor loadings and path diagram, laid out in the table format theses and journals expect. You don't need a licence for any of them, and SPSS-compatible reporting is available too.

My CFA fit indices fell short — what should I do?

Start by locating where the model strains: items with low standardised loadings, item pairs that overlap in content, or a dimensional structure that differs from the one specified. Theoretically grounded adjustments — correlating the errors of two near-identical items, for instance — are defensible; chasing modification indices without justification is the fastest thing for a reviewer to catch. If the structure genuinely differs, going back to EFA and reporting the new structure makes for a stronger paper than a forced model.

Let us run your scale validity analysis end to end

Send us your dataset and we'll handle it all — suitability checks, factor retention and rotation decisions, documented item-removal reasoning, CFA fit indices — and deliver tables you can paste straight into your thesis or manuscript. The initial review is free.

Last updated: July 25, 2026