How to Run a Factor Analysis (EFA and CFA) — A Scale Validity Guide
Factor analysis is the validity analysis that reveals which latent constructs the items of a scale actually measure, and it comes in two forms: exploratory factor analysis (EFA), which discovers the dimensional structure from the data, and confirmatory factor analysis (CFA), which tests whether a pre-specified structure fits the data. If you are developing or adapting a scale, you run EFA first and then CFA on a separate sample; if the scale's dimensional structure is already established in the literature, CFA alone is enough. Before running EFA, check that the data are suitable: KMO should be at least 0.60 (0.80 or above is good) and Bartlett's test of sphericity should be significant (p < 0.05); expect 5-10 respondents per item and typically 200+ in total. Items should load at 0.40 or above, and cross-loading items — those loading on two factors with less than 0.10 difference — need a documented judgement call.
This guide walks through the steps of EFA and CFA, the decision criteria at each stage, and the points reviewers most often flag. If you are running a scale development or adaptation study, send us your data: we handle everything from suitability checks to factor retention, rotation choice, item removal decisions and CFA fit indices, and deliver tables and a path diagram you can paste straight into your thesis or manuscript. The initial review is free.
Who is this guide for?
Master's and doctoral students developing a scale or adapting one to a new language
Researchers who need to establish which dimensions their questionnaire items fall into
Authors whose reviewers noted that construct validity wasn't shown or CFA fit indices weren't reported
Anyone unsure whether their study calls for EFA or CFA
Researchers dealing with scattered loadings and cross-loading items
Exploratory factor analysis (EFA), step by step
EFA derives the dimensional structure from the data. The order of the steps — and the justification recorded at each one — is what makes the result defensible:
- 01
Check that the data are suitable
The KMO measure of sampling adequacy should be at least 0.60 (0.80 and above is considered good), and Bartlett's test of sphericity should be significant (p < 0.05), showing there is enough inter-item correlation to factor. Proceeding without both is not meaningful.
- 02
Choose the extraction method
If your aim is to model a latent construct, use principal axis factoring or maximum likelihood; if it is to reduce the data to fewer components, principal component analysis (PCA) applies. For scale validity in the social sciences, maximum likelihood and principal axis factoring are the theoretically appropriate choices; PCA is not technically factor analysis, and if you use it anyway — as is common in the literature — that choice should be stated explicitly.
- 03
Decide how many factors to retain
Don't rely on the 'eigenvalue greater than 1' rule alone; it systematically over-extracts. Read the scree plot, parallel analysis and — above all — the theoretical interpretability of the dimensions together. Parallel analysis is currently regarded as the most reliable criterion.
- 04
Choose a rotation method
Use orthogonal rotation (varimax) if you expect the factors to be uncorrelated, and oblique rotation (oblimin, promax) if you expect them to correlate. In the social and health sciences, dimensions usually do correlate in practice, so inspect the inter-factor correlations and decide rather than defaulting to varimax.
- 05
Make the item-removal decisions
Items loading below 0.40, and cross-loading items whose two loadings differ by less than 0.10, are removal candidates. But removal is not a purely numerical decision: assess whether the item is needed for the scale's content coverage, and report the reasoning. Re-run the analysis after each removal.
- 06
Evaluate variance explained and name the factors
For multi-factor scales, total variance explained above 50% is the commonly accepted benchmark; lower proportions can be acceptable in unidimensional structures. Finally, name each factor theoretically, based on the shared meaning of the items it contains.
- 07
Add reliability
Compute Cronbach's alpha (or McDonald's omega) for each dimension that emerges. A scale study is incomplete unless validity and reliability are reported together — see our scale reliability analysis guide for the details.
EFA or CFA? What separates them
The two are not alternatives — they are two stages of the same scale validity process:
| Criterion | Exploratory factor analysis (EFA) | Confirmatory factor analysis (CFA) |
|---|---|---|
| Purpose | Discover the dimensional structure from the data | Test a pre-specified structure |
| When | New scale development, item pool of unknown structure | Adaptation studies, structures established in the literature |
| Theoretical input | Items are not assigned to factors in advance | Item-to-factor assignment is specified up front |
| Key outputs | KMO, Bartlett, factor loadings, variance explained | Fit indices (χ²/df, RMSEA, CFI, TLI, SRMR), standardised loadings |
| Sample | 5-10 respondents per item, typically 200+ | Typically 200+; more for complex models |
| Sequence | Comes first in scale development | After EFA, preferably on a separate sample |
EFA decision criteria at a glance
All of these values are expected in the write-up; borderline cases require an explicit justification:
| Criterion | Acceptable value | Note |
|---|---|---|
| KMO sampling adequacy | ≥ 0.60 (0.80+ good) | Below 0.50: factor analysis is not appropriate |
| Bartlett's test of sphericity | p < 0.05 | Non-significant means inter-item correlation is insufficient |
| Sample size | 5-10 per item; generally ≥ 200 | Smaller samples are defensible when loadings are high and factors clean |
| Factor loading | ≥ 0.40 (0.30 as a floor) | Higher loadings mean the item represents the factor strongly |
| Cross-loading | Gap between two loadings ≥ 0.10 | A small gap leaves the item's dimension ambiguous |
| Total variance explained | > 50% in multi-factor structures | Lower is acceptable for unidimensional scales |
| Factor retention criterion | Parallel analysis + scree plot + theory | 'Eigenvalue > 1' alone over-extracts |
CFA fit indices and thresholds
No single index decides model fit in CFA; they are reported and read together:
| Index | Good fit | Acceptable fit |
|---|---|---|
| χ² / df (chi-square / degrees of freedom) | ≤ 3 | ≤ 5 |
| RMSEA | ≤ 0.05 | ≤ 0.08 |
| SRMR | ≤ 0.05 | ≤ 0.08 |
| CFI | ≥ 0.95 | ≥ 0.90 |
| TLI (NNFI) | ≥ 0.95 | ≥ 0.90 |
| Standardised factor loadings | ≥ 0.70 | ≥ 0.50 (significant, p < 0.05) |
Common mistakes
Running EFA and CFA on the same sample — confirming a structure on the very data it was derived from is not confirmation; use a separate sample, or split the sample in half if that isn't possible
Retaining factors on the 'eigenvalue > 1' rule alone, which systematically produces too many factors
Reaching for varimax by reflex; oblique rotation (oblimin/promax) fits better when the dimensions are theoretically expected to correlate
Reporting principal component analysis (PCA) as factor analysis; they are different methods and the one used must be stated
Adding theoretically ungrounded error covariances chased from modification indices to push fit indices over the line — the intervention reviewers question most
Deleting an item purely for a low loading, without ever weighing its contribution to content coverage
Reporting a single Cronbach's alpha for the whole scale before the dimensions are settled; alpha is computed per dimension
Running CFA on a very small sample (say 60-80 respondents) and presenting the resulting fit indices as conclusive
If you are adapting an existing scale
In adaptation studies the statistics are only the final link in the chain: linguistic equivalence is established first through translation and back-translation, content validity is assessed via expert review, and only then is construct validity tested with factor analysis. When the original scale's dimensional structure is known, CFA tests directly whether that structure holds in the new sample; if it doesn't fit, the structure may genuinely differ in this population, so you return to EFA and report the emerging structure with its justification. That is not a failed study — it is an ordinary and publishable finding in adaptation research.
If you plan to compare groups (does the scale measure the same thing for women and men, or across two age groups?), you need measurement invariance testing; mean comparisons can mislead when equivalence across groups hasn't been demonstrated. Convergent and discriminant validity may also call for additional indices such as AVE and CR. During the free initial review we work out which of these your study genuinely needs — without padding it with analyses that add nothing.
Frequently asked questions
Should I run EFA or CFA?
If you are developing the scale yourself, or its dimensional structure is uncertain, start with exploratory factor analysis (EFA); once the structure emerges, test it with confirmatory factor analysis (CFA), preferably on a separate sample. If you are using or adapting a scale whose structure is established in the literature, CFA alone is sufficient — though returning to EFA when CFA fit indices don't hold is a legitimate and common route.
How large a sample does factor analysis need?
The common rule is 5-10 respondents per item, so a 20-item scale calls for roughly 100-200 people. In practice 200 and above is considered safe, and 300+ good. Sample size alone doesn't settle it: with high loadings (0.60+) and cleanly separated dimensions, smaller samples can be defended, while low, scattered loadings won't be rescued even by a large sample.
What KMO value do I need, and what if it's too low?
KMO should be at least 0.60, with 0.80 and above considered good; below 0.50 the data aren't suitable for factor analysis. When it comes out low, the first place to look is the item-level anti-image correlations: identifying and removing items that correlate weakly with the rest usually raises KMO. If the problem persists, revisit the sample size or the theoretical coherence of the item pool.
What factor loading is high enough, and what is a cross-loading?
A loading of 0.40 or above is the widely used criterion (0.30 can serve as a floor). A cross-loading item loads on more than one factor at similar magnitudes; when the gap between the two loadings is under 0.10, the item's dimension is ambiguous and it becomes a removal candidate. Weigh the item's contribution to the scale's content coverage before removing it, and state the reasoning in the write-up.
Do you use AMOS, LISREL or SmartPLS?
No — we run the analyses on our own Python-based stack. The results are identical to what you'd get from those programs: the same fit indices (χ²/df, RMSEA, CFI, TLI, SRMR), standardised factor loadings and path diagram, laid out in the table format theses and journals expect. You don't need a licence for any of them, and SPSS-compatible reporting is available too.
My CFA fit indices fell short — what should I do?
Start by locating where the model strains: items with low standardised loadings, item pairs that overlap in content, or a dimensional structure that differs from the one specified. Theoretically grounded adjustments — correlating the errors of two near-identical items, for instance — are defensible; chasing modification indices without justification is the fastest thing for a reviewer to catch. If the structure genuinely differs, going back to EFA and reporting the new structure makes for a stronger paper than a forced model.
Let us run your scale validity analysis end to end
Send us your dataset and we'll handle it all — suitability checks, factor retention and rotation decisions, documented item-removal reasoning, CFA fit indices — and deliver tables you can paste straight into your thesis or manuscript. The initial review is free.
Last updated: July 25, 2026