JCCES and ACT: A Factor Analysis

3,092 words · 13 min read · 11 references cited

The American College Test (ACT) is one of the two dominant college-admission examinations in the United States, alongside the SAT. Its four sections — English, Mathematics, Reading, and Science Reasoning — collectively claim to measure college readiness across the major academic domains. Whether the ACT also measures something like general cognitive ability — the latent g factor that Spearman (1904) identified as the common variance running through all cognitive tests — is a long-running empirical question, with subsequent work (Deary, Strand, Smith, & Fernandes, 2007; Koenig, Frey, & Detterman, 2008) showing that ACT performance correlates strongly with intelligence-test scores. This study examines that question by factor-analyzing the ACT’s four section scores together with the three subtests of the Jouve Cerebrals Crystallized Educational Scale (JCCES) in the same respondents, and asks not only whether a single factor emerges, but what such a factor can and cannot be taken to mean.

The two batteries and the hypothesis

The JCCES measures crystallized intelligence through three subtests: Verbal Analogies (VA), Mathematical Problems (MP), and General Knowledge (GK). The construct is grounded in Cattell-Horn theory: crystallized intelligence reflects acquired knowledge and skills shaped by education and experience, and the JCCES tasks index its core domains.

The ACT measures college readiness through English (ENG), Mathematics (MATH), Reading (READ), and Science Reasoning (SCIE). Each section has its own content focus, but all four are administered together, in one sitting, on one score scale.

If both batteries tap a common underlying ability — Spearman’s g — then a factor analysis of the seven subtests should recover a single dominant factor on which all seven load substantially. That prediction is easy to state and, as it turns out, easy to satisfy; the harder question, taken up below, is what else could satisfy it. A design with four tasks from one battery and three from another confounds the ability the tasks share with the method they share, and a single factor is the expected result under either.

Method

Fifty-seven participants — high school seniors, college students, and adults — completed the JCCES online under standardized procedures and reported their most recent ACT section scores. ACT scores were self-reported from a prior administration and were not verified against ACT records. Only respondents whose implied testing year fell after the October 1989 introduction of the Enhanced ACT were retained, so that all section scores lie on one comparable scale; complete data on all seven tasks were required, and no other exclusion criterion was applied.

Pearson correlations were computed for all 21 subtest pairs. Sampling adequacy was assessed with the Kaiser-Meyer-Olkin (KMO) measure and Bartlett’s test of sphericity. Internal consistency was assessed with Cronbach’s alpha computed on the correlation matrix: the seven tasks are scored on very different metrics — JCCES raw counts against the ACT’s 1–36 scale, a 2.6-fold spread in standard deviations — and a raw-covariance alpha over such a set measures the metrics as much as the tasks.

Factors were extracted by iterated principal-axis factoring, with squared multiple correlations as initial communalities, following Fabrigar, Wegener, MacCallum, and Strahan’s (1999) recommendations. The number of factors to retain was decided by parallel analysis (Horn, 1965) against 5,000 random data sets of the same size, rather than by the Kaiser (1960) eigenvalue-of-one rule alone; Kaiser’s rule is the least accurate of the common criteria and is defined on the eigenvalues of the unreduced correlation matrix, not on the reduced-matrix eigenvalues a principal-factor extraction produces. Confidence intervals for the loadings come from 2,000 bootstrap resamples of the participants.

Results

Sampling adequacy was good: KMO = 0.818 overall, with per-task values from 0.755 (VA) to 0.851 (SCIE), and Bartlett’s test rejected sphericity, χ²(21) = 266.0, p < .001. Internal consistency across the seven tasks was high, standardized α = 0.905 (the raw-covariance figure, 0.866, is depressed by the metric mismatch and is reported only to make the point that it should not be used here).

All 21 correlations were positive and significant, ranging from 0.39 (MP with READ) to 0.83 (MATH with SCIE) — a textbook positive manifold. The internal structure of that manifold is worth noting before any factor is extracted: the four ACT sections correlate 0.72 with one another on average, the three JCCES subtests 0.57, and the cross-battery pairs 0.51. Tasks from the same battery resemble each other more than tasks from different batteries do, which is exactly the pattern a shared method produces and exactly the pattern a shared ability produces.

How many factors

Principal-axis extraction yielded one large eigenvalue and a series of small ones: 4.312 (61.6% of total variance), 0.642 (9.2%), 0.523 (7.5%), and nothing above 0.16 thereafter. On the unreduced correlation matrix the corresponding values are 4.486, 0.857 and 0.691.

024123ObservedRandom 95thFactorEigenvalue
Figure 1. Only the first factor clears the benchmark from random data of the same size: the observed first eigenvalue (4.312) is more than three times its 95th-percentile benchmark (1.262), while the second (0.642) falls below its own (0.918). One factor is retained.

Parallel analysis is decisive and agrees with the eyeball: the 95th percentile of the first eigenvalue from random data of the same size is 1.262 (reduced matrix) or 1.729 (unreduced), the second is 0.918, and the third is 0.690. The observed first eigenvalue clears its random benchmark by more than threefold; the second falls below it. One factor is retained, on the strictest of the standard criteria as well as on the loosest.

Factor 1 loadings

The single factor loads substantially on all seven tasks (95% bootstrap confidence intervals in brackets):

ACT Science Reasoning0.872 [0.79, 0.93]ACT Mathematics0.868 [0.79, 0.94]ACT English0.800 [0.68, 0.89]ACT Reading0.769 [0.59, 0.89]JCCES Math Problems0.705 [0.50, 0.85]JCCES General Knowledge0.659 [0.52, 0.79]JCCES Verbal Analogies0.648 [0.45, 0.80]00.250.50.751Loading on the common factor (95% CI)
Figure 2. Every task loads substantially on the single factor, but the intervals are wide at this sample size: a loading is pinned to about plus or minus 0.15. The four ACT sections sit above the three JCCES subtests, a gap of 0.157 that reliability cannot explain.
  • JCCES Verbal Analogies (VA): 0.648 [0.45, 0.80]
  • JCCES Mathematical Problems (MP): 0.705 [0.50, 0.85]
  • JCCES General Knowledge (GK): 0.659 [0.52, 0.79]
  • ACT English (ENG): 0.800 [0.68, 0.89]
  • ACT Mathematics (MATH): 0.868 [0.79, 0.94]
  • ACT Reading (READ): 0.769 [0.59, 0.89]
  • ACT Science Reasoning (SCIE): 0.872 [0.79, 0.93]

The mean loading is 0.76, communalities run from 0.42 (VA) to 0.76 (SCIE), and the factor accounts for 58.5% of the total variance in the seven tasks. Those intervals are wide — a subtest loading is estimated to about ±0.15 at this sample size — which is the first thing to keep in mind before reading much into the ordering of individual loadings.

The ACT sections load higher, and the reason matters

The four ACT sections load 0.827 on average, the three JCCES subtests 0.671. That gap of 0.157 is not noise: bootstrapped, it is 95% CI [0.035, 0.276] and positive in 99.6% of resamples.

The tempting reading — the ACT is the better measure of the general factor — is the one the data do not support. Two rival explanations have to be cleared first, and neither can be.

The first is unreliability, which attenuates loadings: a task cannot load above the square root of its reliability. Here it can be bounded rather than assumed. The JCCES subtest reliabilities are published (VA α = .853, MP .914, GK .942), the ACT’s are not entered at all, and the gap can be recomputed under the assumption most generous to this explanation — that the ACT sections are perfectly reliable and every bit of attenuation sits on the JCCES side. Even then the corrected JCCES loadings only rise to 0.706 on average, and a gap of 0.121 survives. Differential reliability cannot account for it.

The second is method variance, and it cannot be cleared, because the design does not allow it. The four ACT sections were taken as one test, in one sitting, on one scale, scored by one organization; the three JCCES subtests likewise. Whatever the four ACT sections share by virtue of being the ACT — format, timing, instructions, the reading load that every section carries, the score-reporting process, and, here, the same act of self-report — enters the common factor along with whatever ability they share, and it inflates their loadings relative to the cross-battery pairs. The within-battery-versus-cross-battery correlation pattern above (0.72 and 0.57 within, 0.51 across) is what that inflation looks like. A factor extracted from four-plus-three tasks cannot separate the two.

What forcing a second factor does

The hypothesis that the two batteries measure separable constructs predicts a two-factor solution splitting along battery lines. Forcing two factors does not produce one. The obliquely rotated solution is improper — Verbal Analogies takes a loading above unity and a communality of 1.03, a Heywood case, the signature of over-extraction — and what structure it does show cuts across the batteries rather than between them: JCCES Mathematical Problems sides with the ACT sections, and the two factors correlate 0.55. Consistent with parallel analysis, there is no second factor here to find, and no bifactor partition of general against specific variance is identified with seven tasks and 57 cases. The honest summary is that this design supports a one-factor description and cannot go further.

What the structure implies

The seven tasks share a great deal: one factor, more than half the total variance, every loading above 0.6. That much is solid, and it is the pattern Spearman’s (1904) general factor predicts, consistent with prior findings that ACT performance correlates strongly with conventional intelligence measures (Koenig, Frey, & Detterman, 2008) and with the wider literature on intelligence and academic achievement (Deary, Strand, Smith, & Fernandes, 2007).

What the factor is remains open, and the analysis cannot close it. The shared variance may be g; it may be crystallized ability specifically, since both batteries are heavily knowledge-loaded; it may be an academic-readiness construct that schooling produces in both; and part of it is certainly method. Disambiguating these requires batteries chosen to pull them apart — a fluid-reasoning measure that does not depend on schooled knowledge, several indicators per domain, and ideally scores that are not self-reported — none of which a two-battery correlational design provides.

The contrast with the JCCES + GAMA pairing, which recovered a two-factor structure, is instructive on exactly this point. GAMA is a nonverbal, figural battery that shares neither content nor format with the JCCES, so the two batteries have little method in common and their shared variance has to be ability. The ACT shares both content and format with the JCCES, and its shared variance therefore cannot be attributed so cleanly. The one-factor result reported here is weaker evidence for a common ability than the two-factor result there is for separable ones.

A multidimensional scaling analysis of these same seven tasks (Jouve, 2016) is informative in the same direction: the tasks arrange themselves by content — verbal-knowledge tasks against quantitative and scientific ones — rather than by the battery they belong to. A single dominant factor and a content-organized configuration are not in conflict. They are two views of the same correlation matrix, and the second says something the first cannot: the shared variance is not so uniform that the instrument of origin is the thing that structures it.

Methodological caveats

Fifty-seven participants is a small sample for factor analysis with seven variables. The KMO is reassuring about sampling adequacy and the bootstrap intervals above are the honest expression of what remains: individual loadings are pinned to about ±0.15, and any claim resting on the ordering of loadings within a battery would be over-reading. The eigenvalue gap is large enough that the one-factor conclusion itself is not in doubt.

The high alpha (0.905) confirms internal consistency but is not independent evidence for one factor. Alpha rises with the number of items and with the average inter-item correlation, and a set with a strong common factor is exactly a set with a high average correlation (Cortina, 1993); it is the same information reported twice, not two findings that agree.

ACT scores were self-reported and unverified. Under a classical model such error would attenuate the ACT’s correlations and loadings — which would mean the true gap in the ACT’s favor is larger still — but the assumption is not safe: respondents may round up, may report a best sitting, or may misreport in ways related to their performance, and error of that kind is neither random nor independent.

No demographic data were analyzed, so measurement invariance across age, education, or background is untested. ACT performance varies with schooling quality and test-preparation access; whether the factor structure holds across those moderators is beyond what this sample can address.

Finally, this is an exploratory analysis, and a confirmatory factor analysis was not run. With seven indicators and 57 cases, a one-factor CFA would have little power to reject a wrong model, and its fit indices would be poorly estimated; the exploratory result, with its intervals, is the more honest summary of what is known.

Implications for practice

For colleges using ACT scores in admissions: the four sections behave, statistically, close to a single dimension, and reporting a composite is defensible. That is a statement about the ACT’s internal structure, not a licence to treat the composite as an intelligence estimate.

For test developers: whether a battery should load on one factor depends on what it is for. A battery meant to measure a unified construct is working when its subtests load uniformly; one meant to measure separable constructs is working when they do not. The JCCES-ACT pairing illustrates the first, the JCCES-GAMA pairing the second — but the comparison is only interpretable because the second pairing does not share a method.

For research on intelligence and academic achievement: the result adds a data point to the case that academic-readiness tests and cognitive batteries share substantial common variance, and a caution about how such data points should be counted. Two batteries, one sitting each, cannot tell a general ability from a shared way of testing for it.

Frequently Asked Questions

How many factors do the JCCES and ACT subtests really have?

One, on this evidence. The first factor has an eigenvalue of 4.312 against a 95th-percentile benchmark of 1.262 from random data of the same size; the second (0.642) falls below its benchmark (0.918). Parallel analysis, the scree pattern, and Kaiser’s rule all agree, and forcing a second factor produces an improper solution.

Why do the ACT sections load higher than the JCCES subtests?

The gap is real (0.157, 95% CI [0.035, 0.276]) but it is not evidence that the ACT measures the general factor better. It is not a reliability artefact — even crediting the ACT with perfect reliability and correcting the JCCES loadings for their published alphas, a gap of 0.121 remains. The likelier contributor is method: the four ACT sections are one test taken in one sitting, and everything they share by being the ACT is folded into the factor along with everything they share by being cognitive tasks.

Does this mean the ACT measures intelligence?

It means the ACT shares a great deal of variance with a crystallized-intelligence battery in this sample. Whether that shared variance is g, crystallized ability, an academic-readiness construct produced by schooling, or partly the two batteries’ own method cannot be settled by a two-battery correlational design. The analysis is consistent with all of them.

How does this compare to the JCCES + GAMA factor analysis?

The JCCES + GAMA factor analysis recovered a two-factor structure — crystallized-verbal against nonverbal-fluid — in respondents who completed both batteries. GAMA shares neither content nor format with the JCCES, so its separation from the JCCES is hard to explain as anything but a difference in ability. The ACT shares both, so its fusion with the JCCES is easier to produce and correspondingly weaker as evidence.

Is a sample of 57 large enough for factor analysis?

It is small. Common rules of thumb ask for several respondents per variable and a few hundred cases for stable solutions, though the sample size actually needed depends on how strong the loadings and communalities are — and here they are strong, which is the condition under which small samples behave best. The bootstrap intervals reported above are the appropriate response: the number of factors is secure, individual loadings are not precise, and replication in a larger sample would sharpen them.

Why is Cronbach’s alpha computed on the correlation matrix?

Because the seven tasks are not on the same scale. JCCES raw scores and ACT 1–36 section scores differ in standard deviation by a factor of 2.6, and a raw-covariance alpha weights each task by its variance — that is, by its metric. The standardized alpha (0.905) weights them equally, which is what a coefficient describing the tasks, rather than their scoring conventions, has to do.

References

Related questions

What is item response theory?

Every time you take a standardized test — an IQ assessment, a college entrance exam, a professional certification — the questions have been calibrated using sophisticated statistical models that most test-takers never learn about. Item Response Theory (IRT) is the mathematical framework behind virtually all modern psychological and educational testing, and understanding its basics illuminates why tests work the way they do. Read more →

What is attenuation-corrected reliability?

Most psychometrics textbooks teach the classical "correction for attenuation" — Spearman's century-old technique for estimating what the correlation between two psychological constructs would be if the tests measuring them were perfectly reliable. The technique is simple: divide the observed correlation by the square root of the product of the two reliabilities. The technique is also limited: it adjusts the relationship between two scales, but assumes the reliability values plugged into the denominator are themselves accurate. A 2022 paper by Jari Metsämuuronen in Applied Psychological Measurement argues that this assumption is broken in practice. Reliability estimates produced by Cronbach's alpha and similar formulas are themselves attenuated by the same mechanical errors that attenuate correlations — and in some datasets, alpha may be deflated by 0.40–0.60 units of reliability. Metsämuuronen's contribution is a class of deflation-corrected reliability estimators that apply the classical attenuation logic inside the reliability formula rather than only to correlations between scales. Read more →

How are JCCES General Knowledge items structured?

The General Knowledge (GK) subtest of the Jouve Cerebrals Crystallized Educational Scale (JCCES) measures factual breadth — the accumulated stock of information about the world that crystallized intelligence theory treats as a core component of acquired cognitive ability. A subtest of this kind has to satisfy two structural requirements: items should span a meaningful range of difficulty, and they should order along a single underlying continuum of factual breadth rather than tapping multiple unrelated dimensions. This study examined the item structure of the JCCES GK subtest using multidimensional scaling (MDS) on response data from 588 respondents, and recovered the empirical signature of a well-ordered unidimensional construct: a horseshoe-shaped scaling pattern that is the canonical evidence of a Guttman simplex — items ordered cleanly along a single difficulty continuum. Read more →

How well does coefficient alpha perform with non-normal data?

Cronbach's coefficient alpha is the most-reported reliability statistic in psychology and educational measurement. It is also one of the most-misunderstood. The classical formula assumes that test items measure a single construct with equal factor loadings (tau-equivalence), uncorrelated errors, and continuously distributed scores. Real psychological measurement rarely meets all three assumptions: most scales use Likert responses (discrete), have items with unequal contributions to the construct (congeneric), and produce score distributions that depart from normality. The natural question is how badly alpha breaks under these violations and which alternatives perform better. A 2023 simulation study by Xiao and Hau in Educational and Psychological Measurement provides a systematic answer, with implications for the routine reliability reporting that fills psychometric methods sections. Read more →

What do the JCCES and ACT measure, and what was hypothesized?

The JCCES measures crystallized intelligence through three subtests: Verbal Analogies, Mathematical Problems, and General Knowledge. The ACT measures college readiness through English, Mathematics, Reading, and Science Reasoning, administered together in one sitting on one score scale. If both batteries tap a common ability, a factor analysis of the seven tasks should recover a single dominant factor on which all seven load. That prediction is easy to satisfy; the harder question is what else could satisfy it, since four tasks from one battery and three from another confound the ability the tasks share with the method they share.

How was the JCCES-ACT study conducted?

Fifty-seven participants completed the JCCES online under standardized procedures and reported their most recent ACT section scores, which were self-reported from a prior administration and not verified against ACT records. Only respondents whose implied testing year fell after the October 1989 introduction of the Enhanced ACT were retained, so that all section scores lie on one comparable scale; complete data on all seven tasks were required, and no other exclusion criterion was applied. Factors were extracted by iterated principal-axis factoring, and the number to retain was decided by parallel analysis against 5,000 random data sets rather than by the eigenvalue-of-one rule alone.