Prenatal Hormonal Milieu as a Common Cause of Correlated Facial and Psychological Phenotype: A Multi-Axis Model and Its First Pre-Registered Falsification Test

Michael Abdo
Independent Researcher · Correspondence: Michael@michaelabdo.com

Theoretical model with a pre-registered falsification test; a refutation, not a confirmation.


Abstract

Prenatal androgen exposure acts as an accepted common cause of co-occurring physical and behavioral traits - digit ratio, spatial cognition, and the anatomical-behavioral profile of congenital adrenal hyperplasia all trace to a shared hormonal gradient. We generalize this single-axis precedent into a multi-axis common-cause model: across critical fetal windows, the intrauterine hormonal milieu (testosterone, cortisol, thyroid, estradiol, and others, varying in dose and timing) simultaneously shapes physical structure, including the face, and neural architecture, so that observable form constitutes a correlated trace of the developmental event that shaped disposition rather than a cause of it. The model yields a falsifiable prediction absent from prior physiognomic claims: the readability of a trait-axis from the face increases monotonically with the strength of that axis’s prenatal hormonal organization. We test this prediction against our own pre-registered, held-out data and report a refutation. A prior, non-pre-registered discovery analysis had produced an encouraging function-axis signal (d ≈ −0.47, sex-residualized −0.45), so the held-out result reads most precisely as a failed replication. We pre-registered the prediction in two registrations (OSF DOI 10.17605/OSF.IO/SMRXT, n = 83; OSF DOI 10.17605/OSF.IO/KJ2RW, n = 549) and ran a frozen estimator on held-out cohorts the discovery analysis never touched. The predicted ordering did not hold: the well-powered test produced sex-modality |d| = 0.39 > letter |d| = 0.27 > function |d| = 0.21, with the flagship function axis failing to separate above chance and failing FDR correction while a letter axis predicted to lie near the null survived. The readability-ordering form of the model is therefore not supported. The value here is as methodological as theoretical: a physiognomy-adjacent claim made precise enough to be killed by its own pre-registered test.


1. Introduction

The idea that the face reveals character boasts a long history, though it has been thoroughly discredited. Classical physiognomy posited a direct body-to-mind causal arrow. This approach collapsed reliably on confounds such as image quality, expression, grooming, demographic correlates, and the self-fulfilling effects of being treated as one’s appearance suggests. Modern machine-learning revivals of the genre inherit the same fatal structure. They predict a socially loaded label from a photograph. These systems cannot separate a genuine developmental signal from the artifacts that co-vary with it.

This paper asks whether a different causal structure can both ground the recurring observation that morphology and disposition co-vary and generate a prediction sharp enough to be falsified. We are talking about a model that explicitly denies the body to mind arrow. We propose such a structure. We derive a falsifiable ordering prediction from it. We then report the result of testing that prediction against our own pre-registered held-out data.

Let us be clear about the paper’s trajectory. We offer two things. First, a theoretical contribution: the model and the falsifiable hierarchy it generates. Second, an empirical contribution: a refutation of that hierarchy’s central claim. Why present both? Because a falsifiable model that fails its first pre-registered test is a more honest scientific object than one decorated with post-hoc confirmation. What follows is simply a falsifiable model and its first failed pre-registered empirical test. This is a held-out replication attempt that did not reproduce an encouraging discovery-cohort signal (§4.1).


2. The Model

2.1 From single-axis precedent to a multi-axis generalization

The organizational hypothesis, which posits that perinatal hormones permanently and jointly shape both soma and brain, finds strong single-axis support in Berenbaum and Beltz (2011). In individuals with excess prenatal androgen due to CAH, virilization of anatomy occurs alongside a shift in behavior, as documented by Berenbaum and Resnick (1997). Similarly, opposite-sex twins display masculinized digit ratios and elevated sensation seeking, according to Cohen-Bendahan et al. (2005). Cord and amniotic testosterone have been shown to predict later language and autistic-trait measures, with Whitehouse et al. (2012) and Auyeung et al. (2009) providing this evidence. The common form across these results is one hormonal gradient yielding two correlated outputs: body and behavior.

We suggest that this structure extends beyond androgens and clinical extremes to encompass the typical case. Multiple hormone axes vary continuously in both concentration and the timing of exposure across critical windows, such as the masculinization programming window (Welsh et al., 2008). These axes jointly specify a high-dimensional developmental trajectory whose outputs include both morphology and disposition. We flag immediately that this generalization runs ahead of the evidence and will return to this point in §5. The strongest single-axis result, 2D:4D, is itself meta-analytically near-null even as a marker of testosterone (Sorokowski & Kowal, 2024; §4.4). Building a multi-axis edifice on this foundation demands caution.

2.2 Two coupled maps from one input

Imagine a fetus’s developmental input as a vector h of hormone exposures, indexed by axis and critical window. The model proposes two coupled maps derived from that same input:

Because M and N share the argument h, their outputs correlate without any causal arrow between body and mind. Personality is not read from the face; both are reads of h. This is the decisive departure from physiognomy, which posits body to mind causation and therefore collapses on confounds.

Two corollaries follow directly and match commonly reported observations:

2.3 The falsifiable prediction

A model earns the right to be taken seriously only if it forbids something. This one does. Define, for each trait-axis a, its hormonal organization O(a) - how strongly axis a is specified by h (strong for sexually dimorphic axes wired by large androgen gradients; weaker for axes with diffuse or late hormonal input). The model predicts:

Readability of axis a from facial morphology increases monotonically in O(a).

This is a hierarchy prediction, not a point prediction, and it is falsifiable three ways: (1) if readability is flat across axes, the common-cause coupling is absent; (2) if the ordering inverts, the model is wrong; (3) if a weakly organized axis reads strongly, an alternative cause (cultural, self-presentation, demographic artifact) is implicated. No prior physiognomic claim made an ordering prediction. They claimed uniform legibility of character, precisely the form that confounds counterfeit. An ordering that tracks independently known hormonal organization resists faking by image-quality or demographic artifacts, because those confounds do not respect O(a).

Operationally, we tiered three families of trait-axis by O(a) from the prior hormonal literature. We note that the high rank assigned to the function axis was not literature-only: a prior discovery analysis on this face dataset had returned an encouraging function-axis effect (§4.1), so the held-out test is a replication attempt, not a first look. The three families:

The predicted sex-residualized readability ordering was frozen before the test.

function ≥ sex-modality > letter axes, with function and sex-modality each separable above chance and the letter axes at or near the null.

We deliberately included the sex-modality axis (an axis essentially saturated by biological sex) as an internal check: the central worry for any face-based claim is that the apparatus merely re-reads sex. The model’s distinctive bet was that the function axis would carry the largest readout after sex is residualized out, i.e., a real, sex-independent developmental signal sitting above the letter axes.

2.4 Scope and self-imposed limits

This model makes a claim about populations and mechanisms, not a license for individual judgment. It predicts a probabilistic prior rather than a verdict. It says nothing about worth, intelligence, criminality, or outcome. The organizational signal is anatomical. It is not a statistical proxy for demographics. This is why demographic controls are mandatory in any test. Any attempt to rank or gate individuals would misuse the model, not apply it. We return to this in §6.


3. The Falsification Test

We publicly pre-registered the prediction from §2.3 and tested it using a frozen estimator on held-out cohorts that the discovery analysis never touched. We report the result verbatim; it is a refutation. The full methodological treatment of this test-the adversarial design, the SHA-256 hash chain, the rare-class power ceiling, and the demographic-shortcut analysis of the surviving signal-is the subject of the companion empirical paper (Abdo, 2026b). We summarize here only what is needed to adjudicate the model’s prediction.

3.1 Design

We registered the prediction on OSF on two occasions, both on 2026-06-15, prior to the confirmatory runs.

Where the larger cohort came from. The two registrations sequester two nested held-out pools, both defined purely by exclusion from the 707-identity discovery manifest. Registration 1 (smrxt) locked a first batch: the 83 identities with a complete 5-coin OPS stack among the 115 OPS-typed folders then identified as absent from the discovery manifest. The smrxt frozen-cohort record osf_prereg_paper3.md registers the count: 115 OPS-typed folders absent from discovery → 83 with a full 5-coin stack, the other 32 dropped for incomplete types. Registration 2 (kj2rw) widened the same exclusion rule to the complete set of OPS-typed identities absent from the 707-identity discovery cohort (every such identity with a full 5-coin type and ≥3 fitted photos), which numbers 549. The additional ~466 identities are not a new construction but the remainder of that same disjoint-by-construction pool, assembled after the smrxt registration. We do not claim the full 549 list was locked at the smrxt stage. It was not. We tie its sequestration to the kj2rw timestamp instead. The powered registration (osf.io/kj2rw §4, §6) records that all 549 identities were frozen feature-blind at registration. Hypothesis-blind 3DMM and pose features extracted for the cohort had no per-coin effect size, Cohen’s d, LDA, or hypothesis test computed at registration time. This binds the frozen cohort list, blind features, and estimator under SHA-256 fingerprints (heldout_large_cohort_frozen.csv e5c5786…; features 2002e174…; estimator c22dcb07…) recorded at the registration timestamp, so any later claim that the cohort was assembled after seeing held-out effects is falsifiable by re-hashing.

Sequestration order (auditable). The held-out identities were sequestered, fixed and disjoint by construction, before any held-out feature was loaded because both cohorts were defined by exclusion from the 707-identity discovery cohort. The estimator was committed as a frozen script, which froze the hypothesis and decision rule. Two distinct fingerprints anchor the two registrations. The smrxt registration sealed the estimator helper confirmatory_analysis.py (SHA-256 578ec8f4…; osf_prereg_paper3.md §6-§7). The kj2rw registration sealed the held-out runner confirmatory_heldout.py (SHA-256 c22dcb07…, §“Data and Code Availability”; osf_prereg_paper3_powered.md §3, §6), which applies that helper unchanged. The runner (c22dcb07…) was applied byte-identically across both cohorts. This means “identical across both runs” refers to the runner being applied unchanged to each cohort, not to a single hash registered at both timestamps. On 2026-06-15 the OSF registration timestamp preceded the confirmatory run. The powered result file records this as a “post-registration timestamp.” We note the registration and run share that single date and present this order as the basis for the “before the confirmatory runs” claim. The OSF and commit timestamps, along with the §6 artifact hashes recorded at registration, are the auditable record.

The estimator remained frozen and identical across both runs. It used the full 40-dimensional 3D-morphable-model (BFM, 3DDFA_V2) per-identity shape centroid, with pose and expression dimensions discarded and no principal-component selection. We calculated covariance-whitened LDA / Mahalanobis separability per axis as a Cohen’s-d-equivalent. Identity-disjoint GroupKFold (k = 5) ensured folds were leakage-free by construction. Sex-residualization served as the primary adjustment, projecting onto the null space of the linear biological-sex direction, with sex read from the annotation database - never from the image. Head pose (yaw/pitch/roll) was regressed out as a confound. We computed bootstrap 95% CIs using 1,000 identity resamples, along with Benjamini-Hochberg FDR (q = 0.05) across axes. The decision rule was frozen in advance. The §2.3 ordering holding with both function and sex-modality surviving FDR confirms. A flat/inverted ordering, or a function axis that fails to separate above chance after sex + pose adjustment, refutes. The registrations committed in advance to reporting the held-out result verbatim, with no substitution of discovery (selected-PC) numbers and no estimator re-selection.

3.2 Result

The well-powered sex-residualized readability ordering involved a sample of n = 549.

Rank Axis d
1 sex-modality 0.39 (AUC 0.609) survives (p = 1.1 × 10⁻⁵)
2 letter (S_N) 0.27 (AUC 0.574) survives (p = 0.0026)
3 function 0.21 (AUC 0.560) fails (p = 0.172)
4 letter (F_T) 0.12 (AUC 0.534) fails (p = 0.169)

The flagship function axis ranked below both the sex-modality axis and a letter axis. Its whole-shape bootstrap CI [−0.377, +0.604] spans zero around the resampled point estimate d = +0.13; note that the d = +0.21 above is the GroupKFold whitened-LDA estimate on the same axis. Because the interval spans zero the axis is not separable above chance. After sex + pose adjustment its effect is only d = +0.15. The sex-residualized AUC is 0.560 and no pose-adjusted AUC was computed for this axis. It fails FDR with p = 0.172. The pre-registered refutation condition is met. This condition is the §3.1 decision rule: a function coin that fails to separate above chance after sex + pose adjustment and corresponds to OSF §5. Meanwhile a letter axis (S_N, predicted near-null) survives FDR. This result is the opposite of the hierarchy.

The sex-modality axis behaved exactly as the model’s high-rank prediction expected. The strongest coin yielded a d of +0.39 using GroupKFold whitened-LDA with an AUC of 0.609, an FDR p of 1.1 × 10⁻⁵, and a bootstrap CI of [+0.003, +0.604] that excluded zero. This CI is around the resampled point estimate d of +0.31. We report this for completeness, but it does not rescue the verdict. This axis is precisely the readout the model needed to exceed, not merely match. A candidate explanation, which we offer as a hypothesis rather than an established interpretation, is that the surviving signal is plausibly dominated by residual sex covariance that survives linear sex-residualization. Sex-residualization projects out only the linear sex direction. We do not treat this as demonstrated. It sits in tension with the independent position that the sex-modality axis encodes cognitive modality rather than biological sex.

The smaller registration (n = 83) also refuted the hypothesis, although it was underpowered on the function axis (10 Se / 12 Te). There, the function effect was inverted (sex-residualized AUC 0.217, below chance), sex-modality collapsed to last (|d| = 0.145), and no coin survived FDR. The two registrations agree on the verdict for different reasons, namely a small-n artifact in the first and a well-powered null in the second, and the function axis fails the confirmation condition in both.

3.3 The verdict, stated plainly

Both registrations return REFUTES according to the frozen pre-registered decision rule. The model’s central falsifiable prediction failed in this first pre-registered test (§4.1). That prediction was that the function axis carries the largest sex-independent prenatal-hormone readout, sitting above the letter axes. Instead, a letter axis the model placed at the null survived. On the evidence we have, the readability-ordering form of the model is not supported.


4. Discussion

4.1 The discovery signal this test failed to replicate

To ensure full transparency regarding the empirical scaffolding behind the function prediction, we note that before the pre-registered held-out test, a non-pre-registered discovery analysis on the same face dataset had produced an encouraging function-axis result. On the 707-identity discovery cohort, a function-axis coin (observer/decider) showed d = −0.47 that survived biological-sex adjustment (sex-residualized d = −0.45). This effect replicated within each sex (males −0.57, females −0.37) and within ethnicity (within-White d = −0.38 at full ethnicity coverage, n = 166). The larger −0.62 came from the low-coverage n = 40 subset before full FairFace labeling, and the source flags the full-coverage figure as resolving that coverage caveat. It was characterized as a “large” discovery effect per the selected-PC discovery comparison in the powered-results file. That discovery signal, not literature alone, is why the model assigned the function axis its high rank. The pre-registered held-out result reported in §3 is therefore most honestly read as a failed replication. On the well-powered held-out cohort the same function axis collapsed to a near-zero, non-significant effect (sex-residualized d = +0.21, pose-adjusted +0.15, bootstrap CI spanning zero, FDR p = 0.172). We disclose the discovery confirmation here precisely so the refutation is not overstated as a first-ever look and the discovery scaffolding is not concealed. The failure of an encouraging discovery signal to replicate on sequestered, pre-registered data is the strongest and most honest form of the result.

4.2 What the refutation kills

The refutation is specific and real: the monotone-in-O(a) readability hierarchy is not supported. After the pre-registered adjustments, the function axis, which is the part of the model that distinguished it from “the face reads sex,” does not separate above chance, while a letter axis that the model demoted to the null does. An honest reading is that the specific ordering the model generated is wrong, at least as operationalized here. We do not get to keep the hierarchy.

This matters because the hierarchy was the model’s entire claim to non-triviality. Without it, what remains?

4.3 What, if anything, survives

Two things are worth separating.

The common-cause logic versus the specific ordering. The model consists of two layers: (a) a structural claim that body and disposition correlate because both read a shared developmental input h (no body→mind arrow), and (b) a quantitative claim that face-readability orders monotonically by hormonal organization. We refute layer (b) here. Layer (a) is not directly tested by this experiment - it is the prior that motivated the ordering, and the failure of (b) neither confirms nor refutes it. What the failure does establish is that, if (a) is true, the mapping from “hormonal organization” to “facial readability” is not the simple monotone function the model assumed: the function axis’s hormonal organization (as we tiered it from the literature) did not translate into facial separability. The model’s weakest link may be the O(a) tiering itself - derived from a contested literature - rather than the common-cause structure.

The surviving signal, soberly. One axis (sex-modality) reads robustly from facial shape after sex-residualization, and one letter axis (S_N) reads weakly but FDR-significantly. That two axes separate above chance at all establishes only that some facial-shape signal exists for some trait-axes - a far weaker statement than the model made, and one fully consistent with mundane explanations (residual sex covariance for the sex-modality axis; demographic or self-presentation correlates for S_N). A further alternative explanation applies to both survivors: a face→label leakage path. The companion empirical paper (Abdo, 2026b) flags that OPS labels are assigned by human typers reasoning over behavioral video - i.e., the typers see the person’s face while assigning the coin (that paper names this “a possible face→label leakage path”). Any such leakage would inflate a positive face-readability estimate, so it is a material alternative explanation for precisely the two positive coins that survive here, further weakening the case that the survivors represent independent confirmation. (It does not bear on the function-axis null: leakage can inflate a positive effect but cannot manufacture a null, so the flagship refutation stands.) We explicitly decline to spin the two surviving coins into a partial confirmation. The pre-registered confirmation required function and sex-modality surviving FDR; function did not.

4.4 The contested mechanism literature

Intellectual honesty demands that we highlight the contested nature of the model’s mechanistic foundation. This reality weakens the prior, regardless of our test.

The refutation comes as no surprise. It aligns with a mechanistic literature that is already weak and contested. We offer the model because its form constitutes a genuine contribution. This form features a falsifiable ordering and an explicit non-physiognomic causal structure. We do not present it because its mechanistic premises are secure.

4.5 Why publish a refuted model

Three reasons. First, the model’s form is the contribution. It takes a vague, confound-prone, physiognomy-adjacent intuition and converts it into an ordering prediction. A held-out, pre-registered test could kill this prediction, and it did. Most claims in this space are never made falsifiable enough to fail. Second, the negative result is informative. It tells the field that the obvious developmental signal beyond sex does not fall out of facial shape under clean adjustment. This should temper the next round of face-based personality claims. Third, transparency about a failed flagship prediction is the only defensible way to discuss a topic with physiognomy’s history. Presenting the model with its refutation is the safeguard.


5. Limitations


6. Ethics

This work engages with the most dangerous claim in applied psychology, the idea that the face reveals the person. It does so to constrain, not license, that claim.


Competing Interests

The author developed, and holds a continuing interest in, the proprietary OPS-style typing system from which the trait-axis labels used here are derived. That system is the locus of the author’s broader applied work and represents a potential commercial stake. This interest is material to a physiognomy-adjacent claim, so we disclose it explicitly; we note that the result reported here is a refutation of the model’s flagship prediction, i.e., a finding that runs against, not toward, any commercial interest in face-based typing. No other competing interests are declared.

Funding

This work received no external funding and was self-funded by the author. No funder had any role in the study design, analysis, interpretation, decision to publish, or preparation of the manuscript.

Research-Conduct and Data Provenance

This study is a secondary analysis of publicly available, third-party, in-the-wild celebrity face images. No images were generated, staged, or solicited for this study. No human-subjects intervention or interaction occurred. No identifiable private information beyond what is already public was collected. No consent was obtained because the analysis uses only pre-existing public images, and individuals are referenced by identity rather than redistributed. On this basis, secondary analysis of public, non-interventional data, and as an independent researcher with no affiliated institution and therefore no IRB of record, the author judges that no IRB review was applicable to this work rather than self-granting an institutional exemption. No IRB review was sought. The trait-axis (typing) labels are author-assigned via the proprietary system described under Competing Interests, not solicited from the depicted individuals.


Data and Code Availability

A frozen, pre-registered estimator was used for the confirmatory analysis. We identify it here to ensure the claim that we did not re-select the estimator is auditable:

The analysis code, the frozen estimator script and hash, and the held-out arrays form the locus of reproducibility, while the OSF registrations serve as the locus for the pre-registered hypotheses and decision rules.


References

Pre-registrations (this work): - OSF DOI 10.17605/OSF.IO/SMRXT - held-out confirmatory test (n = 83), registered 2026-06-15. - OSF DOI 10.17605/OSF.IO/KJ2RW - powered replication (n = 549), registered 2026-06-15.

Companion papers (this series): - Abdo, M. (2026a). The Non-Shared Developmental Residual: What Twin Variance Implies for a Prenatal-Hormone Account of Correlated Phenotype. PsyArXiv. https://doi.org/10.31234/osf.io/qnmtg - Abdo, M. (2026b). Testing a Prenatal-Hormone Effect-Size Hierarchy in Facial Morphology: A Pre-Registered Self-Refutation at Celebrity Scale. PsyArXiv. https://doi.org/10.31234/osf.io/mq4db - Abdo, M. (2026c). Reading the Receipt, Not the Person: An Anti-Physiognomy Framing for Developmental Trace Inference. PsyArXiv. https://doi.org/10.31234/osf.io/2rbqx