Why the sexes are mostly the same and obviously different, and why both sides of the debate have been right
There is a fight that has been running in psychology for thirty years, and it has the peculiar property that neither side has ever landed a decisive blow. One camp, associated with Janet Hyde's gender similarities hypothesis, points out that when you measure men and women trait by trait, the distributions overlap enormously. On most measures, the overlap exceeds 80%, and the honest summary is that the sexes are far more alike than different. The other camp, associated with evolutionary psychology and, more recently, with Marco Del Giudice's multivariate work, points to the domains where the differences are large and cross-culturally stable, and to the fact that many small differences, taken together, can separate two groups almost completely. Del Giudice's 2012 analysis of fifteen personality facets found that when the traits are considered jointly, the overlap between the sexes falls to roughly 10%. Each side has data. Each side has published rebuttals of the other's data. And much of the public, watching from the outside, has largely concluded that the science is politicized because it keeps telling them something that their own lives contradict.
I want to argue that the argument has persisted because both sides are right, and that seeing why they are both right resolves what I will call the primary gender paradox: the sexes are, by careful measurement, largely the same, and by ordinary experience, plainly different. The resolution is not a compromise between the two positions. It is a structural account of why two accurate methods, applied to the same species, return opposite conclusions, and it yields a criterion for predicting which differences should hold across cultures and which should not.
The Shape of The Trait Space
Start with a picture. Take every psychological trait that has been measured—verbal fluency, mental rotation, neuroticism, risk tolerance, interest in mechanical systems, responsiveness to infant distress, sexual variety seeking, choosiness in mate selection, and so on through a few hundred more—and give every person a score on each. Now picture each person as a single data point that carries all of their scores at once, and the whole population as a cloud of those points. Men and women are two clouds.
The first thing to say about the two clouds is that they overlap heavily. There is no region that belongs to one sex alone, and no trait that one sex has and the other lacks; on every trait, both sexes span the full range. If you look at the two clouds through any single trait, they nearly coincide, offset by a small amount, which is why the overlap measured trait by trait typically comes out around eighty percent rather than a hundred. That is the truth the similarities camp has been insisting on, and it is not a small truth. Human sexual dimorphism is modest by primate standards, and human sex differences in psychology are not separate male and female modules but rather different calibrations of shared machinery, largely shaped by hormonal exposure during a few developmental windows. Shared architecture, different tuning. That is why any individual can land anywhere, and why the overlap on any single trait is so large..
The second thing to say is that they are not the same cloud. Each has its own center of density, the place where its points are thickest, and the two centers do not coincide. Most of the displacement between them comes from a small number of traits. On most traits the offset is tiny. On a few it is substantial. And the traits that do the pulling are not scattered at random through the list. They cluster.
Where the Differences Are
The cluster is well documented and was predicted before it was fully documented. In 1995, David Buss laid out the principle that follows from sexual selection theory: the sexes should differ psychologically only in the domains where, over evolutionary time, they faced different adaptive problems, and should be alike in the domains where the problems were the same. Avoiding predators, learning language, reasoning about cause and effect, counting, recalling facts, general intelligence: shared problems, and the sexes are correspondingly alike. Choosing a mate, guarding a mate, competing with rivals of one's own sex, investing in offspring, dividing the labor of provisioning and care: divergent problems, because the reproductive biology of the two sexes made the optimal solutions differ.
John Archer's 2019 review of the entire sex-difference literature confirms the pattern with unusual clarity. The largest and most cross-culturally stable differences fall in exactly these domains: sexual desire and variety seeking, choosiness, the weighting of youth and beauty against status and resources in mate preference, direct physical aggression, dominance seeking, risk tolerance, responsiveness to infants, and the people-versus-things orientation that underlies both occupational interest and, further back, the division of foraging labor. The spatial and targeting differences belong to the same cluster, as products of the division of labor rather than of mating directly. So does a difference that runs the other way and is easy to misfile: women's advantage at remembering where objects are. It looks like an ordinary memory trait, which would put it in the shared domain, until one notices that it is the memory a gatherer needs, the ability to return to a plant that was not ripe last week, and that the reciprocal male advantage in navigating by direction and distance is the memory a hunter needs. The two are one difference seen from both ends of a divided task. Outside this cluster, the differences shrink toward nothing.
So the picture is now sharper. Two overlapping clouds whose centers of density are pulled apart by a compact set of traits corresponding to sexual selection and parental investment, and which are nearly coincident everywhere else. Call that set the divergent domain.
Two Methods, Two Regions
Now look at what each side of the debate is actually doing when it measures.
The similarities method takes traits one at a time and computes an overlap for each. Because most traits lie outside the divergent domain, the typical overlap is high, and averaging across all traits produces a summary in which similarity dominates. This is correct. It is a correct description of the entire trait space, weighted by the number of traits it contains.
The differences method does one of two things. Either it looks specifically inside the divergent domain, where it correctly finds large effects, or it aggregates across many traits at once, asking how far apart the two centers of density are when every trait is counted, rather than how far apart they are on any one trait. Del Giudice's 2012 analysis did the second across fifteen personality facets and found a Mahalanobis distance of about 2.7, which corresponds to roughly ten percent overlap rather than eighty percent. That result was attacked on technical grounds, and some of the attacks have force, but the underlying logic is not in dispute: two groups can overlap heavily on every single trait and still be largely separable when the traits are considered together, because small differences accumulate. Two groups can overlap heavily on every trait and still be nearly separable when the traits are read together, because small differences accumulate. This too is correct. It is a correct description of the distance between the two centers of density.
Neither method is measuring the other's quantity. Overlap on a single trait and separability across the whole profile are different numbers, and a species can have high values of both. The thirty-year argument has been an argument between people measuring different things, each of whom took the other's number as a rival estimate of their own.
Why Experience Sides with Difference
That explains the scientific disagreement. It does not yet explain why ordinary people, who have never computed a Mahalanobis distance (e.g., me), are so confident that the sexes differ.
The answer is that ordinary experience is not a random sample of the trait space. It is a targeted sample of the divergent domain.
Consider where men and women spend their high-intensity time together. Not in the laboratory, and not in the domains of shared adaptive problems. They spend it in courtship, in sex, in the negotiation of commitment, in the raising of children, in the management of jealousy and the division of household effort. Those are the divergent domains, almost by definition; they are the situations that sexual selection shaped differences to handle. A person who has been married for twenty years has spent twenty years observing the other sex precisely on the traits where the two clouds are pulled furthest apart, and has essentially never observed it on the traits where they coincide, because nothing in a marriage makes verbal fluency or mental rotation salient.
Add two amplifiers. The first is that a spouse is not measuring a single trait with a questionnaire but integrating dozens of cues repeatedly over years, which is a multivariate classifier with a very large sample operating on the most diagnostic variables. It is the differences method, run by a nervous system, in the region where the differences method finds the most. The second is that what people notice is tails, not means. A displacement of half a standard deviation between two distributions is a modest overlap figure, but yields a two-to-one ratio at one standard deviation out and a four-to-one ratio at two. The violent man, the obsessive tinkerer, the woman who remembers every slight, the one who weeps at the advertisement: these are tail events, and tails magnify small central shifts into striking asymmetries.
There is a fourth reason: the measurements themselves are taken in the wrong place. A great deal of what is known about sex differences comes from self-report questionnaires administered in neutral settings. But many of the divergent-domain differences are likely state-dependent, triggered by courtship, threat, infants, or conflict rather than constantly present. A questionnaire filled out in a quiet room samples the sexes at their most similar. Marriage samples them at their most divergent. And self-report has a further known bias: people rate themselves against same-sex reference groups, so a woman assertive for a woman and a man assertive for a man both mark the same number, which compresses whatever difference exists.
So the lay perception and the psychometric finding are not in conflict. They are looking at the same underlying pattern from two different vantage points.
From Description to Criterion
Up to this point I have described a structure. The structure becomes a theory when it is combined with a distinction between two kinds of mind.
The adapted mind is the set of mechanisms shaped by selection to solve recurrent ancestral problems. It is conserved across cultures because the problems were universal, and it is resistant to cultural revision because culture did not build it. The adaptive mind is the capacity for flexible learning and cultural acquisition, the machinery by which a human can become a Sumerian scribe or a software engineer. It is malleable by design.
The mapping is now direct. The divergent domain is the domain of sex-differentiated adaptive problems, that is, adapted-mind territory. The traits there should be conserved, generationally regenerated, and present in some form in every culture, resistant to socialization, and visible early in development and in the comparative primate record. The shared domain is where the sexes faced the same problems, and differences there, to the extent they exist at all, should be adaptive-mind products: culturally produced, culturally variable, and responsive to changes in opportunity and norm.
That mapping yields a criterion with predictive content. It says which sex differences should survive a change in culture and which should dissolve. Differences in mating psychology, aggression, parenting orientation, and the people-things dimension should show low cross-cultural variance and should not shrink under egalitarian policy. Differences in, say, mathematics performance, self-reported confidence, or occupational attainment in fields not tied to the people-things dimension should show high cross-cultural variance and should track opportunity closely.
The existing evidence fits. The famous gender-equality paradox, in which occupational and interest differences persist or widen in the most egalitarian societies, is a partial confirmation: those differences load on the people-things dimension, and the criterion says they should not respond to policy. Meanwhile, the mathematics gap, which sits in the shared domain, has narrowed dramatically wherever opportunity has been equalized, which is also what the criterion says.
The Line Between Calibration and Malleability
A reader familiar with the literature will object at this point. Some divergent-domain traits do vary across cultures. Sociosexuality tracks the local sex ratio and pathogen load. Mate-preference weightings shift with economic conditions. If the divergent domain is conserved, why does it move?
The objection conflates two distinct concepts: facultative calibration and cultural malleability.
An adapted mechanism need not produce a constant output. Many are built to read local conditions and adjust. A mechanism that raises sexual restraint when pathogen load is high, or lowers choosiness when eligible partners are scarce, is exhibiting a fixed rule with a variable setting. The rule is conserved; the setting depends on input. That is calibration, and it is adapted-mind behavior through and through. Cultural malleability is different: it is a change in the rule itself, or the acquisition of a disposition the adapted mind never specified, through learning and norms.
The two can be distinguished empirically. Calibration should track ecological variables in the direction predicted by an adaptive analysis, and it should do so in every culture exposed to those variables. Malleability should track cultural variables and be capable of running in directions that make no adaptive sense. The criterion applies to rules, not to settings. A divergent-domain trait that shifts with sex ratio has not been shown to be malleable; it has been shown to be well designed.
The converse also needs to be stated, so that the criterion is not misread. It does not claim that everything in the shared domain is malleable. Much of the shared architecture is shared and fixed. The claim is about where difference is conserved, not about where anything is conserved.
Why This Is the Primary Paradox
Several other things are called gender paradoxes, and it is worth clarifying how they relate to this one.
The gender-equality paradox concerns why occupational and interest differences persist in affluent, egalitarian societies. Under the account given here, it is not a separate puzzle but a derived case: a divergent-domain trait, the people-things dimension, meeting an institutional expectation built on the assumption that all differences are shared-domain and should therefore respond to opportunity. The paradox exists only for someone who has not distinguished the two domains.
The controversies around sex-segregated sport have a similar structure, though here the divergent trait is physical rather than psychological. Athletic categories were built on the observation that differences in strength, speed, and body composition are large and conserved, and that the current disputes arise when a category built for the divergent domain is asked to accommodate a shared-domain conception of identity. Whatever one's view of how those disputes should be settled, the shape of the difficulty is the same shape as the occupational case: an institution designed around one region of the trait space is confronted with claims grounded in another.
The paradox I have described is prior to both. It concerns the structure of the trait space itself, and the others arise where institutions collide with that structure. It also has the widest reach, because it is the one that ordinary people encounter without any institution involved: every person who has lived with the other sex has run into it. That is why I call it primary. The others are its consequences at particular institutional boundaries; this one is the thing they are consequences of.
What Is at Stake in Getting It Right
Two errors follow from missing the structure, and both are currently common.
The first is to take the similarities finding as a claim about the divergent domain, and conclude that the differences people perceive in courtship, sex, and parenting are stereotypes to be corrected. This error underlies a great deal of well-intended advice that does not work, because it treats conserved calibrations as learned habits and expects them to yield to instruction. When they do not yield, the failure is attributed to insufficient effort rather than to a misdiagnosis.
The second is to take the differences finding as a claim about the whole space, and conclude that because the sexes differ in mating psychology they must differ in ability, aptitude, or fitness for roles that lie in the shared domain. This error underlies most of the historical justification for excluding women from work they were fully equipped to do, and it survives in softer forms wherever a divergent-domain difference is quietly extended into a shared-domain conclusion.
The structure defeats both errors at once, without asking either side of the debate to give up its data. The similarities camp is right about the space. The differences camp is right about the region. And the public is right about everyday life. The resolution has been available in the geometry all along; what was missing was the recognition that three correct observations of the same object, taken from three different viewpoints, are not three rival claims.
A Note on Testing
The criterion is falsifiable, and the test is a matter of measuring cross-cultural variance in the two domains and checking whether it sorts as predicted. Much of the necessary data exists in scattered form. A further test is possible through the historical corpus: if the divergent-domain differences are conserved, they should have been described, in the same direction, by every literate tradition, regardless of that tradition's power arrangements and regardless of which sex was doing the describing, while shared-domain claims should flip with era and narrator. That is a different, larger project, but well-suited to the unique access LLMs have to the human narrative corpus, and I plan on doing it.