# What actually predicts a relationship

> Relationship science is consistent: who two people are predicts far less than what happens between them. What that means for anything that claims to match.

Canonical page: https://myorbit.ai/notes/what-predicts-a-relationship

---

> **TL;DR** — Sixty years of relationship research keeps landing on the same awkward result: who two people are barely predicts whether the two of them will work, while what actually happens between them predicts a lot. Similarity drives attraction in the lab and stops mattering once people actually know each other. Machine learning applied to more than a hundred pre-date measures could not predict who would click. What did predict relationship quality, across 43 longitudinal studies, was how each person experienced the relationship they were already in. This note is the evidence, and it is the reason Aura reads relationships rather than scoring personalities.

One thing to declare before the citations start: I build a product that reads relationships. That is exactly why this research is worth publishing instead of hiding. It rules out the version of this product that would have been the easiest to sell.

## What a personality instrument is actually good for

Start with the most respected tool of the lot. The Big Five — openness, conscientiousness, extraversion, agreeableness, negative emotionality — is the dominant structural model of personality, with modern instruments like [the BFI-2 organising the five domains into fifteen facets to improve bandwidth, fidelity and predictive power](https://doi.org/10.1037/pspp0000096).

And it predicts real things. Roberts, Kuncel, Shiner, Caspi and Goldberg reviewed prospective longitudinal evidence and found the effects of personality traits on [mortality, divorce and occupational attainment were "indistinguishable from the effects" of socioeconomic status and cognitive ability](https://doi.org/10.1111/j.1745-6916.2007.00047.x). Personality is not astrology. Traits matter for how a life goes.

Attachment research belongs on the same shelf. The adult attachment literature, [reviewed by Fraley across development, security, and models of thriving through relationships](https://doi.org/10.1146/annurev-psych-010418-102813), describes something real about how people behave when a close bond feels threatened.

But notice what those findings are about. They describe how one person's life tends to go, across many relationships. That is a different question from the one every matching product claims to answer: will these two people work.

## Similarity works in the laboratory and stops working in life

The folk theory is that similar people fit. The measurement is more interesting than that.

Montoya, Horton and Kirchner meta-analysed [460 effect sizes from 313 investigations of actual and perceived similarity](https://doi.org/10.1177/0265407508096700). Actual similarity correlated with attraction at r = .47 and perceived similarity at r = .39 — both large. Then the pattern splits. Actual similarity mattered in no-interaction and short-interaction studies, dropped sharply once people interacted, and "the effect of actual similarity in existing relationships was not significant." Perceived similarity kept predicting attraction at every stage, including in existing relationships.

That finding matters more than any other for anyone building a matching system. **Measured similarity predicts first impressions. Felt similarity predicts relationships.** A system comparing two profiles is describing the four minutes before two people speak.

The same qualification shows up in marriage. Luo and Klohnen studied 291 newlywed couples and found [substantial similarity on attitude-related domains but little on personality-related domains](https://doi.org/10.1037/0022-3514.88.2.304) — and, of the similarity that did predict marital quality, attachment similarity was the strongest.

## The prediction wall

If traits and preferences are the inputs, how far can prediction actually get? This has been tested directly, with the honest method: collect everything, then try.

Joel, Eastwick and Finkel ran two speed-dating studies in which unattached participants completed [more than 100 self-report measures of traits and preferences before meeting anyone](https://doi.org/10.1177/0956797617714580), then met every opposite-sex participant for a four-minute date. Random forests predicted 4% to 18% of actor variance (how romantically interested a person tends to be in general) and 7% to 27% of partner variance (how desirable a person tends to be in general). For relationship variance — the actual compatibility between two specific people — the models "were unable to predict relationship variance using any combination of traits and preferences reported before the dates."

Part of why is that people do not know their own criteria. Eastwick and Finkel found that participants stated the familiar sex-differentiated ideals about attractiveness and earning prospects, but [those stated ideals "failed to predict what inspired their actual desire" at the event itself](https://doi.org/10.1037/0022-3514.94.2.245).

This is also the conclusion of the field's formal review of the industry. Finkel, Eastwick, Karney, Reis and Sprecher's critical analysis of online dating for *Psychological Science in the Public Interest* found [little evidence that matching algorithms can predict whether people are good matches or will have chemistry](https://www.psychologicalscience.org/publications/journals/pspi/online-dating.html), while noting the real benefit these services do provide: access to people you would never otherwise meet.

Access is a solved problem. Compatibility-before-meeting is not, and the people who study it hardest are the most confident about that.

## What does predict: the relationship itself

Now the constructive half, and the most important study in this note.

Joel and 85 co-authors pooled [43 dyadic longitudinal datasets from 29 laboratories and applied machine learning to ask what predicts relationship quality](https://doi.org/10.1073/pnas.1917036117). The top relationship-specific predictors were perceived-partner commitment, appreciation, sexual satisfaction, perceived-partner satisfaction, and conflict. The top individual-difference predictors were life satisfaction, negative affect, depression, attachment avoidance and attachment anxiety.

The numbers matter. Relationship-specific variables predicted up to 45% of variance at baseline and up to 18% at the end of each study; individual differences managed 21% and 12%. And the finding that should reorganise the whole product category: "individual differences and partner reports had no predictive effects beyond actor-reported relationship-specific variables alone."

In plain terms: what you experience inside the relationship carries the signal. Who you are — and even what your partner says — adds nothing on top of it.

That is not a curiosity. It is an instruction. Stop profiling people. Read relationships.

## The popular models, measured

Two frameworks dominate the public conversation, and both deserve a fair hearing and an accurate one.

**Love languages.** Impett, Park and Muise evaluated Chapman's framework — a book that has sold over 20 million copies and been translated into 50 languages — against the empirical record and found [no strong support for its three central assumptions](https://doi.org/10.1177/09637214231217663): that each person has one preferred love language, that there are five, and that couples are more satisfied when partners match. Their proposed replacement is a better metaphor than a ranking: love as a balanced diet, where people need a full range of nutrients — time together, expressed appreciation, affection, practical support — rather than one primary language.

Notice why it became popular anyway. It gives people a vocabulary for needs they could not otherwise name, and a reason to sit down and discuss them. That part is real, and any honest system should keep it.

**Confident prediction numbers.** Work on predicting divorce produced famously high accuracy figures, and the statistics behind them are the part nobody quotes. Heyman and Smith Slep demonstrated with archival data that [a prediction equation correctly classifying 90% of couples in its development sample fell to a positive predictive value of 29% under cross-validation](https://doi.org/10.1111/j.1741-3737.2001.00473.x) — and to about 21% once the actual population base rate of early divorce was applied. Their conclusion was that results without cross-validation "should be interpreted with extreme caution, no matter how impressive the initial results appear to be."

Every product in this space will eventually be tempted to put a confident percentage on a human relationship. This is the paper that explains why that number is usually fiction, and it is the reason a number in this product is never a prediction about how your relationship ends.

## Closeness is made, not matched

If matching cannot manufacture compatibility, what can?

Aron and colleagues built the procedure that answers this: pairs of strangers working through [45 minutes of escalating, reciprocal, personal self-disclosure](https://doi.org/10.1177/0146167297234003) versus comparable small talk. The disclosure condition produced significantly greater closeness. And the details are the finding: matching pairs so they did not disagree on important attitudes made no significant difference, nor did leading them to expect mutual liking, nor did telling them to get close.

The mechanism was never the match. It was the exchange.

That changes the problem. If closeness comes from what two people actually do — opening up, saying thank you, showing up, answering — then the useful help is not better sorting at the front door. It is help with those things, with the people already in your life.

## What we took from all of this

Four design consequences, and they explain what Aura does and, as usefully, what it will not do.

**Read the relationship, not the person.** Aura's read comes from the history the two of you actually have — how you talk, what you have told it, what you have in common — not from a list of traits. Joel's 43 datasets say that is where the signal lives.

**Several kinds of closeness, not one score.** A relationship can be excellent for building something together and thin on everyday company. Averaging those into one compatibility number throws away the only information worth having. So the read has [seven separate views](/notes/how-aura-reads-a-relationship), and they are allowed to disagree.

**No verdicts about the other person.** Aura will not diagnose personality, attachment style or mental health from a profile, and it will not tell you what someone secretly feels about you. The research on how badly stated preferences and observed traits predict a specific pairing is exactly why that restraint is not modesty — it is accuracy.

**Direction, not prediction.** A relationship is not a fixed thing to be measured once. Every new read is saved, so what builds up over time is a record of where something is heading, which is far more honest than a forecast of where it ends up.

The industry's instinct has always been to measure two people and guess. The research says measure what happens between them, and help. That is a harder product to build, a much less impressive demo, and the only one the evidence supports.

## References

- Soto, C. J., & John, O. P. (2017). [The next Big Five Inventory (BFI-2)](https://doi.org/10.1037/pspp0000096). *Journal of Personality and Social Psychology*, 113(1).
- Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., & Goldberg, L. R. (2007). [The power of personality](https://doi.org/10.1111/j.1745-6916.2007.00047.x). *Perspectives on Psychological Science*, 2(4).
- Fraley, R. C. (2019). [Attachment in adulthood: recent developments, emerging debates, and future directions](https://doi.org/10.1146/annurev-psych-010418-102813). *Annual Review of Psychology*, 70.
- Montoya, R. M., Horton, R. S., & Kirchner, J. (2008). [Is actual similarity necessary for attraction? A meta-analysis of actual and perceived similarity](https://doi.org/10.1177/0265407508096700). *Journal of Social and Personal Relationships*, 25(6).
- Luo, S., & Klohnen, E. C. (2005). [Assortative mating and marital quality in newlyweds](https://doi.org/10.1037/0022-3514.88.2.304). *Journal of Personality and Social Psychology*, 88(2).
- Joel, S., Eastwick, P. W., & Finkel, E. J. (2017). [Is romantic desire predictable? Machine learning applied to initial romantic attraction](https://doi.org/10.1177/0956797617714580). *Psychological Science*, 28(10).
- Eastwick, P. W., & Finkel, E. J. (2008). [Sex differences in mate preferences revisited](https://doi.org/10.1037/0022-3514.94.2.245). *Journal of Personality and Social Psychology*, 94(2).
- Finkel, E. J., Eastwick, P. W., Karney, B. R., Reis, H. T., & Sprecher, S. (2012). [Online dating: a critical analysis from the perspective of psychological science](https://www.psychologicalscience.org/publications/journals/pspi/online-dating.html). *Psychological Science in the Public Interest*, 13(1).
- Joel, S., Eastwick, P. W., et al. (2020). [Machine learning uncovers the most robust self-report predictors of relationship quality across 43 longitudinal couples studies](https://doi.org/10.1073/pnas.1917036117). *PNAS*, 117(32).
- Impett, E. A., Park, H. G., & Muise, A. (2024). [Popular psychology through a scientific lens: evaluating love languages from a relationship science perspective](https://doi.org/10.1177/09637214231217663). *Current Directions in Psychological Science*, 33(2).
- Heyman, R. E., & Smith Slep, A. M. (2001). [The hazards of predicting divorce without crossvalidation](https://doi.org/10.1111/j.1741-3737.2001.00473.x). *Journal of Marriage and Family*, 63(2).
- Aron, A., Melinat, E., Aron, E. N., Vallone, R. D., & Bator, R. J. (1997). [The experimental generation of interpersonal closeness](https://doi.org/10.1177/0146167297234003). *Personality and Social Psychology Bulletin*, 23(4).
