Showing posts sorted by relevance for query "james lee". Sort by date Show all posts
Showing posts sorted by relevance for query "james lee". Sort by date Show all posts

Wednesday, May 11, 2016

74 SNP hits from SSGAC GWAS



The SSGAC discovery of 74 SNP hits on educational attainment (EA) is finally published in Nature. Nature News article.

EA was used in order to assemble as large a sample as possible (~300k individuals). Specific cognitive scores are only available for a much smaller number of individuals. But SNPs associated with EA are likely to also be associated with cognitive ability -- see figure above.

The evidence is strong that cognitive ability is highly heritable and highly polygenic. With even larger samples we'll eventually be able to build good genomic predictors for cognitive ability.
Genome-wide association study identifies 74 loci associated with educational attainment A. Okbay et al. Nature http://dx.doi.org/10.1038/nature17671; 2016

Educational attainment is strongly influenced by social and other environmental factors, but genetic factors are estimated to account for at least 20% of the variation across individuals1. Here we report the results of a genome-wide association study (GWAS) for educational attainment that extends our earlier discovery sample1,2 of 101,069 individuals to 293,723 individuals, and a replication study in an independent sample of 111,349 individuals from the UK Biobank. We identify 74 genome-wide significant loci associated with the number of years of schooling completed. Single- nucleotide polymorphisms associated with educational attainment are disproportionately found in genomic regions regulating gene expression in the fetal brain. Candidate genes are preferentially expressed in neural tissue, especially during the prenatal period, and enriched for biological pathways involved in neural development. Our findings demonstrate that, even for a behavioural phenotype that is mostly environmentally determined, a well-powered GWAS identifies replicable associated genetic variants that suggest biologically relevant pathways. Because educational attainment is measured in large numbers of individuals, it will continue to be useful as a proxy phenotype in efforts to characterize the genetic influences of related phenotypes, including cognition and neuropsychiatric diseases.

Here's what I wrote back in September of 2015, based on a talk given by James Lee on this work.
James Lee talk at ISIR 2015 (via James Thompson) reports on 74 hits at genome-wide statistical significance (p < 5E-8) using educational attainment as the phenotype. Most of these will also turn out to be hits on cognitive ability.

To quote James: "Shock and Awe" for those who doubt that cognitive ability is influenced by genetic variants. This is just the tip of the iceberg, though. I expect thousands more such variants to be discovered before we have accounted for all of the heritability.
74 GENOMIC SITES ASSOCIATED WITH EDUCATIONAL ATTAINMENT PROVIDE INSIGHT INTO THE BIOLOGY OF COGNITIVE PERFORMANCE 
James J Lee

University of Minnesota Twin Cities
Social Science Genetic Association Consortium

Genome-wide association studies (GWAS) have revealed much about the biological pathways responsible for phenotypic variation in many anthropometric traits and diseases. Such studies also have the potential to shed light on the developmental and mechanistic bases of behavioral traits.

Toward this end we have undertaken a GWAS of educational attainment (EA), an outcome that shows phenotypic and genetic correlations with cognitive performance, personality traits, and other psychological phenotypes. We performed a GWAS meta-analysis of ~293,000 individuals, applying a variety of methods to address quality control and potential confounding. We estimated the genetic correlations of several different traits with EA, in essence by determining whether single-nucleotide polymorphisms (SNPs) showing large statistical signals in a GWAS meta-analysis of one trait also tend to show such signals in a meta-analysis of another. We used a variety of bio-informatic tools to shed light on the biological mechanisms giving rise to variation in EA and the mediating traits affecting this outcome. We identified 74 independent SNPs associated with EA (p < 5E-8). The ability of the polygenic score to predict within-family differences suggests that very little of this signal is due to confounding. We found that both cognitive performance (0.82) and intracranial volume (0.39) show substantial genetic correlations with EA. Many of the biological pathways significantly enriched by our signals are active in early development, affecting the proliferation of neural progenitors, neuron migration, axonogenesis, dendrite growth, and synaptic communication. We nominate a number of individual genes of likely importance in the etiology of EA and mediating phenotypes such as cognitive performance.
For a hint at what to expect as more data become available, see Five Years of GWAS Discovery and On the genetic architecture of intelligence and other quantitative traits.


What was once science fiction will soon be reality.
Long ago I sketched out a science fiction story involving two Junior Fellows, one a bioengineer (a former physicist, building the next generation of sequencing machines) and the other a mathematician. The latter, an eccentric, was known for collecting signatures -- signed copies of papers and books authored by visiting geniuses (Nobelists, Fields Medalists, Turing Award winners) attending the Society's Monday dinners. He would present each luminary with an ornate (strangely sticky) fountain pen and a copy of the object to be signed. Little did anyone suspect the real purpose: collecting DNA samples to be turned over to his friend for sequencing! The mathematician is later found dead under strange circumstances. Perhaps he knew too much! ...

Saturday, May 06, 2017

More Shock and Awe: James Lee and SSGAC in Oslo, 600 SNP hits


To quote James Lee, the first author listed below: "Shock and Awe" for those who doubt that cognitive ability is influenced by genetic variants.

See work from a year ago: ~100 hits from 300k individuals. Now ~600 hits from 750k. (SNPs associated with EA are likely to also be associated with cognitive ability -- see figure at link above.)
47th Behavior Genetics Annual Meeting, Oslo, Norway

GWAS of Educational Attainment, Phase 3: Biological Findings

Abstract
Genetic factors are estimated to account for at least 20% of the variation across individuals for educational attainment (Rietveld et al., 2013). The results of the latest GWAS for educational attainment identified 74 genome-wide significant loci for educational attainment (Okbay et al., 2016). Here, in one of the largest GWAS to date, we increase our sample to nearly 750,000 individuals, and we identify over 600 genome-wide significant loci associated with the number of years of schooling completed. Note that at the time of presentation, we will likely have updated our meta-analysis to include over 1,000,000 individuals

In this presentation, I will focus on the biological implications of the GWAS results. At the time of writing, 1,656 genes are significantly prioritized, a more than 10-fold increase since our previous report (Okbay et al., 2016). The newly significant results reinforce the biological theme of prenatal brain development and also bring to the foreground new themes that shed light on the biological underpinnings of cognitive performance and other traits affecting educational attainment.

Authors
James Lee (University of Minnesota - Twin Cities), Aysu Okbay (Free University Amsterdam), Robbee Wedow (University of Colorado - Boulder), Edward Kong (Harvard University), Patrick Turley (Broad Institute of MIT and Harvard), Meghan Zacher (Harvard University), Kevin Thom (New York University), Anh Tuan Nguyen Viet (University of Southern California), Omeed Maghzian (Harvard University, NBER), Richard Karlsson Linnér (Vrije Universiteit Amsterdam), Matthew Robinson (The University of Queensland), Social Science Genetic Association Consortium (NA), Peter Visscher (The University of Queensland), Daniel Benjamin (University of Southern California), David Cesarini (New York University)
Note the data here have only been analyzed using summary statistics coming from each sub-cohort. More powerful methods may soon become available:
Penalized regression from summary statistics

One of the difficulties in genomics is that when DNA donors are consented for a study, the agreements generally do not allow sharing (aggregation) of genomic data across multiple studies. This leads to isolated silos of data that can't be fully shared. However, computations can be performed on one silo at a time, with the results ("summary statistics") shared within a larger collaboration. Most of the leading GWAS collaborations (e.g., GIANT for height, SSGAC for cognitive ability) rely on shared statistics. Simple regression analysis (one SNP at a time) can be conducted using just summary statistics, but more sophisticated algorithms cannot. These more sophisticated methods can generate a better phenotype predictor, using less data, than a SNP by SNP analysis.
A successful implementation like the one described at the link above could produce many (several times!) more hits and significantly more variance accounted for by corresponding predictors. Stay tuned!

Note Added: I'm getting lots of questions about how to interpret these results, so here are some comments.

1. I predicted ~10k variants would account for most of the heritability due to common SNPs (i.e., about 50% of total variance; allowing a predictor which correlates ~0.7 with actual cognitive ability). The rate of discovery of genome-wide significant hits and corresponding variance accounted for seems consistent with this prediction. Genetic associations are most easily discovered for variants which are common (e.g., have ~0.5 Minor Allele Frequency, not 0.05) and have large effect sizes. But alleles with this combination of properties are rare. As statistical power increases, one starts to discover (more and more) variants of lower frequency and/or lower effect size. A reasonable guess at the genetic architecture suggests a higher density of such variants, and is consistent with an accelerating rate of discovery of SNP hits (~100 hits from 300k individuals, ~600 hits from 750k). There are more efficient methods that, I believe, would discover nearly all the variants given sample size of ~1M well-phenotyped individuals. But these methods require more than just summary statistics.

I made a similar prediction of ~10k variants for height, and our (unpublished) genomic prediction results make me fairly confident that this will turn out to be correct. We now have moderately good height predictors and they are getting better very fast. That ~10k variants will turn out to be responsible for most of the variation in cognitive ability is still at a somewhat lower confidence level.

2. People are still confused about how many + variants above the mean in the population are required to make a "genius" (or super-genius). I managed to compress the explanation enough to fit in a tweet:
Flip coin 10000 times. 5000 + sqrt(10000)/2 = 5050 heads is +1SD outcome. 5100 is +2SD, etc. sqrt(N) << N for N large. Binomial~Normal dist.
You can see that even if cognitive ability is controlled by ~10k variants, flipping only ~100 of them is enough to cause a big difference in actual intelligence. Flipping a few hundred could get us to super-geniuses beyond anything in human history.

3. If you read press accounts related to our creation of the BGI Cognitive Genomics Lab back in 2011 (at that time there were zero genome-wide significant alleles associated with intelligence), you can find quotes from genomics "experts" asserting that mankind would never discover the genetic architecture of cognitive ability. (Such quotes are easy to obtain even today!) A Bayesian update given what is known in 2017 would call into question the competence of these "experts"!  ;-)

Friday, July 27, 2018

Insight Podcast: James Lee interview on SSGAC EA3



Spencer Wells and Razib Khan interview James Lee (Professor of Psychology, University of Minnesota, BA Berkeley, PhD Harvard) about the recent SSGAC EA3 GWAS.

Comment: James mentions that EA3 may be approaching the GCTA h2 limit (~0.15? so limiting r ~ 0.4) already. But the limit for actual cognitive ability is much higher; with enough data I think we could get to r ~ 0.6 or even r ~ 0.7 eventually for common SNPs -- similar to height.

United Club, HK International Airport



James, me, Chris Chang. (About $1M worth of Illumina HiSeqs in crates behind us?)


Monday, June 15, 2020

Support Freedom of Ideas and Inquiry at MSU

These are letters of support sent on my behalf to the MSU President: presidentstanley@msu.edu

Several are deep, detailed scholarly documents. They firmly rebut the false accusations of the Twitter mob.

Corey Washington's individual letter is over 5000 words long. James Lee's is over 3000 words. The authors have graciously allowed them to be made public.

Corey says:
I have known Steve for 30 years and can attest that he is not a racist or a sexist. Steve is one of the most scrupulously fair people I have ever met and I have seen no evidence that he has ever discriminated against anyone on the basis of their race, sex, or any other status.

... Hsu participated in a 2018 debate at MSU’s Institute for Quantitative Health, Science, and Engineering, at the invitation of Director Chris Contag. The topic was human genetic engineering, and Hsu’s counterpart was MSU bioethicist Len Fleck. A number of people who attended tell me that they found Steve’s view thoughtful and balanced. None that I know of came away from the debate thinking he is a eugenicist in the way that the rest of us are not. It is unclear to me why MSU GEU came to a different conclusion.

Sign the support petition. Email the president: presidentstanley@msu.edu

A (not necessarily up to date) list of signatories, which includes hundreds of professors from MSU and around the world, total ~1500 as of early June 19.


Letter from
Corey Washington, Director of Analytics and Strategic Projects, OSVPRI
Joseph Cesario, Associate Professor, Psychology
Wei Liao, Professor and Director, MSU Anaerobic Digestion Research and Education Center, Biosystems and Agricultural Engineering
John (Xuefeng) Jiang, Professor and Plante Moran Faculty Fellow, Accounting and Information Systems

Letter from Corey Washington, Director of Analytics and Strategic Projects, MSU

Letter from Matt McGue, Regents Professor of Psychology, University of Minnesota

Letter from Russell Warne, Associate Professor of Psychology, Utah Valley University

Letter from Mark Dykman, Professor of Physics, MSU

Letter from Zach Hambrick, Professor of Psychology, MSU

Letter from James Lee, Associate Professor of Psychology, University of Minnesota

Letter from Richard Haier, Emeritus Professor, UC Irvine, author of The Neuroscience of Intelligence (Cambridge)

Letter from Steven Pinker, Johnstone Family Professor of Psychology, Harvard University


Media coverage:

A Twitter Mob Takes Down an Administrator at Michigan State (Wall Street Journal June 25)

Scholar forced to resign over study that found police shootings not biased against blacks (The College Fix)

On Steve Hsu and the Campaign to Thwart Free Inquiry (Quillette)

Michigan State University VP of Research Ousted (Reason Magazine, Eugene Volokh, UCLA)

Research isn’t advocacy (NY Post Editorial Board)

Podcast interview on Tom Woods show (July 2)

College professor forced to resign for citing study that found police shootings not biased against blacks (Law Enforcement Today, July 5)

"Racist" College Researcher Ousted After Sharing Study Showing No Racial Bias In Police Shootings (ZeroHedge, July 6)

Twitter mob: College researcher forced to resign after study finding no racial bias in police shootings (Reclaim the Net, July 8)

Horowitz: Asian-American researcher fired from Michigan State administration for advancing facts about police shootings (The Blaze, July 8)

I Cited Their Study, So They Disavowed It: If scientists retract research that challenges reigning orthodoxies, politics will drive scholarship (Wall Street Journal July 8)

Conservative author cites research on police shootings and race. Researchers ask for its retraction in response (The College Fix, July 8)

Academics Seek to Retract Study Disproving Racist Police Shootings After Conservative Cites It (Hans Bader, CNSNews, July 9)

The Ideological Corruption of Science (theoretical physicist Lawrence Krauss in the Wall Street Journal, July 12)

Foundation for Individual Rights in Education: "chilling academic freedom" (Peter Bonilla, July 22)

Saturday, September 19, 2015

SNP hits on cognitive ability from 300k individuals

James Lee talk at ISIR 2015 (via James Thompson) reports on 74 hits at genome-wide statistical significance (p < 5E-8) using educational attainment as the phenotype. Most of these will also turn out to be hits on cognitive ability.

To quote James: "Shock and Awe" for those who doubt that cognitive ability is influenced by genetic variants. This is just the tip of the iceberg, though. I expect thousands more such variants to be discovered before we have accounted for all of the heritability.
74 GENOMIC SITES ASSOCIATED WITH EDUCATIONAL ATTAINMENT PROVIDE INSIGHT INTO THE BIOLOGY OF COGNITIVE PERFORMANCE 
James J Lee

University of Minnesota Twin Cities
Social Science Genetic Association Consortium

Genome-wide association studies (GWAS) have revealed much about the biological pathways responsible for phenotypic variation in many anthropometric traits and diseases. Such studies also have the potential to shed light on the developmental and mechanistic bases of behavioral traits.

Toward this end we have undertaken a GWAS of educational attainment (EA), an outcome that shows phenotypic and genetic correlations with cognitive performance, personality traits, and other psychological phenotypes. We performed a GWAS meta-analysis of ~293,000 individuals, applying a variety of methods to address quality control and potential confounding. We estimated the genetic correlations of several different traits with EA, in essence by determining whether single-nucleotide polymorphisms (SNPs) showing large statistical signals in a GWAS meta-analysis of one trait also tend to show such signals in a meta-analysis of another. We used a variety of bio-informatic tools to shed light on the biological mechanisms giving rise to variation in EA and the mediating traits affecting this outcome. We identified 74 independent SNPs associated with EA (p < 5E-8). The ability of the polygenic score to predict within-family differences suggests that very little of this signal is due to confounding. We found that both cognitive performance (0.82) and intracranial volume (0.39) show substantial genetic correlations with EA. Many of the biological pathways significantly enriched by our signals are active in early development, affecting the proliferation of neural progenitors, neuron migration, axonogenesis, dendrite growth, and synaptic communication. We nominate a number of individual genes of likely importance in the etiology of EA and mediating phenotypes such as cognitive performance.
For a hint at what to expect as more data become available, see Five Years of GWAS Discovery and On the genetic architecture of intelligence and other quantitative traits.


What was once science fiction will soon be reality.
Long ago I sketched out a science fiction story involving two Junior Fellows, one a bioengineer (a former physicist, building the next generation of sequencing machines) and the other a mathematician. The latter, an eccentric, was known for collecting signatures -- signed copies of papers and books authored by visiting geniuses (Nobelists, Fields Medalists, Turing Award winners) attending the Society's Monday dinners. He would present each luminary with an ornate (strangely sticky) fountain pen and a copy of the object to be signed. Little did anyone suspect the real purpose: collecting DNA samples to be turned over to his friend for sequencing! The mathematician is later found dead under strange circumstances. Perhaps he knew too much! ...

Tuesday, August 09, 2011

Demography and fast evolution

In an earlier post I discussed the population history uncovered by Gregory Clark in his book A Farewell to Alms. By examining British wills, he showed that the rich literally replaced (outreproduced) the poor over a period of several centuries.



The excerpt below is from a review of Greg Clark's book. The review is mostly negative about Clark's big picture conclusions, but does provide some interesting historical information. Note, the reviewer does not seem to understand population genetics (see discussion further below).

The comparison of Beijing nobility and Liaoning peasants is drawn from Lee and Wang’s (1999) survey of Chinese demography, which, in turn, is based on a very detailed investigation of population in Liaoning by Lee and Campbell (1997). In Liaoning, all men had military obligations and were enumerated in the so-called banner roles, which described their families in detail. Individuals’ occupations were also noted, so that fertility can be compared across occupational groups. High status, high income occupations had the most surviving sons: for instance, soldiers aged 46–50 had on average 2.57 surviving sons, artisans had 2.42 sons, and officials had 2.17 sons. In contrast, men aged 46–50 who were commoners had only 1.55 sons on average.

The references cited are

Lee, James Z., and Cameron D. Campbell. 1997. Fate and Fortune in Rural China: Social Organization and Population Behavior in Liaoning 1774–1873. Cambridge and New York: Cambridge University Press.

Lee, James Z., and Feng Wang. 1999. One Quarter of Humanity: Malthusian Mythology and Chinese Realities, 1700–2000. Cambridge and London: Harvard University Press.

So we have at least two documented cases of the descendants of the rich replacing the poor over an extended period of time. My guess is that this kind of population dynamics was quite common in the past. (Today we see the opposite pattern!) Could this type of natural selection lead to changes in quantitative, heritable traits over a relatively short period of time?

Consider the following simple model, where X is a heritable trait such as intelligence or conscientiousness or even height. Suppose that X has narrow sense heritability of one half. Divide the population into 3 groups:

Group 1 bottom 1/6 in X; < 1 SD below average
Group 2 middle 2/3 in X; between -1 and +1 SD
Group 3 highest 1/6 in X; > 1 SD above average

Suppose that Group 3 has a reproductive rate which is 10% higher than Group 2, whereas Group 1 reproduces at a 10% lower rate than Group 2. A relatively weak correlation between X and material wealth could produce this effect, given the demographic data above (the rich outreproduced the poor almost 2 to 1!). Now we can calculate the change in population mean for X over a single generation. In units of SDs, the mean changes by roughly 1/6 ( .1 + .1) 1/2 or about .02 SD. (I assumed assortative mating by group.) Thus it would take roughly 50 generations, or 1k years, under such conditions for the population to experience a 1 SD shift in X.

If you weaken the correlation between X and reproduction rate, or relax the assortative mating assumption, you get a longer timescale. But it's certainly plausible that 10,000 years is more than enough for this kind of evolution. For example, we might expect that the advent of agriculture over such timescales changed humans significantly from their previous hunter gatherer ancestors.

This model is overly simple, and the assumptions are speculative. Nevertheless, it addresses some deep questions about human evolution: How fast did it happen? How different are we from humans who lived a few or ten thousand years ago? Did different populations experience different selection pressures? Amazingly, we may be able to answer some of these questions in the near future.

Thanks to Henry Harpending for reminding me about the Chinese data and about the question of fastest plausible evolution for a quantitative trait.

Monday, July 23, 2018

SSGAC EA3: genomic prediction of educational attainment and related cognitive phenotypes

Years ago I predicted that:

1. Cognitive ability would turn out to be influenced by many thousands of genetic variants, each of small effect.

2. With large enough sample size we would detect these variants and eventually construct genomic predictors.

The Nature Genetics paper below from the SSGAC collaboration takes a significant step in that direction.

Although the study used over a million genotypes, the data had to be aggregated across many sub-cohorts using summary statistics only. This does not permit the L1-penalized optimization we used to build our height predictor.

For out of sample validation of the results below, see this PNAS paper, which (unusually) appeared before the paper on which it is based.

The lead author James Lee is on the left below. Chris Chang, author of Plink 2.0, is on the right. The photo was taken in 2010 at BGI -- they are standing in front of crates of Illumina sequencers.



Article | Published: 23 July 2018

Gene discovery and polygenic prediction from a genome-wide association study of educational attainment in 1.1 million individuals

James J. Lee, Robbee Wedow, […]David Cesarini
Nature Genetics (2018)

Abstract
Here we conducted a large-scale genetic association analysis of educational attainment in a sample of approximately 1.1 million individuals and identify 1,271 independent genome-wide-significant SNPs. For the SNPs taken together, we found evidence of heterogeneous effects across environments. The SNPs implicate genes involved in brain-development processes and neuron-to-neuron communication. In a separate analysis of the X chromosome, we identify 10 independent genome-wide-significant SNPs and estimate a SNP heritability of around 0.3% in both men and women, consistent with partial dosage compensation. A joint (multi-phenotype) analysis of educational attainment and three related cognitive phenotypes generates polygenic scores that explain 11–13% of the variance in educational attainment and 7–10% of the variance in cognitive performance. This prediction accuracy substantially increases the utility of polygenic scores as tools in research.
A nice figure from the paper: Add Health (National Longitudinal Study of Adolescent to Adult Health) and HRS (Health in Retirement Study) are two longitudinal cohorts that have been genotyped; horizontal axis is polygenic score. It appears that individuals with top quintile polygenic scores are about 5 times more likely to complete college than bottom quintile individuals.


Here's a comment on the paper I provided to a journalist:
The EA3 predictor correlates about 0.35 with educational attainment, and slightly less well with measured cognitive ability. While this is far from perfect prediction, it does allow identification of individuals, using DNA alone, who are at unusual risk of being well below average in cognitive ability or struggling in school. Standardized tests, such as SAT, ACT, GRE, LSAT, etc., typically also correlate roughly 0.35 with educational outcomes like grade point average, degree completion, etc. In this sense, the genomic predictor is comparable to widely used tests and it will certainly improve as more data are analyzed. See figure.

Thursday, July 22, 2010

Assortative mating, regression and all that: offspring IQ vs parental midpoint

In an earlier post I did a lousy job of trying to estimate the effect of assortative mating on the far tail of intelligence.

Thankfully, James Lee, a real expert in the field, sent me a current best estimate for the probability distribution of offspring IQ as a function of parental midpoint (average between the parents' IQs). James is finishing his Ph.D. at Harvard under Steve Pinker -- you might have seen his review of R. Nesbitt's book Intelligence and how to get it: Why schools and cultures count.

The results are stated further below. Once you plug in the numbers, you get (roughly) the following:

Assuming parental midpoint of n SD above the population average, the kids' IQ will be normally distributed about a mean which is around +.6n with residual SD of about 12 points. (The .6 could actually be anywhere in the range (.5, .7), but the SD doesn't vary much from choice of empirical inputs.)

So, e.g., for n = 4 (parental midpoint of 160 -- very smart parents!), the mean for the kids would be 136 with only a few percent chance of any kid to surpass 160 (requires +2 SD fluctuation). For n = 3 (parental midpoint of 145) the mean for the kids would be 127 and the probability of exceeding 145 less than 10 percent.

No wonder so many physicist's kids end up as doctors and lawyers. Regression indeed! ;-)

Below are some more details; see here for calculations. In my earlier post I arrived at the same formulae as below, but I had rho = 0.

Assuming bivariate normality (and it appears that IQ has been successfully scaled to produce this), the offspring density function is normal with mean n*h^2 and variance 1-(1/2)(1+rho)h^2, where rho is the correlation between mates attributable to assortative mating and h^2 is the narrow-sense heritability.

I put h^2 between .5 and .7. Bouchard and McGue found a median correlation between husband and wife of .33 in their review many years back, but not all of that may be attributable to assortative mating. So anything in (.20, .25) may be a reasonable guesstimate for rho.

In discussing this topic with smart and accomplished parents (e.g., at foo camp, in academic science, or on Wall Street), I've noticed very strong interest in the results ...

See related posts mystery of non-shared environment , regression to the mean

Note: Some people are confused that the value of h^2 = narrow sense (additive) heritability is not higher than (.5 - .7). You may have seen *broad sense* heritability H^2 estimated at values as large as .8 or .9 (e.g., from twin studies). But H^2 includes genetic sources of variation such as dominance and epistasis (interactions between genes, which violate additivity). Because children are not clones of their parents (they only get half of their genes from each parent, and in a random fashion), the correlation between midparent IQ and offspring IQ is not as large as the correlation between the IQs of identical twins. See here and here for more.

Monday, August 14, 2017

Estimation of genetic architecture for complex traits using GWAS data

These authors extrapolate from existing data to predict sample sizes needed to identify SNPs which explain a large portion of heritability in a variety of traits. For cognitive ability (see red curves in figure below), they predict sample sizes of ~million individuals will suffice.

See also More Shock and Awe: James Lee and SSGAC in Oslo, 600 SNP hits.
Estimation of complex effect-size distributions using summary-level statistics from genome-wide association studies across 32 complex traits and implications for the future

Yan Zhang, Guanghao Qi, Ju-Hyun Park, Nilanjan Chatterjee (Johns Hopkins University)

Summary-level statistics from genome-wide association studies are now widely used to estimate heritability and co-heritability of traits using the popular linkage-disequilibrium-score (LD-score) regression method. We develop a likelihood-based approach for analyzing summary-level statistics and external LD information to estimate common variants effect-size distributions, characterized by proportion of underlying susceptibility SNPs and a flexible normal-mixture model for their effects. Analysis of summary-level results across 32 GWAS reveals that while all traits are highly polygenic, there is wide diversity in the degrees of polygenicity. The effect-size distributions for susceptibility SNPs could be adequately modeled by a single normal distribution for traits related to mental health and ability and by a mixture of two normal distributions for all other traits. Among quantitative traits, we predict the sample sizes needed to identify SNPs which explain 80% of GWAS heritability to be between 300K-500K for some of the early growth traits, between 1-2 million for some anthropometric and cholesterol traits and multiple millions for body mass index and some others. The corresponding predictions for disease traits are between 200K-400K for inflammatory bowel diseases, close to one million for a variety of adult onset chronic diseases and between 1-2 million for psychiatric diseases.


This figure shows predicted effect size distributions for a number of quantitative traits. You can see that height and intelligence are somewhat different, but not dramatically so.

Friday, December 20, 2013

Neanderthals dumb?


This figure is from the Supplement (p.62) of a recent Nature paper describing a high quality genome sequence obtained from the toe of a female Neanderthal who lived in the Altai mountains in Siberia. Interestingly, copy number variation at 16p11.2 is one of the structural variants identified in a recent deCODE study as related to IQ depression; see earlier post Structural genomic variants (CNVs) affect cognition.

From the Supplement (p.62):
Of particular interest is the modern human-specific duplication on 16p11.2 which encompasses the BOLA2 gene. This locus is the breakpoint of the 16p11.2 micro-deletion, which results in developmental delay, intellectual disability, and autism5,6. We genotyped the BOLA2 gene in 675 diverse human individuals sequenced to low coverage as part of the 1000 Genome Project Phase I7 to assess the population distribution of copy numbers in homo-sapiens (Figure S8.3). While both the Altai Neandertal and Denisova individual exhibit the ancestral diploid copy number as seen in all the non-human great apes, only a single human individual exhibits this diploid copy number state.

My recollection from the earlier (less precise) Neanderthal sequences is that the number of bp differences between them and us is few per thousand. Whereas, for modern humans it's 1 per thousand with an additional +/-15% variation due to ethnicity. So, I think it's fair to say that they are qualitatively much more different from us than we (moderns) are from each other. See also The genetics of humanness.

My colleague James Lee (I note he is too modest to list his Harvard Law degree on his faculty page!) describes the current era in genomics as an "age of wonder" :-)  We can anticipate tremendous discoveries in the next decade.

Saturday, January 15, 2022

Manifold Returns!

I'm working on the return of Manifold. I just did the first interview yesterday, with James Lee. I'm not sure when it will be released, but soon I hope.

Thanks to everyone who enjoyed the original show. I received a lot of requests to bring it back over the last 18 months since we went on hiatus. Please make suggestions for guests and show topics!

I'll try to get Dominic Cummings for a long interview. For now, see this great piece in the Guardian about his substack: 

Intoxicating, insidery and infuriating: everything I learned about Dominic Cummings from his £10-a-month blog


Here are the 51 episodes from our first run, with transcripts: https://manifoldlearning.com/


Michigan State University owns the copyright to the old stuff, so we may have to create a new channel and web site. This new project will be entirely mine and I will probably allow comments on the YouTube channel -- we turned those off because the university was/is risk averse about free expression ¯\_(ツ)_/¯ 

I hope to have Corey as a frequent guest on the show, but I will be restarting in solo mode. Here's the first episode where Corey and I introduced ourselves. Very rough around the edges, but still fun for me to listen to again.


Saturday, July 10, 2010

Beyond Bayes: causality vs correlation

A draft paper by Harvard graduate student James Lee (student of Steve Pinker; I'd love to post the paper here but don't know yet if that's OK) got me interested in the work of statistical learning pioneer Judea Pearl. I found the essay Bayesianism and Causality, or, why I am only a half-Bayesian (excerpted below) a concise, and provocative, introduction to his ideas.

Pearl is correct to say that humans think in terms of causal models, rather than in terms of correlation. Our brains favor simple, linear narratives. The effectiveness of physics is a consequence of the fact that descriptions of natural phenomena are compressible into simple causal models. (Or, perhaps it just looks that way to us ;-)

Judea Pearl: I turned Bayesian in 1971, as soon as I began reading Savage’s monograph The Foundations of Statistical Inference [Savage, 1962]. The arguments were unassailable: (i) It is plain silly to ignore what we know, (ii) It is natural and useful to cast what we know in the language of probabilities, and (iii) If our subjective probabilities are erroneous, their impact will get washed out in due time, as the number of observations increases.

Thirty years later, I am still a devout Bayesian in the sense of (i), but I now doubt the wisdom of (ii) and I know that, in general, (iii) is false. Like most Bayesians, I believe that the knowledge we carry in our skulls, be its origin experience, schooling or hearsay, is an invaluable resource in all human activity, and that combining this knowledge with empirical data is the key to scientific enquiry and intelligent behavior. Thus, in this broad sense, I am a still Bayesian. However, in order to be combined with data, our knowledge must first be cast in some formal language, and what I have come to realize in the past ten years is that the language of probability is not suitable for the task; the bulk of human knowledge is organized around causal, not probabilistic relationships, and the grammar of probability calculus is insufficient for capturing those relationships. Specifically, the building blocks of our scientific and everyday knowledge are elementary facts such as “mud does not cause rain” and “symptoms do not cause disease” and those facts, strangely enough, cannot be expressed in the vocabulary of probability calculus. It is for this reason that I consider myself only a half-Bayesian. ...

Friday, April 11, 2014

Human Capital, Genetics and Behavior

See you in Chicago next week :-)
HCEO: Human Capital and Economic Opportunity Global Working Group

Conference on Genetics and Behavior

April 18, 2014 to April 19, 2014

This meeting will bring together researchers from a range of disciplines who have been exploring the role of genetic influences on socioeconomic outcomes. The approaches taken to incorporating genes into social science models differ widely. The first goal of the conference is to provide a forum in which alternative frameworks are discussed and critically evaluated. Second, we are hopeful that the meeting will trigger extended interactions and even future collaboration. Third, the meeting will help focus future genetics-related initiatives by the Human Capital and Economic Opportunity Global Working Group, which is pursuing the study of inequality and social mobility over the next several years.

PROGRAM

9:00 to 11:00
Genes and Socioeconomic Aggregates
Gregory Cochran University of Utah
Steven Durlauf University of Wisconsin–Madison
Henry Harpending University of Utah
Aldo Rustichini University of Minnesota
Enrico Spolaore Tufts University

11:30 to 1:30
Population-Based Studies
Sara Jaffee University of Pennsylvania
Matthew McGue University of Minnesota
Peter Molenaar
Jenae Neiderhiser

2:30 to 4:30
Genome-Wide Association Studies (GWAS)
Daniel Benjamin Cornell University
David Cesarini New York University
Dalton Conley New York University/NBER
Jason Fletcher University of Wisconsin–Madison
Philipp Koellinger University of Amsterdam

APRIL 19, 2014

9:00 to 11:00
Neuroscience
Paul Glimcher New York University
Jonathan King National Institute on Aging
Aldo Rustichini University of Minnesota

11:30 to 1:30
Intelligence
Stephen Hsu Michigan State University
Wendy Johnson University of Edinburgh
Rodrigo Pinto The University of Chicago

2:30 to 4:30
Role of Genes in Understanding Socioeconomic Status
Gabriella Conti University College London
Steven Durlauf University of Wisconsin–Madison
Felix Elwert University of Wisconsin–Madison
James Lee University of Minnesota

Friday, July 26, 2019

RadioLab on embryo selection in IVF



I'm in this RadioLab podcast covering genetic selection of embryos in IVF. Apologies to SSGAC, Robert Plomin, Ian Deary, James Lee, Tom Bouchard, and countless other dedicated scientists for the impression given that progress in genomics of cognitive ability is largely my work. See last paragraph below.

This is the email I sent to RadioLab this morning:
Hi Pat and Michelle,

Congratulations on a high quality podcast. I thought you were admirably fair and balanced. I also thought the production (esp. the music) was excellent.

My main comment is that the juxtaposition between my remarks and Benjamin's is misleading: when he says 60-40 or 55% chance of rank ordering properly, that is a very different question than identifying an outlier who is, say, among the 1% highest in risk. We are not trying to rank order embryos, but to warn against unusual risk of a medical condition.

To use the SAT analogy, given two kids with scores 1250 and 1200, only some of the time does the 1250 kid end up with a higher GPA. (You can't predict rank order very well.) But if the engineering dean admits an SAT 770 kid (i.e., a negative outlier compared to the average score of, say, 1300 among engineers) in his freshman class, he knows the likelihood is high that the kid will struggle. Benjamin is talking about the first scenario, and I am talking about the second.

Finally, I realize that to hook listeners you had to make me the focus of the episode. But I want to make clear that many scientists contribute to this work, which I feel will ultimately be beneficial to our species and civilization. I am just a small part of a worldwide research endeavor.

Best wishes,
Steve
For more on recent progress in genomic prediction, see The Diffusion of Knowledge.

Wednesday, January 26, 2022

Friday, August 03, 2012

Correlation, Causation and Personality

A new paper from my collaborator James Lee. Ungated copy here, including commentary from other researchers including Judea Pearl.
Correlation and Causation in the Study of Personality 
Abstract: Personality psychology aims to explain the causes and the consequences of variation in behavioural traits. Because of the observational nature of the pertinent data, this endeavour has provoked many controversies. In recent years, the computer scientist Judea Pearl has used a graphical approach to extend the innovations in causal inference developed by Ronald Fisher and Sewall Wright. Besides shedding much light on the philosophical notion of causality itself, this graphical framework now contains many powerful concepts of relevance to the controversies just mentioned. In this article, some of these concepts are applied to areas of personality research where questions of causation arise, including the analysis of observational data and the genetic sources of individual differences.
From the conclusions:
... This article is in part an effort to unify the contributions of three innovators in causal reasoning: Ronald Fisher, Sewall Wright, and Judea Pearl 
Fisher began his career at a time when the distinction between correlation and causation was poorly understood and indeed scorned by leading intellectuals. Nevertheless, he persisted in valuing this distinction. This led to his insight that randomization of the putative cause—whether by the deliberate introduction of ‘error’, as his biologist colleagues thought of it, or ‘beautifully . . . by the meiotic process’—in fact reveals more than it obscures. His subsequent introduction of the average excess and average effect is perhaps the first explicit use of the distinction between correlation and causation in any formal scientific theory. 
Structural equation modelers will know Wright—Fisher’s great rival in population genetics—as the ingenious inventor of path analysis. Wright’s diagrammatic approach to cause and effect serves as a conceptual bridge toward Pearl’s graphical formalization, which has greatly extended the innovations developed by both of the population-genetic pioneers. 
The fruitfulness of Pearl’s graphical framework when applied to the problems discussed in this article bear out its utility to personality psychology. Perhaps the most surprising instance of the theory’s fruitfulness concerns the role of colliders. Although obscure before Pearl’s seminal work, this role turns out to be obvious in retrospect and a great aid to the understanding of covariate choice, assortative mating, selection bias, and a myriad of other seemingly unrelated problems. This article has surely only scratched the surface of the ramifications following from our recognition of colliders.  
Conspicuous from these accolades by his absence is Charles Spearman—the inventor of factor analysis and thereby a founder of personality psychology. Spearman (1927) did conceive of his g factor as a hidden causal force. However, new and brilliant ideas are often only partially understood, even by their authors. After a century of theoret- ical scrutiny and empirical applications, common factors appear to be more plausibly defended as mild formalizations of folk-psychological terms than as causal forces uncovered by matrix algebra. I have thus advocated a sharp distinction between the measurement of personality traits (factor analysis) and the study of their causal relations (graphical SEM).  
... The puzzle is that by using common factors in our causal explanations, we seem to be retreating from this reductionistic approach. A single node called g sending an arrow to a single node called liberalism is surely an approximation to the true and extraordinarily more complicated graph entangling the various physical mechanisms that underlie mental characteristics. Why this compromise? Is it sensible to test models of ethereal emergent properties shoving and being shoved by corporeal bits of matter—or, perhaps even worse, by other emergent properties? If we are committing to a calculus of causation, should we not also discard the convenient fictions of folk psychology? 
The answer to this puzzle may be that reductionistic decomposition is not always the royal road to scientific understanding. ... [[In physics we refer to "effective descriptions" or "effective degrees of freedom" appropriate to a particular scale of dynamics or organization -- no need to invoke quarks to explain the mechanics of a baseball.]]
From the author's response to commentary:
... It was the genius of Darwin to realize the power of explanation (4): phenotypes and environments cohere in such an uncanny way because nature is a statistician who has allowed only a subset of the logically possible combinations to persist over time. 
Although phenotypes are what nature selects, it cannot be phenotypes alone that preserve the record of natural selection. Phenotypes typically lack the property that variations in them are replicated with high fidelity across an indefinite number of generations. DNA, however, does have this property— hence the memorable phrase “the immortal replicator” (Dawkins, 1976). If DNA is furthermore causally efficacious, such that the possession of one variant rather than another has phenotypic consequences that are reasonably robust, then we have the potential for natural selection to bring about a lasting correlation between environmental demands and the causes of adaptation to those very same demands. 
When statistically controlling fitness, nature does not actually use the [causal] average effect of any allele. If an allele has a positive average excess in [is correlated with] fitness, for any reason whatsoever, it will tend to displace its alternatives. Nevertheless, it seems to be the case that nature correctly picks out alleles for their effects often enough; the results are evident in the living world all around us. Davey Smith and I are confident that where nature has succeeded, patient and ingenious human scientists will be able to follow.
For more on Judea Pearl's work, see the earlier post: Beyond Bayes: causality vs correlation.

Wednesday, October 30, 2013

Nabokov on teaching


Nabokov was professor of literature at Cornell from 1948-1959. The excerpt below is from a 1964 Playboy interview, reproduced at longform.org (a site I highly recommend).
Nabokov: I gave up teaching—that’s about all in the way of change. Mind you, I loved teaching, I loved Cornell, I loved composing and delivering my lectures on Russian writers and European great books. But around 60, and especially in winter, one begins to find hard the physical process of teaching, the getting up at a fixed hour every other morning, the struggle with the snow in the driveway, the march through long corridors to the classroom, the effort of drawing on the blackboard a map of James Joyce’s Dublin or the arrangement of the semi-sleeping car of the St. Petersburg-Moscow express in the early 1870s—without an understanding of which neither Ulysses nor Anna Karenina, respectively, makes sense. For some reason my most vivid memories concern examinations. Big amphitheater in Goldwin Smith. Exam from 8 a.m. to 10:30. About 150 students—unwashed, unshaven young males and reasonably well-groomed young females. A general sense of tedium and disaster. Half-past eight. Little coughs, the clearing of nervous throats, coming in clusters of sound, rustling of pages. Some of the martyrs plunged in meditation, their arms locked behind their heads. I meet a dull gaze directed at me, seeing in me with hope and hate the source of forbidden knowledge. Girl in glasses comes up to my desk to ask: “Professor Kafka, do you want us to say that…? Or do you want us to answer only the first part of the question?” The great fraternity of C-minus, backbone of the nation, steadily scribbling on. A rustic arising simultaneously, the majority turning a page in their bluebooks, good teamwork. The shaking of a cramped wrist, the failing ink, the deodorant that breaks down. When I catch eyes directed at me, they are forthwith raised to the ceiling in pious meditation. Windowpanes getting misty. Boys peeling off sweaters. Girls chewing gum in rapid cadence. Ten minutes, five, three, time’s up.
The first paragraph of Lolita, one of my favorites in all of literature:
Lolita, light of my life, fire of my loins. My sin, my soul. Lo-lee-ta: the tip of the tongue taking a trip of three steps down the palate to tap, at three, on the teeth. Lo. Lee. Ta.

Friday, January 01, 2016

GCTA, Missing Heritability, and All That

New Update: Yang, Visscher et al. respond here.

Update: see detailed comments and analysis here and here by Sasha Gusev. Gusev claims that the problems identified in Figs 4,7 are the result of incorrect calculation of the SE (4) and failure to exclude related individuals in the Framingham data (7).


Bioinformaticist E. Stovner asked about a recent PNAS paper which is critical of GCTA. My comments are below.

It's a shame that we don't have a better online platform (e.g., like Quora or StackOverflow) for discussing scientific papers. This would allow the authors of a paper to communicate directly with interested readers, immediately after the paper appears. If the authors of this paper want to correct my misunderstandings, they are welcome to comment here!
I took a quick look at it. My guess is that Visscher et al. will respond to the paper. It has not changed my opinion of GCTA. Note I have always thought the standard errors quoted for GCTA are too optimistic, as the method makes strong assumptions (e.g., fits a model with Gaussian random effect sizes). But directionally it is obvious that total h2 accounted for by all common SNPs is much larger than what you get from only genome wide significant hits obtained in early studies. For example, the number of genome wide significant hits for some traits (e.g., height) has been growing steadily, along with h2 accounted for using just those hits, eventually approaching the GCTA prediction. That is, even *without* GCTA the steady progress of GWAS shows that common variants account for significant heritability (amount of "missing" heritability steadily declines with GWAS sample size), so the precise reliability of GCTA becomes less important.

Regarding this paper, they make what sound like strong theoretical points in the text, but the simulation results don't seem to justify the aggressive rhetoric. The only point they really make in figs 4,7 is that the error estimate from GCTA in the case where SNP coverage is inadequate (i.e., using 5k out of 50k SNPs) are way off. But this doesn't correspond to any real world study that we care about. Real world results show that as you approach ~few x 100k SNPs used the h2 result asymptotes (approaches its limiting value), because you have enough coverage of common variants. The authors of the paper seem confused about this point -- see "Saturation of heritability estimates" section.

What they should do is simulate repeatedly with multiple disjoint populations (using good SNP coverage) and see how the heritability results fluctuate. But I think that kind of calculation has been done by other people and does not show large fluctuations in h2.

Well, since you got me to write this much already I suppose I should promote this to an actual blog post at some point ... Please keep in mind that I've only given the paper a quick read so I might be missing something important. Happy New Year!
Here is the paper:
Limitations of GCTA as a solution to the missing heritability problem
http://www.pnas.org/content/early/2015/12/17/1520109113

The genetic contribution to a phenotype is frequently measured by heritability, the fraction of trait variation explained by genetic differences. Hundreds of publications have found DNA polymorphisms that are statistically associated with diseases or quantitative traits [genome-wide association studies (GWASs)]. Genome-wide complex trait analysis (GCTA), a recent method of analyzing such data, finds high heritabilities for such phenotypes. We analyze GCTA and show that the heritability estimates it produces are highly sensitive to the structure of the genetic relatedness matrix, to the sampling of phenotypes and subjects, and to the accuracy of phenotype measurements. Plausible modifications of the method aimed at increasing stability yield much smaller heritabilities. It is essential to reevaluate the many published heritability estimates based on GCTA.
It's important to note that although GCTA fits a model with random effects, it purports to estimate the heritability of more realistic genetic architectures with some other (e.g., sparse) distribution of effect sizes (see Lee and Chow paper at bottom of this post). The authors of this PNAS paper seem to take the random effects assumption more seriously than the GCTA originators themselves. The latter fully expected a saturation effect once enough SNPs are used; the former seem to think it violates the fundamental nature of the model. Indeed, AFAICT, the toy models in the PNAS simulations assume all 50k SNPs affect the trait, and they run simulations where only 5k at a time are included in the computation. This is likely the opposite of the real world situation, in which a relatively small number (e.g., ~10k SNPs) affect the trait, and by using a decent array with > 200k SNPs one already obtains sensitivity to the small subset.

One can easily show that genetic architectures of complex traits tend to be sparse: most of the heritability is accounted for by a small subset of alleles. (Here "small" means a small fraction of ~ millions of SNPs: e.g., 10k SNPs.) See section 3.2 of On the genetic architecture of intelligence and other quantitative traits for an explanation of how to roughly estimate the sparsity using genetic Hamming distances. In our work on Compressed Sensing applied to genomics, we showed that much of the heritability for many complex traits can be recovered if sample sizes of order millions are available for analysis. Once these large data sets are available, this entire debate about missing heritability and GCTA heritability estimates will recede in importance. (See talk and slides here.)

For more discussion, see Why does GCTA work?
This paper, by two of my collaborators, examines the validity of a recently introduced technique called GCTA (Genome-wide Complex Trait Analysis). GCTA allows an estimation of heritability due to common SNPs using relatively small sample sizes (e.g., a few thousand genotype-phenotype pairs). The new method is independent of, but delivers results consistent with, "classical" methods such as twin and adoption studies. To oversimplify, it examines pairs of unrelated individuals and computes the correlation between pairwise phenotype similarity and genotype similarity (relatedness). It has been applied to height, intelligence, and many medical and psychiatric conditions.

When the original GCTA paper (Common SNPs explain a large proportion of the heritability for human height) appeared in Nature Genetics it stimulated quite a lot of attention. But I was always uncertain of the theoretical justification for the technique -- what are the necessary conditions for it to work? What are conservative error estimates for the derived heritability? My impression, from talking to some of the authors, is that they had a mainly empirical view of these questions. The paper below elaborates significantly on the theory behind GCTA. 
Conditions for the validity of SNP-based heritability estimation

James J Lee, Carson C Chow
doi: 10.1101/003160

...

Sunday, March 30, 2014

Why does GCTA work?

This paper, by two of my collaborators, examines the validity of a recently introduced technique called GCTA (Genome-wide Complex Trait Analysis). GCTA allows an estimation of heritability due to common SNPs using relatively small sample sizes (e.g., a few thousand genotype-phenotype pairs). The new method is independent of, but delivers results consistent with, "classical" methods such as twin and adoption studies. To oversimplify, it examines pairs of unrelated individuals and computes the correlation between pairwise phenotype similarity and genotype similarity (relatedness). It has been applied to height, intelligence, and many medical and psychiatric conditions.

When the original GCTA paper (Common SNPs explain a large proportion of the heritability for human height) appeared in Nature Genetics it stimulated quite a lot of attention. But I was always uncertain of the theoretical justification for the technique -- what are the necessary conditions for it to work? What are conservative error estimates for the derived heritability? My impression, from talking to some of the authors, is that they had a mainly empirical view of these questions. The paper below elaborates significantly on the theory behind GCTA.
Conditions for the validity of SNP-based heritability estimation

James J Lee, Carson C Chow
doi: 10.1101/003160

ABSTRACT

The heritability of a trait ($h^2$) is the proportion of its population variance caused by genetic differences, and estimates of this parameter are important for interpreting the results of genome-wide association studies (GWAS). In recent years, researchers have adopted a novel method for estimating a lower bound on heritability directly from GWAS data that uses realized genetic similarities between nominally unrelated individuals. The quantity estimated by this method is purported to be the contribution to heritability that could in principle be recovered from association studies employing the given panel of SNPs ($h^2_\textrm{SNP}$). Thus far the validity of this approach has mostly been tested empirically. Here, we provide a mathematical explication and show that the method should remain a robust means of obtaining $h^2_\textrm{SNP}$ under circumstances wider than those under which it has so far been derived.

Blog Archive

Labels