Showing posts sorted by date for query "james lee". Sort by relevance Show all posts
Showing posts sorted by date for query "james lee". Sort by relevance Show all posts

Wednesday, January 26, 2022

Saturday, January 15, 2022

Manifold Returns!

I'm working on the return of Manifold. I just did the first interview yesterday, with James Lee. I'm not sure when it will be released, but soon I hope.

Thanks to everyone who enjoyed the original show. I received a lot of requests to bring it back over the last 18 months since we went on hiatus. Please make suggestions for guests and show topics!

I'll try to get Dominic Cummings for a long interview. For now, see this great piece in the Guardian about his substack: 

Intoxicating, insidery and infuriating: everything I learned about Dominic Cummings from his £10-a-month blog


Here are the 51 episodes from our first run, with transcripts: https://manifoldlearning.com/


Michigan State University owns the copyright to the old stuff, so we may have to create a new channel and web site. This new project will be entirely mine and I will probably allow comments on the YouTube channel -- we turned those off because the university was/is risk averse about free expression ¯\_(ツ)_/¯ 

I hope to have Corey as a frequent guest on the show, but I will be restarting in solo mode. Here's the first episode where Corey and I introduced ourselves. Very rough around the edges, but still fun for me to listen to again.


Sunday, March 21, 2021

The Contribution of Cognitive and Noncognitive Skills to Intergenerational Social Mobility (McGue et al. 2020)

If you have the slightest pretension to expertise concerning social mobility, meritocracy, inequality, genetics, psychology, economics, education, history, or any related subjects, I urge you to carefully study this paper.
The Contribution of Cognitive and Noncognitive Skills to Intergenerational Social Mobility  
(Psychological Science https://doi.org/10.1177/0956797620924677)
Matt McGue, Emily A. Willoughby, Aldo Rustichini, Wendy Johnson, William G. Iacono, James J. Lee 
We investigated intergenerational educational and occupational mobility in a sample of 2,594 adult offspring and 2,530 of their parents. Participants completed assessments of general cognitive ability and five noncognitive factors related to social achievement; 88% were also genotyped, allowing computation of educational-attainment polygenic scores. Most offspring were socially mobile. Offspring who scored at least 1 standard deviation higher than their parents on both cognitive and noncognitive measures rarely moved down and frequently moved up. Polygenic scores were also associated with social mobility. Inheritance of a favorable subset of parent alleles was associated with moving up, and inheritance of an unfavorable subset was associated with moving down. Parents’ education did not moderate the association of offspring’s skill with mobility, suggesting that low-skilled offspring from advantaged homes were not protected from downward mobility. These data suggest that cognitive and noncognitive skills as well as genetic factors contribute to the reordering of social standing that takes place across generations.
From the paper:
We believe that a reasonable explanation of our findings is that the degree to which individuals are more or less skilled than their parents contributes to their upward or downward mobility. Behavioral genetic and genomic research has established the heritability of social achievements (Conley, 2016) as well as the skills thought to underlie them (Bouchard & McGue, 2003). Nonetheless, these associations may be due to passive gene–environment correlation, whereby high-achieving parents both transmit genes and provide a rearing environment that promotes their children’s social success (Scarr & McCartney, 1983). Our within-family design controlled for passive gene–environment correlation effects. Although offspring inherit all of their genes from their parents, they inherit a random subset of parental alleles because of meiotic segregation. Consequently, some offspring inherit a favorable subset of their parents’ alleles, whereas others inherit a less favorable subset. We found, as did previous researchers (Belsky et al., 2018), that the inheritance of a favorable subset of alleles was associated with an increased likelihood of upward mobility... 
...In summary, our analysis of intergenerational social mobility in a sample of 2,594 offspring from 1,321 families found that (a) most individuals were educationally and occupationally mobile, (b) mobility was predicted by offspring–parent differences in skills and genetic endowment, and (c) the relationship of offspring skills with social mobility did not vary significantly by parent social background. In an era in which there is legitimate concern over social stagnation, our findings are noteworthy in identifying the circumstances when parents’ educational and occupational success is not reproduced across generations.

See also Game Over: Genomic Prediction of Social Mobility (PNAS July 9, 2018: 201801238). Both papers provide out of sample validation of polygenic predictors for cognitive ability, specifically of the relationship to intergenerational social mobility.


Monday, June 15, 2020

Support Freedom of Ideas and Inquiry at MSU

These are letters of support sent on my behalf to the MSU President: presidentstanley@msu.edu

Several are deep, detailed scholarly documents. They firmly rebut the false accusations of the Twitter mob.

Corey Washington's individual letter is over 5000 words long. James Lee's is over 3000 words. The authors have graciously allowed them to be made public.

Corey says:
I have known Steve for 30 years and can attest that he is not a racist or a sexist. Steve is one of the most scrupulously fair people I have ever met and I have seen no evidence that he has ever discriminated against anyone on the basis of their race, sex, or any other status.

... Hsu participated in a 2018 debate at MSU’s Institute for Quantitative Health, Science, and Engineering, at the invitation of Director Chris Contag. The topic was human genetic engineering, and Hsu’s counterpart was MSU bioethicist Len Fleck. A number of people who attended tell me that they found Steve’s view thoughtful and balanced. None that I know of came away from the debate thinking he is a eugenicist in the way that the rest of us are not. It is unclear to me why MSU GEU came to a different conclusion.

Sign the support petition. Email the president: presidentstanley@msu.edu

A (not necessarily up to date) list of signatories, which includes hundreds of professors from MSU and around the world, total ~1500 as of early June 19.


Letter from
Corey Washington, Director of Analytics and Strategic Projects, OSVPRI
Joseph Cesario, Associate Professor, Psychology
Wei Liao, Professor and Director, MSU Anaerobic Digestion Research and Education Center, Biosystems and Agricultural Engineering
John (Xuefeng) Jiang, Professor and Plante Moran Faculty Fellow, Accounting and Information Systems

Letter from Corey Washington, Director of Analytics and Strategic Projects, MSU

Letter from Matt McGue, Regents Professor of Psychology, University of Minnesota

Letter from Russell Warne, Associate Professor of Psychology, Utah Valley University

Letter from Mark Dykman, Professor of Physics, MSU

Letter from Zach Hambrick, Professor of Psychology, MSU

Letter from James Lee, Associate Professor of Psychology, University of Minnesota

Letter from Richard Haier, Emeritus Professor, UC Irvine, author of The Neuroscience of Intelligence (Cambridge)

Letter from Steven Pinker, Johnstone Family Professor of Psychology, Harvard University


Media coverage:

A Twitter Mob Takes Down an Administrator at Michigan State (Wall Street Journal June 25)

Scholar forced to resign over study that found police shootings not biased against blacks (The College Fix)

On Steve Hsu and the Campaign to Thwart Free Inquiry (Quillette)

Michigan State University VP of Research Ousted (Reason Magazine, Eugene Volokh, UCLA)

Research isn’t advocacy (NY Post Editorial Board)

Podcast interview on Tom Woods show (July 2)

College professor forced to resign for citing study that found police shootings not biased against blacks (Law Enforcement Today, July 5)

"Racist" College Researcher Ousted After Sharing Study Showing No Racial Bias In Police Shootings (ZeroHedge, July 6)

Twitter mob: College researcher forced to resign after study finding no racial bias in police shootings (Reclaim the Net, July 8)

Horowitz: Asian-American researcher fired from Michigan State administration for advancing facts about police shootings (The Blaze, July 8)

I Cited Their Study, So They Disavowed It: If scientists retract research that challenges reigning orthodoxies, politics will drive scholarship (Wall Street Journal July 8)

Conservative author cites research on police shootings and race. Researchers ask for its retraction in response (The College Fix, July 8)

Academics Seek to Retract Study Disproving Racist Police Shootings After Conservative Cites It (Hans Bader, CNSNews, July 9)

The Ideological Corruption of Science (theoretical physicist Lawrence Krauss in the Wall Street Journal, July 12)

Foundation for Individual Rights in Education: "chilling academic freedom" (Peter Bonilla, July 22)

Friday, July 26, 2019

RadioLab on embryo selection in IVF



I'm in this RadioLab podcast covering genetic selection of embryos in IVF. Apologies to SSGAC, Robert Plomin, Ian Deary, James Lee, Tom Bouchard, and countless other dedicated scientists for the impression given that progress in genomics of cognitive ability is largely my work. See last paragraph below.

This is the email I sent to RadioLab this morning:
Hi Pat and Michelle,

Congratulations on a high quality podcast. I thought you were admirably fair and balanced. I also thought the production (esp. the music) was excellent.

My main comment is that the juxtaposition between my remarks and Benjamin's is misleading: when he says 60-40 or 55% chance of rank ordering properly, that is a very different question than identifying an outlier who is, say, among the 1% highest in risk. We are not trying to rank order embryos, but to warn against unusual risk of a medical condition.

To use the SAT analogy, given two kids with scores 1250 and 1200, only some of the time does the 1250 kid end up with a higher GPA. (You can't predict rank order very well.) But if the engineering dean admits an SAT 770 kid (i.e., a negative outlier compared to the average score of, say, 1300 among engineers) in his freshman class, he knows the likelihood is high that the kid will struggle. Benjamin is talking about the first scenario, and I am talking about the second.

Finally, I realize that to hook listeners you had to make me the focus of the episode. But I want to make clear that many scientists contribute to this work, which I feel will ultimately be beneficial to our species and civilization. I am just a small part of a worldwide research endeavor.

Best wishes,
Steve
For more on recent progress in genomic prediction, see The Diffusion of Knowledge.

Friday, July 27, 2018

Insight Podcast: James Lee interview on SSGAC EA3



Spencer Wells and Razib Khan interview James Lee (Professor of Psychology, University of Minnesota, BA Berkeley, PhD Harvard) about the recent SSGAC EA3 GWAS.

Comment: James mentions that EA3 may be approaching the GCTA h2 limit (~0.15? so limiting r ~ 0.4) already. But the limit for actual cognitive ability is much higher; with enough data I think we could get to r ~ 0.6 or even r ~ 0.7 eventually for common SNPs -- similar to height.

United Club, HK International Airport



James, me, Chris Chang. (About $1M worth of Illumina HiSeqs in crates behind us?)


Monday, July 23, 2018

SSGAC EA3: genomic prediction of educational attainment and related cognitive phenotypes

Years ago I predicted that:

1. Cognitive ability would turn out to be influenced by many thousands of genetic variants, each of small effect.

2. With large enough sample size we would detect these variants and eventually construct genomic predictors.

The Nature Genetics paper below from the SSGAC collaboration takes a significant step in that direction.

Although the study used over a million genotypes, the data had to be aggregated across many sub-cohorts using summary statistics only. This does not permit the L1-penalized optimization we used to build our height predictor.

For out of sample validation of the results below, see this PNAS paper, which (unusually) appeared before the paper on which it is based.

The lead author James Lee is on the left below. Chris Chang, author of Plink 2.0, is on the right. The photo was taken in 2010 at BGI -- they are standing in front of crates of Illumina sequencers.



Article | Published: 23 July 2018

Gene discovery and polygenic prediction from a genome-wide association study of educational attainment in 1.1 million individuals

James J. Lee, Robbee Wedow, […]David Cesarini
Nature Genetics (2018)

Abstract
Here we conducted a large-scale genetic association analysis of educational attainment in a sample of approximately 1.1 million individuals and identify 1,271 independent genome-wide-significant SNPs. For the SNPs taken together, we found evidence of heterogeneous effects across environments. The SNPs implicate genes involved in brain-development processes and neuron-to-neuron communication. In a separate analysis of the X chromosome, we identify 10 independent genome-wide-significant SNPs and estimate a SNP heritability of around 0.3% in both men and women, consistent with partial dosage compensation. A joint (multi-phenotype) analysis of educational attainment and three related cognitive phenotypes generates polygenic scores that explain 11–13% of the variance in educational attainment and 7–10% of the variance in cognitive performance. This prediction accuracy substantially increases the utility of polygenic scores as tools in research.
A nice figure from the paper: Add Health (National Longitudinal Study of Adolescent to Adult Health) and HRS (Health in Retirement Study) are two longitudinal cohorts that have been genotyped; horizontal axis is polygenic score. It appears that individuals with top quintile polygenic scores are about 5 times more likely to complete college than bottom quintile individuals.


Here's a comment on the paper I provided to a journalist:
The EA3 predictor correlates about 0.35 with educational attainment, and slightly less well with measured cognitive ability. While this is far from perfect prediction, it does allow identification of individuals, using DNA alone, who are at unusual risk of being well below average in cognitive ability or struggling in school. Standardized tests, such as SAT, ACT, GRE, LSAT, etc., typically also correlate roughly 0.35 with educational outcomes like grade point average, degree completion, etc. In this sense, the genomic predictor is comparable to widely used tests and it will certainly improve as more data are analyzed. See figure.

Monday, August 14, 2017

Estimation of genetic architecture for complex traits using GWAS data

These authors extrapolate from existing data to predict sample sizes needed to identify SNPs which explain a large portion of heritability in a variety of traits. For cognitive ability (see red curves in figure below), they predict sample sizes of ~million individuals will suffice.

See also More Shock and Awe: James Lee and SSGAC in Oslo, 600 SNP hits.
Estimation of complex effect-size distributions using summary-level statistics from genome-wide association studies across 32 complex traits and implications for the future

Yan Zhang, Guanghao Qi, Ju-Hyun Park, Nilanjan Chatterjee (Johns Hopkins University)

Summary-level statistics from genome-wide association studies are now widely used to estimate heritability and co-heritability of traits using the popular linkage-disequilibrium-score (LD-score) regression method. We develop a likelihood-based approach for analyzing summary-level statistics and external LD information to estimate common variants effect-size distributions, characterized by proportion of underlying susceptibility SNPs and a flexible normal-mixture model for their effects. Analysis of summary-level results across 32 GWAS reveals that while all traits are highly polygenic, there is wide diversity in the degrees of polygenicity. The effect-size distributions for susceptibility SNPs could be adequately modeled by a single normal distribution for traits related to mental health and ability and by a mixture of two normal distributions for all other traits. Among quantitative traits, we predict the sample sizes needed to identify SNPs which explain 80% of GWAS heritability to be between 300K-500K for some of the early growth traits, between 1-2 million for some anthropometric and cholesterol traits and multiple millions for body mass index and some others. The corresponding predictions for disease traits are between 200K-400K for inflammatory bowel diseases, close to one million for a variety of adult onset chronic diseases and between 1-2 million for psychiatric diseases.


This figure shows predicted effect size distributions for a number of quantitative traits. You can see that height and intelligence are somewhat different, but not dramatically so.

Saturday, May 06, 2017

More Shock and Awe: James Lee and SSGAC in Oslo, 600 SNP hits


To quote James Lee, the first author listed below: "Shock and Awe" for those who doubt that cognitive ability is influenced by genetic variants.

See work from a year ago: ~100 hits from 300k individuals. Now ~600 hits from 750k. (SNPs associated with EA are likely to also be associated with cognitive ability -- see figure at link above.)
47th Behavior Genetics Annual Meeting, Oslo, Norway

GWAS of Educational Attainment, Phase 3: Biological Findings

Abstract
Genetic factors are estimated to account for at least 20% of the variation across individuals for educational attainment (Rietveld et al., 2013). The results of the latest GWAS for educational attainment identified 74 genome-wide significant loci for educational attainment (Okbay et al., 2016). Here, in one of the largest GWAS to date, we increase our sample to nearly 750,000 individuals, and we identify over 600 genome-wide significant loci associated with the number of years of schooling completed. Note that at the time of presentation, we will likely have updated our meta-analysis to include over 1,000,000 individuals

In this presentation, I will focus on the biological implications of the GWAS results. At the time of writing, 1,656 genes are significantly prioritized, a more than 10-fold increase since our previous report (Okbay et al., 2016). The newly significant results reinforce the biological theme of prenatal brain development and also bring to the foreground new themes that shed light on the biological underpinnings of cognitive performance and other traits affecting educational attainment.

Authors
James Lee (University of Minnesota - Twin Cities), Aysu Okbay (Free University Amsterdam), Robbee Wedow (University of Colorado - Boulder), Edward Kong (Harvard University), Patrick Turley (Broad Institute of MIT and Harvard), Meghan Zacher (Harvard University), Kevin Thom (New York University), Anh Tuan Nguyen Viet (University of Southern California), Omeed Maghzian (Harvard University, NBER), Richard Karlsson Linnér (Vrije Universiteit Amsterdam), Matthew Robinson (The University of Queensland), Social Science Genetic Association Consortium (NA), Peter Visscher (The University of Queensland), Daniel Benjamin (University of Southern California), David Cesarini (New York University)
Note the data here have only been analyzed using summary statistics coming from each sub-cohort. More powerful methods may soon become available:
Penalized regression from summary statistics

One of the difficulties in genomics is that when DNA donors are consented for a study, the agreements generally do not allow sharing (aggregation) of genomic data across multiple studies. This leads to isolated silos of data that can't be fully shared. However, computations can be performed on one silo at a time, with the results ("summary statistics") shared within a larger collaboration. Most of the leading GWAS collaborations (e.g., GIANT for height, SSGAC for cognitive ability) rely on shared statistics. Simple regression analysis (one SNP at a time) can be conducted using just summary statistics, but more sophisticated algorithms cannot. These more sophisticated methods can generate a better phenotype predictor, using less data, than a SNP by SNP analysis.
A successful implementation like the one described at the link above could produce many (several times!) more hits and significantly more variance accounted for by corresponding predictors. Stay tuned!

Note Added: I'm getting lots of questions about how to interpret these results, so here are some comments.

1. I predicted ~10k variants would account for most of the heritability due to common SNPs (i.e., about 50% of total variance; allowing a predictor which correlates ~0.7 with actual cognitive ability). The rate of discovery of genome-wide significant hits and corresponding variance accounted for seems consistent with this prediction. Genetic associations are most easily discovered for variants which are common (e.g., have ~0.5 Minor Allele Frequency, not 0.05) and have large effect sizes. But alleles with this combination of properties are rare. As statistical power increases, one starts to discover (more and more) variants of lower frequency and/or lower effect size. A reasonable guess at the genetic architecture suggests a higher density of such variants, and is consistent with an accelerating rate of discovery of SNP hits (~100 hits from 300k individuals, ~600 hits from 750k). There are more efficient methods that, I believe, would discover nearly all the variants given sample size of ~1M well-phenotyped individuals. But these methods require more than just summary statistics.

I made a similar prediction of ~10k variants for height, and our (unpublished) genomic prediction results make me fairly confident that this will turn out to be correct. We now have moderately good height predictors and they are getting better very fast. That ~10k variants will turn out to be responsible for most of the variation in cognitive ability is still at a somewhat lower confidence level.

2. People are still confused about how many + variants above the mean in the population are required to make a "genius" (or super-genius). I managed to compress the explanation enough to fit in a tweet:
Flip coin 10000 times. 5000 + sqrt(10000)/2 = 5050 heads is +1SD outcome. 5100 is +2SD, etc. sqrt(N) << N for N large. Binomial~Normal dist.
You can see that even if cognitive ability is controlled by ~10k variants, flipping only ~100 of them is enough to cause a big difference in actual intelligence. Flipping a few hundred could get us to super-geniuses beyond anything in human history.

3. If you read press accounts related to our creation of the BGI Cognitive Genomics Lab back in 2011 (at that time there were zero genome-wide significant alleles associated with intelligence), you can find quotes from genomics "experts" asserting that mankind would never discover the genetic architecture of cognitive ability. (Such quotes are easy to obtain even today!) A Bayesian update given what is known in 2017 would call into question the competence of these "experts"!  ;-)

Wednesday, May 11, 2016

74 SNP hits from SSGAC GWAS



The SSGAC discovery of 74 SNP hits on educational attainment (EA) is finally published in Nature. Nature News article.

EA was used in order to assemble as large a sample as possible (~300k individuals). Specific cognitive scores are only available for a much smaller number of individuals. But SNPs associated with EA are likely to also be associated with cognitive ability -- see figure above.

The evidence is strong that cognitive ability is highly heritable and highly polygenic. With even larger samples we'll eventually be able to build good genomic predictors for cognitive ability.
Genome-wide association study identifies 74 loci associated with educational attainment A. Okbay et al. Nature http://dx.doi.org/10.1038/nature17671; 2016

Educational attainment is strongly influenced by social and other environmental factors, but genetic factors are estimated to account for at least 20% of the variation across individuals1. Here we report the results of a genome-wide association study (GWAS) for educational attainment that extends our earlier discovery sample1,2 of 101,069 individuals to 293,723 individuals, and a replication study in an independent sample of 111,349 individuals from the UK Biobank. We identify 74 genome-wide significant loci associated with the number of years of schooling completed. Single- nucleotide polymorphisms associated with educational attainment are disproportionately found in genomic regions regulating gene expression in the fetal brain. Candidate genes are preferentially expressed in neural tissue, especially during the prenatal period, and enriched for biological pathways involved in neural development. Our findings demonstrate that, even for a behavioural phenotype that is mostly environmentally determined, a well-powered GWAS identifies replicable associated genetic variants that suggest biologically relevant pathways. Because educational attainment is measured in large numbers of individuals, it will continue to be useful as a proxy phenotype in efforts to characterize the genetic influences of related phenotypes, including cognition and neuropsychiatric diseases.

Here's what I wrote back in September of 2015, based on a talk given by James Lee on this work.
James Lee talk at ISIR 2015 (via James Thompson) reports on 74 hits at genome-wide statistical significance (p < 5E-8) using educational attainment as the phenotype. Most of these will also turn out to be hits on cognitive ability.

To quote James: "Shock and Awe" for those who doubt that cognitive ability is influenced by genetic variants. This is just the tip of the iceberg, though. I expect thousands more such variants to be discovered before we have accounted for all of the heritability.
74 GENOMIC SITES ASSOCIATED WITH EDUCATIONAL ATTAINMENT PROVIDE INSIGHT INTO THE BIOLOGY OF COGNITIVE PERFORMANCE 
James J Lee

University of Minnesota Twin Cities
Social Science Genetic Association Consortium

Genome-wide association studies (GWAS) have revealed much about the biological pathways responsible for phenotypic variation in many anthropometric traits and diseases. Such studies also have the potential to shed light on the developmental and mechanistic bases of behavioral traits.

Toward this end we have undertaken a GWAS of educational attainment (EA), an outcome that shows phenotypic and genetic correlations with cognitive performance, personality traits, and other psychological phenotypes. We performed a GWAS meta-analysis of ~293,000 individuals, applying a variety of methods to address quality control and potential confounding. We estimated the genetic correlations of several different traits with EA, in essence by determining whether single-nucleotide polymorphisms (SNPs) showing large statistical signals in a GWAS meta-analysis of one trait also tend to show such signals in a meta-analysis of another. We used a variety of bio-informatic tools to shed light on the biological mechanisms giving rise to variation in EA and the mediating traits affecting this outcome. We identified 74 independent SNPs associated with EA (p < 5E-8). The ability of the polygenic score to predict within-family differences suggests that very little of this signal is due to confounding. We found that both cognitive performance (0.82) and intracranial volume (0.39) show substantial genetic correlations with EA. Many of the biological pathways significantly enriched by our signals are active in early development, affecting the proliferation of neural progenitors, neuron migration, axonogenesis, dendrite growth, and synaptic communication. We nominate a number of individual genes of likely importance in the etiology of EA and mediating phenotypes such as cognitive performance.
For a hint at what to expect as more data become available, see Five Years of GWAS Discovery and On the genetic architecture of intelligence and other quantitative traits.


What was once science fiction will soon be reality.
Long ago I sketched out a science fiction story involving two Junior Fellows, one a bioengineer (a former physicist, building the next generation of sequencing machines) and the other a mathematician. The latter, an eccentric, was known for collecting signatures -- signed copies of papers and books authored by visiting geniuses (Nobelists, Fields Medalists, Turing Award winners) attending the Society's Monday dinners. He would present each luminary with an ornate (strangely sticky) fountain pen and a copy of the object to be signed. Little did anyone suspect the real purpose: collecting DNA samples to be turned over to his friend for sequencing! The mathematician is later found dead under strange circumstances. Perhaps he knew too much! ...

Friday, January 01, 2016

GCTA, Missing Heritability, and All That

New Update: Yang, Visscher et al. respond here.

Update: see detailed comments and analysis here and here by Sasha Gusev. Gusev claims that the problems identified in Figs 4,7 are the result of incorrect calculation of the SE (4) and failure to exclude related individuals in the Framingham data (7).


Bioinformaticist E. Stovner asked about a recent PNAS paper which is critical of GCTA. My comments are below.

It's a shame that we don't have a better online platform (e.g., like Quora or StackOverflow) for discussing scientific papers. This would allow the authors of a paper to communicate directly with interested readers, immediately after the paper appears. If the authors of this paper want to correct my misunderstandings, they are welcome to comment here!
I took a quick look at it. My guess is that Visscher et al. will respond to the paper. It has not changed my opinion of GCTA. Note I have always thought the standard errors quoted for GCTA are too optimistic, as the method makes strong assumptions (e.g., fits a model with Gaussian random effect sizes). But directionally it is obvious that total h2 accounted for by all common SNPs is much larger than what you get from only genome wide significant hits obtained in early studies. For example, the number of genome wide significant hits for some traits (e.g., height) has been growing steadily, along with h2 accounted for using just those hits, eventually approaching the GCTA prediction. That is, even *without* GCTA the steady progress of GWAS shows that common variants account for significant heritability (amount of "missing" heritability steadily declines with GWAS sample size), so the precise reliability of GCTA becomes less important.

Regarding this paper, they make what sound like strong theoretical points in the text, but the simulation results don't seem to justify the aggressive rhetoric. The only point they really make in figs 4,7 is that the error estimate from GCTA in the case where SNP coverage is inadequate (i.e., using 5k out of 50k SNPs) are way off. But this doesn't correspond to any real world study that we care about. Real world results show that as you approach ~few x 100k SNPs used the h2 result asymptotes (approaches its limiting value), because you have enough coverage of common variants. The authors of the paper seem confused about this point -- see "Saturation of heritability estimates" section.

What they should do is simulate repeatedly with multiple disjoint populations (using good SNP coverage) and see how the heritability results fluctuate. But I think that kind of calculation has been done by other people and does not show large fluctuations in h2.

Well, since you got me to write this much already I suppose I should promote this to an actual blog post at some point ... Please keep in mind that I've only given the paper a quick read so I might be missing something important. Happy New Year!
Here is the paper:
Limitations of GCTA as a solution to the missing heritability problem
http://www.pnas.org/content/early/2015/12/17/1520109113

The genetic contribution to a phenotype is frequently measured by heritability, the fraction of trait variation explained by genetic differences. Hundreds of publications have found DNA polymorphisms that are statistically associated with diseases or quantitative traits [genome-wide association studies (GWASs)]. Genome-wide complex trait analysis (GCTA), a recent method of analyzing such data, finds high heritabilities for such phenotypes. We analyze GCTA and show that the heritability estimates it produces are highly sensitive to the structure of the genetic relatedness matrix, to the sampling of phenotypes and subjects, and to the accuracy of phenotype measurements. Plausible modifications of the method aimed at increasing stability yield much smaller heritabilities. It is essential to reevaluate the many published heritability estimates based on GCTA.
It's important to note that although GCTA fits a model with random effects, it purports to estimate the heritability of more realistic genetic architectures with some other (e.g., sparse) distribution of effect sizes (see Lee and Chow paper at bottom of this post). The authors of this PNAS paper seem to take the random effects assumption more seriously than the GCTA originators themselves. The latter fully expected a saturation effect once enough SNPs are used; the former seem to think it violates the fundamental nature of the model. Indeed, AFAICT, the toy models in the PNAS simulations assume all 50k SNPs affect the trait, and they run simulations where only 5k at a time are included in the computation. This is likely the opposite of the real world situation, in which a relatively small number (e.g., ~10k SNPs) affect the trait, and by using a decent array with > 200k SNPs one already obtains sensitivity to the small subset.

One can easily show that genetic architectures of complex traits tend to be sparse: most of the heritability is accounted for by a small subset of alleles. (Here "small" means a small fraction of ~ millions of SNPs: e.g., 10k SNPs.) See section 3.2 of On the genetic architecture of intelligence and other quantitative traits for an explanation of how to roughly estimate the sparsity using genetic Hamming distances. In our work on Compressed Sensing applied to genomics, we showed that much of the heritability for many complex traits can be recovered if sample sizes of order millions are available for analysis. Once these large data sets are available, this entire debate about missing heritability and GCTA heritability estimates will recede in importance. (See talk and slides here.)

For more discussion, see Why does GCTA work?
This paper, by two of my collaborators, examines the validity of a recently introduced technique called GCTA (Genome-wide Complex Trait Analysis). GCTA allows an estimation of heritability due to common SNPs using relatively small sample sizes (e.g., a few thousand genotype-phenotype pairs). The new method is independent of, but delivers results consistent with, "classical" methods such as twin and adoption studies. To oversimplify, it examines pairs of unrelated individuals and computes the correlation between pairwise phenotype similarity and genotype similarity (relatedness). It has been applied to height, intelligence, and many medical and psychiatric conditions.

When the original GCTA paper (Common SNPs explain a large proportion of the heritability for human height) appeared in Nature Genetics it stimulated quite a lot of attention. But I was always uncertain of the theoretical justification for the technique -- what are the necessary conditions for it to work? What are conservative error estimates for the derived heritability? My impression, from talking to some of the authors, is that they had a mainly empirical view of these questions. The paper below elaborates significantly on the theory behind GCTA. 
Conditions for the validity of SNP-based heritability estimation

James J Lee, Carson C Chow
doi: 10.1101/003160

...

Saturday, September 19, 2015

SNP hits on cognitive ability from 300k individuals

James Lee talk at ISIR 2015 (via James Thompson) reports on 74 hits at genome-wide statistical significance (p < 5E-8) using educational attainment as the phenotype. Most of these will also turn out to be hits on cognitive ability.

To quote James: "Shock and Awe" for those who doubt that cognitive ability is influenced by genetic variants. This is just the tip of the iceberg, though. I expect thousands more such variants to be discovered before we have accounted for all of the heritability.
74 GENOMIC SITES ASSOCIATED WITH EDUCATIONAL ATTAINMENT PROVIDE INSIGHT INTO THE BIOLOGY OF COGNITIVE PERFORMANCE 
James J Lee

University of Minnesota Twin Cities
Social Science Genetic Association Consortium

Genome-wide association studies (GWAS) have revealed much about the biological pathways responsible for phenotypic variation in many anthropometric traits and diseases. Such studies also have the potential to shed light on the developmental and mechanistic bases of behavioral traits.

Toward this end we have undertaken a GWAS of educational attainment (EA), an outcome that shows phenotypic and genetic correlations with cognitive performance, personality traits, and other psychological phenotypes. We performed a GWAS meta-analysis of ~293,000 individuals, applying a variety of methods to address quality control and potential confounding. We estimated the genetic correlations of several different traits with EA, in essence by determining whether single-nucleotide polymorphisms (SNPs) showing large statistical signals in a GWAS meta-analysis of one trait also tend to show such signals in a meta-analysis of another. We used a variety of bio-informatic tools to shed light on the biological mechanisms giving rise to variation in EA and the mediating traits affecting this outcome. We identified 74 independent SNPs associated with EA (p < 5E-8). The ability of the polygenic score to predict within-family differences suggests that very little of this signal is due to confounding. We found that both cognitive performance (0.82) and intracranial volume (0.39) show substantial genetic correlations with EA. Many of the biological pathways significantly enriched by our signals are active in early development, affecting the proliferation of neural progenitors, neuron migration, axonogenesis, dendrite growth, and synaptic communication. We nominate a number of individual genes of likely importance in the etiology of EA and mediating phenotypes such as cognitive performance.
For a hint at what to expect as more data become available, see Five Years of GWAS Discovery and On the genetic architecture of intelligence and other quantitative traits.


What was once science fiction will soon be reality.
Long ago I sketched out a science fiction story involving two Junior Fellows, one a bioengineer (a former physicist, building the next generation of sequencing machines) and the other a mathematician. The latter, an eccentric, was known for collecting signatures -- signed copies of papers and books authored by visiting geniuses (Nobelists, Fields Medalists, Turing Award winners) attending the Society's Monday dinners. He would present each luminary with an ornate (strangely sticky) fountain pen and a copy of the object to be signed. Little did anyone suspect the real purpose: collecting DNA samples to be turned over to his friend for sequencing! The mathematician is later found dead under strange circumstances. Perhaps he knew too much! ...

Friday, March 13, 2015

The Fourth Law of Behavior Genetics?


I believe the law stated below almost follows from the observation that humans brains are complex machines: hence the DNA blueprint has many components, and variance is spread over these components  :^)

However, note the evidence for discrete genetic modules of large effect in other species: Discrete genetic modules can control complex behavior (burrowing behavior in cute mouse in picture at bottom), As flies to wanton boys are we to the gods (discrete genetic controls on drosophila behavior).

THE FOURTH LAW OF BEHAVIOR GENETICS

Christopher F. Chabris, Union College
James J. Lee, University of Minnesota Twin Cities
David Cesarini, New York University
Daniel J. Benjamin, Cornell University and University of Southern California
David I. Laibson, Harvard University

Abstract
Behavior genetics is the study of the relationship between genetic variation and psychological traits. Turkheimer (2000) proposed “Three Laws of Behavior Genetics” based on empirical regularities observed in studies of twins and other kinships. On the basis of molecular studies that have measured DNA variation directly, we propose a Fourth Law of Behavior Genetics: “A typical human behavioral trait is associated with very many genetic variants, each of which accounts for a very small percentage of the behavioral variability.” This law explains several consistent patterns in the results of gene discovery studies, including the failure of candidate gene studies to robustly replicate, the need for genome-wide association studies (and why such studies have a much stronger replication record), and the crucial importance of extremely large samples in these endeavors. We review the evidence in favor of the Fourth Law and discuss its implications for the design and interpretation of gene-behavior research.

Wednesday, October 01, 2014

Adventures in the high dimensional space of genomes


2000+ views in 4 months is not bad considering that this is a genomics paper but uses terms like phase transition, sparsity, L1-penalized regression, Gaussian random matrices, etc. I wish I knew how many views the arXiv and BioMed Central versions of the paper have received. Related posts.
Dear Dr Hsu,

We thought you might be interested to know how many people have read your article:

Applying compressed sensing to genome-wide association studies
Shashaank Vattikuti, James J Lee, Christopher C Chang, Stephen D H Hsu and Carson C Chow
GigaScience, 3:10 (16 Jun 2014)
http://www.gigasciencejournal.com/content/3/1/10

Total accesses to this article since publication: 2266

This figure includes accesses to the full text, abstract and PDF of the article on the GigaScience website. It does not include accesses from PubMed Central or other archive sites (see http://www.biomedcentral.com/about/archive). The total access statistics for your article are therefore likely to be significantly higher. ...
My guess is still that it will take some time before these methods become widely understood in genomics.
Crossing boundaries: ... In a similar way Turing found a home in Cambridge mathematical culture, yet did not belong entirely to it. The division between 'pure' and 'applied' mathematics was at Cambridge then as now very strong, but Turing ignored it, and he never showed mathematical parochialism. If anything, it was the attitude of a Russell that he acquired, assuming that mastery of so difficult a subject granted the right to invade others.

PS I will be at the ASHG meeting in San Diego later this month along with (I think) all of the other authors of the paper. Vattikuti will be giving a poster session.

Friday, April 11, 2014

Human Capital, Genetics and Behavior

See you in Chicago next week :-)
HCEO: Human Capital and Economic Opportunity Global Working Group

Conference on Genetics and Behavior

April 18, 2014 to April 19, 2014

This meeting will bring together researchers from a range of disciplines who have been exploring the role of genetic influences on socioeconomic outcomes. The approaches taken to incorporating genes into social science models differ widely. The first goal of the conference is to provide a forum in which alternative frameworks are discussed and critically evaluated. Second, we are hopeful that the meeting will trigger extended interactions and even future collaboration. Third, the meeting will help focus future genetics-related initiatives by the Human Capital and Economic Opportunity Global Working Group, which is pursuing the study of inequality and social mobility over the next several years.

PROGRAM

9:00 to 11:00
Genes and Socioeconomic Aggregates
Gregory Cochran University of Utah
Steven Durlauf University of Wisconsin–Madison
Henry Harpending University of Utah
Aldo Rustichini University of Minnesota
Enrico Spolaore Tufts University

11:30 to 1:30
Population-Based Studies
Sara Jaffee University of Pennsylvania
Matthew McGue University of Minnesota
Peter Molenaar
Jenae Neiderhiser

2:30 to 4:30
Genome-Wide Association Studies (GWAS)
Daniel Benjamin Cornell University
David Cesarini New York University
Dalton Conley New York University/NBER
Jason Fletcher University of Wisconsin–Madison
Philipp Koellinger University of Amsterdam

APRIL 19, 2014

9:00 to 11:00
Neuroscience
Paul Glimcher New York University
Jonathan King National Institute on Aging
Aldo Rustichini University of Minnesota

11:30 to 1:30
Intelligence
Stephen Hsu Michigan State University
Wendy Johnson University of Edinburgh
Rodrigo Pinto The University of Chicago

2:30 to 4:30
Role of Genes in Understanding Socioeconomic Status
Gabriella Conti University College London
Steven Durlauf University of Wisconsin–Madison
Felix Elwert University of Wisconsin–Madison
James Lee University of Minnesota

Sunday, March 30, 2014

Why does GCTA work?

This paper, by two of my collaborators, examines the validity of a recently introduced technique called GCTA (Genome-wide Complex Trait Analysis). GCTA allows an estimation of heritability due to common SNPs using relatively small sample sizes (e.g., a few thousand genotype-phenotype pairs). The new method is independent of, but delivers results consistent with, "classical" methods such as twin and adoption studies. To oversimplify, it examines pairs of unrelated individuals and computes the correlation between pairwise phenotype similarity and genotype similarity (relatedness). It has been applied to height, intelligence, and many medical and psychiatric conditions.

When the original GCTA paper (Common SNPs explain a large proportion of the heritability for human height) appeared in Nature Genetics it stimulated quite a lot of attention. But I was always uncertain of the theoretical justification for the technique -- what are the necessary conditions for it to work? What are conservative error estimates for the derived heritability? My impression, from talking to some of the authors, is that they had a mainly empirical view of these questions. The paper below elaborates significantly on the theory behind GCTA.
Conditions for the validity of SNP-based heritability estimation

James J Lee, Carson C Chow
doi: 10.1101/003160

ABSTRACT

The heritability of a trait ($h^2$) is the proportion of its population variance caused by genetic differences, and estimates of this parameter are important for interpreting the results of genome-wide association studies (GWAS). In recent years, researchers have adopted a novel method for estimating a lower bound on heritability directly from GWAS data that uses realized genetic similarities between nominally unrelated individuals. The quantity estimated by this method is purported to be the contribution to heritability that could in principle be recovered from association studies employing the given panel of SNPs ($h^2_\textrm{SNP}$). Thus far the validity of this approach has mostly been tested empirically. Here, we provide a mathematical explication and show that the method should remain a robust means of obtaining $h^2_\textrm{SNP}$ under circumstances wider than those under which it has so far been derived.

Friday, December 20, 2013

Neanderthals dumb?


This figure is from the Supplement (p.62) of a recent Nature paper describing a high quality genome sequence obtained from the toe of a female Neanderthal who lived in the Altai mountains in Siberia. Interestingly, copy number variation at 16p11.2 is one of the structural variants identified in a recent deCODE study as related to IQ depression; see earlier post Structural genomic variants (CNVs) affect cognition.

From the Supplement (p.62):
Of particular interest is the modern human-specific duplication on 16p11.2 which encompasses the BOLA2 gene. This locus is the breakpoint of the 16p11.2 micro-deletion, which results in developmental delay, intellectual disability, and autism5,6. We genotyped the BOLA2 gene in 675 diverse human individuals sequenced to low coverage as part of the 1000 Genome Project Phase I7 to assess the population distribution of copy numbers in homo-sapiens (Figure S8.3). While both the Altai Neandertal and Denisova individual exhibit the ancestral diploid copy number as seen in all the non-human great apes, only a single human individual exhibits this diploid copy number state.

My recollection from the earlier (less precise) Neanderthal sequences is that the number of bp differences between them and us is few per thousand. Whereas, for modern humans it's 1 per thousand with an additional +/-15% variation due to ethnicity. So, I think it's fair to say that they are qualitatively much more different from us than we (moderns) are from each other. See also The genetics of humanness.

My colleague James Lee (I note he is too modest to list his Harvard Law degree on his faculty page!) describes the current era in genomics as an "age of wonder" :-)  We can anticipate tremendous discoveries in the next decade.

Wednesday, October 30, 2013

Nabokov on teaching


Nabokov was professor of literature at Cornell from 1948-1959. The excerpt below is from a 1964 Playboy interview, reproduced at longform.org (a site I highly recommend).
Nabokov: I gave up teaching—that’s about all in the way of change. Mind you, I loved teaching, I loved Cornell, I loved composing and delivering my lectures on Russian writers and European great books. But around 60, and especially in winter, one begins to find hard the physical process of teaching, the getting up at a fixed hour every other morning, the struggle with the snow in the driveway, the march through long corridors to the classroom, the effort of drawing on the blackboard a map of James Joyce’s Dublin or the arrangement of the semi-sleeping car of the St. Petersburg-Moscow express in the early 1870s—without an understanding of which neither Ulysses nor Anna Karenina, respectively, makes sense. For some reason my most vivid memories concern examinations. Big amphitheater in Goldwin Smith. Exam from 8 a.m. to 10:30. About 150 students—unwashed, unshaven young males and reasonably well-groomed young females. A general sense of tedium and disaster. Half-past eight. Little coughs, the clearing of nervous throats, coming in clusters of sound, rustling of pages. Some of the martyrs plunged in meditation, their arms locked behind their heads. I meet a dull gaze directed at me, seeing in me with hope and hate the source of forbidden knowledge. Girl in glasses comes up to my desk to ask: “Professor Kafka, do you want us to say that…? Or do you want us to answer only the first part of the question?” The great fraternity of C-minus, backbone of the nation, steadily scribbling on. A rustic arising simultaneously, the majority turning a page in their bluebooks, good teamwork. The shaking of a cramped wrist, the failing ink, the deodorant that breaks down. When I catch eyes directed at me, they are forthwith raised to the ceiling in pious meditation. Windowpanes getting misty. Boys peeling off sweaters. Girls chewing gum in rapid cadence. Ten minutes, five, three, time’s up.
The first paragraph of Lolita, one of my favorites in all of literature:
Lolita, light of my life, fire of my loins. My sin, my soul. Lo-lee-ta: the tip of the tongue taking a trip of three steps down the palate to tap, at three, on the teeth. Lo. Lee. Ta.

Wednesday, October 09, 2013

The human genome as a compressed sensor



Compressed sensing (see also here) is a method for efficient solution of underdetermined linear systems: y = Ax + noise , using a form of penalized regression (L1 penalization, or LASSO). In the context of genomics, y is the phenotype, A is a matrix of genotypes, x a vector of effect sizes, and the noise is due to nonlinear gene-gene interactions and the effect of the environment. (Note the figure above, which I found on the web, uses different notation than the discussion here and the paper below.)

Let p be the number of variables (i.e., genetic loci = dimensionality of x), s the sparsity (number of variables or loci with nonzero effect on the phenotype = nonzero entries in x) and n the number of measurements of the phenotype (i.e., the number of individuals in the sample = dimensionality of y). Then  A  is an  n x p  dimensional matrix. Traditional statistical thinking suggests that  n > p  is required to fully reconstruct the solution  x  (i.e., reconstruct the effect sizes of each of the loci). But recent theorems in compressed sensing show that  n > C s log p  is sufficient if the matrix A has the right properties (is a good compressed sensor). These theorems guarantee that the performance of a compressed sensor is nearly optimal -- within an overall constant of what is possible if an oracle were to reveal in advance which  s  loci out of  p  have nonzero effect. In fact, one expects a phase transition in the behavior of the method as  n  crosses a critical threshold given by the inequality. In the good phase, full recovery of  x  is possible.

In the paper below, available on arxiv, we show that

1. Matrices of human SNP genotypes are good compressed sensors and are in the universality class of random matrices. The phase behavior is controlled by scaling variables such as  rho = s/n  and our simulation results predict the sample size threshold for future genomic analyses.

2. In applications with real data the phase transition can be detected from the behavior of the algorithm as the amount of data  n  is varied. A priori knowledge of  s  is not required; in fact one deduces the value of  s  this way.

3.  For heritability h2 = 0.5 and p ~ 1E06 SNPs, the value of  C log p  is ~ 30. For example, a trait which is controlled by s = 10k loci would require a sample size of n ~ 300k individuals to determine the (linear) genetic architecture.
Application of compressed sensing to genome wide association studies and genomic selection          
http://arxiv.org/abs/1310.2264
Authors: Shashaank Vattikuti, James J. Lee, Stephen D. H. Hsu, Carson C. Chow
Categories: q-bio.GN
Comments: 27 pages, 4 figures; Supplementary Information 5 figures

We show that the signal-processing paradigm known as compressed sensing (CS)
is applicable to genome-wide association studies (GWAS) and genomic selection
(GS). The aim of GWAS is to isolate trait-associated loci, whereas GS attempts
to predict the phenotypic values of new individuals on the basis of training
data. CS addresses a problem common to both endeavors, namely that the number
of genotyped markers often greatly exceeds the sample size. We show using CS
methods and theory that all loci of nonzero effect can be identified (selected)
using an efficient algorithm, provided that they are sufficiently few in number
(sparse) relative to sample size. For heritability h2 = 1, there is a sharp
phase transition to complete selection as the sample size is increased. For
heritability values less than one, complete selection can still occur although
the transition is smoothed. The transition boundary is only weakly dependent on
the total number of genotyped markers. The crossing of a transition boundary
provides an objective means to determine when true effects are being recovered.
For h2 = 0.5, we find that a sample size that is thirty times the number
of nonzero loci is sufficient for good recovery.

Blog Archive

Labels