Showing posts sorted by relevance for query replication. Sort by date Show all posts
Showing posts sorted by relevance for query replication. Sort by date Show all posts

Sunday, February 26, 2017

Perverse Incentives and Replication in Science

Here's a depressing but all too common pattern in scientific research:
1. Study reports results which reinforce the dominant, politically correct, narrative.

2. Study is widely cited in other academic work, lionized in the popular press, and used to advance real world agendas.

3. Study fails to replicate, but no one (except a few careful and independent thinkers) notices.
For numerous examples, see, e.g., any of Malcolm Gladwell's books :-(

A recent example: the idea that collective intelligence of groups (i.e., ability to solve problems and accomplish assigned tasks) is not primarily dependent on the cognitive ability of individuals in the group.

It seems plausible to me that by adopting certain best practices for collaboration one can improve group performance, and that diversity of knowledge base and personal experience could also enhance performance on certain tasks. But recent results in this direction were probably oversold, and seem to have failed to replicate.

James Thompson has given a good summary of the situation.

Parts 1 and 2 of our story:
MIT Center for Collective Intelligence: ... group-IQ, or “collective intelligence” is not strongly correlated with the average or maximum individual intelligence of group members but is correlated with the average social sensitivity of group members, the equality in distribution of conversational turn-taking, and the proportion of females in the group.
Is it true? The original paper on this topic, from 2010, has been cited 700+ times. See here for some coverage on this blog when it originally appeared.

Below is the (only independent?) attempt at replication, with strongly negative results. The first author is a regular (and very insightful) commenter here -- I hope he'll add his perspective to the discussion. Have we reached part 3 of the story?
Smart groups of smart people: Evidence for IQ as the origin of collective intelligence in the performance of human groups

Timothy C. Bates a,b,⁎, Shivani Gupta a
a Department of Psychology, University of Edinburgh
b Centre for Cognitive Ageing and Cognitive Epidemiology, University of Edinburgh

What allows groups to behave intelligently? One suggestion is that groups exhibit a collective intelligence accounted for by number of women in the group, turn-taking and emotional empathizing, with group-IQ being only weakly-linked to individual IQ (Woolley, Chabris, Pentland, Hashmi, & Malone, 2010). Here we report tests of this model across three studies with 312 people. Contrary to prediction, individual IQ accounted for around 80% of group-IQ differences. Hypotheses that group-IQ increases with number of women in the group and with turn-taking were not supported. Reading the mind in the eyes (RME) performance was associated with individual IQ, and, in one study, with group-IQ factor scores. However, a well-fitting structural model combining data from studies 2 and 3 indicated that RME exerted no influence on the group-IQ latent factor (instead having a modest impact on a single group test). The experiments instead showed that higher individual IQ enhances group performance such that individual IQ determined 100% of latent group-IQ. Implications for future work on group-based achievement are examined.


From the paper:
Given the ubiquitous importance of group activities (Simon, 1997) these results have wide implications. Rather than hiring individuals with high cognitive skill who command higher salaries (Ritchie & Bates, 2013), organizations might select-for or teach social sensitivity thus raising collective intelligence, or even operate a female gender bias with the expectation of substantial performance gains. While the study has over 700 citations and was widely reported to the public (Woolley, Malone, & Chabris, 2015), to our knowledge only one replication has been reported (Engel, Woolley, Jing, Chabris, & Malone, 2014). This study used online (rather than in-person) tasks and did not include individual IQ. We therefore conducted three replication studies, reported below.

... Rather than a small link of individual IQ to group-IQ, we found that the overlap of these two traits was indistinguishable from 100%. Smart groups are (simply) groups of smart people. ... Across the three studies we saw no significant support for the hypothesized effects of women raising (or men lowering) group-IQ: All male, all female and mixed-sex groups performed equally well. Nor did we see any relationship of some members speaking more than others on either higher or lower group-IQ. These findings were weak in the initial reports, failing to survive incorporation of covariates. We attribute these to false positives. ... The present findings cast important doubt on any policy-style conclusions regarding gender composition changes cast as raising cognitive-efficiency. ...

In conclusion, across three studies groups exhibited a robust cognitive g-factor across diverse tasks. As in individuals, this g-factor accounted for approximately 50% of variance in cognition (Spearman, 1904). In structural tests, this group-IQ factor was indistinguishable from average individual IQ, and social sensitivity exerted no effects via latent group-IQ. Considering the present findings, work directed at developing group-IQ tests to predict team effectiveness would be redundant given the extremely high utility, reliability, validity for this task shown by individual IQ tests. Work seeking to raise group-IQ, like re- search to raise individual IQ might find this task achievable at a task- specific level (Ritchie et al., 2013; Ritchie, Bates, & Plomin, 2015), but less amenable to general change than some have anticipated. Our attempt to manipulate scores suggested that such interventions may even decrease group performance. Instead, work understanding the developmental conditions which maximize expression of individual IQ (Bates et al., 2013) as well as on personality and cultural traits supporting cooperation and cumulation in groups should remain a priority if we are to understand and develop cognitive ability. The present experiments thus provide new evidence for a central, positive role of individual IQ in enhanced group-IQ.
Meta-Observation: Given the 1-2-3 pattern described above, one should be highly skeptical of results in many areas of social science and even biomedical science (see link below). Serious researchers (i.e., those who actually aspire to participate in Science) in fields with low replication rates should (as a demonstration of collective intelligence!) do everything possible to improve the situation. Replication should be considered an important research activity, and should be taken seriously.

Most researchers I know in the relevant areas have not yet grasped that there is a serious problem. They might admit that "some studies fail to replicate" but don't realize the fraction might be in the 50 percent range!

More on the replication crisis in certain fields of science.

Tuesday, March 08, 2016

One funeral at a time?


A new scientific truth does not triumph by convincing its opponents and making them see the light, but rather because its opponents eventually die, and a new generation grows up that is familiar with it. -- Max Planck
I'm at the annual AAU meeting of research Vice-Presidents, and the "reproducibility crisis" (in some fields) is one of the topics on the agenda. Today I heard a nice talk by Brian Nosek on the Reproducibility Project and Open Science Framework. Afterwards I asked him whether other social psychologists were really absorbing the implications of the low reproducibility rate in their field. Nosek said that attitudes were changing rapidly and that in 10 years the field would be very different. If so, progress can happen much faster than in Max Planck's pessimistic quote. Let's hope Nosek is right!

As an example of soul searching (learning can hurt!) Nosek pointed me to this blog post, by Professor Michael Inzlicht of Toronto. To his credit, Inzlicht is taking seriously the replication difficulties of two effects he has worked on in the past: ego depletion (i.e., willpower fatigue) and stereotype threat. Anyone who has looked carefully at the literature (positive effects from small sample sizes, but failed replication or very small effect size results from much larger samples) has to question whether these much-hyped phenomena are real. There is also lots of evidence for publication bias.
Inzlicht: ... I have spent nearly a decade working on the concept of ego depletion, including work that is critical of the model used to explain the phenomenon. I have been rewarded for this work, and I am convinced that the main reason I get any invitations to speak at colloquia and brown-bags these days is because of this work. The problem is that ego depletion might not even be a thing. By now, many people are aware that a massive replication attempt of the basic ego depletion effect involving over 2,000 participants found nothing, nada, zip. Only three of the 24 participating labs found a significant effect, but even then, one of these found a significant result in the wrong direction!

There is a lot more to this registered replication than the main headline, and deep in my heart, it is hard to believe that fatigue is not a real phenomenon. I promise to get to it in a later post. But for now, we are left with a sobering question: If a large sample pre-registered study found absolutely nothing, how has the ego depletion effect been replicated and extended hundreds and hundreds of times? More sobering still: What other phenomena, which we now consider obviously real and true, will be revealed to be just as fragile?

As I said, I’m in a dark place. I feel like the ground is moving from underneath me and I no longer know what is real and what is not.

I edited an entire book on stereotype threat, I have signed my name to an amicus brief to the Supreme Court of the United States citing stereotype threat, yet now I am not as certain as I once was about the robustness of the effect. I feel like a traitor for having just written that; like, I’ve disrespected my parents, a no no according to Commandment number 5. But, a meta-analysis published just last year suggests that stereotype threat, at least for some populations and under some conditions, might not be so robust after all. P-curving some of the original papers is also not comforting. Now, stereotype threat is a politically charged topic and I really really want it to be real. I think a lot more pain-staking work needs to be done before I stop believing (and rumor has it that another RRR of stereotype threat is in the works), but I would be lying if I said that doubts have not crept in. ...
See also Is science self-correcting? Figure at top is from this meta-analysis of stereotype threat.

Tuesday, December 10, 2013

Is science self-correcting?


More fun from our man Ioannidis. See earlier posts Medical science? , NIH discovers reproducibility and Bounded cognition.

A toy model of the dynamics of scientific research, with probability distributions for accuracy of experimental results, mechanisms for updating of beliefs by individual scientists, crowd behavior, bounded cognition, etc. can easily exhibit parameter regions where progress is limited (one could even find equilibria in which most beliefs held by individual scientists are false!). Obviously the complexity of the systems under study and the quality of human capital in a particular field are important determinants of the rate of progress and its character.

In physics it is said that successful new theories swallow their predecessors whole. That is, even revolutionary new theories (e.g., special relativity or quantum mechanics) reduce to their predecessors in the previously studied circumstances (e.g., low velocity, macroscopic objects). Swallowing whole is a sign of proper function -- it means the previous generation of scientists was competent: what they believed to be true was (at least approximately) true. Their models were accurate in some limit and could continue to be used when appropriate (e.g., Newtonian mechanics).

In some fields (not to name names!) we don't see this phenomenon. Rather, we see new paradigms which wholly contradict earlier strongly held beliefs that were predominant in the field* -- there was no range of circumstances in which the earlier beliefs were correct. We might even see oscillations of mutually contradictory, widely accepted paradigms over decades.

It takes a serious interest in the history of science (and some brainpower) to determine which of the two regimes above describes a particular area of research. I believe we have good examples of both types in the academy.

* This means the earlier (or later!) generation of scientists in that field was incompetent. One or more of the following must have been true: their experimental observations were shoddy, they derived overly strong beliefs from weak data, they allowed overly strong priors to determine their beliefs.

Why Science Is Not Necessarily Self-Correcting 
(DOI: 10.1177/1745691612464056)

John P. A. Ioannidis
Stanford Prevention Research Center, Department of Medicine and Department of Health Research and Policy, Stanford University School of Medicine, and Department of Statistics, Stanford University School of Humanities and Sciences

The ability to self-correct is considered a hallmark of science. However, self-correction does not always happen to scientific evidence by default. The trajectory of scientific credibility can fluctuate over time, both for defined scientific fields and for science at-large. History suggests that major catastrophes in scientific credibility are unfortunately possible and the argument that “it is obvious that progress is made” is weak. Careful evaluation of the current status of credibility of various scientific fields is important in order to understand any credibility deficits and how one could obtain and establish more trustworthy results. Efficient and unbiased replication mechanisms are essential for maintaining high levels of scientific credibility. Depending on the types of results obtained in the discovery and replication phases, there are different paradigms of research: optimal, self-correcting, false nonreplication, and perpetuated fallacy. In the absence of replication efforts, one is left with unconfirmed (genuine) discoveries and unchallenged fallacies. In several fields of investigation, including many areas of psychological science, perpetuated and unchallenged fallacies may comprise the majority of the circulating evidence. I catalogue a number of impediments to self-correction that have been empirically studied in psychological science. Finally, I discuss some proposed solutions to promote sound replication practices enhancing the credibility of scientific results as well as some potential disadvantages of each of them. Any deviation from the principle that seeking the truth has priority over any other goals may be seriously damaging to the self-correcting functions of science

Sunday, May 03, 2015

Replication is hard; understanding what that means is even harder



Bad news for psychology -- only 39 of 100 published findings were replicated in a recent coordinated effort.
Nature | News: An ambitious effort to replicate 100 research findings in psychology ended last week — and the data look worrying. Results posted online on 24 April, which have not yet been peer-reviewed, suggest that key findings from only 39 of the published studies could be reproduced. ...
The article goes on:
But the situation is more nuanced than the top-line numbers suggest (See graphic, 'Reliability test'). Of the 61 non-replicated studies, scientists classed 24 as producing findings at least “moderately similar” to those of the original experiments, even though they did not meet pre-established criteria, such as statistical significance, that would count as a successful replication.  [ Yeah, right. ]
This makes me suspect bounded cognition -- humans trusting their post hoc stories and intuition instead of statistical criteria chosen before planned replication attempts.

The most tragic thing about Ioannidis's work on low replication rates and wasted research funding is that while medical researchers might pay lip service to his results (which are highly cited), they typically have not actually grasped the implications for their own work. In particular, they typically have not updated their posteriors to reflect the low reliability of research results, even in the top journals.


Wednesday, March 08, 2017

Examples of Perverse Incentives and Replication in Science

In an earlier post (Perverse Incentives and Replication in Science), I wrote:
Here's a depressing but all too common pattern in scientific research:

1. Study reports results which reinforce the dominant, politically correct, narrative.
2. Study is widely cited in other academic work, lionized in the popular press, and used to advance real world agendas.
3. Study fails to replicate, but no one (except a few careful and independent thinkers) notices.
This seems to have hit a nerve, as many people have come forward with their own examples of this pattern.

From a colleague at MSU:

Parts 1 and 2: Green revolution in Malawi from Farm Input Subsidy Program? Hurrah! Gushing coverage in the NYTimes, Jeffrey Sachs claiming credit, etc.

http://www.nytimes.com/2012/04/20/opinion/how-malawi-fed-its-own-people.html
http://www.nytimes.com/2007/12/02/world/africa/02malawi.html
http://www.nytimes.com/slideshow/2007/12/01/world/20071202MALAWI_index.html


Part 3: Failed replication? No actual green revolution? Will anyone notice?
Re-evaluating the Malawian Farm Input Subsidy Programme (Nature Plants)

Joseph P. Messina1*†, Brad G. Peter1† and Sieglinde S. Snapp2†

Abstract: The Malawian Farm Input Subsidy Programme (FISP) has received praise as a proactive policy that has transformed the nation’s food security, yet irreconcilable differences exist between maize production estimates distributed by the Food and Agriculture Organization of the United Nations (FAO), the Malawi Ministry of Agriculture and Food Security (MoAFS) and the National Statistical Office (NSO) of Malawi. These differences illuminate yield-reporting deficiencies and the value that alternative, politically unbiased yield estimates could play in understanding policy impacts. We use net photosynthesis (PsnNet) as an objective source of evidence to evaluate production history and production potential under a fertilizer input scenario. Even with the most generous harvest index (HI) and area manipulation to match a reported error, we are unable to replicate post-FISP production gains. In addition, we show that the spatial delivery of FISP may have contributed to popular perception of widespread maize improvement. These triangulated lines of evidence suggest that FISP may not have been the success it was thought to be. Lastly, we assert that fertilizer subsidies may not be sufficient or sustainable strategies for production gains in Malawi.

Introduction: Input subsidies targeting agricultural production are frequent and contentious development strategies. The national scale FISP implemented in Malawi has been heralded as an ‘African green revolution’ success story1. The programme was developed by the Malawian government in response to long-term recurring food shortages, following the notably poor maize harvest of 2005; the history of FISP is well described by Chirwa and Dorward2. Scholars and press sources alike commonly refer to government statistics regarding production and yields as having improved significantly. Reaching widespread audiences, Sachs broadcasted that “production doubled within one harvest season” following its deployment3. The influential policy paper by Denning et al. opened with the statement that the “Government of Malawi implemented one of the most ambitious and successful assaults on hunger in the history of the African continent”4. The Malawi success narrative has certainly influenced global development agencies, resulting in increased support for agricultural input subsidies; Tanzania, Zambia, Kenya and Rwanda have all followed suit and implemented some form of input subsidy programme. There has been mild economic criticism of the subsidy implementation process, including disruption of private fertilizer distribution networks within the policy’s first year5. Moreover, the sustainability of subsidies in Malawi has been debated6,7, yet crop productivity gains from subsidies have gone largely unquestioned. As Sanchez commented, “in spite of criticisms by donor agencies and academics, the seed and fertilizer subsidies provided food security to millions of Malawians”1. This optimistic assessment of potential for an “African green revolution” must be tempered by the fact that the Malawian production miracle appears, in part, to be a myth. ...
For more on the 1-2-3 pattern and replication, see this blog post by economist Douglas Campbell, and discussion here.

For more depressing narrative concerning the reliability of published results (this time in cancer research), see this front page NYTimes story from today.

Wednesday, May 03, 2017

Replication and the "Crises of Confidence" in Science

One of the authors pointed me to the interesting paper below, which contains a proposal meant to improve the reliability of scientific research (specifically in some areas such as social science or biomedicine). Could this proposal work if, say, it were strongly supported by scholarly associations and funding agencies?

Note the authors' use of "Crises of Confidence" in their title. I believe that certain areas of science really should be experiencing a crisis of confidence, if only their practitioners were smarter and more honest. But in fact I suspect it's mostly business as usual, even in the research areas with the lowest (and best documented as lowest) replication rates.
An Economic Approach to Alleviate the Crises of Confidence in Science: With An Application to the Public Goods Game

Luigi Butera, John A. List  (University of Chicago)

April 5, 2017

... This paper proposes and puts into practice a novel and simple mechanism that allows mutually beneficial gains from trade between original investigators and other researchers. In our mechanism, the original investigators, upon completing their initial study, write a working paper version of their study. While they do share their working paper online, they do however commit not to submit it to any journal for publication, ever. The original investigators instead offer co-authorship of a second paper to other researchers who are willing to independently replicate the experimental protocol in their own research facilities.2 Once the team is established, but before beginning replications, the replication protocol is pre-registered at the AEA experimental registry, and referenced in the first working paper. This is to guarantee that all replications, both successful and failed, are properly accounted for, eliminating any concerns about publication biases. The team of researchers composed by the original investigators and the other scholars will then write and coauthor a second paper, which will reference the original unpublished working paper, and submit it to an academic journal. Under such an approach, the original investigators accept to publish their work with several other coauthors, a feature that is typically unattractive to economists, but in turn gain a dramatic increase in the credibility and robustness of their results, should they replicate. Further, the referenced working paper would provide a credible signal about the ownership of the initial research design and idea, a feature that is particularly desirable for junior scholars. On the other hand, other researchers would face the monetary cost of replicating the original study, but would in turn benefit from coauthoring a novel study, and share the related payoffs. Overall, our mechanism could critically strengthen the reliability of novel experimental results and facilitate the advancement of scientific knowledge.

Thursday, September 22, 2016

Annals of Reproducibility in Science: Social Psychology and Candidate Gene Studies

Andrew Gelman offers a historical timeline for the reproducibility crisis in Social Psychology, along with some juicy insight into the one funeral at a time manner in which academic science often advances.
OK, that was a pretty detailed timeline. But here’s the point. Almost nothing was happening for a long time, and even after the first revelations and theoretical articles you could still ignore the crisis if you were focused on your research and other responsibilities. ...

Then, all of a sudden, the world turned upside down.

If you’d been deeply invested in the old system, it must be pretty upsetting to think about change. Fiske is in the position of someone who owns stock in a failing enterprise, so no wonder she wants to talk it up. The analogy’s not perfect, though, because there’s no one for her to sell her shares to. What Fiske should really do is cut her losses, admit that she and her colleagues were making a lot of mistakes, and move on. She’s got tenure and she’s got the keys to PPNAS, so she could do it. Short term, though, I guess it’s a lot more comfortable for her to rant about replication terrorists and all that.

... Why do I go into all this detail? Is it simply mudslinging? Fiske attacks science reformers, so science reformers slam Fiske? No, that’s not the point. The issue is not Fiske’s data processing errors or her poor judgment as journal editor; rather, what’s relevant here is that she’s working within a dead paradigm. A paradigm that should’ve been dead back in the 1960s when Meehl was writing on all this, but which in the wake of Simonsohn, Button et al., Nosek et al., is certainly dead today. It’s the paradigm of the open-ended theory, of publication in top journals and promotion in the popular and business press, based on “p less than .05” results obtained using abundant researcher degrees of freedom. It’s the paradigm of the theory that in the words of sociologist Jeremy Freese, is “more vampirical than empirical—unable to be killed by mere data.”

... In her article that was my excuse to write this long post, Fiske expresses concerns for the careers of her friends, careers that may have been damaged by public airing of their research mistakes. Just remember that, for each of these people, there may well be three other young researchers who were doing careful, serious work but then didn’t get picked for a plum job or promotion because it was too hard to compete with other candidates who did sloppy but flashy work that got published in Psych Science or PPNAS. It goes both ways. ...
An old timer who has seen it all before comments.
ex-social psychologist says:
September 21, 2016 at 5:36 pm

Former professor of social psychology here, now happily retired after an early buyout offer. If not so painful, it would almost be funny at how history repeats itself: This is not the first time there has been a “crisis” in social psychology. In the late 1960s and early 1970s there was much hand-wringing over failures of replication and the “fun and games” mentality among researchers; see, for example, Gergen’s 1973 article “Social psychology as history” in JPSP, 26, 309-320, and Ring’s (1967) JESP article, “Experimental social psychology: Some sober questions about some frivolous values.” It doesn’t appear that the field ever truly resolved those issues back when they were first raised–instead, we basically shrugged, said “oh well,” and went about with publishing by any means necessary.

I’m glad to see the renewed scrutiny facing the field. And I agree with those who note that social psychology is not the only field confronting issues of replicability, p-hacking, and outright fraud. These problems don’t have easy solutions, but it seems blindingly obvious that transparency and open communication about the weaknesses in the field–and individual studies–is a necessary first step. Fiske’s strategy of circling the wagons and adhering to a business-as-usual model is both sad and alarming.

I took early retirement for a number of reasons, but my growing disillusionment with my chosen field was certainly a primary one.
Geoffrey Miller also contributes
Geoffrey Miller says:
September 21, 2016 at 8:43 pm

There’s also a political/ideological dimension to social psychology’s methodological problems.

For decades, social psych advocated a particular kind of progressive, liberal, blank-slate ideology. Any new results that seemed to support this ideology were published eagerly and celebrated publicly, regardless of their empirical merit. Any results that challenged it (e.g. by showing the stability or heritability of individual differences in intelligence or personality) were rejected as ‘genetic determinism’, ‘biological reductionism’, or ‘reactionary sociobiology’.

For decades, social psychologists were trained, hired, promoted, and tenured based on two main criteria: (1) flashy, counter-intuitive results published in certain key journals whose editors and reviewers had a poor understanding of statistical pitfalls, (2) adherence to the politically correct ideology that favored certain kinds of results consistent with a blank-slate, situationist theory of human nature, and derogation of any alternative models of human nature (see Steven Pinker’s book ‘The blank slate’).

Meanwhile, less glamorous areas of psychology such as personality, evolutionary, and developmental psychology, intelligence research, and behavior genetics were trundling along making solid cumulative progress, often with hugely greater statistical power and replicability (e.g. many current behavior genetics studies involve tens of thousands of twin pairs across several countries). But do a search for academic positions in the APS job ads for these areas, and you’ll see that they’re not a viable career path, because most psych departments still favor the kind of vivid but unreplicable results found in social psych and cognitive neuroscience.

So, we’re in a situation where the ideologically-driven, methodologically irresponsible field of social psychology has collapsed like a house of cards … but nobody’s changed their hiring, promotion, or tenure priorities in response. It’s still fairly easy to make a good living doing bad social psychology. It’s still very hard to make a living doing good personality, intelligence, behavior genetic, or evolutionary psychology research.

In the title of this post I mention Candidate Gene Studies. Forget, for the moment, about goofy Social Psychology experiments conducted on undergraduates. Much more money was wasted in the early 21st century on under-powered genomics studies that looked for gene-trait associations using small samples. Researchers, overconfident in their vaunted biological or biochemical intuition, performed studies using p < 0.05 thresholds that produced (ultimately false) associations between candidate genes and a variety of traits. According to Ioannidis, almost none of these results replicate (more). When I first became aware of GWAS almost a decade ago, the field was in disarray, with some journals still publishing results at the p < 0.05 threshold, whereas others having adopted the corrected p < 5E-08 = 0.05 x 1E-06 "genome wide significance" threshold (based on multiple testing correction for 1E06 SNPs)! The latter results routinely replicate, as expected.

Clearly, many researchers fundamentally misunderstood basic statistics, or at least were grossly overconfident in their priors for no good reason. But as of today, genomics has corrected its practices and although no one wants to dwell on the 5+ years worth of non-replicable published results, science is at least moving forward. I hope Social Psychology and other problematic areas (such as in biomedical research) can self-correct their practices as genomics has.

See also One funeral at a time?


Bonus Feature!
Empirical assessment of published effect sizes and power in the recent cognitive neuroscience and psychology literature

Denes Szucs, John PA Ioannidis
doi: http://dx.doi.org/10.1101/071530

We have empirically assessed the distribution of published effect sizes and estimated power by extracting more than 100,000 statistical records from about 10,000 cognitive neuroscience and psychology papers published during the past 5 years. The reported median effect size was d=0.93 (inter-quartile range: 0.64-1.46) for nominally statistically significant results and d=0.24 (0.11-0.42) for non-significant results. Median power to detect small, medium and large effects was 0.12, 0.44 and 0.73, reflecting no improvement through the past half-century. Power was lowest for cognitive neuroscience journals. 14% of papers reported some statistically significant results, although the respective F statistic and degrees of freedom proved that these were non-significant; p value errors positively correlated with journal impact factors. False report probability is likely to exceed 50% for the whole literature. In light of our findings the recently reported low replication success in psychology is realistic and worse performance may be expected for cognitive neuroscience.
From the paper. FRP = False Report Probability = the probability that the null hypothesis is true when we get a statistically significant finding.
... In all, the combination of low power, selective reporting and other biases and errors that we have documented in this large sample of papers in cognitive neuroscience and psychology suggest that high FRP are to be expected in these fields. The low reproducibility rate seen for psychology experimental studies in the recent Open Science Collaboration (Nosek et al. 2015a) is congruent with the picture that emerges from our data. Our data also suggest that cognitive neuroscience may have even higher FRP rates, and this hypothesis is worth evaluating with focused reproducibility checks of published studies. Regardless, efforts to increase sample size, and reduce publication and other biases and errors are likely to be beneficial for the credibility of this important literature.

Monday, February 11, 2013

On the verge

This paper is based on a combined sample of 18k individuals, taken from several European longitudinal studies of child development. Note early childhood intelligence is not as heritable as adult intelligence, according to classical (twins, adoption) methods.
Nature Molecular Psychiatry (29 January 2013) | doi:10.1038/mp.2012.184

Childhood intelligence is heritable, highly polygenic and associated with FNBP1L

Intelligence in childhood, as measured by psychometric cognitive tests, is a strong predictor of many important life outcomes, including educational attainment, income, health and lifespan. Results from twin, family and adoption studies are consistent with general intelligence being highly heritable and genetically stable throughout the life course. No robustly associated genetic loci or variants for childhood intelligence have been reported. Here, we report the first genome-wide association study (GWAS) on childhood intelligence (age range 6–18 years) from 17 989 individuals in six discovery and three replication samples. Although no individual single-nucleotide polymorphisms (SNPs) were detected with genome-wide significance, we show that the aggregate effects of common SNPs explain 22–46% of phenotypic variation in childhood intelligence in the three largest cohorts (P=3.9 × 10−15, 0.014 and 0.028). FNBP1L, previously reported to be the most significantly associated gene for adult intelligence, was also significantly associated with childhood intelligence (P=0.003). Polygenic prediction analyses resulted in a significant correlation between predictor and outcome in all replication cohorts. The proportion of childhood intelligence explained by the predictor reached 1.2% (P=6 × 10−5), 3.5% (P=10−3) and 0.5% (P=6 × 10−5) in three independent validation cohorts. Given the sample sizes, these genetic prediction results are consistent with expectations if the genetic architecture of childhood intelligence is like that of body mass index or height. Our study provides molecular support for the heritability and polygenic nature of childhood intelligence. Larger sample sizes will be required to detect individual variants with genome-wide significance.

Compare to this figure for height and BMI from an earlier post: Five years of GWAS discovery. It appears we may be close to the threshold required to find the first genome-wide significant hits.



Wednesday, June 17, 2015

Hopfield on physics and biology

Theoretical physicist John Hopfield, inventor of the Hopfield neural network, on the differences between physics and biology. Hopfield migrated into biology after making important contributions in condensed matter theory. At Caltech, Hopfield co-taught a famous course with Carver Mead and Richard Feynman on the physics of computation.
Two cultures? Experiences at the physics-biology interface

(Phys. Biol. 11 053002 doi:10.1088/1478-3975/11/5/053002)

Abstract: 'I didn't really think of this as moving into biology, but rather as exploring another venue in which to do physics.' John Hopfield provides a personal perspective on working on the border between physical and biological sciences.

... With two parents who were physicists, I grew up with the view that science was about understanding quantitatively how things worked, not about collecting details and categorizing observations. Their view, though not so explicitly stated, was certainly that of Rutherford: 'all science is either physics or stamp collecting.' So, when selecting science as a career, I never considered working in biology and ultimately chose solid state physics research.

... I attended my first biology conference in the summer of 1970 at a small meeting with the world's experts on the hemoglobin molecule. It was held at the Villa Serbelloni in Bellagio, in sumptuous surroundings verging on decadence as I had never seen for physics meetings. One of the senior biochemists took me aside to explain to me why I had no place in biology. As he said, gentlemen did not interpret other gentlemen's data, and preferably worked on different organisms. If you wish to interpret data, you must get your own. Only the experimentalist himself knows which of the data points are reliable, and so only he should interpret them. Moreover, if you insist on interpreting other people's data, they will not publish their best data. Biology is very complicated, and any theory with mathematics is such an oversimplification that it is essentially wrong and thus useless. And so on... On closer examination, this diatribe chiefly describes differences between the physics and biology paradigms (at the time at least) for engaging in science. Physics papers use data points with error bars; biology papers lacked them. Physics was based on the quantitative replication of experiments in different laboratories; biology broadened its fact collecting by devaluing replication. Physics education emphasized being able to look at a physical system and express it in mathematical terms. Mathematical theory had great predictive power in physics, but very little in biology. As a result, mathematics is considered the language of the physics paradigm, a language in which most biologists could remain illiterate. Time has passed, but there is still an enormous difference in the biology and physics paradigms for working in science. Advice? Stick to the physics paradigm, for it brings refreshing attitudes and a different choice of problems to the interface. And have a thick skin. ...
Also by Hopfield: Physics, Computation, and Why Biology Looks so Different and Whatever happened to solid state physics?

See also In search of principles: when biology met physics (Bill Bialek), For the historians and the ladiesAs flies to wanton boys are we to the gods and Prometheus in the basement.

Tuesday, November 06, 2018

1 In 4 Biostatisticians Surveyed Say They Were Asked To Commit Scientific Fraud


In the survey reported below, about 1 in 4 biostatisticians were asked to commit scientific fraud. I don't know whether this bad behavior was more prevalent in industry as opposed to academia, but I am not surprised by the results.

I do not accept the claim that researchers in data-driven areas can be ignorant of statistics. It is common practice to outsource statistical analysis to people like the "consulting biostatisticians" surveyed below. But scientists who do not understand statistics will not be effective in planning future research, nor in understanding the implications of results in their own field. See the candidate gene and missing heritability nonsense the field of genetics has been subject to for the last decade.

I cannot count the number of times, in talking to a scientist with limited quantitative background, that I have performed -- to their amazement -- a quick back of the envelope analysis of a statistical design or new results. This kind of quick estimate is essential to understand whether the results in question should be trusted, or whether a prospective experiment is worth doing. The fact that they cannot understand my simple calculation means that they literally do not understand how inference in their own field should be performed.
Researcher Requests for Inappropriate Analysis and Reporting: A U.S. Survey of Consulting Biostatisticians

(Annals of Internal Medicine 554-558. Published: 16-Oct-2018. DOI: 10.7326/M18-1230)

Results:
Of 522 consulting biostatisticians contacted, 390 provided sufficient responses: a completion rate of 74.7%. The 4 most frequently reported inappropriate requests rated as “most severe” by at least 20% of the respondents were, in order of frequency, removing or altering some data records to better support the research hypothesis; interpreting the statistical findings on the basis of expectation, not actual results; not reporting the presence of key missing data that might bias the results; and ignoring violations of assumptions that would change results from positive to negative. These requests were reported most often by younger biostatisticians.
This kind of behavior is consistent with the generally low rate of replication for results in biomedical science, even those published in top journals:
What is medicine’s 5 sigma? (Editorial in the Lancet)... much of the [BIOMEDICAL] scientific literature, perhaps half, may simply be untrue. Afflicted by studies with small sample sizes, tiny effects, invalid exploratory analyses, and flagrant conflicts of interest, together with an obsession for pursuing fashionable trends of dubious importance, [BIOMEDICAL] science has taken a turn towards darkness. As one participant put it, “poor methods get results”. The Academy of Medical Sciences, Medical Research Council, and Biotechnology and Biological Sciences Research Council have now put their reputational weight behind an investigation into these questionable research practices. The apparent endemicity of bad research behaviour is alarming. In their quest for telling a compelling story, scientists too often sculpt data to fit their preferred theory of the world. ...
More background on the ongoing replication crisis in certain fields of science. See also Bounded Cognition.

Monday, July 20, 2015

What is medicine’s 5 sigma?

Editorial in the Lancet, reflecting on the Symposium on the Reproducibility and Reliability of Biomedical Research held April 2015 by the Wellcome Trust.
What is medicine’s 5 sigma?

... much of the [BIOMEDICAL] scientific literature, perhaps half, may simply be untrue. Afflicted by studies with small sample sizes, tiny effects, invalid exploratory analyses, and flagrant conflicts of interest, together with an obsession for pursuing fashionable trends of dubious importance, [BIOMEDICAL] science has taken a turn towards darkness. As one participant put it, “poor methods get results”. The Academy of Medical Sciences, Medical Research Council, and Biotechnology and Biological Sciences Research Council have now put their reputational weight behind an investigation into these questionable research practices. The apparent endemicity of bad research behaviour is alarming. In their quest for telling a compelling story, scientists too often sculpt data to fit their preferred theory of the world. ...

One of the most convincing proposals came from outside the biomedical community. Tony Weidberg is a Professor of Particle Physics at Oxford. ... the particle physics community ... invests great effort into intensive checking and rechecking of data prior to publication. By filtering results through independent working groups, physicists are encouraged to criticise. Good criticism is rewarded. The goal is a reliable result, and the incentives for scientists are aligned around this goal. Weidberg worried we set the bar for results in biomedicine far too low. In particle physics, significance is set at 5 sigma—a p value of 3 × 10–7 or 1 in 3·5 million (if the result is not true, this is the probability that the data would have been as extreme as they are). The conclusion of the symposium was that something must be done ...
I once invited a famous evolutionary theorist (MacArthur Fellow) at Oregon to give a talk in my institute, to an audience of physicists, theoretical chemists, mathematicians and computer scientists. The Q&A was, from my perspective, friendly and lively. A physicist of Hungarian extraction politely asked the visitor whether his models could ever be falsified, given the available field (ecological) data. I was shocked that he seemed shocked to be asked such a question. Later I sent an email thanking the speaker for his visit and suggesting he come again some day. He replied that he had never been subjected to such aggressive and painful attack and that he would never come back. Which community of scientists is more likely to produce replicable results?

See also Medical Science? and Is Science Self-Correcting?

To answer the question posed in the title of the post / editorial, an example of a statistical threshold which is sufficient for high confidence of replication is the p < 0.5 x 10^{-8} significance requirement in GWAS. This is basically the traditional p < 0.05 threshold corrected for multiple testing of 10^6 SNPs. Early "candidate gene" studies which did not impose this correction have very low replication rates. See comment below for what this implies about the validity of priors based on biological intuition.

I discuss this a bit with John Ioannidis in the video below.


Saturday, April 02, 2016

Jonathan Haidt and Tyler Cowen




Highly recommended: a great conversation (transcript) between Tyler Cowen and NYU psychology professor Johnathan Haidt. More Haidt.

The transformation of the Academy and the two universities:
COWEN: But is it at least possibly the case that we’re seeing the greatest threat to intellectual diversity in some of the areas which matter least, and when the stakes are high we overcome it. Physics looks pretty good, computer science looks pretty good.

HAIDT: No, it’s not — there are two universities now, but it’s not which ones matter more and which ones matter less. It’s what is the sacred value. The sacred value of universities from sometime in the 19th century through maybe the 1980s was truth. Now it was not perfect, but we all talked that way. Look at the mottos of Harvard and Yale — Veritas, Lux et Veritas, it’s right there on the motto, veritas, truth.

We made a big show — it was largely true — of saying this is what we’re here for, we’re here to find truth. But in the 1970s and ’80s as we had a big influx of baby boomers who were involved in social protest, who were fighting for very good causes, civil rights, women’s rights — they flood into the academy in ’70s and ’80s, they get tenure in the ’80s and ’90s, but also in the 1990s, the Greatest Generation begins to retire. There were a lot of Republicans who became professors after World War II.

But the ’90s is the decade where everything flips. At the start of the 1990s, the overall left‑right ratio of the academy, taking all departments, was two to one, just twice as many people on the left as right. That’s fine, that’s not a problem. But by 2005, it had gone to five to one, five people on the left for every one on the right. Those people on the right are mostly engineering, nursing, things like that. If you look at the core — the humanities and the social sciences, other than economics, it’s closer to 10 to 1 or 20 to 1.

In other words, right‑wing, or libertarian, or social conservative voices have basically vanished between 1995 and 2005. This has made us unfunctional, but it’s in the social sciences and humanities where the sacred value has become social justice and the protection of victims. That’s the division. One university of the sciences still pursues truth, the other university in the social sciences and humanities pursues social justice.
The Replication Crisis (see also One funeral at a time?):
HAIDT: ... I think Brian Nosek, who’s been leading the charge on the problems in psychology, is largely right. That our methods have been sloppy, which has allowed us to engage in practices where we’re just more likely than we should be to get a significant result. And then of course, that’s more likely to get published.

Given that we find the same problem in cancer research and biomedical research — in almost every field where it’s been looked at — I think that the replication crisis is very real. It should be a top priority for science.

A lot of my work is on how we are not fully rational creatures. We are deeply emotional and tribal creatures. If you have this idealized view of researchers and our null hypothesis significant testing is based on idealized view of researchers who are basically testing samples honestly.

“Well, this could only happen 1 in 20 times by chance,” but we’re not those creatures. We want certain outcomes to happen. We make certain choices unconsciously. We all have to up our game. I don’t think there’s anything special about social psychology. It’s no worse than other fields. But we have been the leaders at actually addressing it, and saying, “Why are we not able to replicate each other’s work so much?”

I actually am impressed that the young generation has really embraced this and simply committing to making your data available — if you know that other people are going to get access to your SPSS file, or whatever, your data file, and they’re going to be looking it over, boy, you’re going to be a lot more careful.

I think just by raising the crisis, raising the alarm last year, the quality of our work is going to go substantially up. I’m really excited by this.
Social Psychology IS worse than some other fields, when it comes to reproducibility. First, it is in the wrong (SJW) part of the two universities Haidt describes in the earlier excerpt. Secondly, along with biomedical research, it is in the part of the university where most researchers lack a deep understanding of statistics and quantitative inference. See What is medicine's 5 sigma?

Thursday, April 07, 2016

GWAS of cognitive function using UK Biobank data

This paper is based on analysis of UK Biobank data. The phenotypes (cognitive scores) were obtained via brief on-screen tests. Although there is significant noise in the scores obtained (see test-retest correlations in the table at bottom), there was enough signal to obtain a number of genome-wide significant SNP hits.
Genome-wide association study of cognitive functions and educational attainment in UK Biobank (N=112,151)

Nature Molecular Psychiatry 5 April 2016 doi: 10.1038/mp.2016.45

People’s differences in cognitive functions are partly heritable and are associated with important life outcomes. Previous genome-wide association (GWA) studies of cognitive functions have found evidence for polygenic effects yet, to date, there are few replicated genetic associations. Here we use data from the UK Biobank sample to investigate the genetic contributions to variation in tests of three cognitive functions and in educational attainment. GWA analyses were performed for verbal–numerical reasoning (N=36 035), memory (N=112 067), reaction time (N=111 483) and for the attainment of a college or a university degree (N=111 114). We report genome-wide significant single-nucleotide polymorphism (SNP)-based associations in 20 genomic regions, and significant gene-based findings in 46 regions. These include findings in the ATXN2, CYP2DG, APBA1 and CADM2 genes. We report replication of these hits in published GWA studies of cognitive function, educational attainment and childhood intelligence. There is also replication, in UK Biobank, of SNP hits reported previously in GWA studies of educational attainment and cognitive function. GCTA-GREML analyses, using common SNPs (minor allele frequency>0.01), indicated significant SNP-based heritabilities of 31% (s.e.m.=1.8%) for verbal–numerical reasoning, 5% (s.e.m.=0.6%) for memory, 11% (s.e.m.=0.6%) for reaction time and 21% (s.e.m.=0.6%) for educational attainment. Polygenic score analyses indicate that up to 5% of the variance in cognitive test scores can be predicted in an independent cohort. The genomic regions identified include several novel loci, some of which have been associated with intracranial volume, neurodegeneration, Alzheimer’s disease and schizophrenia.


Discussion

The results of the present study make novel contributions to three scientific aims of GWAS: helping towards identifying specific mechanisms of genomic variation; describing the genetic architecture of complex traits; and predicting phenotypic variation in independent samples. The most important novel contribution of the present study is the discovery of many new genome-wide significant genetic variants associated with reasoning ability, cognitive processing speed and the attainment of a college or university degree. The study provided robust estimates of the SNP-based heritability of the four cognitive variables and their genetic correlations. The study makes important steps toward genetic consilience, because several of the genomic regions identified by the present analyses have previously been associated in GWASs of general cognitive function, executive function, educational attainment, intracranial volume, neurodegenerative disorders and Alzheimer’s disease. The study was successful in using the GWAS results from UK Biobank to predict cognitive variation in new samples. ...

Tuesday, September 19, 2017

Accurate Genomic Prediction Of Human Height

I've been posting preprints on arXiv since its beginning ~25 years ago, and I like to share research results as soon as they are written up. Science functions best through open discussion of new results! After some internal deliberation, my research group decided to post our new paper on genomic prediction of human height on bioRxiv and arXiv.

But the preprint culture is nascent in many areas of science (e.g., biology), and it seems to me that some journals are not yet fully comfortable with the idea. I was pleasantly surprised to learn, just in the last day or two, that most journals now have official policies that allow online distribution of preprints prior to publication. (This has been the case in theoretical physics since before I entered the field!) Let's hope that progress continues.

The work presented below applies ideas from compressed sensing, L1 penalized regression, etc. to genomic prediction. We exploit the phase transition behavior of the LASSO algorithm to construct a good genomic predictor for human height. The results are significant for the following reasons:
We applied novel machine learning methods ("compressed sensing") to ~500k genomes from UK Biobank, resulting in an accurate predictor for human height which uses information from thousands of SNPs.

1. The actual heights of most individuals in our replication tests are within a few cm of their predicted height.

2. The variance captured by the predictor is similar to the estimated GCTA-GREML SNP heritability. Thus, our results resolve the missing heritability problem for common SNPs.

3. Out-of-sample validation on ARIC individuals (a US cohort) shows the predictor works on that population as well. The SNPs activated in the predictor overlap with previous GWAS hits from GIANT.
The scatterplot figure below gives an immediate feel for the accuracy of the predictor.
Accurate Genomic Prediction Of Human Height
(bioRxiv)

Louis Lello, Steven G. Avery, Laurent Tellier, Ana I. Vazquez, Gustavo de los Campos, and Stephen D.H. Hsu

We construct genomic predictors for heritable and extremely complex human quantitative traits (height, heel bone density, and educational attainment) using modern methods in high dimensional statistics (i.e., machine learning). Replication tests show that these predictors capture, respectively, ∼40, 20, and 9 percent of total variance for the three traits. For example, predicted heights correlate ∼0.65 with actual height; actual heights of most individuals in validation samples are within a few cm of the prediction. The variance captured for height is comparable to the estimated SNP heritability from GCTA (GREML) analysis, and seems to be close to its asymptotic value (i.e., as sample size goes to infinity), suggesting that we have captured most of the heritability for the SNPs used. Thus, our results resolve the common SNP portion of the “missing heritability” problem – i.e., the gap between prediction R-squared and SNP heritability. The ∼20k activated SNPs in our height predictor reveal the genetic architecture of human height, at least for common SNPs. Our primary dataset is the UK Biobank cohort, comprised of almost 500k individual genotypes with multiple phenotypes. We also use other datasets and SNPs found in earlier GWAS for out-of-sample validation of our results.
This figure compares predicted and actual height on a validation set of 2000 individuals not used in training: males + females, actual heights (vertical axis) uncorrected for gender. For training we z-score by gender and age (due to Flynn Effect for height). We have also tested validity on a population of US individuals (i.e., out of sample; not from UKBB).


This figure illustrates the phase transition behavior at fixed sample size n and varying penalization lambda.


These are the SNPs activated in the predictor -- about 20k in total, uniformly distributed across all chromosomes; vertical axis is effect size of minor allele:


The big picture implication is that heritable complex traits controlled by thousands of genetic loci can, with enough data and analysis, be predicted from DNA. I expect that with good genotype | phenotype data from a million individuals we could achieve similar success with cognitive ability. We've also analyzed the sample size requirements for disease risk prediction, and they are similar (i.e., ~100 times sparsity of the effects vector; so ~100k cases + controls for a condition affected by ~1000 loci).


Note Added: Further comments in response to various questions about the paper.

1) We have tested the predictor on other ethnic groups and there is an (expected) decrease in correlation that is roughly proportional to the "genetic distance" between the test population and the white/British training population. This is likely due to different LD structure (SNP correlations) in different populations. A SNP which tags the true causal genetic variation in the Euro population may not be a good tag in, e.g., the Chinese population. We may report more on this in the future. Note, despite the reduction in power our predictor still captures more height variance than any other existing model for S. Asians, Chinese, Africans, etc.

2) We did not explore the biology of the activated SNPs because that is not our expertise. GWAS hits found by SSGAC, GIANT, etc. have already been connected to biological processes such as neuronal growth, bone development, etc. Plenty of follow up work remains to be done on the SNPs we discovered.

3) Our initial reduction of candidate SNPs to the top 50k or 100k is simply to save computational resources. The L1 algorithms can handle much larger values of p, but keeping all of those SNPs in the calculation is extremely expensive in CPU time, memory, etc. We tested computational cost vs benefit in improved prediction from including more (>100k) candidate SNPs in the initial cut but found it unfavorable. (Note, we also had a reasonable prior that ~10k SNPs would capture most of the predictive power.)

4) We will have more to say about nonlinear effects, additional out-of-sample tests, other phenotypes, etc. in future work.

5) Perhaps most importantly, we have a useful theoretical framework (compressed sensing) within which to think about complex trait prediction. We can make quantitative estimates for the sample size required to "solve" a particular trait.

I leave you with some remarks from Francis Crick:
Crick had to adjust from the "elegance and deep simplicity" of physics to the "elaborate chemical mechanisms that natural selection had evolved over billions of years." He described this transition as, "almost as if one had to be born again." According to Crick, the experience of learning physics had taught him something important — hubris — and the conviction that since physics was already a success, great advances should also be possible in other sciences such as biology. Crick felt that this attitude encouraged him to be more daring than typical biologists who tended to concern themselves with the daunting problems of biology and not the past successes of physics.

Monday, May 22, 2017

NYTimes: In ‘Enormous Success,’ Scientists Tie 52 Genes to Human Intelligence


The Nature Genetics paper below made a big splash in today's NYTimes: In ‘Enormous Success,’ Scientists Tie 52 Genes to Human Intelligence. The picture above is of a UK Biobank storage facility for blood (DNA) samples.

The results are not especially surprising to people who have been following the subject, but this is the largest sample of genomes and cognitive scores yet analyzed (~80k individuals). SSGAC has assembled a much larger dataset (~750k, soon to be over 1M; over 600 genome-wide significant SNP hits), but are working with a proxy phenotype for cognitive ability: years of education.
Genome-wide association meta-analysis of 78,308 individuals identifies new loci and genes influencing human intelligence

Nature Genetics (2017) doi:10.1038/ng.3869
Received 10 January 2017 Accepted 24 April 2017 Published online 22 May 2017

Intelligence is associated with important economic and health-related life outcomes1. Despite intelligence having substantial heritability2 (0.54) and a confirmed polygenic nature, initial genetic studies were mostly underpowered3, 4, 5. Here we report a meta-analysis for intelligence of 78,308 individuals. We identify 336 associated SNPs (METAL P < 5 × 10−8) in 18 genomic loci, of which 15 are new. Around half of the SNPs are located inside a gene, implicating 22 genes, of which 11 are new findings. Gene-based analyses identified an additional 30 genes (MAGMA P < 2.73 × 10−6), of which all but one had not been implicated previously. We show that the identified genes are predominantly expressed in brain tissue, and pathway analysis indicates the involvement of genes regulating cell development (MAGMA competitive P = 3.5 × 10−6). Despite the well-known difference in twin-based heratiblity2 for intelligence in childhood (0.45) and adulthood (0.80), we show substantial genetic correlation (rg = 0.89, LD score regression P = 5.4 × 10−29). These findings provide new insight into the genetic architecture of intelligence.
Perhaps the most interesting aspect of this study is the further evidence it provides that many (the vast majority?) of the hits discovered by SSGAC are indeed correlated with cognitive ability (as opposed to other traits such as Conscientiousness, which might influence educational attainment without affecting intelligence):
To examine the robustness of the 336 SNPs and 47 genes that reached genome-wide significance in the primary analyses, we sought replication. Because there are no reasonably large GWAS for intelligence available and given the high genetic correlation with educational attainment, which has been used previously as a proxy for intelligence7, we used the summary statistics from the latest GWAS for educational attainment21 for proxy-replication (Online Methods). We first deleted overlapping samples, resulting in a sample of 196,931 individuals for educational attainment. Of the 336 top SNPs for intelligence, 306 were available for look-up in educational attainment, including 16 of the independent lead SNPs. We found that the effects of 305 of the 306 available SNPs in educational attainment were sign concordant between educational attainment and intelligence, as were the effects of all 16 independent lead SNPs (exact binomial P < 10−16; Supplementary Table 14). ...
Carl Zimmer did a good job with the Times story. The basic ideas, that
0. Intelligence is (at least crudely) measurable
1. Intelligence is highly heritable (much of the variance is determined by DNA)
2. Intelligence is highly polygenic (controlled by many genetic variants, each of small effect)
3. Intelligence is going to be deciphered at the molecular level, in the near future, by genomic studies with very large sample size 
are now supported by overwhelming scientific evidence. Nevertheless, they are and have been heavily contested by anti-Science ideologues.

For further discussion of points (0-3), see my article On the genetic architecture of intelligence and other quantitative traits.

Wednesday, June 10, 2015

Replication and cumulative knowledge in life sciences

See Ioannidis at MSU for video discussion of related topics with the leading researcher in this area, and also Medical Science? Is Science Self-Correcting?
The Economics of Reproducibility in Preclinical Research (PLoS Biology)

Abstract: Low reproducibility rates within life science research undermine cumulative knowledge production and contribute to both delays and costs of therapeutic drug development. An analysis of past studies indicates that the cumulative (total) prevalence of irreproducible preclinical research exceeds 50%, resulting in approximately US$28,000,000,000 (US$28B)/year spent on preclinical research that is not reproducible—in the United States alone. We outline a framework for solutions and a plan for long-term improvements in reproducibility rates that will help to accelerate the discovery of life-saving therapies and cures.
From the introduction:
Much has been written about the alarming number of preclinical studies that were later found to be irreproducible [1,2]. Flawed preclinical studies create false hope for patients waiting for lifesaving cures; moreover, they point to systemic and costly inefficiencies in the way preclinical studies are designed, conducted, and reported. Because replication and cumulative knowledge production are cornerstones of the scientific process, these widespread accounts are scientifically troubling. Such concerns are further complicated by questions about the effectiveness of the peer review process itself [3], as well as the rapid growth of postpublication peer review (e.g., PubMed Commons, PubPeer), data sharing, and open access publishing that accelerate the identification of irreproducible studies [4]. Indeed, there are many different perspectives on the size of this problem, and published estimates of irreproducibility range from 51% [5] to 89% [6] (Fig 1). Our primary goal here is not to pinpoint the exact irreproducibility rate, but rather to identify root causes of the problem, estimate the direct costs of irreproducible research, and to develop a framework to address the highest priorities. Based on examples from within life sciences, application of economic theory, and reviewing lessons learned from other industries, we conclude that community-developed best practices and standards must play a central role in improving reproducibility going forward. ...

Friday, November 10, 2006

Hedge fund clones

The Economist discusses some proposals for cheap replication of hedge fund strategies. The first paper mentioned below is by Andy Lo of MIT. Recent innovations like ETFs and other narrowly focused instruments allow cheaper exposure to well defined types of risk -- the cost of placing a bet on a particular strategy is lower than ever before. However, this begs the question of how one decides which bet to make, and when. I don't think the difficult part of generating alpha is the mechanics of making an investment (at least, not anymore), but rather the decision making.

Economist: ...But financial scholars are beginning to demystify hedge funds. They think they can replicate their performance using garden-variety financial products. The result could be a cheap competitor for the hedge-fund titans, akin to the index-tracking funds that have eaten into the market shares of active fund managers.

Replication is possible because hedge-fund managers are not as distinctive as they claim. They say their returns are based on skill, or “alpha”, but in fact their performance is largely derived from market movements. A recent paper* by two academics at the Massachusetts Institute of Technology breaks down the returns of 1,610 funds from 1986 to 2005. It finds that six common factors, such as the change in the S&P 500 index and the return on corporate bonds, explained a significant part of hedge-fund returns.

Investors can gain exposure to these factors through widely available liquid instruments. Thus it should be possible to build “clone” portfolios that resemble hedge funds. Such portfolios would not only avoid hedge-fund fees, but would also escape the risk of backing a mismanaged fund, such as Amaranth, which lost 65% of its value in September.

The authors suggest cloning a fund by dissecting its performance over the past year or two. One could sift and sort the factors behind its success, and allocate the clone's money accordingly. A back-tested clone portfolio returned an average of 12.8% a year over nearly 20 years compared with 14.2% for the typical hedge fund. And the copycat portfolio offered investors many of the same benefits of diversification as the fund it mimicked.

Not every academic is impressed by this approach, however. Harry Kat, at the Cass Business School, says that such “multi-factor” models fail to explain a large proportion of hedge-fund returns. But Mr Kat proposes his own cheap alternative†. It may be impossible to know the particular plays hedge-fund managers make. But, he says, you can devise a formulaic trading strategy in the futures markets that would duplicate the overall shape of their returns. His strategies would give investors two of the three things they look for from a hedge: a low correlation with their existing portfolio, at a level of volatility they can tolerate. The return would be out of their hands, but tests suggest profits can be decent: 10% a year in one example.

Saturday, August 27, 2005

Dynamical hedging and Black-Scholes

Derman and Taleb claim that you can derive Black-Scholes without assuming instantaneous replication of the option using cash and stock. (This assumption has some practical limitations in the real world.) I always thought the replication insight very important for justifying the use of the riskless rate to discount cash flows. As the authors note, formulae equivalent to B-S were derived by others such as Samuelson, but leaving the discount rate as an unknown parameter. Once you know the option can be perfectly hedged, it becomes straightforward to price it given any model in which you can compute the expected return.

For elementary discussion, see here.

Sunday, February 07, 2010

The new dating game



In case you are unfamiliar with terms like (no, this has nothing to do with portfolio theory): alpha, beta, neg, PUA, AFC, and chick crack, read the excerpt below. The photo above is just one of many from the site Hot chicks with douchebags. More details in this Wikipedia entry.

I spent my late teen years at an approximately all-male university near Los Angeles, so I endured way too much time at bars talking to women like the ones described in the article below (in case you are wondering, I had a very good fake ID, but that's another story). I remember a weeknight (happy hour!) at a club in Glendale, with a French guy (grad student, I knew him from the gym) who is now a professor of bioinformatics. I was just a kid -- all the women there were much older than I was. Pierre, I'll call him, had just finished dancing with a modestly attractive blonde and sat down at the bar with me. Are you really interested in her? I asked*. He winked at me and mouthed a single word: Practice :-)

The evo-psych explanations given below date back at least to Caltech guys (anthropologists of the LA singles scene) of the 1980s, and probably much earlier.

* Modern lingo: Would you really hit that?

Weekly Standard: ... In the late 1990s, Mystery developed a precise and exacting “algorithm” of moves and routines—pre-scripted lines to be practiced in the field—that are virtually guaranteed (according to Mystery at least) to lure a female into your bed after just seven hours in her company from a cold turkey meeting in a public place. ... The fundamental strategy is to “demonstrate higher value” (DHV, another Mystery acronym), to appear so fascinating that the woman will want to prove her worthiness to you, not the other way around. You don’t buy her a drink; you offer to let her buy you one. You don’t give her your phone number; you get her to give you hers, in what Mystery calls a “number closing.” If she asks you what you do for a living, you don’t mention the drone desk job that you actually hold down; you tell her you “repair disposable razors” (the choice of a Mystery disciple). You “peacock” (yet another Mystery coinage), which means donning outlandish, attention-grabbing attire. Mystery’s signature peacocking wardrobe includes a black fur bucket hat and matching black nail polish and eyeliner. On The Pickup Artist, he sported a seemingly inexhaustible supply of exotic headgear and man-baubles.

...

If it all sounds cheesy, tedious, manipulative, obvious, condescending to women, maybe kind of gay, it’s because it is. But here’s the rub: This stuff works. If you think men who peacock look ridiculous and unmanly, click onto the photo-website Hot Chicks With Douchebags, where spectacular-looking babes hang on the pecs of preening rednecks and “Jersey Shore”-style guidos sporting chest-baring shirts and product-stiffened fauxhawks. Watch the video “Learn Enough Guitar to Get Laid” on YouTube (three chords, max). In June 2005, Craig Malisow, a reporter for the Houston Press, trailed 24-year-old Bashev, a Bulgarian-born graduate student in engineering at Rice University and self-styled pickup expert, to a series of bars and clubs in Houston. Bashev had no intention of telling the 20-something HBs he met that his day job consisted of working with multivariable calculus. Instead he pointed to his shoes and informed them that he was a “foot model.” Then he launched into his canned opener: Did they think reality shows were “really real”? Sure, two groups of females on whom Bashev tried that line rolled their eyes and smirked, but three bars (and the same routine) later, he was relaxing in a lounge chair reading a shapely brunette’s palm (chick crack plus “kino,” a Mystery-ism that refers to getting a woman to crave your touch), and soon enough “her fingers were gently grasping the backs of his wrists,” Malisow observed. Within minutes, Bashev had not only number-closed but gotten a date for the following Wednesday.

Pickup mentors are relying, consciously or sub, on the principles of evolutionary psychology, which uses Darwinian theory to account for human traits and practices. Robert Wright introduced the reading public to evolutionary psychology in his 1994 book, The Moral Animal: Why We Are the Way We Are. He summarized what biologists had observed in the field: that among animals—and especially among our closest relatives, the great apes—males often fight each other for females and so the most dominant, or “alpha,” male has access to the most desirable, and perhaps all, of the females. But it’s the female of the species who ultimately makes the choice as to which member of the pack she will deem the alpha male. “Females are choosy in all the great ape species,” Wright wrote. He also noted that, for example, a female gorilla will be faithful—forced into fidelity, actually—to a single dominant male, but she will willingly desert him for a rival male who impresses her with his superior dominance by fighting with her mate. That’s because, as Darwin postulated, evolution isn’t merely a matter of survival of the fittest but also of the replication of the fittest, “selfish genes,” in the words of neo-Darwinian Richard Dawkins. Driven by instinctual desire for offspring, male primates chase fertile females so they can replicate themselves, while female primates choose strong males on the basis of survival traits to be passed on to young ones.

Evolutionary psychologists like David Buss in The Evolution of Desire (1994) and Geoffrey Miller in The Mating Mind (2000) have elaborated on these theories, arguing that the human brain itself, with its capacity for consciousness, reasoning, and artistic creation, evolved as an entertainment device for male hominids competing to impress the females in the pack. Dennis Dutton’s new book, The Art Instinct, makes much the same argument. Evolutionary psychologists postulate that the same physical and psychological drives prevail among modern humans: Men, eager for replication, are naturally polygamous, while women are naturally monogamous—but only until a man they perceive as of higher status than their current mate comes along. Hypergamy—marrying up, or, in the absence of any constrained linkage between sex and marriage, mating up—is a more accurate description of women’s natural inclinations. Long-term monogamy—one spouse for one person at one time—may be the most desirable condition for ensuring personal happiness, accumulating property, and raising children, but it is an artifact of civilization, Western civilization in particular. In the view of many evolutionary psychologists, long-term monogamy is natural for neither men nor women.

...

Evolutionary psychology also provides support for a truth universally denied: Women crave dominant men. And it seems that where men are forbidden to dominate in a socially beneficial way—as husbands and fathers, for example—women will seek out assertive, self-confident men whose displays of power aren’t so socially beneficial. This game of sexual Whack-a-Mole is played regularly these days in a culture that, starting with children’s schoolbooks and moving up through films and television, targets as oppressors and mocks as bumblers the entire male sex.

...

Living in the New Paleolithic can be hard on women, many of whom party on merrily until they reach age 30 and then panic. “They’re at the peak of their beauty in their early 20s—they’re luscious—but the guys their age don’t look as good, so they say to themselves: ‘Why do I want to get married?,’ ” notes Kay Hymowitz, a contributing editor to the Manhattan Institute’s City Journal, who is writing a book about the singles crisis. “Then they get to age 28, 29, and their fertility goes down and they’re not quite so luscious. But the guys their age are starting to make money, they look better, they’ve got self-assurance, and they’ve also got the pick of the 23-year-olds.”

Some argue, though, that it is actually beta men who are the greatest victims of the current mating chaos: the ones who work hard, act nice, and find themselves searching in vain for potential wives and girlfriends among the hordes of young women besotted by alphas. That is the underlying message of what is undoubtedly the most deftly written and also the darkest of the seduction-community websites, the blog Roissy in DC. Unlike his confreres, Roissy does not sell books or boot camps, and his site carries no ads. He also blogs anonymously, or at least tries to. (Purported photos of Roissy circulating on the Internet show a tall unshaven man in his late 30s with piercing blue eyes and good, if somewhat dissolute, looks.) The pseudonym Roissy derives from the chateau that was the setting for sadomasochistic orgies in The Story of O, the French pornographic classic of the 1960s which featured a beautiful young woman who couldn’t get enough of being violated and flagellated by masterful men. Roissy maintains that he is not an S&M-fetishist but picked the pseudonym because “chicks dig power.”

...

“The sexual revolution in America was an attempt by women to realize their own [hypergamous] utopia, not that of men,” Devlin wrote. Beta men become superfluous until the newly liberated women start double-clutching after years in the serial harems of alphas who won’t “commit,” lower their standards, and “settle.” During this process, monogamy as a stable and civilization-maintaining social institution is shattered. “Monogamy is a form of sexual optimization,” Devlin told me. “It allows as many people who want to get married to do so. Under monogamy, 90 percent of men find a mate at least once in their life.” This isn’t necessarily so anymore in today’s chaotic combination of polygamy for lucky alphas, hypergamy in varying degrees for females depending on their sex appeal, and, at least in theory, large numbers of betas left without mates at all—just as it is in baboon packs. The aim of Mystery-style game is to give those betas better odds. ...


Related: NYTimes on dating and gender imbalances on campus.

Saturday, August 31, 2013

Genetic architecture of schizophrenia and related psychiatric disorders

This recent paper estimates that 10k common SNPs account for most of the heritability of schizophrenia. I'd guess there are some rare variants of large effect around as well. Where have I seen numbers like 10k causal variants (see also here) before? Just a few years ago, those were crazy numbers.
Genome-wide association analysis identifies 13 new risk loci for schizophrenia (Nature Genetics)

Schizophrenia is an idiopathic mental disorder with a heritable component and a substantial public health impact. We conducted a multi-stage genome-wide association study (GWAS) for schizophrenia beginning with a Swedish national sample (5,001 cases and 6,243 controls) followed by meta-analysis with previous schizophrenia GWAS (8,832 cases and 12,067 controls) and finally by replication of SNPs in 168 genomic regions in independent samples (7,413 cases, 19,762 controls and 581 parent-offspring trios). We identified 22 loci associated at genome-wide significance; 13 of these are new, and 1 was previously implicated in bipolar disorder. Examination of candidate genes at these loci suggests the involvement of neuronal calcium signaling. We estimate that 8,300 independent, mostly common SNPs (95% credible interval of 6,300–10,200 SNPs) contribute to risk for schizophrenia and that these collectively account for at least 32% of the variance in liability. Common genetic variation has an important role in the etiology of schizophrenia, and larger studies will allow more detailed understanding of this disorder.

Also of interest:
Genetic relationship between five psychiatric disorders estimated from genome-wide SNPs (Nature Genetics)

Most psychiatric disorders are moderately to highly heritable. The degree to which genetic variation is unique to individual disorders or shared across disorders is unclear. To examine shared genetic etiology, we use genome-wide genotype data from the Psychiatric Genomics Consortium (PGC) for cases and controls in schizophrenia, bipolar disorder, major depressive disorder, autism spectrum disorders (ASD) and attention-deficit/hyperactivity disorder (ADHD). We apply univariate and bivariate methods for the estimation of genetic variation within and covariation between disorders. SNPs explained 17–29% of the variance in liability. The genetic correlation calculated using common SNPs was high between schizophrenia and bipolar disorder (0.68 ± 0.04 s.e.), moderate between schizophrenia and major depressive disorder (0.43 ± 0.06 s.e.), bipolar disorder and major depressive disorder (0.47 ± 0.06 s.e.), and ADHD and major depressive disorder (0.32 ± 0.07 s.e.), low between schizophrenia and ASD (0.16 ± 0.06 s.e.) and non-significant for other pairs of disorders as well as between psychiatric disorders and the negative control of Crohn's disease. This empirical evidence of shared genetic etiology for psychiatric disorders can inform nosology and encourages the investigation of common pathophysiologies for related disorders.


Blog Archive

Labels