Showing posts with label GWAS. Show all posts
Showing posts with label GWAS. Show all posts

Thursday, February 17, 2022

Another ADHD "Meta-Analysis" Makes Genetic Claims for the Disorder, But Shows the Opposite.

 The latest ADHD GWAS is available in pre-print:

Genome-wide analyses of ADHD identify 27 risk loci, refine the genetic architecture and implicate several cognitive domains

It is now formulaic to perform a GWAS "meta-analysis," rather than independently examining a new dataset. I put meta-analysis in quotes, because this is not really even what we have, since this new data which makes up half the data in the study has not been in a previous study. As I have noted repeatedly, this is problematic and I will touch on why in this critique. Let's get to the claims. 

 The meta-analysis identified 32 lead variants (r2 < 0.1) located in 27 genome-wide significant loci (Figure 1; Table 1, locus plots in Supplementary Data 1), including 21 novel loci. No statistically significant heterogeneity was observed between cohorts 

The first question you might ask is why these 21 novel loci were not noted in the previous GWAS for ADHD? The argument is that when you increase the number of cases, working with a higher N, you are more likely to pick up smaller correlations. The problem with that argument can be seen by the fact that there were 12 loci found significant previously and now only 6 of them are still significant. If we were talking about two entirely different studies, where the larger one picked up 6 out of 12 loci from the previous study, you might make some claims of a modest success and the authors seem to imply exactly this: 

Six of the previously identified 12 loci in the ADHD2019 study14 were significant in the present study (Table 1), and the remaining six loci demonstrated P-values < 8x10-4 

The problem here is that the data from the ADHD2919 study referenced above WAS INCLUDED IN THE CURRENT STUDY. It makes up about half the data, in fact. Thus we are not talking about independent replication, which apparently was not even attempted (or at least no such results were included). If you make the argument that increasing the case numbers identifies more significant loci, then why wouldn't you expect the previous 12 loci to be confirmed? Without even considering population stratification issues, if you have 12 loci with low p values for correlation, you are bolstering the dataset. The fact that half the loci did not retain significance should sound alarm bells. 

Similarly, it is assumed that increasing case size would increase the identified h2 heritability related to genes. Let's see how that turns out:

The SNP heritability (h2 SNP) was estimated to 0.14 (s.e. = 0.01), which is lower than the previously reported h2 SNP of 0.2214. The h2 SNP for iPSYCH (h2 SNP = 0.23; s.e. = 0.01) was in line with the previous finding, but lower h2 SNP was observed for PGC (h2 SNP = 0.12; s.e. = 0.03) and deCODE (h2 SNP = 0.081; s.e. = 0.014). Between-cohort heterogeneity in h2 SNP is not unusual and has been observed for other disorders like e.g. MDD <Major Depressive Disorder>.

One interpretation of this finding, apparently not occurring to the authors, is that the positive findings they have are little more than population stratification, and even in relatively homogenous (white European) cohorts, such pop strat loses its strength from one study to the next. It is a bit amusing that the counter to this is that it was observed in MDD, circularly assuming that both are valid. In other words, getting contradictory results for other diagnoses validates that it should be expected for ADHD. They, in fact, double down on this dubious argument:

The observation that previously identified loci may not reach genome-wide significance in a subsequent larger GWAS, has also been seen for other psychiatric disorders, e.g. bipolar disorder, where eight out of 19 loci were significant in a subsequent larger study.

It's hard not to laugh, and I'll point out that the "larger" GWAS for other disorders like bipolar disorder also had this contradiction even though they were also using data from the studies that first "discovered" the loci.  

Much of the rest of the study involved "enrichment," statistics, making the argument that cognitive related genes are more common among the significant loci. This is impossible to critique without access to the methods used. However, I would ask the authors to consider whether the 6 loci that did not remain significant were claimed to be enriched in previous studies? Is this an indication for the loci being valid, or is this an indication that these enrichment statistics are misguided?


 

 

 

Wednesday, December 9, 2020

Wake Up Call for Insomnia GWAS

Here is another GWAS, this time for insomnia, that I think buries the lead:

Genome-wide meta-analysis of insomnia in over 2.3 million individuals implicates involvement of specific biological pathways through gene-prioritization

Here's an alternate title:

Based on 1.3 million GWAS, the maximum variance explained was 2.6% and based on 2.3 million individuals the maximum variance explained seems to be only 2% !

- (Hat tip to Veera M. Rajagopal, twitter handle: @doctorveera, who might not really appreciate the hat tip)

Obviously, there is a problem here, when, even when finding novel loci by expanding your dataset, you are getting getting worse "variance explained" from your PRS. I think this suggests that they have already reached their peak, which seems to run in the 2 to 3% range for most behavioral traits. I will once again point out that even this number is suspect, since it is not compared to any null trait. They try to rationalize it by suggesting that that the added data (from 23andMe) was less stringently phenotyped, but you can't have it both ways. Expanding datasets does not appear to give us any more real insight. It just bumps up the number of loci meeting significance, which arguably just a collection of false positives.

As the datasets expands beyond just white Europeans, I suspect this will onlly further water down the success of these studies, since they will not be able to rely as much on pop strat to get correlations.

Bipolar Genetics Makes No Progress

Nice critique by Peter Simons of a genetic study for bipolar disorder among Han Chinese with the diagnosis  of Bipolar Disorder (original study here). A couple of excerpts:

The researchers analyzed thousands of Han Chinese people and found that genetics explained just 2.3% of whether they received a diagnosis of bipolar disorder (BD) or not...

However, it is unclear how tiny correlations like this—which affect but a tiny sample of the population studied and explain less than 3% of the risk for a diagnosis—could help researchers understand the supposed “biological etiology” of bipolar disorder. In fact, they rather show that more than 97% of the reason that someone gets a diagnosis is explained by factors other than biology. 

As I like to point out, even the 3% is quite suspect and is arguably noise and should be tested against a null trait to establish that the 3% is not the null.

Tuesday, December 1, 2020

The Unembarrassed Bot

 A shortcutting of the usual GWAS is a bot that simply cranks out a Manhattan plot with no further analysis. While, those who do traditional GWAS downplay it, there is really little difference between what they are doing and what the bot is doing other than some shoddy speculation and perhaps a bit of data cleaning, but the real issue is that the bot does GWAS that most would be too embarrassed to publish and these get a lot of hits. Take this one, for example, that ironically, without embarrassment, finds genetic variants for "worrying too long after embarrassment":


This should be a clear indication that silly false positives can be produced from anything you can ask on a questionnaire. In addition to the likelihood of some massive pop strat dependent on particular cultural backgrounds, what exactly is meant by "too long"? Is this a subjective opinion of the person or is it a specific amount of time? 


Sunday, September 20, 2020

GWAS Meta-analyis for Bipolar Disorder Gives Glowing Analysis, but is impossible to Interpret (Again)

 A brief review of this GWAS for Bipolar Disorder:

Genome-wide association study of over 40,000 bipolar disorder cases provides novel biological insights (Mullins et al. )

Like almost all the behavioral genetic GWAS studies, this one uses a meta-analysis, despite having new data added to previous data and the new data was never assessed (at least in print) independently. Thus, it is difficult to assess statistically what is success and what is failure, although it is filled with the usual accolades:

This GWAS provides the best-powered BD polygenic scores to date, when applied in both European and diverse ancestry samples. Together, these results advance our understanding of the biological etiology of BD, identify novel therapeutic leads and prioritize genes for functional follow-up studies.

 Well, the best and the only, really. But, of course, I have a lot of questions. The first is related to their significant loci count, and for which I needed partial clarification from one of the authors, as I will discuss after the fold (click "read more" to continue).

Monday, August 17, 2020

My Four Laws of the Behavioral Genetics Fallacy

 I discussed these in more length, here as a response to Eric Turkheimer's Three Laws of Behavior Genetics. But just wanted to lay them out in one short post (credit Turkheimer for the second, which is his third).

My Four Laws of the Behavioral Genetics Fallacy:

1. Any behavioral trait studied within a society will be correlated genetically to specific subpopulations, regardless of whether these genetic correlations are directly related to the trait.

2. A substantial portion of the variation in complex human behavioral traits is not accounted for by the effects of genes or families.

3. Differences in human behavior, intelligence and personality are not accounted for by structural or functional differences in the brain.

4. Advancements in understanding human behavior and psychology require inner exploration from the scientist, the subject or both.

Wednesday, August 12, 2020

Yet More UK BioBank Pop Strat Issues Noted.

 There are so many studies coming out noting population stratification issues that it is hard to keep track. This is an interesting preprint looking at CAD and BMI:

Fine-scale population structure confounds genetic risk scores in the ascertainment population

From the Abstract:

we investigated the accuracy of two different GRS across population strata of the UK Biobank, separated along principal component (PC) axes, considering different approaches to account for social and environmental confounders. We found that these scores did not predict the real differences in phenotypes observed along the first principal component, with evidence of discrepancies on axes as high as PC45. These results demonstrate that the measures currently taken for correcting for population structure are not sufficient, and the need for social and environmental confounders to be factored into the creation of GRS. 

One interesting aspect of this study, I think, is that it highlights how it can be necessary to have a good working knowledge of the population you are studying.  This plot is striking in that respect:

This was confined only to white European descent, but still had this kind of stratification. A larger point here is that more and more pop/strat issues arise, many of which were not accounted for in earlier studies and perhaps should lead to corrections. Moreover, for those doing GWAS in the future, particularly in the UK BioBank, it is worth having a bit of skepticism that at least some of what you are seeing is pop/strat that has yet to be recognized.

 

Wednesday, July 22, 2020

Another Paper Related to Pop Strat issues for GWAS

Another study discussing pop strat issues:

Demographic history impacts stratification in polygenic scores

Points out more issues with population stratification:
We show that when population structure is recent, it cannot be fully corrected using principal components based on common variants—the standard approach—because common variants are uninformative about recent demographic history.
They further note some limitations with sibling based studies:
While sibling-based association tests are immune to stratification, the hybrid approach of ascertaining variants in a standard GWAS and then re-estimating effect sizes in siblings reduces but does not eliminate bias. 
As I've argued previously, the "immune to stratification" point is not necessarily true secondary to factors like varying ages of the siblings and selections biases of the databases. Nonetheless, if using sibling studies  "reduces but does not eliminate bias," and they are bringing the variance explained down to 2 or 3 %, then arguably they are scraping along near the null. So, far from showing that some of the variance explained is retained in sibling studies, it might suggest that there is no real genetic component found.

Finally, it's worth pointing out that despite the growing number of studies showing pop/strat issues in the UK Biobank and other such databases, no one has taken it upon themselves to reevaluate their previous, published GWAS results in light of this. It's as if they are grandfathered in.

Friday, June 19, 2020

Genes for Substance Abuse Has Made No Progress, but Unjustified Optimism Continues

Yet another genetic study of substance abuse:

Using polygenic scores for identifying individuals at increased risk of substance use disorders in clinical and population samples
Highlights:
These PRSs explain ~2.5–3.5% of the variance in AUD (across FT12 and COGA) when all PRSs are included in the same model. 
...usefulness for identifying those at increased risk in their current form is modest, at best 
This was from an all white European sample, ftr, with the assumption that pop strat is accounted for. One can assume, as has been the case, that such pop strat will be found and water this down to next to nothing. That said, is the null 0% or is 2 or 3 % about as low as you can get? I'd be happy to see an example in which a PRS does worse than this.

So is the conclusion that perhaps we are barking up the wrong tree? Of course not:
 Improvement in predictive ability will likely be dependent on increasing the size of well-phenotyped discovery samples. 

The shell game continues...

Friday, December 20, 2019

Perhaps We Have a Use for These GWAS, Afterall

In my last post, I briefly critiqued this study absurdly correlating genetics to income and offered a challenge to the authors. My opinion of GWAS is obvious to anyone who skims through this blog, but it occurs to me that perhaps we might have a use for GWAS, afterall, as I recently tweeted:
Here’s a different take: The extent to which you can correlate genes to income in a society, is a direct measure of the unfairness and class stratification of that society. 
If we assume, as I do, that most genetic correlations in the behavioral genetics realm are due to population stratification, then we know that any genetic correlations would demonstrate ways in which the society is stratified. This could be in obvious ways such as racial delineations, but might also include more subtle classist issues (He/She is not from the right family...) and would be an even better way to measure more covert discrimination. By the way, I think this is provable in the sense that other societies will have entirely different loci correlated to income, a fact that will cause a lot of mental gymnastics to explain away.
If we can't prove the causality of the genes flagged in such studies, shouldn't we assume that they are an indication, of an unfair stratification of the society? If we could rid ourselves of all such genetic commonalities, wouldn't that lead us to a true meritocracy? Therefore, wouldn't it make sense and be more fair to give job and college admission preferences to those with the LOWEST polygenic scores for income? As the very "not racist" individuals who embraced this study and took me to task on Twitter pointed out, shouldn't we pursue the truth wherever it happens to lead?

Monday, December 16, 2019

A Challenge to the "Income" Genes Clan

I was going to do a longer critique of this study:


Genome-wide analysis identifies molecular systems and 149 genetic loci associated with income

However, I do a lot of these and it seems like I  go after each head of a hydra, only to be met with one more absurd than the last.  Instead I am going to say a couple of things and offer a challenge to the authors. This study claimed to have found 30 loci associated with income (29 novel). They then went digging around the UK BioBank, for which many of the authors in this study are all too familiar, and used MTAG for other dubious phenotypes, like "Educational Attainment", and "intelligence" to crank up another 120 associations. Imagine the assumptions of the authors regarding income, our economic system, the illusion of meritocracy, IQ, the primacy of income, and the fantasy that you can add up a bunch of genetic variants and determine the likelihood of a high or low income for a person.
A couple of points about the 30 loci noted above. First, there were only 2 previous loci correlated with this trait in the past. So one of the two was not significant. You might think, well, at least they replicated one loci. However, the previous 2 loci come from a smaller version of the same damn UK BioBank dataset. Thus, even though they were using some of the same data as their last study (yes, same authors, same database), one of the loci didn't reach significance with additional data. So what we really have are 30 "novel" loci, that have never been replicated. This, in my opinion, is what one might view as a "screening study." So, good, go and do a GWAS of an INDEPENDENT dataset that you haven't been turning upside down for the past 5 years and see if any of these same loci meet significance instead of trying every kind of gymnastic exercise to correlate these likely false positives into something meaningful.
In fact, I challenge the authors of this study to do a GWAS for "income" on an independent dataset other than the UK BioBank, which at this point is like playing poker when you can see what's in everyone else's hand, and see if you can even replicate a single one of these loci. I'll even handicap you and say you can use all white people again, but somewhere other than the UK. 
If you can't replicate any of these loci, then admit you are playing a shell game, pack up your shit and go find a real, honest job, instead of fueling the prejudices of Charles Murray and Quillette.

Thursday, November 21, 2019

Cross Ancestry Study of Schizophrenia puts out its best face (only)

I wanted to make a few quick points about this study:
Comparative genetic architectures of schizophrenia in East Asian and European populations
I tried to ask a few questions to one of the authors promoting it on Twitter, but he did not respond, so if I am incorrect about any fact, leave a comment here and I will update. Let's start with the Abstract, which is below in full:
Schizophrenia is a debilitating psychiatric disorder with approximately 1% lifetime risk globally. Large-scale schizophrenia genetic studies have reported primarily on European ancestry samples, potentially missing important biological insights. Here, we report the largest study to date of East Asian participants (22,778 schizophrenia cases and 35,362 controls), identifying 21 genome-wide-significant associations in 19 genetic loci. Common genetic variants that confer risk for schizophrenia have highly similar effects between East Asian and European ancestries (genetic correlation = 0.98 ± 0.03), indicating that the genetic basis of schizophrenia and its biology are broadly shared across populations. A fixed-effect meta-analysis including individuals from East Asian and European ancestries identified 208 significant associations in 176 genetic loci (53 novel). Trans-ancestry fine-mapping reduced the sets of candidate causal variants in 44 loci. Polygenic risk scores had reduced performance when transferred across ancestries, highlighting the importance of including sufficient samples of major ancestral groups to ensure their generalizability across populations.

The reason I am showing the entire abstract is to point out what it doesn't say: That the study apparently failed to replicate any of the previous significant loci for schizophrenia (as far as I can tell). The authors simply ignore this, almost as if it is expected, yet expend a lot of time trying to make lemonade out of a lemon without telling us it was a lemon, trying to justify why schizophrenia would present in the same way in different cultures, when it is presumably due to entirely different gene sets.
In my view, you would not expect any of the loci to match between the two studies because the loci are generally false positives, probably enhanced by population stratification issues that are going to be different in these two different populations. Let me go over some of the findings and why I believe they are consistent with pop/strat, false positives after the fold:

Thursday, October 10, 2019

PTSD and the GWAS Hype Machine

A new PTSD GWAS makes a few bold claims.  I think it's a good example of the kind of hype that these studies, which show next to nothing, crank out to hype their results. In this puff piece related to the study, they start with:
Large study reveals PTSD has strong genetic component like other psychiatric disorders
Which 1. It does not and 2. Is not really shown to be true of other psychiatric disorders, either, except in the same hyped fashion as this study. Now let's look at this from the same puff piece:
The study team also reports that, like other psychiatric disorders and many other human traits, PTSD is highly polygenic, meaning it is associated with thousands of genetic variants throughout the genome, each making a small contribution to the disorder. Six genomic regions called loci harbor variants that were strongly associated with disease risk, providing some clues about the biological pathways involved in PTSD.
 If it is highly polygenic, on what basis are they saying this if only 6 loci were strongly associated with disease risk (this is not even accurate, as I'll discuss in moment)? "Genome-wide, a substantial number of variants had some level of association with PTSD, showing the disorder to be highly polygenic," What this is saying is that there are other loci (presumed genetic variants) that did not reach significance, but they include through the subterfuge of "polygenic scores." There is no basis, other than the hopefulness of those doing these studies, that these below significant findings are anything more than non-significant findings. I'll  also note that none of these 6 loci were found in previous studies. Thus, this is an entirely unreplicated study. Now, let's take a look at the loci they did claim to find:

Thursday, September 12, 2019

Depression and Bipolar: Looking at the Positive While Inadvertantly Demonstrating the Negative

I wanted to critique this study which I admittedly struggled to get my head around, so I needed to get help from one of the authors on Twitter. In short, it takes data from two previous studies of depression and bipolar disorder and recombines them. Here is his given explanation:
The combination of the MDD and the bipolar data (which have not been combined in this way before). That is, we are seeing some loci have statistical evidence for "MDD or bipolar" versus control individuals that we haven't seen when looking at either individually so far.
I'm not really sure if that is what they established even on its face since, as I understand it, they simply combine the data from the two studies and perform new GWAS's for both Bipolar Disorder and Depression, creating a new case vs. control for both (I welcome the authors giving a better explanation than I'm putting forth, lest I be accused of creating a straw man. I really just don't fully understand the underlying premise). In doing so, they came up with 15 new loci related to these disorders without using any new data. I believe the point here is to show that bipolar disorder and depression have some genetic commonalities that were demonstrated. They go on to assess these further, but I suggest maybe the lede was buried here and that the study demonstrated another, perhaps more plausible, conclusion: That the original significant loci were false positives, as are these. Let me explain below the fold:

Wednesday, May 1, 2019

Is This a Successful Study for Bipolar Genetics? That's how they are billing it.

A "new" GWAS came out for Bipolar Disorder. As yet, I have only seen the abstract, but I wonder whether I need to see more? Let me comment on a few things:
"Eight of the 19 variants that were genome-wide significant (P < 5 × 10−8) in the discovery GWAS were not genome-wide significant in the combined analysis, consistent with small effect sizes and limited power but also with genetic heterogeneity."
Is that really what it's consistent with? If you have variants that were found to be significant in previous studies and you include the data from those studies in your current study, even if the effect size was small (and, the power now increased), you should expect most of them to retain significance, even if they weren't significant in the new data set independently. The fact that half of them have lost significance is a good indication that most or all of them were false positives to begin with. Moreover, once again, why not do an independent GWAS (I'm assuming they did not) of the new data and compare it to the old data?
Now let's look at the very next sentence:

Wednesday, January 23, 2019

The UK BioBank: The Beast of Pop/Strat

Here is yet another study looking at population stratification issues related to GWAS studies and polygenic score results: Apparent latent structure within the UK Biobank sample has implications for epidemiological analysis.
They looked at geographic structure and found that the UK Biobank is subject to a lot of stratification in that regard. They looked at BMI (body mass index), household income, and educational attainment and found all of them to be subject to geographic population stratification, even with principle component analysis.  First they looked at a smaller subset of genetic data from a previous study (ALSPAC)
...we anticipate that the educational attainment of people who migrate for economic reasons differs from people who do not. Educational attainment is therefore aligned to subtle genetic differences even in this apparently geographically and ethnically homogenous population and this is co-incident with axes of ancestry.
They move on to the beast, the UK Biobank:

Wednesday, January 2, 2019

Genes for Ice Cream Flavor preference...

Yes, the bar gets lowered once again as it approaches a Coke vs. Pepsi gene.  This time we have a "study" that purports to find genes for a preference for chocolate vs. vanilla ice cream (and strawberry, of course).  I know, you are thinking I'm making this up, so here it is.  I'm pretty sure, back in the day, I considered using exactly this possibility to mock these studies, but opted on "finding raisins to be tasty."  I'll get a little bit into the study, but first, let me ask anyone reading this to try to catch yourself in that moment between when something absurd was stated and when you convince yourself that it is somehow valid because, you know, it's science and all.  I like to call that the "Emperor has no clothes!" moment.  Maybe that moment has already passed you by, so try, really try, to remember how you felt in those few seconds before you had to tuck it away.  That moment is important.  It's the brief time when one can see, no matter how much they've been inundated with "scientific" pronouncements to the contrary, that this entire field of study might just have a kind of absurdity to it.
Can't go there?  Well, then you have to believe that there are genetic variants for preferring chocolate or strawberry ice cream over vanilla.  Lot's of them, in fact.  Or you have to explain why this study is not valid and other GWAS studies are.  Let's go through this important study:

Thursday, October 11, 2018

New Depression Study Finding 102 Variants. What is Replication?

A new Depression study claiming 102 genetic variants has just come out (pre-publish).  I don't want to do an extended critique, so I will stick to a few main points:

Wednesday, September 5, 2018

More Risk-Taking genetics

I will make a quick point related to this study:
Genetics of self-reported risk-taking behaviour, trans-ethnic consistency and relevance to brain gene expression (Strawbridge, et al.).


The study notes 8 novel loci for this so-called "risk-taking behavior" (diagnosed by asking people one question: Would you describe yourself as someone who takes risks?”), as well as noting "...two replicated previous findings."  My quick point is that the 2 "replicated" SNP's were from a previous study by the same author using the same UK Biobank dataset (which has expanded since the last study from a few months ago).  Obviously, you would expect some "replication" when using overlapping datasets.
In short, no independently replicated SNP's from previous studies of this ilk, and some of the SNP's, even bolstered by using some of the same data, were not replicated.  Moreover, no independent analysis of the new data, which was simply folded into the old study in a meta-analysis type of format.  Again, in an attempt to bolster N, the study did not look at the new data independently.  
I might add to this critique when I've had more time to examine it in detail. 

Wednesday, July 4, 2018

Genes for loneliness, health club attendance, bar hopping and churchgoing, all in one study!

I am attempting to critique this study:

Elucidating the genetic basis of social interaction and isolation (Day et al.)

There are many directions I can go with such a critique.  The most appealing and easiest, would be to mock it with a couple of quick quotes and be done with it.  Then, I think to myself, are there people out there that take a study like this seriously?  And, of course there are a lot of people who take a study like this seriously.  
I'm hoping, though, that there are a few scientists who have been holding onto these GWAS studies as some sort of proof of all kinds of mental constructs, who might have a bit of a crisis of confidence when reading something like this.  Perhaps they would like to dismiss this as an outlier, or misguided in some way.
Here's where they have a problem.  Because, this study was done by the book.  It has all the elements used to prove that there are genetic associations for these traits, just as studies are done to find associations for IQ, mental disorders, and personality traits.  So if you are touting GWAS studies related to any of these traits, you need to explain why your study is good and this one is ridiculous, or you need to embrace this ridiculous study.  There is no in between.
With that in mind, I will go through how this study follows the same formula as your cherished studies and you can decide which side of the health club attendance gene fence you sit on.