Discussion about this post

User's avatar
John K.'s avatar

This item 'may' pertain to me, but is Way over my head. But thanks for the shrouded Rocky Horror reference. Gave yourself away. (...pation)

Rex's avatar

The pair of them both look pretty embarrassing to me.

When you think about taking action to solve a problem, you have four questions. (1) How _much_ will it help? (2) How _reliably_ will it help? (3) How sure am I that I didn't just fool myself? (4) What are we even measuring?

Question number one is a question of effect size. Question number two is a question of variance explained. And question number three is a question of p-value, confidence interval, log odds ratio, and more esoteric things that statisticians get very animated about before going to war with pitchfork and torches about frequentist vs. Bayesian interpretations.

R^2 is a measure of variance explained. It's asking about reliability. And, indeed, your deciles are quite reliably correlated with CDR. For doctors treating deciles (in this case, a "decile" is a group of ~160 people), this is good news! You have found something that reveals a relationship between two variables. Unfortunately, while you know the drug changes one variable, you _don't_ know that the drug application preserves the relationship. But, anyway, reducing noise is _exactly how block averaging works_.

But let's ask: is it _true_ that there's a relationship? If we look at the unbinned data, the CI on the R^2 value is 0.02-0.05. That means _R^2 is not random in the original_. (And, since the original analysis showed a small significant effect, that's what you'd suspect, though there are ways it could not come out that way). In the reanalysis, it is 0.7-0.97. Still significant! Significance doesn't tell you how big, though! So the binning didn't change the significance conclusion.

Now let's ask how big the average effect is by looking at the average spread in amyloid levels. If you just eyeball it, amyloid levels vary on average by maybe 50-75. That corresponds to an improvement of about 0.5 points, maybe a bit more, according to the decile table. And if you try to eyeball the full graph, lo and behold you get about 0.5 points, maybe a bit more. The reanalysis did not change the effect size!

Now, this doesn't tell you whether the drug did anything, because the analysis throws away labels (that's what we're actually measuring here). This was genuinely dumb (or diabolical) of the reanalysis paper. The underlying relationship could have a negative slope, the drug could keep the same relationship but push the population left on the X-axis, and you'd have a completely useless drug that, if you forgot you threw away your labels, would have a Simpson's paradox-like structure. And the reply paper mentioned in text that something could happen but _didn't provide a clear demonstration_. (It is trivial to demonstrate, like with the R^2-binning!) In this case, in fact, it didn't seem to happen (the original paper _already showed a small effect_, roughly consistent). And the reply paper either didn't notice or didn't deign to note that in fact, in this case, despite it being a stupid thing to do, it wasn't actually a problem.

The entire affair seems incredibly embarrassing. That the reanalysis was published at all, despite being a classic invitation to Simpson's paradox and failing to answer the critical question of whether the drug was beneficial OR whether amyloid predicts decline rates because the two were scrambled together, speaks poorly of the review process. (There are statistical methods to unscramble them.) That the reply didn't even _demonstrate that effect_, and instead focused on the statistical trivia that if you pour patients into differently-shaped cups, even though the noise level is different, the total amount of water (effect size / significance) is the same. And the conclusion wasn't "so don't fret about the shape of your cups, just keep it in mind". No--it was that oh, how dare you, because we love our R^2 to be corrupted by all the confounds it normally is like difference between clinical outcome and assessment variable, measurement noise, and so on, but not changes due to batching!

This is in a journal with an impact factor of 24.

I'm quite appalled.

2 more comments...

No posts

Ready for more?