Showing posts with label Reviews & meta-analysis. Show all posts
Showing posts with label Reviews & meta-analysis. Show all posts

Monday, November 21, 2022

Some studies are MONSTERS!

 

Cartoon of 2 small studies on one side of a meta-analysis, with a very big 3rd study on the other side pulling the studies' combined result over to his side. One of the little studies is thinking "That jerk is always throwing his weight around!"

On the plus side, this jerk explains a lot about the data in a meta-analysis!

This cartoon is a forest plot, a style of data visualization for meta-analysis results. Some people call them "blobbograms". Each of these horizontal lines with a square in the middle represents the results of a different study. The length of that horizontal line represents the length of the confidence interval (CI). That gives you an estimate of how much uncertainty there is around that result - the shorter it is, the more confident we can be about the result. (Statistically Funny explainer here.)

The square is called the point estimate - the study's "result" if you like. Often, it's sized according to how much weight the study has in the meta-analysis. The bigger it is, the more confident we can be about the result.

The size of the point estimate is echoing the length of the confidence interval. They are two perspectives on the same information. Small square and long line provides less confidence than a big square with a short line.



Cartoon showing a big smirking cartoon dragging the summary estimate diamond over to his side of the meta-analysis


The diamond here is called the summary estimate. It represents the summary of the results from the 3 studies combined. It doesn't just add up the 3 results then divide them by 3. It's a weighted average. Bigger studies with more events count for more. (More on that later.)

The left and right tips of the diamond are the two ends of the confidence interval. With each study that gets added to the plot, those tips will get closer together, and it will move left or right if a study's result tips the scales in one direction.

The vertical line in the center is the "line of no effect". If a result touches or crosses it, then the result is not statistically significant. (That's a tricky concept: my explainer here.)

In biomedicine, forest plots are the norm. But in other fields, like psychology, the results of meta-analyses are often presented as tables of data. That means that each data point - the start and end of each confidence interval, and so on - are numbers in a column instead of plotted on a graph. (Here's a study that does that.)

So what about that jerk? He carries so much weight not just because the study has a lot of participants in it. What's called a study's precision depends on the number of "events" in the study, too. 

Say the event you’re interested in is heart attacks – and you are investigating a method for reducing them. But for whatever reason, not a single person in the experimental or control group has a heart attack even though the study was big enough for you to have expected several. That study would have less ability to detect any difference your method could have made, so the study would have less weight.

It's very common for a study, or a couple of them, to carry most of the weight in a meta-analysis. A study by Paul Glasziou and colleagues found that the trial with the most precision carried an average of 51% of the whole result. When that's the case, you really want to understand that study.

Some studies are such whoppers that they overpower all other studies – no matter how many of them there are. They may never be challenged, just because of their sheer size: No one might ever do a study that large on the same question again.

The size of the point estimate and length of the line around it are clues to the weight of the study. The meta-analysis might also include the percentages of weight for each study.

Like to know more? This is a shorter version of one of the tips in my post at Absolutely Maybe5 Tips for Understanding Data in Meta-Analyses. Check it out for a more in-depth example of looking at the weight of a study and 4 more key tips!

Hilda

Monday, October 31, 2022

Researching our way to better research?

 

Cartoon: I do research on research. Person 2: Terrific! I research the research of research


Here we see an expert in evidence synthesis meet a metascientist!

Evidence synthesis is an umbrella term for the work of finding and making sense of a body of research – methods like systematic reviews and meta-analysis. And metascience is studying the methods of science itself. It includes studying the way science is published – see for example my posts on peer review research. And yes, there's metascience on evidence synthesis, too – and syntheses of metascience!

The terms metascience and metaresearch haven't been tossed around for all that long, compared to other types of science. Back when I took my first steps down this road in the early 1990s, in my neck of the science woods we called people who did this methodologists. A guiding light for us was the statistician and all-round fantastic human Doug Altman (1948-2018). He wrote a rousing editorial in 1994 called "The scandal of poor medical research," declaring "We need less research, better research, and research done for the right reasons." Still true, of course.

Altman and colleague, Iveta Simera, chart the early history of metascience over at the James Lind Library. Box 1 in that piece has a collection of scathing quotes about poor research methodology, starting in 1917 with this one on clinical evidence: "A little thought suffices to show that the greater part cannot be taken as serious evidence at all."

The first piece of research on research that they identified was published – with only the briefest of detail, unfortunately – by Halbert Dunn in 1929. He analyzed 200 quantitative papers, and concluded, "About  half of the papers should never have been published as they stood." (It's on the second page here.)

The first detailed report came in 1966, by a statistician and medical student. They reckoned over 70% of the papers they examined should either have been rejected or revised before being published. A few years after that, the methods for evidence synthesis took an important step forward when Richard Light and Paul Smith published their "procedures for resolving contradictions among different research studies" (Light and Smith, 1971.)

Evidence synthesis and metascience have proliferated wildly since the 1990s. And there's lots of the better research that Altman hoped for, too. Unfortunately, though, it's still in the minority – even in evidence synthesis. Sigh! Will more research on research help? Someone should do research on that!

Hilda Bastian


Sunday, June 30, 2013

Goldilocks and the three reviews



Goldilocks is right: that review is FAR too complicated. The methods section alone is 652 pages long! Which wouldn't be too bad, if it weren't that it is a few years out of date. It took so long to do this review and go through rigorous enough quality review, it was already out of date the day it was released. Something that happens often enough to be rather disheartening.

When methodology for systematic reviewing gets overly rococo, the point of diminishing returns will be passed. That's a worry, for a few reasons. For one, it's inefficient and more reviews could be done with the resources. Secondly, more complex methodology can both be daunting, and it can be hard for researchers to accomplish with consistency. Thirdly, when a review gets very elaborate, reproducing or updating it isn't going to be easy either.

It's unavoidable for some reviews to be massive and complex undertakings, though, if they're going to get to the bottom of massive and complex questions. Goldilocks is right about review number 2, as well: that one is WAY too simple. And that's a serious problem, too.

Reviewing evidence needs to be a well-conducted research exercise. A great way to find out more about what goes wrong when it's not, is reading Testing Treatments. And see more on this here at Statistically Funny, too.

You need to check the methods section of every review before you take its conclusions seriously - even when it claims to be "evidence-based" or systematic. People can take far too many shortcuts. Fortunately, it's not often that a review gets as bad as the second one Goldilocks encountered here. The authors of that review decided to include only one trial for each drug "in order to keep the tables and figures to a manageable size." Gulp!

Getting to a good answer also quite simply takes some time and thought. Making real sense of evidence and the complexities of health, illness and disability is often just not suited to a "fast food" approach. As the scientists behind the Slow Science Manifesto point out, science needs time for thinking and digesting.

To cover more ground, people are looking for reasonable ways to cut corners, though. There are many kinds of rapid review, including reliance on previous systematic reviews for new reviews. These can be, but aren't always, rigorous enough for us to be confident about their conclusions.

You can see this process at work in the set of reviews discussed at Statistically Funny a few cartoons ago. Review number 3 there is in part based on review number 2 - without re-analysis. And then review number 4 is based on review number 3.

So if one review gets it wrong, other work may be built on weak foundations. Li and Dickersin suggest this might be a clue to the perpetuation of incorrect techniques in meta-analyses: reviewers who got it wrong in their review, were citing other reviews that had gotten it wrong, too. (That statistical technique, by the way, has its own cartoon.)

Luckily for Goldilocks, the bears had found a third review. It had sound methodology you can trust. It had been totally transparent from the start - included in PROSPERO, the international prospective register for systematic reviews. Goldilocks can get at the fully open review and its data are in the Systematic Review Data Repository, open to others to check and re-use. Ahhh - just right!


PS:

I'm grateful to the Wikipedians who put together the article on Goldilocks and the three bears. That article pointed me to the fascinating discussion of "the rule of three" and the hold this number has on our imaginations.

Tuesday, May 21, 2013

He said, she said, then they said...



Conflicting studies can make life tough. A good systematic review could sort it out. It might be possible for the studies to be pooled into a meta-analysis. That can show you the spread of individual study results and what they add up to, at the same time.

But what about when systematic reviews disagree? When the "he said, she said" of conflicting studies goes meta, it can be even more confusing. New layers of disagreement get piled onto the layers from the original research. Yikes! This post is going to be tough-going...

A group of us defined this discordance among reviews as: the review authors disagree about whether or not there is an effect, or the direction of effect differs between reviews. A difference in direction of effect can mean one review gives a "thumbs up" and another a "thumbs down."

Some people are surprised that this happens. But it's inevitable. Sometimes you need to read several systematic reviews to get your head around a body of evidence. Different groups of people approach even the same question in different but equally legitimate ways. And there are lots of different judgment calls people can make along the way. Those decisions can change the results the systematic review will get.

When and how they searched for studies - and what type and subject - means that it's not at all unusual for groups of reviewers to be looking at different sets of studies for much the same question.

After all that, different groups of people can interpret evidence differently. They often make different judgments about the quality of a study or part of one - and that could dramatically affect its value and meaning to them.

It's a little like watching a game of football where there are several teams on the field at once. Some of the players are on all the teams, but some are playing for only one or two. Each team has goal posts in slightly different places - and each team isn't necessarily playing by the same rules. And there's no umpire.

Here's an example of how you can end up with controversy and people taking different positions even when there's a systematic review. The area of some disagreement in this subset of reviews is about psychological intervention after trauma to prevent post-traumatic stress disorder (PTSD) or other problems:

Published in 2002Published in 2005Published in 2005Published in 2010; Published in 2012Published in 2013.

The conclusions range from saying debriefing has a large benefit to saying there is no evidence of benefit and it seems to cause some PTSD. Most of the others, but not all, fall somewhere in between, leaning to "we can't really be sure". Most are based only on randomized trials, but one has none, and one has a mixture of study types.

The authors are sometimes big independent national or international agencies. A couple of others include authors of the studies they are reviewing. The definition of trauma isn't the same - they may or may not include childbirth, for example. The interventions aren't the same.

The quality of evidence is very low. And the biggest discordance - whether or not there is evidence of harm - hinges mostly on how much weight you put on one trial.

It's about debriefing. The debriefing group is much bigger than the control group because they stopped the trial early, and while it's complicated, that can be a source of bias.

The people in the debriefing group were at quite a lot higher risk of PTSD in the first place. Data for more than 20% of the people randomized is missing - and that biases the results too (it's called attrition bias). You can't be sure those people didn't return because they were depressed, for example. If so, that could change the results.

It's no wonder there's still a controversy here.


See also my 5 tips for understanding data in meta-analysis.

Links to key papers about this in my comment at PubMed Commons (archived here).


If you want to read more about debriefing, here's my post in Scientific American: Dissecting the controversy about early psychological response to disasters and trauma.


Friday, October 12, 2012

The Forest Plot Trilogy - a gripping thriller concludes



Forest plots, funnel plots - and what's with the mysterious diamond symbol, lurking like a secret sign, in meta-analyses? Meta-analysis is a statistical technique for combining the results of studies. It is often used in systematic reviews (and in non-systematic reviews, too).

A forest plot is a graphical way of presenting the results of each individual study and the combined result. The diamond is one way of showing that combined result. Here's a representation of a forest plot, with 4 trials (a line for each). The 4th trial finds the treatment better than what it's compared to: the other 3 had equivocal results because they're crossing the vertical line of no effect.



A funnel plot is one way of exploring for publication bias: whether or not there may be unpublished studies. Funnel plots can look kind of like the sketches below. The first shows a pretty normal distribution of studies - each blob is a study. It's roughly symmetrical: small under-powered studies spread around, with both positive and negative results.



This second one is asymmetrical or lopsided, suggesting there might be some studies that didn't show the treatment works - but they weren't published:


        Gaping hole where negative studies should be



(This post uses snapshots from slides I'll be using to explain systematic reviews at the 2012 NIH Medicine in the Medicine course that's starting this weekend. It's several days of in-depth training in evidence and statistics for journalists. This year it's being held at Potomac, just near Washington. And here's a post on the start of the course that I wrote for Scientific American online.)