Showing posts sorted by date for query Kirsch. Sort by relevance Show all posts
Showing posts sorted by date for query Kirsch. Sort by relevance Show all posts

Monday, 24 January 2011

Is low dose citalopram as effective as other antidepressants?

From this post at the 21st Floor I found out about a study of citalopram dosing (Aursnes et al 2010) that is a sort of pre-print (apparently Webmedcentral is a kind of post-publication peer review model). It suggests that citalopram in lower doses (below 45mg per day) might not be as effective as comparator antidepressants. Citalopram is used in doses from 20-60mg for depression and an average dose is probably around 30mg so this is a significant claim*. It is actually titled "Are Regular Doses of Citalopram for Depression only Placebos? Meta-analysis And Meta-regression Analysis Of Pre-registration Clinical Trial Data" which is pretty tendentious since the study only includes two placebo controlled trials, both of which are in high doses of citalopram. Whether citalopram is or isn't a placebo at lower doses is not going to be answered by this study since, even if lower doses of citalopram are less effective than other antidepressants we don't know how they'd compare to placebo using this data.


The paper found a statistically significant relationship (p=.034) between mean citalopram dose and effect size (see Figure 1 below) and went on to split the data into doses above and below 45mg and found that doses below 45mg (but not above) had a statistically significantly smaller effect size than the control antidepressants.

Figure 1. Regression from Aursnes et al 2010

It is very unclear from the paper what methods they used - it appears that they used standardised mean difference (SMD) in change scores but it is very unclear how they computed the change scores from the data they had available. The data they show for most studies does not report change scores and you wouldn't normally be able to estimate the standard deviation of this measure (which you need for the meta-analysis) unless the study reported the correlation between baseline and final score (which is very unlikely).***

Naturally I wanted to have a look at the data and regular readers of this blog will know that baseline severity is an important predictor of antidepressant efficacy in trials and I was interested to see what effect that would have in this study. Some of the source data is conveniently provided, it was obtained from Danish medicine licensing applications for citalopram so likely to be less subject to publication bias (in a similar way to the Kirsch et al data). I was able to get final score data for the studies using the Hamilton Rating Scale for Depression (HRSD-17) as outcome but had to impute standard deviations for those using the Montgomery-Åsberg Depression Rating Scale (MADRS). I could obtain change score SMD data from their figures but I think this is a little unreliable***. I used all the studies with an active comparator (usually a tricyclic antidepressant**). All analyses were random effects unless otherwise noted, and everything was done in R.

So the first thing I wanted to do was reproduce the correlation between mean dose of citalopram and effect size. Using the data I extracted (see Figure 2 below) my meta-regression showed a slope of .038 of SMD against citalopram dose (p=.13). Then I used their numbers (SMD and standard error) and performed another meta-regression. This showed a regression slope of .037 (similar, but not the same as their .038) but it was not statistically significant (p=.20) even if I used a fixed effects regression (p=.19). This is surprising as the authors report statistical significance at p=.034! The only way I can get their data to give a statistically significant regression is to include the placebo studies so that the slope is now .041 with p=.033 but this would be completely wrong. Having placebo studies included only for high doses of 50mg and 60mg will overestimate the effect of citalopram at high doses since we expect citalopram to be better than placebo but no better or worse than another antidepressant as control.

Performing multiple regression with dose and baseline score against weighted mean difference in HRSD score and MADRS score separately**** showed slopes of .73 (p=.10) and -.047 (p=.76) for baseline severity and mean dose respectively for HRSD scores and 1.06 (p=.30) and .25 (p=.75) for MADRS scores. So no statistically significant effect of baseline severity. That isn't necessarily surprising because it may be that both citalopram and other antidepressants are all equally affected by increasing efficacy with increasing severity of depression.


Figure 2. The data I have extracted presented as weighted mean differences in RevMan 5.0

So, in summary, there does not seem to be a robust relationship between the effect size of citalopram and the mean dose used in trials - I do not think a discrepancy of this magnitude is likely to due to methodological differences and I hypothesise that Aursnes et al (2010) have made a calculation error. Therefore it does not make sense to divide the studies based on citalopram dose. If we combine all studies using the SMD as outcome (see Figure 3 below) there is not a statistically significant difference between the citalopram and active control arms of trials although the difference is borderline significant (p=.057) and includes differences up to -.46 which corresponds to scores on the HRSD up to 4 points*****. It is worth bearing in mind that not a single trial had a mean dose of citalopram less than 40mg so this analysis cannot tell us much about lower doses of 20-30mg of citalopram and, if the correlation with dose is not real, then it can tell us nothing about these lower doses. In some ways this resembles the study by Kirsch et al where the authors used a regression analysis to make claims about patients with low severity by extrapolating the line into a region where there were not actually any studies.******

Figure 3. All active control studies combined using SMD as outcome, to the left of 0.0 favours control and to the right favours citalopram


* No study had a mean dose of citalopram less than 40mg.
** Two amitriptyline, two clomipramine, three mianserin, two maprotiline, nortripyline, and imipramine. The paper seems to misreport this but looking at the original data sheets they get their numbers from I think my numbers are right and their report of five studies using mianserin is wrong.
*** This is what the authors say about their methods:
These data were fed into a tailor-made program for performing meta-analysis (Comprehensive Meta-analysis Version 2, from Biostat, Englewood, USA), which uses standard statistical procedures [9]. We filled in columns for the mean score, its standard deviation, and the number of patients in the citalopram group and in the comparator group. We added a column for correlation between before and after, Pearson’s coefficient of correlation found to be 0.48, and standardized the effect analysis with standard deviations of the differences between values before and after. We found I2 to be 73.5 and performed random effect meta-analyses with the effects weighted with the inverse of their variances.
And I think they've probably committed an error in their methodology because I'm not sure what they think that correlation is for - they can't use that (the between study correlation) to estimate the correlation between baseline and final scores within studies and then use that to calculate change score standard deviations. That would be completely wrong. It is also noteable that the authors do an 'intention to treat' analysis in addition to 'last observation carried forward' but this is a flawed approach for continuous measures if you don't have access to the original data and it inflates the apparent sample size and does not capture any useful information regarding drop-outs that 'intention to treat' does with dichotomous data.
**** They can't be combined together because the baseline scores are in HRSD or MADRS units respectively. 
***** You can see from Figure 2 that the studies using the HRSD did show a statistically significant decrease in the citalopram group.
****** If the regression slope were to be taken seriously an arbitrary 'clinically significant' difference of .35 (around 3 points on the HRSD) is reached around a dose of 40mg.

Thursday, 6 January 2011

Lamotrigine: an exciting new treatment for acute bipolar depression?

I've talked a fair bit about the efficacy of antidepressants (or lack thereof) in treating major depression. I won't go into that again, but I did want to discuss something that has been neglected in that debate.

Bipolar depression
It's estimated that some 10% of people with a major depressive episode have underlying bipolar disorder - that is they'll go on to have a manic or hypomanic episode (if they haven't had one already) - and if the efficacy of antidepressants in straight major depression is controversial then bipolar depression is a minefield.

It is certainly considered that antidepressant usage alone gives a big risk of provoking a swing from a depressive to a manic episode so they would usually be used alongside a 'mood stabiliser' like lithium or an anticonvulsant, and even then there is considerable disagreement as to how much they help.

Lamotrigine in acute bipolar depression
Something that has got a lot of attention in recent years is the anticonvulsant lamotrigine. There is good evidence for the efficacy of lamotrigine in preventing further depressive episodes in people with bipolar disorder. But recently there has been much interest in its use for treating an acute episode of depression in bipolar disorder, and this is despite the fact that it takes a considerable time to titrate up the dose to therapeutic levels (if you go by the BNF it takes 5 weeks to get to the usual dose of 200mg).

A major paper influencing people's thinking came out of Oxford by Geddes et al in 2009. This was a meta-analysis of trials of lamotrigine in acute bipolar depressive episodes and had a considerable impact. The Canadian Network for Mood and Anxiety Treatments (CANMAT) guidelines now recommend lamotrigine as a first-line treatment for acute bipolar depression largely on the strength of this analysis.

What they did was, apparently, contact GSK (who make lamotrigine as 'Lamictal') and get hold of all the individual patient level data from all five trials performed by the drug company and used it to perform a meta-analysis. They also identified two other studies not by GSK but didn't combine them because they didn't have any individual subject data and the trials were a crossover trial (which is difficult to correctly combine in a meta-analysis) and the other used lamotrigine as add-on to lithium therapy. They excluded data from one of the trials which had used a 50mg dose as this is generally considered subtherapeutic.

I'll concentrate on their findings using the Hamilton Rating Scale for Depression (HRSD, 17-item version) which is very widely used in antidepressant trials (I've mentioned it before) which has a maximum of 54 points and a score greater than 18 is usually needed to be recruited into a trial as 'moderately' depressed. Again, I'll focus on two measures of outcome, 'response rate' (the proportion of patients in each arm who achieve a 50% or more reduction in their initial HRSD score) and mean difference in the final HRSD score.
If we look at the mean difference in the final HRSD score (this was adjusted for baseline severity in a regression) it was –1.01 (–2.17 to 0.14). That mean difference is not statistically significant (although using a different measure, the Montgomery–Åsberg Depression Rating Scale, they did find a statistically significant effect), nor is it even suggesting a particularly large effect is possible (with the upper limit of the effect size around 2 points on the HRSD). This is smaller than the 1.8 point effect size reported by Irving Kirsch's meta-analysis of 'new' antidepressants in major depression and when I reanalysed Kirsch's data properly I (and others) found an effect size of 2.7 points on the HRSD, just short of NICE's (arbitrary) 3 point threshold for 'clinical significance'.

So a mean improvement of 1.0 points on the HRSD is not exactly impressive - certainly it wouldn't be very good if that was a uniform single point improvement across every patient. But potentially it could represent a really big, 'clinically significant' improvement for a subset of patients - and we'd be interested in a drug that could do that.

So this is why we look at 'response rates' - what proportion of patients got 'clinically significantly' better, or 'remission rates', the proportion who score sufficiently low to count as being better. Commonly the former is defined as a 50% reduction in the score on the HRSD (or other symptom scale). We can see below (Figure 1) that significantly more patients responded in the lamotrigine group than the placebo group with a risk ratio of 1.27 (95% CI: 1.09-1.47), that is 27% more patients in the lamotrigine group showed an improvement in HRSD of 50% or more - that implies 11 patients need to be treated for one additional 'response'.

Figure 1. Figure from Geddes et al - Meta-analysis of HRSD 50% response rates using individual patient data from GSK trials
As has been pointed out before, response rates in depression trials are slippery beasts. By calling a 50% reduction in HRSD score a 'response' to treatment you actually need to improve by less points on the HRSD if you are less depressed (and have a lower baseline score) so an improvement on exactly the same items of the HRSD could be classified as 'response' or 'no response' depending on that patient's initial severity. This also means that this measure is very vulnerable to small differences in baseline severity between the arms of a trial.*  In practice response rates based on thresholding continuous variables like the HRSD (such as a 50% improvement threshold, or a threshold of 8 points for 'remission') are vulnerable to artefactual non-linear effects where very small improvements tip a few people over the threshold (something to particularly worry about if the threshold used seems a bit arbitrary anyway as you can easily pick one after the fact that amplifies any effect).  

The response rates in this study are around 35% of patients in the placebo arm and 45% of those taking lamotrigine - so an increase of 10 percentage points due to lamotrigine. If we consider that, for the average patient, 50% improvement implies a minimum change of around 12 points on the HRSD the actual mean improvement of 1.0 points (and the standard deviation) seems a little small (in a back of the envelope calculation you could say that you would expect an average of 1.2 point improvement - thats 10% getting 12 points averaged over all the group). So a bit of a mixed bag I'd say.

Traditional meta-analysis versus individual patient data
A question that occurs to me is what extra information we gain from having the individual patient data? It is pretty rare to get hold of individual patient data when doing a meta-analysis, usually all we have is the overall results for each trial. Sometimes this doesn't make much difference, comparing the individual level meta-analysis by Fournier et al of antidepressants with Kirsch et al's meta-analysis did not reveal any major differences (and these studies looked at fairly different sets of antidepressants).

I had a quick go at performing a study level meta-analysis of the GSK lamotrigine data** and found that the response rate had a relative risk of 1.22 (1.06-1.41) (Figure 2 below) which is pretty similar to the individual level data 1.27 risk ratio.

Figure 2. Meta-analysis of HRSD 50% response rates from GSK per trial data
Looking at the mean difference in the change in HRSD scores I found an effect size of -.86 points (-1.84-.12) which is similar to the -1.0 effect using individual level data and is similarly also only borderline statistically significant (Figure 3 below).

Figure 3. Meta-analysis of mean difference in HRSD scores from GSK per trial data
This is pretty reassuring as it suggests the majority of the information is present in the trial data so we wouldn't lose too much or miss any important effect if we didn't have the individual patient data available. You have to ask why GSK didn't re-analyse the data themselves, this would have been easily done (it took me a few minutes). I think there are a few reasons for this, one intimated in the paper is that regulatory authorities do not accept the results of meta-analyses, they require two large positive trials for licensing a drug, and all but one of the GSK trials was negative so a meta-analysis would have done them no good with licensing. However, they could still have influenced the scientific literature or provided justification for a further larger trial and I wonder whether they didn't in fact do the same analysis as either me or Geddes et al and realise that this was probably a small and fairly dubious effect and reckon it wouldn't do them any favours in the long run.

Antidepressant response and the severity of depression
Like many previous studies of antidepressants Geddes et al also find that there is an interaction between the size of the antidepressant effect of lamotrigine and how severely depressed the patients in the study were at baseline. They found a significant interaction on their ANCOVA analysis between baseline HRSD severity and final HRSD score with a regression coefficient of .30 (p=.04). They go on to comment that:
"Thus, the interaction by severity was because of a higher response rate in the moderately ill placebo-treated group, rather than, for example, a higher response rate in the severely ill lamotrigine-treated group."
This statement is very redolent of Kirsch et al's claim:
"The relationship between initial severity and antidepressant efficacy is attributable to decreased responsiveness to placebo among very severely depressed patients, rather than to increased responsiveness to medication."
Kirsch et al got a lot of stick for saying this, not least from yours truly, so I was interested in what a set of rather more mainstream figures in the psychiatric world (two Oxford academic psychiatrists and a psychiatrist from the GSK advisory board) were trying to say, and how this relates the Kirsch's work. Maybe I was being unfair to Kirsch?

Kirsch et al argued that the apparent increase in antidepressant efficacy with increasing baseline severity in trials was due to decreasing placebo response in more severe trials (see Figure 4 below). They further argued that this increasing efficacy was therefore only 'apparent' because the response to antidepressant was the same. I've discussed before how, on many levels, it is meaningless to claim that this effect is only 'apparent'. 

Figure 4. Figure from Kirsch et al - Regression of baseline severity against standardised mean difference of HRSD score improvement for antidepressant and placebo groups using per trial data
However, when I re-analysed their data it was their finding of decreased placebo response that was actually only 'apparent', and was due to needlessly normalising the raw HRSD data using the standardised mean difference (see Figure 5 below) and in fact placebo responses remained fairly static with increasing severity while antidepressant responses increased.

Figure 5. Regression of baseline severity against HRSD score improvement for antidepressant and placebo groups using per trial data from Kirsch et al

Of course these correlations are only looking at the average baseline severity between each trial and doesn't tell us whether the relationship between baseline severity and HRSD improvement holds true within each trial for the individual patients in that trial. Fournier et al used individual level data to look at this relationship and found increases in both placebo and antidepressant responses, with the greater gradient in the latter leading to increased efficacy of antidepressants overall (see Figure 6 below) so there is actually not great evidence for a decreased placebo response with increasing baseline severity in straight trials of antidepressants in major depression.

Figure 6. Figure from Fournier et al - Regression of baseline severity against HRSD score improvement for antidepressant and placebo groups using individual patient data

So what did Geddes et al find? Well, unlike Kirsch et al and Fournier et al they primarily looked at response rates rather than mean change in HRSD, obviously you can't have an individual patient's response rate (an individual either does, or does not respond) so they divided the subjects into two groups, those with a baseline severity below 24 points, and those above. They found that only those in the 'severe' severity (>24 point) group had a statistically significant response rate greater than placebo with 46% of those on lamotrigine responding compared to 30% on placebo. In the 'moderate' (<= 24 point) group response rates were 48% for lamotrigine and 45% for placebo. Geddes and others have argued that this increased placebo response at moderate severity is due to something like inflation of baseline severity at trial recruitment (that is doctors subconsciously inflate severity for those around the lower threshold for trial recruitment and they then regress to the 'real' lower score they would have had anyway when assessed blindly as part of the trial while this doesn't happen for the more severe patients).***

Now I've outlined some of my reservations about 50% reduction in HRSD score as a measure of 'response rate' and I don't think that these numbers necessarily show what Geddes et al think it does. Let's consider a simple model of how antidepressants might work, let's say they can be modelled simply by saying that an antidepressant reduces the baseline HRSD score by X HRSD points which is the simple sum of the placebo effect (P) and a 'true' antidepressant effect (A):

X = P + A

Based on this model, as discussed above, we can see how patients in the 'moderate' severity group could be more likely to respond even if X is the same for both severity groups. This is because a lower baseline severity means less HRSD points need to be lost to reach a 50% reduction. If we take this observation it is conceivable that the lower placebo response in the 'severe' group could be, at least partly, due to pure artefact. Given that those treated with lamotrigine in the 'severe' group had a larger response rate than those on placebo you might then go on to posit that maybe P is larger for the more severely depressed patients.

The way we ideally would want to answer this question would be to look at the mean HRSD scores as we did above for the data from Kirsch et al and Fournier et al. The response rate figures alone would still be completely consistent with a relationship like that shown in Figure 5 above and suggest that the claim that "the interaction by severity was because of a higher response rate in the moderately ill placebo-treated group" is false.

Unfortunately Geddes et al don't show the mean HRSD data by baseline severity nor do they report the correlations between mean HRSD score and baseline severity for the lamotrigine and placebo groups separately, either of which might help to answer this question. This is a bit odd as this is one of the main areas where the individual patient data would prove very useful and answer questions that the trial only data cannot. So we'll never know whether they do or don't show that placebo responses are constant, decreased, or increased with greater severity of depression. I've reproduced the data presenting each trial below (Figures 7 & 8) but these can't really answer this question, for that we need the within trial data.

Figure 7. Simple regression of mean baseline severity against 50% response rates from GSK per trial data split by lamotrigine (blue) and placebo (red)

Figure 8. Simple regression of mean baseline severity against mean change in HRSD score from GSK per trial data split by lamotrigine (blue) and placebo (red)
Amusingly, if you consider Figure 8 in the same way as my Figure 5 and attempt to determine a severity 'threshold' above which the NICE 3 point 'clinical significance' criteria is reached then you find that you never actually reach it.


Summary
So, what should we conclude? Well a few things:
  • Lamotrigine is not very effective for acute bipolar depression
  • Individual patient data in these trials doesn't tell us much more than a classical meta-analysis
  • The effect of lamotrigine, like other antidepressants in major depression, increases with greater baseline severity of depression, but never reaches the NICE 'clinical significance' criteria
  • It is unclear exactly why antidepressant efficacy increases in this way but it is far from established that it is due to "higher response rate in the moderately ill placebo-treated [patients], rather than, for example, a higher response rate in the severely ill lamotrigine-treated [patients]"


So what next?
So what medication should be used for acute treatment of bipolar depression? Well I think the data from quetiapine is pretty promising and certainly a lot more convincing and impressive than for lamotrigine monotherapy. Perhaps lamotrigine will be synergistic when added to quetiapine and the first author, John Geddes, is heading up the CEQUEL trial looking at just this question. 




* I think a better measure might be a fixed improvement in HRSD score, call it a 'clinically significant response', this would avoid the assumption that antidepressants somehow cause an X% reduction in HRSD score rather than say improving Y number of symptoms, and thus the problem I've mentioned above about how mildly depressed patients need less improvement to 'respond' than more severely depressed. 

** I got the response rate data from the Geddes et al paper and the HRSD mean difference data from the GSK trials register (since it wasn't presented in the paper). I used mean change scores rather than final HRSD scores (although this shouldn't make much difference), I had to estimate standard deviations for trial SCA40910 and the means for trial SCA30924 were already adjusted for baseline severity. Data are for fixed effect models but random effects makes minimal difference to the results.

*** Interestingly this explanation would only be tenable if trials of greater severity showed the same within-trial effect as trials of lower severity, since this effect should take place at the recruitment threshold irrespective of what that threshold is. So therefore within a trial the less severe subjects around the recruitment threshold should show a greater placebo response whereas there would not necessarily be any relationship between average baseline severity and average response between trials. This means that my regression data from Figure 5 above doesn't necessarily argue against this model - what we need to know is whether this relationship holds within the trials, something that even the individual subject data from Figure 6 doesn't rule out because we'd ideally have data separately plotted for each trial since the recruitment thresholds for each trial may have differed.   

Friday, 26 March 2010

Treating depression in general medical patients

As a doctor with an interest in psychiatry currently working in general medicine the issue of depression in general medical patients is one that interests me. We commonly see overdoses secondary to depressive illness and depression in patients with terminal diagnoses but also in many other conditions, particularly chronic disease. While we have access to specialist psychiatric or palliative care services for the former conditions that still leaves a substantial number of depressed patients to care for, and that is something of a treatment dilemma. 

Physical illness is strongly associated with depression and some 10-20% of general medical inpatients or outpatients have a depressive disorder. This is particularly marked for people with chronic disease where rates range from 11% of diabetics to 20% of people after a heart attack.  Depression is a a risk factor for poor prognosis in physical disease, being associated with worse mortality, at least partly mediated via decreased adherence to treatment.  Yet evidence shows that physically ill patients receive less antidepressant prescriptions than other depressed patients.

There are some specific challenges in recognising and treating depression in general medicine, early in an admission somatic symptoms of depression can difficult to distinguish from symptoms of physical illness and later on during treatment low mood can be considered 'understandable', with a natural resistance on the part of clinicians to medicalise normal emotional reactions. The inpatient environment is also unusual and stressful and it is unclear whether patients will maintain a low mood or improve when discharged home.  Practically, the onset of antidepressants is generally believed to be delayed over two weeks which means that any effect may not be seen during an acute admission and psychological therapies such as CBT are just not available in general medicine.

In recent years the risks of self-harm and discontinuation syndromes with antidepressants have received significant coverage and since Irving Kirsch's 2008 paper much doubt has been raised about overall antidepressant efficacy in any other than the most severe patients.  A recent Cochrane Review has addressed the question of antidepressant usage for depression in physically ill patients:

Rayner et al 'Antidepressants for depression in physically ill patients' Cochrane Database of Systematic Reviews 2010, Issue 3.


They looked at studies of depression quite broadly defined (major depressive disorder, adjustment disorder, dysthymia) and found 51 studies (mostly in SSRIs but also in tricyclics and a few less common antidepressants), with fluoxetine (Prozac) the drug most commonly studied (12 trials).  The physical diseases studied included stroke (11 studies), HIV (7), Parkinson's disease (6), cancer (4), COPD (chronic bronchitis and emphysema; 3), diabetes (3) , heart attacks (2), and renal failure (2). At the two follow-up periods of most interest (6-8 wks and 9-18 wks) there were around 1,000-1,500 subjects included in the analysis.

Overall they found that antidepressants were similarly effective at all follow-up durations (ranging from 4 to greater than 18 weeks) as seen in the summary figures on the right. We can see an odds ratio of around 2, that is antidepressants roughly double the chance of a 50% improvement in outcome score (most studies used the Hamilton Rating Scale for Depression) or showed a standardised mean difference of around 0.5*

Looking at other aspects they found that there were more people dropped out of the study from the antidepressant group than the placebo group (this was marginally significant) with an odds-ratio of 1.3 (95% confidence interval 1.0-1.8). Looking at side-effects, dry mouth and sexual dysfunction were both significantly more likely to be reported by those in the antidepressant group, the latter being primarily driven by those taking SSRIs.  So overall antidepressants had side-effects sufficiently bad to lead more people to drop out of the study.

The study didn't find any striking differences between the efficacy of SSRIs and other antidepressants, nor differences between taking a narrow (major depressive disorder only) or broad definition of depression.

Looking at the I-squared statistic for trial hetrogeneity we can see that for dichotomous outcomes differences between trials were not very large but for the mean difference outcomes there was very large heterogeneity. However, this seems to be due to two studies with stonkingly big effect sizes (improvements greater than 10 points on the HRSD) and excluding these from analyses drops the I-squared right down without massively affecting the results.

Overall this is quite an interesting finding and it suggests that antidepressants can be really very effective for depression in physically ill patients. But there are some limitations to bear in mind:
  • Most studies were pretty small, almost all with less than 100 subjects and we know that small studies are more likely to overestimate the size of the beneficial effect
  • Trial quality was actually pretty low, and low quality trials are known to overestimate effect sizes (more on this below)
  •  Publication bias was apparent in the studies (more below)
  • The effect of baseline severity has become a hot topic since Kirsch et al and this study didn't look at this (more below)
  • They looked only at the 10 most common side-effects but not overall adverse event rates, or specifically serious adverse events, and this prevents detection of less common but serious complications (stuff like death or suicide)
  • No subgroup analyses were performed to see if antidepressants were more effective in specific physical illnesses (say in stroke rather than HIV)
  • They did not look at studies with co-morbid psychiatric illness, this is important because mixed disorders, particularly with aspects of both depression and anxiety, are very common
Looking at a funnel plot from the study we can see apparent publication bias (see right), the gap at the bottom left of the pyramid represents missing small trials (or rather, trials with a large standard error) with a large negative effect of antidepressants. This suggests that some negative trials (which we would have predicted would exist based on the effect size we are finding) are missing from the literature included in the review. This is an example of how small positive trials are much more likely to get published than small negative trials which disappear into the file drawer.  Publication bias is a known problem in antidepressant trials. When Turner et al analysed data submitted to the FDA before approval** they found that 37/38 positive trials were published but only 14/36 negative trials were published, and 11 of these actually claimed a positive result!

Trial quality was disturbingly low, the authors used the 'Risk of Bias' table from the Cochrane Handbook to score as 'low risk', 'unclear risk', or 'high risk' of bias on six items:
  • Sequence generation
  • Allocation concealment
  • Blinding
  • Incomplete outcome data
  • Selective outcome reporting
  • Other issues
Only three studies scored as 'low risk' of bias on four or more items*** and only something like 13 on three or more items. The authors find that by looking only at these 13 odd studies the effect size is not grossly different to looking at all the studies.  If we just look at the three best quality studies (see right, data from 9-18wks) there is a large effect that is not statistically significant for the mean difference in HRSD scores (but it is significant looking at SMD) that suggests that it isn't purely low quality trials driving the beneficial effect of antidepressants seen in this study.

Not looking at the effect of baseline severity in the wake of Kirsch et al and its widespread impact is curious. Kirsch et al, looking at the same FDA data as Turner et al, found that the NICE threshold for 'clinical significance' (an improvement of 3 points on the HRSD or 0.5 SMD) was reached around a baseline severity (as measured by the HRSD) of 26 points, which is classified as 'very severe' by NICE and the American Psychiatric Association (see right). Similar results were found by Fournier et al looking at individual subject level data.

I made a back of the envelope attempt to plot the meta-analysis data against baseline severity**** and we find that the NICE threshold is actually reached at quite low baseline severity (18.5-21.5) which falls in the upper range of moderate through to severe severity.
In summary, studies of antidepressant use in physical illness indicate a large effect size that is 'clinically significant' in the 'severe' depression range, and there is a disparity between the large effects sizes in this review and in other studies of depression.  Although I have some criticisms of Kirsch et al it seems most likely that this disparity is due to publication bias in the Cochrane meta-analysis.  There are some interesting issues regarding the way that studies in general depression usually have a more severe major depression population and any extrapolation to less severe patients is on the basis of few studies whereas the Cochrane review includes a number of less severe conditions and it is possible that this makes it therefore more sensitive to beneficial effects at the lesser degrees of severity.  It is also possible that physically ill patients may be more responsive to antidepressants but I'm unconvinced.

This study looked at largely outpatient populations with chronic illness and it isn't clear whether the results are directly applicable to inpatient populations but it certainly supports the use of antidepressants in inpatient depression and suggests that at the very least they are likely to be as effective in this population as in the general population of depressed patients.

Finally it is worth noting that NICE has a guideline on treating depression in chronic physical illness which makes recommendations which are broadly similar to those they make for depression in general:
  • For low persistent subthreshold depressive symptoms or mild to moderate depression:
    • Low intensity psychosocial intervention (e.g. computerised CBT etc.)
    • If symptoms persist, previous severe depression, or symptoms compromising care consider either:
      • SSRI (citalopram or sertraline first line)
      • High intensity psychosocial intervention (e.g. individual CBT etc.)
  • Severe depression
    • Antidepressant and individual CBT
  • Be aware of drug interactions


* Standardised mean difference is the difference between the mean outcome scores for the antidepressant and placebo groups divided by the standard deviation, this corresponds to something like a difference of 4 points on the HRSD. Since many studies don't report dichotomous 'improvement' measures, or use different definitions, and these have to be 'imputed' using the mean difference data (making assumptions about how the data is distributed),  I prefer mean difference data, ideally using the raw HRSD figures rather than the standardised mean difference (since this can create odd distortions in the data, e.g. in Kirsch et al's study).  Almost all studies use the original 17-item HRSD but the few studies that instead use, say, the Montgomery-Åsberg Depression Rating Scale means that the authors have used the SMD so that this data can be combined (the SMD is supposed to allow you to combine data from different scales that are intended to measure the same thing). 
** This data should be free from publication bias because the FDA legally mandates the pharmaceutical companies to supply all studies performed on the drug.
*** Cochrane actually discourage adding these up to produce a scale but I can't think of a better way to see which studies are more or less biased.
**** Only including those studies with HRSD data, and those trials where I could access the article and extract it.  Since I didn't try too hard to check everything it is quite possible some scoring from scales other than the 17-item HRSD crept in there. 


UPDATE
In response to neuroskeptic in the comments, here's the baseline severity data split by antidepressant and placebo groups (as seen in Kirsch et al's analysis), the regression is weighted by sample size, the baseline severity is mean HRSD score, the improvement is mean baseline severity minus mean HRSD score at 6-8 weeks. We can see that increasing baseline severity leads to increasing response to antidepressant with placebo response fairly flat. This was pretty much what we found when we looked at HRSD outcome data from the Kirsch et al study.

Sunday, 31 January 2010

The drugs do work?

An interesting article in the NY Times by psychiatrist Richard Friedman:
Last week, The Journal of the American Medical Association published a study questioning the effectiveness of antidepressant drugs. The drugs are useful in cases of severe depression, it said. But for most patients, those with mild to moderate cases, the most commonly used antidepressants are generally no better than a placebo.
...the authors of the new analysis gave themselves an additional handicap: they decided to exclude a whole class of studies, those that tried to correct for the so-called placebo response.
...
Another drawback of the study is that its conclusions are based on studies that included only two antidepressants — when there are 25 or so on the market. By contrast, when the Food and Drug Administration wanted to investigate the safety of antidepressants, it analyzed data from some 300 clinical trials, with nearly 80,000 patients, involving about a dozen antidepressants.
...
Every once in a while, a landmark study comes along and overturns everyone’s cherished ideas about a particular treatment. But the current study is not one of them. So it would be a shame if it discouraged depressed patients from taking antidepressants.
Neuroskeptic has blogged about this study before (Fournier et al 2010 JAMA 303(1)), but it is worth noting that, despite my criticisms of Irving Kirsch's meta-analysis of the FDA data on antidepressant efficacy, even when I reanalysed the data I found that the NICE threshold for 'clinical significance' was met at around a baseline severity of 26 points on the Hamilton scale. In this study by Fournier et al the threshold was met around a baseline severity of 25 points.

My analysis could only correct for baseline severity on a per trial basis whereas the above study was a patient level meta-analysis which is a better approach when available. So while the above study did include only two antidepressants (one of which was the older tricyclic class, although these are thought to be similarly effective to the newer SSRIs, just with more side-effects) it is consistent with the study by Kirsch et al, even given the criticisms of it I've previously raised.

So I don't think Friedman's criticism holds up, I think a more sensible attack is that the NICE threshold is entirely arbitrary, an argument I made at the time of Kirsch's paper.

Saturday, 7 March 2009

Are the new antidepressants any better?

You may recall that the Kirsch et al study showed a trend (that was not statistically significant) for venlafaxine to be superior to other antidepressants. But there are a number of newer antidepressants, how do they compare to the vanilla SSRIs?

A new meta-analysis in the Lancet tried to find out:
"Mirtazapine, escitalopram, venlafaxine, and sertraline were significantly more efficacious than duloxetine (odds ratios [OR] 1·39, 1·33, 1·30 and 1·27, respectively), fluoxetine (1·37, 1·32, 1·28, and 1·25, respectively), fluvoxamine (1·41, 1·35, 1·30, and 1·27, respectively), paroxetine (1·35, 1·30, 1·27, and 1·22, respectively), and reboxetine (2·03, 1·95, 1·89, and 1·85, respectively). Reboxetine was significantly less efficacious than all the other antidepressants tested. Escitalopram and sertraline showed the best profile of acceptability, leading to significantly fewer discontinuations than did duloxetine, fluvoxamine, paroxetine, reboxetine, and venlafaxine."
I haven't read it yet, so I can't speak for its veracity.

Tuesday, 23 September 2008

Yet more antidepressants

Thanks to paul in the comments below here's a link to the Maudsley debate on antidepressant efficacy, 'This House Believes Antidepressants are no Better than Placebo', featuring Irving Kirsch, Joanna Moncrief (for), Guy Goodwin, and Lewis Wolpert (against).

Interestingly Kirsch turns the question around to argue that there is no evidence that antidepressants are clinically significantly better than placebo (as you may be aware, I've addressed the Kirsch et al PLoS paper before).

Both Kirsch and Moncrief make some good points, but I was struck by Moncrief's claim that because we don't know that antidepressants act specifically against 'the' biological cause of depression, and in fact may have rather more non-specific effects that help to ameliorate the symptoms of depression, we therefore should not use them.* She goes on to argue a completely contradictory point at the end of her speech to claim that because antidepressants are (according to her) no better than placebo they are therefore harmful because they are not inert. Yet her argument that antidepressants are not superior to placebo hinges on the claim that their apparent superority hinges on their active side effects in clinical trials.

Goodwin and Wolpert make some fairly pedestrian counterarguments, some fallacious, largely anecdotal in Wolpert's case. The comments from the floor included some 'service users' ranting and anti-psychiatry, which is quite common at psychiatric talks.


* I've long been drawn to the idea (can't remember who first proposed it) that serotonergic antidepressants flatten emotional responses and noradrenergic antidepressants are activating. But both these effects would seem to me to be very useful in helping to relieve symptoms of depression and facilitate true recovery. Comments from the floor point out that in the rest of medicine we don't abandon treatments proven to work in clinical trials because we don't know their mechanism of action. In fact, it is pretty hard to see what Moncrief would advocate to treat depression instead of antidepressants, surely we can't know that physical exercise or CBT are definitely treating the underlying physical abnormality?

Sunday, 4 May 2008

Kirsch et al reply again

Huedo-Medina, Johnson, and Kirsch have submitted a further response on PLOS Medicine in reply to further comments by myself and others, after their last reply.*


Placebo response and severity

Interestingly, while I observed that they needed to:

"clarify their position on the claim that placebo response decreases with increasing baseline severity, since this appears to be an artefact"
Rather than address this observation they instead repeat the claim, saying that it is the 'unique' contribution of their article:

"without the within-group analyses it would not have been possible to conclude that placebo responses were lessening as initial severity of depression increased (whereas drug response remained constant; see our article’s Figures 2 and 3). This unique contribution of our article contradicts Wohlfarth’s conclusion that it contained “nothing new.”"

Flawed meta-analytic methods

Further, they defend their bizarre and biased analytical method:

"One of the main concerns in the new commentaries centred on one of our main analyses, which evaluated change for drug and placebo groups without taking a direct difference between them. Thus, effect sizes were calculated separately for each group for this analysis, though the analysis combined them. Leonard regarded this practice as “unorthodox” and Wohlfarth regarded it as “erroneous because the effect size in an RCT is defined as the difference between the effect of active compound and placebo.” First, these concerns ignore the fact that our article’s between-group analyses confirmed the major trends present in the analyses that considered within-group change. Specifically, both sets of analyses concluded that antidepressants’ efficacy was greater at higher initial severity, attaining clinical significance standards only for samples with extremely severe initial depression. Second, although the commentators may be correct that our within-group analyses are relatively innovative in this literature, it does not mean that they were wrong. To the contrary, these statistics are in conventional usage elsewhere (e.g., 3, 4, 5), as Waldman’s commentary implies...Finally, the analyses did incorporate a direct contrast between drug and placebo (see Table 2, and Model 2c, for example).

...

Although alternative weighting strategies may yield somewhat different results, the choices converge well both for the overall mean difference and for analyses of the trends across the literature. As an example, Leonard (04 March 2008) reported replicating our meta-regression patterns using alternative precision weights.

Importantly, as our article documented (Figures 2 & 3), the size of the difference between drug and placebo grows as the samples’ initial severity increases to extremely severe depression (but is very small at lower observed levels of initial severity). Because the overall differences between drug and placebo depended on initial severity, it is misleading to consider the overall difference in isolation."

But this does not address my objection that:

"This is not an acceptable analytic technique because it ignores that there is a relationship between the improvement in placebo and drug groups from the same study, but that the placebo and drug groups from any given study can have grossly different weightings when considered separately (e.g. there would be half as much weighting to the results from the fluoxetine trials in the drug analysis as the placebo analysis, the result of, for example, different sample sizes between the experimental arms)."
And as I say about the more conventional analysis they claim supports their 'unorthodox' analysis:


"I note that Figure 4 in the paper of Kirsch et al is actually more consistent with my finding of 'clinical significance' at a baseline of 26 (this threshold is found both by regression on the difference scores, or separate regressions for each group's change score) than their suggestion of 28 points..."

And Robert notes:

"The available unbiased estimate of the overall average benefit of NDA’s is equal to 2.65 HRSD units, which is considerably higher than Kirsch et al’s biased estimate [of 1.8]."

So while my and Robert's analyses confirm that the effect size of antidepressants increases with increasing baseline severity, they also show that their claim that placebo responses decrease with baseline severity of depression are false, and that Kirsch et al report effect sizes that are considerably biased downwards.

It is worth thinking about the references they give to support their analytical method (numbers 3,4, & 5, notice they are either in psychology or education journals), the most recent (and thus most easily available) is reference 5, Morris & DeShon (2002) 'Combining Effect Size Estimates in Meta-Analysis With Repeated Measures and Independent-Groups Designs' in Psychological Methods 7(1) 105-25. It is about combining the results from repeated measures designs and independent group designs, concentrating on training effectiveness, organizational development, and psychotherapy, and is not an article about medical meta-analysis:

"The issue of combining effect sizes across different research designs is particularly important when the primary research literature consists of a mixture of independent-groups and repeated measures designs. For example, consider two researchers attempting to determine whether the same training program results in improved outcomes (e.g., smoking cessation, job performance, academic achievement). One researcher may choose to use an independent-groups design, in which one group receives the training and the other group serves as a control. The difference between the groups on the outcome measure is used as an estimate of the treatment effect. The other researcher may choose to use a single-group pretest-posttest design, in which each individual is measured before and after treatment has occurred, allowing each individual to be used as his or her own control.1 In this design, the difference between the individuals’ scores before and after the treatment is used as an estimate of the treatment effect."

I hope you can already see why this situation is not comparable to a meta-analysis of double blind randomised placebo controlled drug trials because these repeated measures designs would not be appropriate (you can have within-subjects cross-over designs but that is not what is being discussed here) because we know placebo effects are very important in drug trials so we require the use of placebo control arms. Therefore the Kirsch et al meta-analysis only involved independent groups and there is no need to worry about combining repeated measures and independent groups, and, as Morris & DeShon say:

"When the research base consists entirely of independent-groups designs, the calculation of effect sizes is straightforward and has been described in virtually every treatment of meta-analysis"

That is there is no need to use this unusual method because perfectly good methods already exist for analysing this data.

So Morris & DeShon are concerned with what to do when you have no control group for some of your studies - which is not the case in double blind RCTs because a study without a control group is considered an invalid measure of drug effects. The other two references, number 4, Gibbons et al (1993), and number 3, Becker (1988), also emphasise this aspect of the method (I haven't read these studies):

"With this approach, data from studies using different designs may be compared directly and studies without control groups do not need to be omitted."

But we have no need to do this, so we have no need for the analytical method used by Kirsch et al, and we have no need for this method precisely because medical meta-analysis consider only double blind RCTs and explicitly rejects studies without control groups precisely because an estimate of placebo responses in each trial is considered essential.

So we have no reason to use the method of Kirsch et al, but what reasons do we have for not using these "innovative...statistics...in conventional usage elsewhere"? Well what do Morris & DeShon have to say? Well obviously they're concerned about when it is acceptable to combine studies with and without controls, and conclude that ideally, if you intend to do this, there oughtn't to be a change in the control group with time, i.e. there should be no placebo effect in the control group. But they do refer to Becker for a meta-analytic method proposed for use when there is a placebo effect:

"Becker (1988) described two methods that can be used to integrate results from single-group pretest-posttest designs with those from independent-groups pretest-posttest designs. In both cases, meta-analytic procedures are used to estimate the bias due to a time effect. The methods differ in whether the correction for the bias is performed on the aggregate results or separately for each individual effect size. The two methods are briefly outlined below, but interested readers should refer to Becker (1988) for a more thorough treatment of the issues.

An important assumption of this method is that the source of bias (i.e., the time effect) is constant across studies. This assumption should be tested as part of the initial meta-analysis used to estimate the pretest-posttest change in the control group. If effect sizes are heterogeneous, the investigator should explore potential moderators, and if found, separate time effects could be estimated for subsets of studies." [my emphasis]

That is, if you can't be sure that the placebo effect is constant across studies, you shouldn't combine studies using this method. And, of course, this is precisely the objection that I and others have raised to this method - because we already know that placebo responses can vary between trials - that is why we have placebo control arms in randomised controlled trials!

So Huedo-Medina, Johnson, and Kirsch are advocating the rejection of the usual meta-analytic techniques used in medical research where the highest standards are required and control groups considered very important, in favour of adopting a methods from psychology and education that is only used when two different designs, one of which is rejected in medical research, need to be combined, and where placebo effects are downplayed, a method that even its advocates recognise is unsuitable with heterogeneous placebo responses between studies.

This is quite some defence when you look at the scatter on the placebo responses in the Kirsch et al meta-analysis, that's about as heterogeneous as it gets, and it isn't explained by baseline severity of depression - so the assumptions underlying the meta-analytic method used by Kirsch et al are violated, even according to the citations they refer to in justifying their approach! Even the original study (with standardised mean differences rather than raw change scores) showed great heterogeneity in the placebo arm:

"The amounts of change for...placebo groups varied widely around their respective means, Q(34)s = ... 74.59, p-values [less than] 0.05, and I2s = ... 54.47"

Precision versus bias

Huedo-Medina et al also completely misunderstand the objections of Robert Waldmann and others by saying:

"Waldman argued that our estimates of the overall difference between drug and placebo was conservatively biased (i.e., too small) because of assumptions present in our estimates of precision for each effect size. It is of course not possible to be certain that one has completely removed error from any measurement, or for that matter, to do so in an analysis of measures from independent trials. As Young noted, there are uncontrolled measurement errors or artefacts that necessitate the use of a control group and the randomised controlled trial design.
...
The calculation of a weighted effect size by using the inverse of each within-subjects variance is more precise than a sample-size weighted average (9), contrary to the Waldman’s assertion."
When, of course, his assertion is that their estimates may be precise but are biased:

"In each case, Kirsch chose a method which, under strong assumptions, gives an efficient and unbiased estimate of the true overall average benefit. In each case there are alternative approaches which are less efficient under those assumptions but which are unbiased not only when the Kirsch et al estimates are unbiased, but also for many cases in which the Kirsch et al estimates are biased. That is they are less efficient under the null but more robust. In each case the null hypothesis that the Kirsch et al estimator is unbiased has been tested and overwhelmingly rejected. The available unbiased estimate of the overall average benefit of NDA’s is equal to 2.65 HRSD units, which is considerably higher than Kirsch et al’s biased estimate."

* UPDATE
PJ Leonard replies to Huedo-Medina et al on PLoS Medicine.

Monday, 7 April 2008

More or Less Depression?

Thanks to lemmuslemmus (and also via the badscience forums) here's a BBC Radio 4 episode of More or Less talking about and to Kirsch et al. Nothing new here to be honest, More or Less was better under Andrew Dilnot.

Sunday, 30 March 2008

Kirsch et al update

Following on from the reply by Johnson et al, my re-analyses and Robert Waldmann's work of the last few weeks there's been a nice discussion of the paper by Nick Barrowman on Log base 2 which refers to both my, and Robert's reservations.

Robert has submitted a response to PLoS and develops his argument and analysis in this post. Amongst the other responses to PLoS biostatistician Jim Young develops a similar line of argument:

"...To see what might be regression to the mean with increasing initial disease severity, one would need to plot raw improvement against initial disease severity.

Most of the authors’ analyses seem to represent each trial as two group means, not as a single difference between groups. It is not possible to estimate the pure effect of treatment and of placebo within a single trial of this sort. Each effect is confounded with other effects – such as regression to mean and spontaneous improvement. If it were possible to measure separate effects of treatment and placebo within each trial, then there would be no need for the placebo group at all. It would be much more efficient to run trials with only a treatment group [1].

The authors’ write “Finally…we calculated the difference between the change for the drug group minus the change for the placebo group, leaving the difference in raw units and deriving its analytic weight from its standard error.” This is the only analysis that makes any sense because it is an analysis of the difference between groups in each trial. I think this corresponds to Models 3a and 3b in Table 2 but, as PJ Leonard notes, Model 3b is best because it’s sensible to drop out the one study where patients had only moderate depression.

The authors first conclude that “Drug–placebo differences in antidepressant efficacy increase as a function of baseline severity”. I have no problem with this – that’s what Figure 4 shows, at least in patients with more than moderate depression. But the authors go on to conclude that “The relationship between initial severity and antidepressant efficacy is attributable to decreased responsiveness to placebo among very severely depressed patients, rather than to increased responsiveness to medication.” This second conclusion requires the assumption that other effects (such as regression to the mean and spontaneous improvement or deterioration) are the same in each trial. Because of measurement error, it’s logical to expect regression to the mean to be greater in trials recruiting patients with more severe disease or greater in trials with longer follow up. Likewise it’s logical to expect spontaneous improvement or deterioration to differ with length of follow up. Even if the authors are happy to make the assumption that these other effects are the same in each trial, I think they should have made this assumption explicit, because I would not want to assume this myself.

Even if willing to make this assumption, why would you base this second conclusion on Figure 3? Why plot the standardised difference (between initial and final measurements in each group); why not just plot the raw difference? In meta-analysis, the only reason for standardising is to convert measurements on different scales into a common metric so that one can compare them. But if measurements are already on the same scale in each trial, why standardise them? It’s more difficult to interpret and requires stronger assumptions (“that variation between standard deviations reflects only differences in measurement scales and not differences in the reliability of outcome measures or variability among trial populations” [2]). Figure 3 may mislead because standard deviations will vary from trial to trial (these’s an order of magnitude difference in sample size between the smallest and largest trials). PJ Leonard says that if you use the raw differences, the “placebo response does not in fact decrease with increasing baseline severity”. If he’s right, then the authors’ second conclusion just looks like wishful thinking."

Saturday, 15 March 2008

Kirsch et al reply

Blair Johnson and the other authors of the Kirsch et al paper in PLoS Medicine reply to the responses to their paper here. Some relevant parts for the discussions here:

"...as we reported in our results, this difference was more apparent than real, disappearing when we controlled for baseline severity. It is worth noting that Turner et al. (2008) found between-group effect size (d) estimates of 0.40 for venlafaxine and 0.26 for nefazodone, both of which are close to the mean of 0.40 for all 12 newer antidepressants and are identical to those for fluoxetine (0.26) and paroxetine (0.42)."

"Leonard took the trouble of re-analyzing the data from our Table 1 and concluded that a clinically significant difference emerged at a lower point of severity than we concluded in our article (i.e., 26 vs. 28). We are grateful that his work confirms our major conclusion, which is that the efficacy of anti-depressants depends on the initial severity of depression. Unfortunately, however, his estimates of the standard deviation underlying each effect size relied on between-subjects’ rather than within-subjects’ formulations. In examining improvement in response to drug or placebo, individual trials conventionally control for the correlation between the HRSD scores at baseline. We adopted this convention in our analyses of drug and placebo improvement. Reassuringly, the analyses at the end of our Results section pertaining to each trial’s drug vs. placebo comparison also used a between-subjects variance formulation and confirmed that clinical significance emerges in the vicinity of an HRSD score of 28."

"We found a nonsignificant benefit of drug compared to placebo for moderately depressed patients. Yet, consistent with our other conclusions, the difference between drug and placebo grows at higher levels of depression. Davies commented on the fact that there were few samples with scores below the category of very severe depression on the Hamilton Rating Scale of Depression (HRSD), a limitation that our Discussion mentioned. "

I note that they don't engage with the finding by 'Leonard' (that's me that is) that there is no real decrease in placebo response with increasing severity, nor do they address my concerns that their use of the measure 'd' (mean change divided by SD of the change) biases the effect size (expressed in HRSD change scores), nor that looking at raw HRSD changes suggests that paroxetine and venlafaxine exceed the NICE 'clinical significance' criteria. I'm not quite sure what they mean by referring to within-subjects variance versus between-subjects variance (since I've changed the analysis based on Robert Waldmann's findings I don't know which analysis they looked at), they could be referring to normalising to the change score SD, which makes little difference compared to my previous analyses, or to analysing the drug group and placebo groups separately, which is just plain statistically wrong (and seems to be what they did, note that my analysis of separate regression lines produces the same results as looking at the between-subjects regression). They refer to their analyses at the end of their results section as confirming their 'within-subjects' results, I wonder if they mean their Figure 4 (repeated here), you might want to compare that to my regression (and their Figure 2) - and decide for yourself whether that confirms that the threshold for 'clinical significance' of 3 HRSD points difference is at baseline HRSD of 28 points as they claim, or 26 as I find.

They also don't really seem sufficiently contrite over their claim that in 'moderate' depression antidepressants should be avoided, given that it was based on a single study plus extrapolating a regression line.
The only real finding, that is robust, is that the difference between placebo and antidepressant response seems to increase with baseline HRSD severity. Although Kirsch et al emphasise that the level at which this difference becomes 'clinically significant' is in severe depression, it is worth noting that in fact the level at which it is significant (around 26 according to both my and their analysis of raw HRSD figures) is pretty much the middle of the pack in terms of the baseline severity of the studies (which were pretty much all in the 'very severe' range over 23 HRSD points - see that figure). [Their finding that the differences between the drugs may be largely explained by the differing baseline of the studies is not unreasonable].

UPDATE
'PJ Leonard' has submitted a response, titled 'Analytical differences', pretty much repeating what I said above:
"It is good of Johnson et al to reply to the responses here. However, I do not think they have sufficiently dealt with some of the reservations concerning their paper.

In particular, I do not think that they have engaged with my finding that using the raw HRSD change scores reveals that the placebo response does not in fact decrease with increasing baseline severity on the HRSD.

I am not clear exactly what they mean when they say that I have used between-subjects analyses to suggest that the effect size (when analysing the raw HRSD change scores) is larger than presented in their paper, whereas they have used within-subjects analyses.

My analyses utilise conventional methods for meta-analysis where the effect size in each study is analysed directly, whereas it seems likely that the low estimated effect size in HRSD units in this study is the result of carrying out the meta-analytic weighting on the drug and placebo groups separately (a 'within subjects' analysis?), and then comparing the effect sizes thus obtained (which would explain the lack of forest plots in the paper).

This is not an acceptable analytic technique because it ignores that there is a relationship between the improvement in placebo and drug groups from the same study, but that the placebo and drug groups from any given study can have grossly different weightings when considered separately (e.g. there would be half as much weighting to the results from the fluoxetine trials in the drug analysis as the placebo analysis, the result of, for example, different sample sizes between the experimental arms).

Normalising the HRSD change to the change standard deviation in each group separately is also unnaceptable because a larger change in HRSD score in the drug group could be associated with a greater variance, although this does not appear to be the case in this study.

Robert Waldmann estimates that there is more bias in analytical method in this paper than publication bias present in the data itself:

http://rjwaldmann.blogspot.com/2008/03/just-cant-let-it-go.html

I note that Figure 4 in the paper of Kirsch et al is actually more consistent with my finding of 'clinical significance' at a baseline of 26 (this threshold is found both by regression on the difference scores, or separate regressions for each group's change score) than their suggestion of 28 points, this difference is undoubtedly because this figure looks at raw HRSD scores, as did my analyses, and because the NICE 'clinical significance' threshold of d > .5 is actually stricter than the NICE threshold of an HRSD difference > 3.

I concur that there is a relationship between baseline HRSD severity and effect size but it is worth noting that almost all studies examined had baselines over 23 points (and were thus in APA/NICE categories of 'very severe' depression) so the threshold of 26 points is a fairly average baseline severity for the studies analysed in this paper (as can be seen from my regression plots or their Figure 4). Any generalisation to less severe categories of depression is unwarranted given that it would depend on extrapolating the regression line to a region with only a single study."

Tuesday, 11 March 2008

Statistics and depression

Robert Waldmann has some statistical thoughts on the Kirsch et al meta-analysis of anti-depressants:

Just can't let it go. SSRI Meta-analysis meta-addiction

Caveat lector

Caveat Lector II

The Simplest Meta-Analysis Problem

That Hideous Strength

Prozac Fan Talks Back

Personal chat with pj

In particular, in response to my confusion about where they got an effect size of 1.8 from:

"Actually I think I understand how Kirsch et al got their results. I get a weighted average difference of change of 1.

notation: changeij is the average change in HRSD of patients in study i who got the SSRI (j=1) or the placebo (j=0). Nij is the sample size of patients in trial i who get j pills.

In one calculation I used changeij/dij as the standard deviation of the change and thus (changeij/dij)^2/Nij as the estimated variance of the average change. The *separately* for drug and placebo data, I calculated the precision weighted average over i of changeij. This gave me an average change of 7.809 for the placebo and 9.592 for the SSRI treated for a difference of 1.78.

I guess this is what they did. I the confidence intervals are screwy and d is as described in the paper."

He also finds that it looks like Kirsch et al did indeed divide the change score by the standard deviation of the change score to obtain their 'd' measure - the wide confidence intervals that I thought argued against this seem to be a particular idiosyncrasy of their study. Fortunately this makes very little difference to my previous analyses (I've updated the effect sizes in this study, but the only real impact is on calculating proper SMD effect measures where these are larger, because the estimated SD is smaller).

I'll repeat one of my comments here:

"Hmm, that would be very annoying if someone had based their analyses on the confidence intervals being, you know, normal confidence intervals.

Looking back at the data it seems you're right that it is by essentially carrying out the meta-analysis on two entirely separate populations, the drug changes, and the placebo changes, and then subtracting one from the other, that they get their very low estimate of HRSD change.

That is a very odd way of doing things indeed, it is basically assuming that each study is really two separate and entirely unrelated studies, one on how people improve with drugs, and one on how they improve with placebo, so the way to analyse them is to ignore the study design and just try and estimate the pooled effect size for each group (drug and placebo) as if they were unrelated. It partly has an effect because the SDs depend on response, and because sample sizes are skewed towards drug groups in some studies (so the placebo group is much smaller than the drug group).

Taking your SD = change/d approach, and just plugging it into a meta-analysis program (SE weighting, fixed effects) giving an overall effect size of 1.9, it is interesting to note that the fluoxetine trials contribute half as much to the drug analysis (in terms of weighting) compared to the placebo analysis!

But as before, it is also interesting to see that segregating by drug gives effect sizes from 3.6 to .6 (or, given the silly form of this analysis, comparing individual drug groups to pooled placebo subjects 3.8 to -.2)."

I'll expand on that last bit - basically if Robert is right, and it is the best explanation I've found (looking back at the paper there are tantalising suggestions that it is correct because they report model statistics separately), then they have assumed that there are two entirely separate populations, the drug group and the placebo group, and that each trial is simply an attempt to estimate the size of the improvement in HRSD score within each group, ignoring any information about which placebo group went with which drug group in any particular trial (this is an approach that chimes with their regression analysis approach looking at each group separately).

When I attempted to replicate this sort of analysis as I mention above I find that the effect sizes are 9.6 and 7.7 (difference of 1.9) with the drug groups paroxetine, fluoxetine, nefazadone, and venlafaxine 9.6, 7.5, 10.6, and 11.5 respectively, making differences (to overall placebo) of 2.0, -.2, 3.0, and 3.8, although compared to their relevant placebo groups these differences are 3.0, .6, 1.8, and 3.6, giving you an idea of how the placebo groups vary by the drug study they are in.

Robert also finds a particularly telling aspect of the study:
"In my view in passing from the publication biased 3.23 to the final 1.78 only 0.6 of the change is due to removing the publication bias and 0.85 is due to inefficient and biased meta analysis.if the subsample of studies with references (I guess published studies) is analyzed with the method of Kirsch et al the weighted average improvement with SSRI is 9.63 and the weighted average improvement with placebo is 7.37 so the added improvement with SSRI is 2.26.If I have correctly inferred which studies were publicly available before Kirsch et al's FOIA request, I conclude that they would have argued that the effect of SSRI's is not clinically significant based on meta analysis of only published studies."
UPDATE
As mentioned in the comments, here's a pretty graph showing the effect size adjusted by regression on the baseline HRSD scores (to a baseline severity of 26 points) giving an overall effect of 3.0 (the grey line, as we'd expect from the regression lines which reach 'clinical significance' at baseline = 26) we get 2.4, 2.9, 3.2, and 3.7 for the effect sizes in nefazadone, flouxetine, paroxetine, and venlafaxine respectively, although the differences between drugs doesn't seem to be stastically significant (the closest is nefazadone versus venlafaxine).