コーエンの


12

Cohenのは、エフェクトのサイズを測定する最も一般的な方法の1つです(Wikipediaを参照)。プールされた標準偏差に関して2つの平均間の距離を単純に測定します。Cohenのdの分散推定の数式をどのように導出できますか? dd

2015年12月編集:この質問に関連するのは、dの周りの信頼区間を計算するという考え方です。この記事ではd

σd2=n+n×+d22n+

ここで、は2つのサンプルサイズの合計であり、n ×は2つのサンプルサイズの積です。n+n×

この式はどのように導出されますか?


@Clarinetist:他の人の質問を編集して、より多くの実質とより多くの質問を追加することは(文言を改善するのではなく)幾分物議を醸しています。私はあなたの編集を承認するために自由を取りました(あなたが寛大な賞金を置いて、あなたの編集が質問を改善すると思うと)、他の人はロールバックすることを決めるかもしれません。
— アメーバは、モニカを復活させる

1
@amoeba問題ありません。式がためにそこにある長いほど(前無かった)、それは私たちが式の数学的な導出を探していることは明らかだ、の罰金という。σd2
— クラリネット奏者

2番目の分数の分母はなければなりません。以下の私の答えをご覧ください。2(n+−2)

回答:


15

質問の分散式は近似値であることに注意してください。 Hedges(1981)は、一般的な設定(すなわち、複数の実験/研究)で大きなサンプル分散と近似を導き出しました。そして、私の答えは、論文の派生物をほとんど通り抜けます。d

まず、利用する前提は次のとおりです。

2つの独立した治療グループ、(治療)とC(対照)があるとします。ましょうY T IとY C jはどの被験者からのスコア/応答/ BE I群におけるT被写体J群でC、夫々 。TCYTiYCjiTjC

応答は正規分布であり、治療群と対照群は共通の分散を共有していると仮定します。

YTi∼N(μT,σ2),i=1,…nTYCj∼N(μC,σ2),j=1,…nC

私たちはそれぞれの研究で推定することに興味が効果の大きさは、。私たちが使用します効果の大きさの推定がある D= ˉ Y T- ˉ Y Cδ=μT−μCσ ここで、S2kはグループkの不偏サンプル分散です。

d=Y¯T−Y¯C(nT−1)ST2+(nC−1)SC2nT+nC−2
Sk2k

大標本プロパティを考えてみましょう。 d

まず、ノートその: 及び()私の表記とルーズである: (N T - 1 )S 2 T

Y¯T−Y¯C∼N(μT−μC,σ2nT+nCnTnC)
および (NC-1)S 2 C
(1)(nT−1)ST2σ2(nT+nC−2)=1nT+nC−2(nT−1)ST2σ2∼1nT+nC−2χnT−12
(2)(nC−1)SC2σ2(nT+nC−2)=1nT+nC−2(nC−1)SC2σ2∼1nT+nC−2χnC−12

1σ2(nT−1)ST2+(nC−1)SC2nT+nC−2∼1nT+nC−2χnT+nC−22

今、いくつかの巧妙な代数:

d=Y¯T−Y¯C(nT−1)ST2+(nC−1)SC2nT+nC−2=(σnT+nCnTnC)−1(Y¯T−Y¯C)(σnT+nCnTnC)−1(nT−1)ST2+(nC−1)SC2nT+nC−2=(Y¯T−Y¯C)−(μT−μC)σnT+nCnTnC+μT−μCσnT+nCnTnC(nT+nCnTnC)−1(nT−1)ST2+(nC−1)SC2σ2(nT+nC−2)=nT+nCnTnC(θ+δnTnCnT+nCVν)
where θ∼N(0,1), V∼χν2, and ν=nT+nC−2. Thus, d is nT+nCnTnC times a variable which follows a non-central t-distribution with nT+nC−2 degrees of freedom and non-centrality parameter of δnTnCnT+nC.

Using the moment properties of the non-central t distribution, it follows that:

(3)Var(d)=(nT+nC−2)(nT+nC−4)(nT+nC)nTnC(1+δ2nTnCnT+nC)−δ2b2
where
b=Γ(nT+nC−22)nT+nC−22Γ(nT+nC−32)≈1−34(nT+nC−2)−1

So Equation (3) provides the exact large sample variance. Note that an unbiased estimator for δ is bd, with variance:

Var(bd)=b2(nT+nC−2)(nT+nC−4)(nT+nC)nTnC(1+δ2nTnCnT+nC)−δ2

For large degrees of freedom (i.e. large nT+nC−2), the variance of a non-central t variate with ν degrees of freedom and non-centrality parameter p can be approximated by 1+p22ν (Johnson, Kotz, Balakrishnan, 1995). Thus, we have:

Var(d)≈nT+nCnTnC(1+δ2(nTnCnT+nC)2(nT+nC−2))=nT+nCnTnC+δ22(nT+nC−2)

Plug in our estimator for δ and we're done.


Very, very nice derivation. Just a few questions: 1) could you clarify what the notation Y¯iT−Y¯iC means (I know it's something to do with difference of sample means, but how can they both have the same index?)? 2) could you clarify how the approximation for b is done (I don't need all of the details, a source is fine and maybe a brief explanation)? Otherwise, I'm quite pleased with this. (+1) This also agrees with the observation that I've made that d doesn't follow a normal distribution, contrary to the explanation in the linked article in the OP.
— Clarinetist

@Clarinetist Thanks! 1) How can they have the same index? Typo, that's how! :P They're an artifact of my first draft of the answer. I'll fix that. 2) I pulled it from the Hedges paper -- don't know its derivation at the moment but will think about it some more.

I'm looking into the derivation now, but FYI, the numerator of b should be Γ(nT+nC−22).
— Clarinetist

Derivation provided for reference: math.stackexchange.com/questions/1564587/… . Turns out there's likely a sign error.
— Clarinetist

@mike : very impressing answer. Thanks for taking the time to share it with us.
— Denis Cousineau
弊社のサイトを使用することにより、あなたは弊社のクッキーポリシーおよびプライバシーポリシーを読み、理解したものとみなされます。
Licensed under cc by-sa 3.0 with attribution required.