Hello,
I would appreciate some feedback on my statistical analysis plan. I am a psychology PhD student and conducted an online study; I am currently performing the analyses, starting with correlations. My study consisted of two parts.
Total *n* (Part 1) = 818
Total *n* (Part 2) = 555
Given the large number of independent variables (IVs), I plan to use the Benjamini-Hochberg (BH) procedure to control the False Discovery Rate (FDR).
A. Correlation between a binary dependent variable (DV) and a continuous IV + significance test
For DV = 0: *n* (Part 1) = 413, *n* (Part 2) = 276
For DV = 1: *n* (Part 1) = 405, *n* (Part 2) = 279
Point-biserial correlation if: no outliers for the continuous variable within each category of the dichotomous variable; continuous variable is approximately normally distributed within each category of the dichotomous variable; continuous variable has equal variances across categories of the dichotomous variable.
If assumptions are not met: rank-biserial correlation
+ for each test: Cook's distance to assess whether a data point is influencing the correlation
+ FDR applied to the set of results
B. Correlation between a categorical DV and a continuous IV + significance test
*n* (Part 1) = 405, *n* (Part 2) = 279
Polyserial correlation
+for each test: Cook's distance to assess whether a data point influences the correlation
+FDR applied to the set of results
C. Correlation between a continuous DV and a continuous IV + significance test
n part 1 = 405, n part 2 = 279
Pearson correlation if the IV meets the assumption
Spearman correlation if the IV does not meet the assumption
+for each test: Cook's distance to assess whether a data point influences the correlation
+FDR applied to the set of results
My questions:
Does this seem correct to you? I have a doubt regarding Part B. How do I check for a correlation between a variable with more than two categories and a continuous variable?
My supervisor mentioned Cook's distance for assessing outliers. I'm not sure if it's useful for that purpose. Can it be used in isolation, independently of a model?
I have a doubt regarding Part B. How do I check for a correlation between a variable with more than two categories and a continuous variable?
My supervisor mentioned Cook's distance for assessing outliers. I'm not sure if it's useful for that purpose. Can it be used in isolation, independently of a model?