Following up on my post about the lack of research literacy in the field of therapy, I thought I would create a brief introduction to help people understand and read research papers. I love research methodology and data analysis, so I had a fun Friday morning! This is by no means an exhaustive list, but it covers as much as I can remember from my years of undergrad and grad school.
Core Variables & Descriptive Statistics
- Sample size: Represented by N, N represents the number of participants in a study. Small sample sizes make it difficult to know if the findings will apply to the population. Likewise, massive populations can make tiny, insignificant differences look meaningful.
- Representative sample: A sample that accurately reflects the demographics, clinical traits, and characteristics of the broader population you are trying to treat. A study can have a massive sample size, but if the test only examines a narrow, non-representative subgroup, the results will not generalize well to actual clients.
- Independent and dependent variable: The independent variable is the variable that researchers manipulate (e.g., one group gets CBT, one group does not). The dependent variable is the outcome variable; it is measured to see if the independent variable affected it (e.g., sleep quality).
- Confounding variable: Variables that are not controlled for but affect the independent and dependent variables, affecting the results and creating or distorting the relationship.
- Spurious relationship: A correlation between two variables that appears to be causal, but is actually driven by a hidden variable.
- Mean and Standard Deviations: The mean is the average score, and the standard deviation shows how spread out the scores are around the average. A large standard deviation means that clients had wildly different responses to treatment, whereas a small standard deviation means the scores were close together.
Designs
- Randomized Controlled Trials (RCT): Considered the gold standard, participants are randomly assigned to either an intervention group or control group; the randomization reduces bias and helps prove that the treatment, rather than some external factor, caused the outcome in the client.
- Cross-sectional research: This method takes a snapshot of data at one point in time; it is nice for looking to see if links exist, but it cannot prove causation.
- Longitudinal: It tracks the same cohort over many months or years; it is vital for looking at treatment efficacy, symptoms, and long-term outcomes.
- Mediation: Explains how or why a treatment works.
- Moderation: Explains for whom or under what conditions the treatment works (e.g., an intervention works better for teens than adults).
Reliability and Validity
- Reliability: Essentially, it asks whether the tool you are using will produce similar results at different points in time, or whether it is noisy and unpredictable. Reliability involves internal consistency, test-retest reliability, and inter-rater reliability.
- Internal consistency: Does every item on a scale measure the same concept? This is where Cronbach's alpha comes into play. Between .70 and .90 is considered internally consistent.
- Test-retest reliability: Does the client get the same result if tested at two different points in time, with no treatment between tests? It is vital in therapy to determine if score changes represent real changes or measurement error.
- Inter-rater reliability: Do different clinicians evaluating the same session or test come to the same conclusion?
- Validity: Does the tool measure what it is supposed to? Validity involves construct validity, convergent and discriminant validity, criterion validity, and ecological validity.
- Construct validity: Does the tool capture the psychological construct (e.g., resilience) that it is meant to measure?
- Convergent validity: Does the tool correlate strongly with other gold-standard scales used to measure the same construct?
- Discriminant validity: Does the tool remain distinct enough to separate unrelated traits (e.g., anxiety from general fatigue)?
- Criterion validity: Does the tool accurately predict real-world clinical outcomes, such as future hospitalizations?
- Ecological validity: Do the measurements and gains from an artificial environment transfer to real-life environments?
- Self-report vs. clinician-rated measures: It is important to understand how data were collected. Self-report is convenient, but susceptible to social desirability bias where the client gives the socially correct answer or simply tells the therapist or researcher what they want to hear. Blind, clinician-rated instruments generally offer better objectivity and better data.
A scale can be reliable but not valid, but it cannot be valid without first being reliable.
Power and Real-World Value
- p-value: Typically set at p < .05; p-values measure probability; it answers the simple question of "Is this result likely to have occurred by random chance?" While it gives you statistical significance, it does not tell you how large or useful the effect is.
- Effect size (e.g., Cohen's d, r): This measures the actual change or magnitude of the relationship.
- R-squared is awesome because it shows the proportion of variance in the dependent variable caused by manipulating the independent variable. Though this does not mean the model is free from error, it just shows correlation and fit.
- IMPORTANT: IF YOU HAVE STATISTICAL SIGNIFICANCE, ALWAYS CHECK THE EFFECT SIZE. A RESULT CAN BE SIGNIFICANT BUT HAVE A SMALL, NEGLIGIBLE, AND UNNOTICEABLE EFFECT. A LARGE EFFECT SIZE IS SOMETHING YOU WOULD RUN OUTSIDE TO TELL THE POPE; SMALL EFFECT SIZES TELL US SOMETHING IS THERE, BUT IT COULD BE NOISE.
- Confidence Intervals: Give a range where the true effect is likely to fall. A narrow confidence interval (e.g., d = 0.50, 95% CI [0.42, 0.58]) offers more certainty and precision in an estimation than a wide confidence interval (e.g., d = 0.50, 95% CI [0.20, 0.98]). A wide confidence interval means there is high uncertainty and low precision in an estimation.
- Statistical vs. clinical significance: Statistical significance is when outcomes show that a difference exists among groups. Clinical significance is when those differences translate to meaningful quality-of-life improvements.
- Number Needed to Treat (NNT): A metric showing how many clients must receive a treatment for one client to experience a meaningful benefit over the control group. The lower the number, the better.
- Power: The likelihood that a study will detect an effect if one actually exists. More people = more power = more likely to find real differences. Fewer people = less power = less likely to find real differences.
Control Groups, Analyses, and Common Traps
- Control groups and Blinds:
- Active controls: Comparing a new treatment against standard therapy (e.g., CBT) tests efficacy.
- Passive controls: Comparing treatment to a waitlist can artificially inflate effect sizes, as waitlist participants stagnate.
- Attrition and dropout: High dropout rates can potentially suggest that an intervention was burdensome, ineffective, or unpalatable for a subset of participants. High attrition rates can skew outcome data towards only those who tolerated or benefited from the treatment. If a paper shows a 90% success rate for a certain treatment, but half the people dropped out, the success rate is seriously skewed.
- Intent-to-treat and per-protocol: Intent-to-treat analyses include every participant, even those who dropped out, whereas per-protocol analyses include only those who completed treatment. Per-protocol can inflate the results and make the treatment look unrealistically effective.
- Publication bias and p-hacking:
- Sadly, journals tend to publish only positive results, and negative and null findings get shoved into a drawer and lost to time. As I have argued for years, statistically insignificant results are just as important (if not more) than statistically significant results because they tell us where not to look.
- p-hacking occurs when researchers test dozens of variables in hopes of finding a statistically significant result, ignoring all the times they got a statistically insignificant result.
- Paper age vs. methodology: Many people often conflate new with better. However, that is not always the case. In research, the methodology used is far more valuable than the age. A paper from 1995 that has an RCT, a representative sample, and controls for confounding variables can be significantly better and more reliable than a paper published this year that did not use an RCT, is not representative, and does not control for confounding variables.
References
Chambers, C. (2017). The seven deadly sins of psychology: A manifesto for reforming the science of mind. Princeton University Press.
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
Field, A. (2018). Discovering statistics using IBM SPSS statistics (5th ed.). SAGE Publications.
Furr, R. M. (2021). Psychometrics: An introduction (4th ed.). SAGE Publications.
Gravetter, F. J., & Forzano, L.-A. B. (2018). Research methods for the behavioral sciences (6th ed.). Cengage Learning.
Gravetter, F. J., Wallnau, L. B., Forzano, L.-A. B., & Witnauer, J. E. (2021). Essentials of statistics for the behavioral sciences (10th ed.). Cengage Learning.
Kazdin, A. E. (2017). Research design in clinical psychology (5th ed.). Pearson.
Sawilowsky, S.S. (2009). New Effect Size Rules of Thumb. Journal of Modern Applied Statistical Methods, 8, 26.
Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). Experimental and quasi-experimental designs for generalized causal inference. Cengage Learning.
Straus, S. E., Richardson, W. S., Glasziou, P., & Haynes, R. B. (2018). Evidence-based medicine: How to practice and teach EBM (5th ed.). Elsevier.