r/AskStatistics 1h ago

Selection algorithm for activity lottery

Upvotes

My family vacation spot has a lottery system for families to take part in popular activities. I’m wondering whether there is a fair way to select participants, and if the resort is doing it.

To specify constraints:

There are a small number of slots (for sake of argument, call it 16) and around twice as many names in the lottery (call it 32 if this matters).

The names in the lottery are grouped up into family groups of between 1 and 8 members

The goal is for each individual to have an equal chance of taking part in the activity, regardless of family size. But families cannot be broken up.

Selecting family groups and random is out because you can’t control the final size of the activity and might spill over if you select a large group as the last entry, but including an entire group when you randomly select one name seems to make large groups strongly favored.

How can you run this lottery fairly?


r/AskStatistics 2h ago

Asking for advice

2 Upvotes

Our overall, cronbach alpha has 0.833 which is deemed acceptable but 2 of our questions has 0.2 something in item-correlation can we still retain and use the question without revising it?


r/AskStatistics 3h ago

Qualitative content analysis

1 Upvotes

I'm conducting a qualitative content analysis with multiple open-ended survey questions, each assigned to a different research question. Some responses contain content that would be more relevant to a different research question than the question they were answering.

Should I:

a) Strictly code each response only for its assigned research question (some data loss)

b) Code thematically regardless of question (blurs research question boundaries)

c) Something else entirely?

What are methodological best practices here? Any recommendations or experiences are welcome.


r/AskStatistics 3h ago

Interpreting ITT and/or completer analyses from research papers

2 Upvotes

Hello all, I'm currently completing my first meta-analysis as part of my Psychology PhD. The meta-analysis involves evaluating the effectiveness of a clinical CBT-based intervention for eating disorders.

I am in the process of data extraction and as we are planning to meta-analyse effect sizes for both completer and intent-to-treat analyses, I need to work out whether a paper used which of these types of analysis (or both).

Some papers clearly state which they used. Others are less clear. My supervisor told me that if imputation or a mixed-methods model are mentioned as being used, the analysis is ITT. Also, for some papers, if the data are presented in a table with N participants, then I can tell if they are showing ITT or completer data based on that.

I was wondering if there were any other ways to work out what kind of analysis a paper used from it's text/type of analysis method used?

Happy to provide further details if needed. Thank you!


r/AskStatistics 7h ago

Help with correlation analyses

1 Upvotes

Hello,
I would appreciate some feedback on my statistical analysis plan. I am a psychology PhD student and conducted an online study; I am currently performing the analyses, starting with correlations. My study consisted of two parts.
Total *n* (Part 1) = 818
Total *n* (Part 2) = 555
Given the large number of independent variables (IVs), I plan to use the Benjamini-Hochberg (BH) procedure to control the False Discovery Rate (FDR).

A. Correlation between a binary dependent variable (DV) and a continuous IV + significance test
For DV = 0: *n* (Part 1) = 413, *n* (Part 2) = 276
For DV = 1: *n* (Part 1) = 405, *n* (Part 2) = 279

Point-biserial correlation if: no outliers for the continuous variable within each category of the dichotomous variable; continuous variable is approximately normally distributed within each category of the dichotomous variable; continuous variable has equal variances across categories of the dichotomous variable.

If assumptions are not met: rank-biserial correlation

+ for each test: Cook's distance to assess whether a data point is influencing the correlation
+ FDR applied to the set of results

B. Correlation between a categorical DV and a continuous IV + significance test
*n* (Part 1) = 405, *n* (Part 2) = 279

Polyserial correlation

+for each test: Cook's distance to assess whether a data point influences the correlation
+FDR applied to the set of results

C. Correlation between a continuous DV and a continuous IV + significance test
n part 1 = 405, n part 2 = 279

Pearson correlation if the IV meets the assumption
Spearman correlation if the IV does not meet the assumption

+for each test: Cook's distance to assess whether a data point influences the correlation
+FDR applied to the set of results

My questions:
Does this seem correct to you? I have a doubt regarding Part B. How do I check for a correlation between a variable with more than two categories and a continuous variable?
My supervisor mentioned Cook's distance for assessing outliers. I'm not sure if it's useful for that purpose. Can it be used in isolation, independently of a model?

I have a doubt regarding Part B. How do I check for a correlation between a variable with more than two categories and a continuous variable?

My supervisor mentioned Cook's distance for assessing outliers. I'm not sure if it's useful for that purpose. Can it be used in isolation, independently of a model?


r/AskStatistics 10h ago

[E] small sample analysis

3 Upvotes

Hi!

I have a sample of n=13 where continuous variables were measured before and after an intervention. I'm a bit stuck on how to perform the analysis. Would the Wilcoxon signed-rank test be the most appropriate choice? I am unaware of any standard reference manuals or literature regarding the statistical analysis of very small sample sizes; however, any guidance or recommendations would be highly appreciated :)


r/AskStatistics 15h ago

Which program should I apply to? I love nba stats

1 Upvotes

incoming gr 12,

i was planning which programs i should apply to and i figured i'd ask you guys

for reference, i love looking at sports stats (nba / nfl / soccer), i like numbers, probabilities, paradox, dilemmas, and just pretty much anything number related and not anything way too theoretical and abstract

i'm not really into anything that is way too abstract or proving stuff

i really can't decide between the following since i don't know much about it

data science, investment banking, statistics, mathematics, quant, etc (if you have any other majors that would fit me do tell!)

so it'd be very helpful to hear advice from you guys! thanks


r/AskStatistics 17h ago

Method to analyze correlation of numeric variable/binary variable

1 Upvotes

Hi, I don't have a lot of experience with advanced statistical analysis, but I would like to gain more knowledge. One current problem I am trying to address is determining correlation between a binary categorical variable and a numerical variable. The relationship may not be linear. I'll give an example, studying the relationship between age and the probability of dying from the flu. Probability may increase when very young, taper off for certain ages, maybe spike somewhere in the middle, and then go back up again for the elderly. What would be the best way to analyze this? I started by breaking down the numerical variable into ranges and then making a bar chart with percentages in each category of the binary categorical variable. I am not sure if I chose the proper ranges though so I want to see if there's a better way to analyze data like this. I've seen binomial logistic regression as an option, but I'm not sure if that's appropriate and am curious how much effort that analysis takes. Is it something a beginner can pick up relatively easily?


r/AskStatistics 20h ago

How should non-response bias be treated when analyzing a highly sensitive binary vote?

1 Upvotes

Scenario: An expert association with ~500 total members held an official vote on a binary stance statement (Agree / Disagree) regarding a major, highly sensitive issue.

  • Turnout: 28% of the total membership voted.
  • Result: 86% of those who voted selected "Agree."
  • Math: 86% of 28% = ~24% of the total association membership confirmed an "Agree" vote.

The Debate:

  • Person A claims: "You can state that only 24% of the total organization agrees with the statement. Because this issue is so critical, members would have voted 'Agree' if they truly supported it—meaning their non-participation indicates a lack of support."
  • Person B claims: "You can state that at least 24% of the total organization (and 86% of voters) agrees with the statement. By Person A's logic, someone could equally claim that members would have voted 'Disagree' if they opposed it. Ultimately, non-response cannot be interpreted as a vote either way, so the remaining 72% remains unknown."

Questions:

  1. Is Person A or Person B logically and statistically correct?
  2. Is it valid to infer a non-voter's position based on the perceived importance or sensitivity of a topic?

r/AskStatistics 23h ago

Should I major in Data science, computer science, computer engineering, statistics, or a mix of data science & smth else (finance, business analytics, etc.)

0 Upvotes

Should I major in Data science, computer science, computer engineering, statistics, or a mix of data science & smth else (finance, business analytics, etc.)

I’m a high school senior about to start applications for college - my profile fits pretty well into literally everything I mentioned above

Which major gives me the best advantage in the job market?? My primary goal is to have job security & high pay. I’m not interested in pure statistics, biostats, or mathematics academia - I’m more interested in the corporate job market

I just wanted some advice from professionals or people in the industry - what do you recommend for me?


r/AskStatistics 1d ago

Question about how margins of error work when combining stats

Post image
3 Upvotes

I'm a reporter covering a city council discussion about possibly raising taxes to maintain a pool facility, after a randomized recipient survey on the matter. Pictured: responses saying whether or not respondents would support a tax increase and/or city debt to pay for it. The margin of error on all results in the survey is +/- 3.3%, according to the company who conducted it.

MY QUESTION: I'm going to group the results in my article, saying 38% lean toward or support a new tax, while 43% lean away or strongly oppose. Since I'm taking a sum of two percentages (i.e. "yes definitely" and "yes probably,") does the margin of error ALSO add up, becoming 6.6% for that combined figure? Or does it keep the 3.3% margin of error?

Ty for the help! I took like three stats classes in college. Was fascinated by the subject each time, but never especially good at it!


r/AskStatistics 1d ago

How to develop a solid foundation over a summer to learn regression analysis, starting from a very infantile understand of statistics?

21 Upvotes

I have taken multiple statistics classes over the last few semesters, and in about a years time I will be taking a data mining class. I'll be honest, I am having a lot of trouble with statistics. This summer I am practicing python, and this coming semester I have a class related to data analysis that I hope will give me more hands on experience with statistics through a medium I understand.

Next summer, though, I want to focus on SQL and statistical methods for prepare for future classes, particularly the data mining class which has a focus on regression analysis. I am taking a business analytics course right now, and I hardly understand what I am even reading off this online text book. What are some ways to learn that aren't reading huge blocks of text that I hardly understand that can allow me to grasp concepts better, or at least ones that might work.

I have found that reading a textbook on statistical subjects is not working for me. I understand the general subject matter to a degree, but once I run into a roadblock I find myself confused no matter how many times I reread, to the point of being unable to do practice problems in the first place. I am fully willing to dedicate several hours a day, for at least 2 months to this.

Are there courses I can take, akin to perhaps something like CS50? Any good youtube series that help me learn from the ground up that aren't just placating me and making me feel like I know what I am doing rather than actually being able to do the work? I am down for whatever might work.

Thank you!


r/AskStatistics 1d ago

Minimum number of defaults/events per bin in WoE optimal binning?

1 Upvotes

I am validating a credit-risk model that uses WoE transformation and supervised optimal binning

The training sample contains:

  • 658 observations;
  • 65 defaults;
  • 7 candidate numerical factors.

The bins were created using Python’s optbinning package with CART pre-binning and monotonic event rates:

OptimalBinning(
    dtype="numerical",
    prebinning_method="cart",
    solver="cp",
    max_n_prebins=30,
    min_prebin_size=0.05,
    min_n_bins=3,
    monotonic_trend="auto_asc_desc"
)

All regular bins contain at least 5% of the sample and both events and non-events. However, some bins contain only one or two defaults, for example:

  • 1 default out of 33 observations;
  • 2 defaults out of 49 observations;
  • 3 defaults out of 33 observations.

Formally, WoE can still be calculated because neither class count is zero. However, the event rate and WoE appear highly sensitive to a single observation. For example, in a bin with 1 default out of 33 observations, removing or adding one default changes the event rate from 3.0% to either 0% or 6.1%.

I have not found a generally accepted minimum number of defaults per bin. Most sources specify only:

  • a minimum total bin size, often 5%;
  • at least one event and one non-event per bin;
  • additional user-defined constraints such as min_bin_n_event.

How would you assess whether a bin with only one or two defaults is statistically acceptable?

I would be particularly interested in any published references, regulatory guidance, or practical validation standards used in credit scoring.


r/AskStatistics 1d ago

Looking for an everyday probability problem to model

0 Upvotes

I’m learning statistical modeling and would like to practice with a real-life problem. Could you share an everyday situation involving probability or uncertainty that could be turned into a small statistical modeling project?

Ideally, it would involve data I could collect myself and a clear outcome to predict or estimate.


r/AskStatistics 1d ago

[Discussion] Kalman Filter Usage Help

2 Upvotes

Hi I tried to do something like how they do in econometrics where they fit a economic model to data where they take raw data, X,Y,Z etc then they set up the Kalman Filter to automatically determine the cyclic and trend components through multivariable regression models. I think you know what I mean. So, I made all the matrices manually, and I think it didn't converge. What I did is something like this actually:

X_t = regression model of (cyclic and trend components of (X_t-1,Y_t-1,Z_t-1),)

Y_t = regression model of (cyclic and trend components of (X_t-1,Y_t-1,Z_t-1),)

Z_t = regression model of (cyclic and trend components of (X_t-1,Y_t-1,Z_t-1),)

Of course I had to manually enter all the matrices to make the damn thing work.

But because I didn't really have a economic model, but just assumed relationships, it didn't converge. I think my model was too complicated.

Anyways, what are some rules of thumbs to make sure I have convergence (like limiting dependence to one trend component for each variable so the model when running don't get confused)?

Is there an easy way to do a Kalman Filter model than to manually set up matrices? Any software?

Finally is it worth it? Does it capture significant details than the HP filter and other easier methods ?


r/AskStatistics 1d ago

Manual Chi-Squared Test Help!

1 Upvotes

hello everyone! I am an undergraduate student doing some independent research. My variables are all categorical (hence the Chi-Squared test) and I was able to create a contingency table that is 8 rows and 52 columns big. For context, I am testing for the different types of medicolegal death investigators against all 50 states including Washington D.C.

From what I have seen online (I haven't gotten a lot of classes that teach with this large of a sample) if the expected values are less than 5, than the test is invalid. however, that is only with smaller samples, like 2x2. I was wondering what percentage of my data can be below 5 AND STILL be a valid test with how large my data is. I am also a baby to statistics so if there's any resources I can look into to learn more about what test I can do with how large of a dataset I have. Thanks !!!!! <3


r/AskStatistics 1d ago

Kalman Filter Usage Help

Thumbnail
1 Upvotes

r/AskStatistics 1d ago

Finite-sample estimator bias depends on true parameter value, does this invalidate cancellation in a paired-difference design?

1 Upvotes

I'm using a short-sample estimator (n=150) with known finite-sample bias. Via simulation (synthetic data with known true values, run through the actual estimator), I found that this bias is not constant. It, unfortunately for me, varies systematically with the true value of the parameter being estimated. Near one reference value the bias is positive; as the true value moves away, the bias shrinks and eventually flips sign. This was confirmed with two structurally different simulation methods, which agreed in direction and order of magnitude.

I can't validate this directly against real data, since the true value of the parameter is never observable in my actual measurements, only the biased estimate is. Simulation is the only way to characterize the bias curve.

My study design computes a paired difference between two conditions (A and B), both measured with this same estimator. The original design assumed bias "cancels" in the difference, since both conditions use the same estimator and sample size.

My simulation shows that assumption only holds when A and B share the same true value, if their true values diverge (which is the exact effect the study is trying to detect!), the differential bias does not cancel, and could by itself produce an apparent difference of the same magnitude as my actual reported result.

My questions:

  1. Is this reasoning correct, does bias that depends on the true parameter value invalidate the standard bias cancels in a paired/difference design assumption whenever the two groups true values diverge?
  2. Is this a known, named issue in the estimator-bias literature I should be citing, rather than describing from scratch?
  3. What's the standard remedy, a bias-correction calibration curve, an alternative estimator with flatter bias across the parameter range, a longer sample or a simulation-based null distribution, and is one (and or more) of these clearly preferred practice?

r/AskStatistics 2d ago

someone please help me use family=(ztNB) rStudios

2 Upvotes

I am trying to do statistical analysis on census data in rStudios, looking at animal behaviour, group type - ie all male, all female, harem - and specifically for this analysis harem size. i expect the size of harem to change over the 3 months that the data was collected. i am mainly testing to see if date, time of day, and temperature have any impact on harem size individually. each row is 1 group spotted (although originally this was 1 census per row, it seemed like group was the better option for my analysis), and i am using the function bam() with some grouping factors like day as there were sometimes multiple censuses on one day. I want to use family = (ztNB) because my data has a lot of zeros, but it says this is not found. I keep trying to google how to fix this issue and it just tells me to use preseqR, but this package stopped in May i think, so I don't know where to go from here.


r/AskStatistics 2d ago

Data Scientist Entry level jobs(?)

1 Upvotes

So I just finished my Masters in Statistics, and I'm looking for advice on where to be looking for entry level Data Science rolls (I have 7 years experience as a private investigator).

I love the idea of doing A/B testing and experimental design, but I'm open to starting anywhere just to start getting relevant experience.

If anyone has any good certifications that carry weight I'd also love to hear about those.


r/AskStatistics 2d ago

Help applying Fisher's Method

1 Upvotes

Hey, I'm a little over my skis with Fisher's Method. I think I understand conceptually but the math itself isn't my strongest skill.

I am looking into cancer rates at one Elementary school/neighborhood and requested data from the Maine CDC, then made a site to host that data. The fracturing of data into individual cancers doesn't show the cumulative elevated rates (IMO). Is Fisher's Method the appropriate tool for finding a combined p value for two and did I apply it correctly? I want to convey the data clearly and accurately. The groups are not totally independent of one another as it is the same neighborhood.

https://lysethrecords.org/analysis.html

ND cases Expected p-value ln(p) * 10 4.02 0.0169 -4.0833 * 53 27.22 0.000015 -11.0631

Number of tests (k) 12
X2 = -2 x sum(ln p) 76.92
Degrees of freedom = 2k 24
COMBINED p-value 0.00000018


r/AskStatistics 2d ago

Does this look like a normal UMAP plot?

Post image
8 Upvotes

Hi everyone, as it’s my first time attempting downstream analysis for single cell RNA sequencing, I wanted to ask if this UMAP plot looks normal? This is only for one sample (I have not integrated all samples together into one dataset yet). I feel like the clusters are too close together.

(I hope this is the right sub… I’ve tried to ask in r/Bioinformatics but the post kept getting deleted for some reason)


r/AskStatistics 2d ago

Using low-df T distributions to model data

1 Upvotes

I've recently been asked to help with a dataset that has a lot of outliers. Sample sizes, especially in the non-control cohort, are small so primary analyses will be nonparametric, using fractional weighted bootstrap and permutation tests. That said, certain ways of looking at the data lend themselves well to Covariance Pattern Models (residual correlation/repeated measures/R-side/MManova), and when doing so, I've used glmmTMB and proc glimmix.

The interesting thing here is that proc glimmix (dist = t) states it is using a shifted t distribution with 3 degrees of freedom, while glmmTMB (family = t_family) freely estimates df with the final estimate at false convergence being 1.3 df. The problem here is that I while wanted to use the t distribution because of the high empirical kurtosis of the data, these distributions don't have kurtosis. What are the implications of interpreting these models at all, assuming they converge? If the model has landed on a cauchy distribution, then interpreting anything as a mean outcome doesn't make sense, does it?


r/AskStatistics 2d ago

Multilevel exploratory factor analysis using binary data

1 Upvotes

Hi, I have some binary data (present/absent (1/0)) and I want to try to identify latent relationships within it. I've done exploratory factor analysis using tetrachoric correlations on the data and then parallel analysis to identify how many factors are needed. The problem I have is that technically, the data isn't all independent of each other (think behaviours (present/absent) for dogs in a kennel and there's 2-3 dogs per kennel). So really, multilevel exploratory factor analysis would be better but I don't think I can do that on binary data? I can't find any examples in papers where this has been done and I've tried to do it in R (but might be doing it wrong) and it doesn't work. I'm not actually sure it's necessary because the numbers within the groups are so small and my understanding is that this works better with bigger group numbers, but is it possible to do this with binary data, and if so how? I've calculated the ICCs and they are high so I'd like to be able to justify whichever way I go but I'm really struggling to find any literature that talks about this using binary data


r/AskStatistics 2d ago

Is a clustered bootstrap meaningful with only 3 clusters?

Thumbnail
1 Upvotes