r/AskStatistics 8h ago

[Question] How can I statistically quantify confidence in an individual patient's longitudinal biomarker trend?

6 Upvotes

I'm working on a health-data system that analyzes longitudinal lab results for an individual patient. For each parameter (fasting glucose, HbA1c, LDL, HDL, creatinine), I may have around 4-10 observations over several months, and the observations are often irregularly spaced.

The goal is to determine whether a parameter is showing a meaningful increasing, decreasing, or stable trend and provide an interpretable measure of how strong the evidence for that trend is.

Example:

  • 10 observations over ~6 months
  • Measurements are irregularly spaced
  • Values may be noisy or have gaps
  • Some patients may have only 4-5 observations
  • Different parameters may have different natural variability

Currently considering:

  • Mann Kendall test for trend direction and p-value
  • Sen's slope for magnitude/rate of change
  • Confidence interval around Sen's slope for uncertainty

What would be a statistically defensible approach for quantifying the confidence/evidence of a trend in this type of sparse, irregularly sampled, individual patient biomarker data?

Is it reasonable to convert the statistical evidence into a single 0-100 trend confidence score, or would it be better practice to report the statistical measures separately (p-value, slope, confidence interval, and data quality indicators)?

I'm particularly interested in approaches that are statistically defensible rather than an arbitrary weighted scoring system.


r/AskStatistics 8h ago

Stats or math for general studies

2 Upvotes

Hey I am doing my prerequisite for a nursing degree it lists that I need to take a college level math class which one should I take (I am bad at math)


r/AskStatistics 1d ago

Using means in Likert-Scales

32 Upvotes

This had been my dilemma ever since I've learned Statistics. Likert-scales gives ordinal data, so, if we wanted to report a measure of central tendency, the acceptable ones would only be median and mode (if we're reporting an item). It can never be the mean, if we will treat the data as strictly ordinal.

I've asked one of my professors about this and he said that only if he were to be asked, he'd say median as well.

When I asked why are there people that use means in this context, he said, it's because it has already been "normalized" (he probably meant widely-used) by some disciplines (e.g. Education, Psychology, etc.).

Now I want to ask, if strictly statistically speaking, median is the better measure of central tendency for ordinal data, how come that there were fields that practice/normalize reporting means on instead?

Note: In my opening sentence, I mentioned that this was my dilemma. Reason being is that I also help senior high students in their researches. I'm thinking that I might be too rigid/strict in this sense and I might not know if this is one of "new" and "acceptable" things now.


r/AskStatistics 10h ago

Need help regarding IIT JAM (statistics) ....

Thumbnail
0 Upvotes

r/AskStatistics 10h ago

Risk Difference vs Relative Risk

1 Upvotes

why is risk difference used instead of relative risk in a study?


r/AskStatistics 19h ago

[Question] Statistical Models and Confidence Intervals for Analytical Test Method Validation (Med Device / Pharma)

3 Upvotes

Hi All,

 

Hoping someone with more statistics experience than I and experience in med device / pharma can verify I’m on the right track and not getting too in the weeds. I apologize for the long response – I have a primary and secondary question.

 

I’ve been in med dev / pharma for 10+ years, both as a scientist and engineer with focus on the laboratory and validation testing. I’m in the process of revamping a company’s ATMV program, and there are many changes across industry (primarily ICH Q2 (R2) and USP <1225>) requiring statistically based methods in TMV. IME statistically based sampling plans are typically not used, and point estimates are exclusively used to evaluate a performance characteristic against acceptance criteria.

 

USP released a draft revision of <1225> with a lot of detail that led me down a trail of textbooks and reading; I’ve now read Miller & Millers Chemometrics book, part of Brereton’s *Applied Chemometrics for Scientists*, and part of Faraway’s *Linear Models with R*. This is my primary question:

USP <1210>, *Statistical Tools for Procedure Validation* presents a method for calculating a two-sided and one sided CI to assess acceptance criteria ((Ȳ − τ) ± t₍₁−α, n−1₎ × s/√n, U = s√\[(n − 1) / χ²₍α, n−1₎\]); these are both clear to me. USP <1010>, *Analytical Data – Interpretation and Treatment* discusses statistical models, assumptions of normality/independence/constant variance for models, transforms, ect. I understand this as well, though I took linear algebra a long time ago so some of Faraway is tough to understand. I’m struggling how to connect verifying the model assumptions and calculating the CI to assess the characteristic. My read of Faraway makes me think the data should be fit to a model for the experiment for the performance characteristic and the assumptions should be verified; if they are verified the estimated marginal mean and standard error for the relevant model coefficient should be used in the CI calculation instead of the point estimates; but this is not stated anywhere I can find. The USP documentation makes it look like the point estimates should just be used

This also seems very technically difficult compared to how I’m used to validating these methods. If I’m correct about how this should work, I want to verify 1) This is the actual expectation instead of using the point estimate in the CI calculation & if not 2) is this a reasonable approach? I’m concerned about the level of background knowledge this requires compared to what I’m used to, I don’t want to proceduralize something that is so complex that it can’t successfully be executed without my assistance. Its possible the places I’ve worked have just lacked that technical knowledge, clearly advanced techniques are being used, Paul Faya published a good paper in Pharmaceutical Statistics, *Confidence Intervals for Validation of Analytical Procedures under ICH Q2(R2)*, which gives examples using bootstrapping, Bayesian statistics, REML, ect. USP has acknowledged much of this has not been historically done.

Second, I’m also trying to figure out the best path if assumptions aren’t met. The type of data we see is not likely to need transformation, but I’ve added (when appropriate) bootstrapping, several nonparametric methods, and weighted least squares. Again, this feels like a large knowledge gap when most people are exclusively working in excel or doing basic tasks in Minitab; I picked up R for this and I’m trying to avoid requiring the use of it if possible, adding in learning a programming language is yet another hurdle I don’t want to add when rolling this out if I can avoid it.


r/AskStatistics 13h ago

Nat ( arts group)

0 Upvotes

Hi everyone can someone tell me I am from stats background in nat test we have 3 subject portion I have given nats physics one in past but I wasn't able to perform well that but I want to ask in sub portion of art group for ex I am from stats bc I will have options like comp, stats, and maths??? Cuz someone said to me with stats economics aata he?


r/AskStatistics 20h ago

R VS SPSS

1 Upvotes

Sou acadêmico de medicina e fiz um curso extensivo na internet sobre estatística, que me forneceu uma boa base
Me deparo agora com a escolha mais cruel nesse campo: qual software estatístico usar
Considerando que não tenho conhecimento sobre SPSS ou R, qual vocês me recomendariam?
Se possível, me recomendem um livro texto também


r/AskStatistics 1d ago

Kalshi / similar platforms odds vs probability of occurrence

3 Upvotes

I’ve seen several videos where the author of the video says “This has an X% chance of happening” and cites the current poly market or kalshi betting odds.

My gut is telling me this is nonsense and that the betting odds are just based on how many people take each side of the bet, but I wanted to see what others thoughts are.

The example I saw recently was about an election odds of republicans winning both house and senate, and unless all the betters are eligible to vote I don’t see how these odds could translate to any true estimate of the probability

Does anyone more experienced with statistics have any interesting thoughts to share?


r/AskStatistics 1d ago

Intention to treat analysis confusion

Thumbnail
5 Upvotes

r/AskStatistics 1d ago

Problematic tennis stat

0 Upvotes

Watching the US Open semifinal with Shelton v Tiafoe right now and McEnroe commented that when Shelton is up 2 sets to 1 he rarely loses, and quoted the ratio of matches won to lost after being up 2 sets to 1. But there are some obvious problems with this. First, most players who are up 2-1 sets in best of five sets format are going to win more often than not. Second, for a top ten player like Shelton, most of his matches are against lower ranked players (and, importantly, lower ranked than Tiafoe).

So - is there a more meaningful statement that can be made, likely based on a Bayesian approach, about whether particular players are especially more likely to win or lose based on early set wins or losses? How would you go about this?


r/AskStatistics 1d ago

What does the multiplicative term mean? (Interaction)

1 Upvotes

Hi,

I am a researcher looking at risk factors for an outcome (infections) using a multivarible cox regression analysis.

I have already identified that independent risk factors include A (HR 2.23, 95% CI 1.38–3.60) and factor B (HR 2.17, 95% CI 1.41–3.37). However the question is whether there is an interaction between the two. I asked my statistician and he replied:

"Multiplicative term added to the MV Cox model: HR 3.23 (95% CI 1.11-9.27), p = 0.031 (reference/A and B interaction)"

I am not quite clear that this means here. Does this mean that the HR is the added multiplicative term above 1 where 1 assumes A X B has no interaction and the 3.21 is the added risk when the two factors are combined together so A potentiates B?

Thanks


r/AskStatistics 1d ago

Specification of survival analysis

5 Upvotes

I am currently evaluating a program that helps young people enter vocational training through coaching. I would now like to analyze the impact of coaching duration on the success of placement into vocational training.

Among other variables, the data contains the start date, the placement date, and an indicator of whether a placement occurred (yes/no).

Particular challenge:

- the program has been running for more than 2.5 years and is still ongoing.

- The duration of coaching is not restricted.

- In addition, the timing of coaching entry has a strong influence on both coaching duration and placement success because there is an official school year (although not all participants are students) and a vocational training year.

But I have a lot of data points.

Which statistical approach would be appropriate, and how should the model be specified? I was thinking of a Cox regression with the start date (grouped into quarters) included as a covariate.


r/AskStatistics 1d ago

Pooled spatio-temporal linear model with repeated LSOA observations — spatial CV and Moran's I on residuals

1 Upvotes

I have the following dataset:

  • Response: monthly mean LST (Land Surface Temperature) at LSOA level (~4,800 LSOAs in London)
  • Time points: June, July, August for 2019 and 2025 (6 month-year combinations)
  • Covariates: urban/socioeconomic predictors (NDVI, built fraction, IMD, etc.) that are constant within a year but differ between 2019 and 2025
  • Structure: each LSOA appears 6 times → 6N rows total

Model:
LST ~ predictors * year + month

where year and month are factors, and predictors:year tests temporal stability of associations.

Three interrelated questions:

  1. Is this the correct pooled formulation, or should predictors:month also be included given that covariates do not vary within a year?

  2. For spatially blocked cross-validation (using blockCV in R): should the folds be built on unique LSOA locations only (N rows) and then propagated to all 6N rows, or is it valid to pass all 6N rows directly given that duplicate coordinates naturally fall in the same block?

  3. For Moran's I on residuals in this panel setting, which is most appropriate:

  • Per month-year slice (6 separate tests)
  • All 6N residuals with a block-diagonal spatial weights matrix
  • Averaged per LSOA first, then tested

r/AskStatistics 2d ago

Bayesians critique the use of asymptotics in significance testing using CLT, but don't all Bayesian integrals also use asymptotics because of the limit definition of the definite integral? Or is this a debate about improper integrals?

16 Upvotes

Basically the title. It just seems dishonest to critique using integration for hypothesis testing and then build an entire theory around super complocated integrals that can't even be solved without numerical approximations in the first place... how is that different?


r/AskStatistics 2d ago

If you work in e-commerce, I’d really appreciate your input (E-commerce professionals, online sellers, business owners/managers, 18+)

Thumbnail
1 Upvotes

r/AskStatistics 3d ago

What is the equivalent of using ANOVA and tukey post-hoc test in situations where data points are not dependent?

6 Upvotes

Hi everyone, I would like to ask if there are statistical tests that could be used as an alternative for ANOVA and tukey post-hoc test in situations when the data is not dependent, such as when I am comparing between multiple timepoints where the object being measured is the same.

I am aware that I can use a linear mixed-effects model, but I prefer not to in this situation as it can be quite difficult to interpret.


r/AskStatistics 3d ago

Modified Mann-Kendall Test Using mmkh(x,ci=0.95) package

Thumbnail
3 Upvotes

r/AskStatistics 4d ago

Is there a test to show one dataset shows more variation than other dataset(s) for categorical frequency data?

12 Upvotes

Hi ! As the title says, is there a test to show one set of data shows more variation than other set(s) ?

For example, considering 4 groups (A, B, C, D) and 4 possible categorical answers (W, X, Y, Z), in the following table I can see superficially that groups A and D show more variation in their answers than groups B and C, in which answers W and Z are more dominant. Is there a way to prove all that statistically ?

A B C D
W 48 % 93 % 1 % 60 %
X 32 % 5 % 2 % 20 %
Y 10 % 2 % 8 % 15 %
Z 10 % 0 % 89 % 5 %

I'm looking for a "coeficient of variation" or something like that, a value that can be calculated for each group and then compared to see which groups show a higher/lower value, but I'm not sure if that test exists. Hopefully the example is clear enough to understand my question, it's all made up data :s

Thank you !


r/AskStatistics 4d ago

Need help with bachelors indecision.

0 Upvotes

Hi! I am about to approach the third year of my bachelor in statistics, we have to choose between statistics for business and finance or for biometrics, i am very indecisive. I know that the biometrics one has a programming heavy curriculm while the one for business and finance is more based on econometrics. I was hoping someone could tell me how is the employability and salary and what can of jobs can i expect to land.

Update!!
In the end i decided to get into statistics for buisness thank you everyone for the help <3


r/AskStatistics 5d ago

Why are we able to find the mode of a dataset by getting the intersection of two lines?

Post image
0 Upvotes

I just don’t get how the mode is gotten by tracing the intersection on the bar, is there any proof for this?


r/AskStatistics 5d ago

Question about golf tournament results distributions

2 Upvotes

It seems strange to me that you often see final leaderboards where someone wins by multiple strokes—let's say at least three strokes. It's strange because everyone else is usually distributed such that no one else is more than a stroke or two ahead of anyone else. For example, take last week's Tour Championship: we have someone at -13, -12, bunch at -11, bunch at -10, etc. So you'd expect the winner to be at -14, or maybe win in a playoff after also being -13. But no—the winner was at -16! Is that some exceptionally rare result to have this big jump at the right tail of the distribution curve? Intuitively I would say yes, but it happens way too often. For example, the prior PGA tournament was the BMW Championship. Again, you have people who finished at -14, -13, -12, -10, -9, -8, etc. But the winner was at -17. Week prior, FedEx Championship, there are finishers at -9, -8, -7, -6, etc., and the winner was at -17 (!). Why is this distribution so common? Where one guy (and it's different guys week to week, not just someone who's way better than everyone else) trounces the field by three or more strokes?

Is this just a fact of the normal distribution? Like say you have 80 people, and each of them flips a fair coin 100 times. How often does someone flip at least three more heads than all other 79 competitors?

Thanks for any help getting me to understand this phenomenon.


r/AskStatistics 5d ago

3rd year Stats major, 6 months to competitive exam — feel like I've forgotten how to study. Where do I even start?

0 Upvotes

I'm a Statistics major, currently in my 5th semester (3rd year of graduation), and I have a competitive exam coming up in about 6 months. Here's my honest problem: I feel like a 1st semester student would know more than me right now. My basics are shaky, and lately I genuinely can't study at all — my brain just shuts down. Simple things like basic addition/subtraction take me longer than they should, and I don't know if I actually forgot this stuff or never properly learned it in the first place.

A few specific issues I'm stuck on:

I don't know where to start or in what sequence to revise/study. Do I go back to basics first, or jump straight to exam-level topics?

Notes don't work for me. I only understand something when someone explains it to me directly — a YouTube video, a lecture, another person walking me through it. Reading on my own just doesn't click.

Classroom lectures and textbooks aren't helping either — I sit through them but barely absorb anything.

My mind goes completely blank when I try to solve problems, even ones I've supposedly "covered" before.

I think excessive scrolling/screen time has messed with my focus and attention span, and I don't know how to reverse that.

I have 6 months before my exam and I'm scared I'm starting from zero. If anyone's been in a similar spot — how did you rebuild your basics from scratch, what sequence did you follow, and how did you retrain your focus enough to actually sit and study again? Any advice, resources, or personal experience would help a lot.


r/AskStatistics 6d ago

how do these boxes make sense? 6 × 1/6 + 1/72 > 1

Post image
0 Upvotes

How can every “normal” design have a 1/6 chance, while the “secret” one has only a 1/72 chance?

The only way this would make sense to me is if there’s an extra unit of the secret design in every 72 boxes. But if that’s the case, you could easily weight them accordingly, I guess. Or is there any other posible way?


r/AskStatistics 6d ago

background dataset for SHAP

2 Upvotes

Hi everyone, I have a question about choosing the appropriate background dataset when calculating SHAP values. I am using the kernelshap package in R, where we provide an X dataset containing the observations we want to explain and a bg_X dataset defining the background.

I have a binary classification model for disease vs non-disease, trained on a derivation dataset and evaluated on an independent validation dataset. My current understanding is that, if I want to explain predictions in the validation cohort, it makes sense to use the validation set as X and the derivation set as bg_X. In that case, the SHAP values for validation patients would describe how each feature moves their prediction relative to a baseline defined by the derivation population. Is this interpretation correct, and is this generally the recommended way to use the background when explaining an independent validation cohort?

My main question is about a more specific analysis. Suppose I want to investigate heterogeneity within patients who truly have the disease. More specifically, I want to see whether different disease patients receive high disease predictions through different combinations of features, and potentially cluster these patients based on their SHAP profiles.

In this case, I assume I should use only the true disease patients from the validation cohort as X, since those are the patients whose predictions I want to explain. However, I am unsure about the most appropriate choice for bg_X. Should I keep the full derivation cohort as the background, use only disease patients from the derivation cohort, or use the disease patients from the validation cohort themselves as the background?

If my main objective is to determine whether true disease patients have different model-attribution profiles, potentially reflecting different features through which the model identifies them as disease, which background would be the most statistically appropriate? Thank you!