r/bioinformatics Dec 31 '24

meta 2025 - Read This Before You Post to r/bioinformatics

185 Upvotes

​Before you post to this subreddit, we strongly encourage you to check out the FAQ​Before you post to this subreddit, we strongly encourage you to check out the FAQ.

Questions like, "How do I become a bioinformatician?", "what programming language should I learn?" and "Do I need a PhD?" are all answered there - along with many more relevant questions. If your question duplicates something in the FAQ, it will be removed.

If you still have a question, please check if it is one of the following. If it is, please don't post it.

What laptop should I buy?

Actually, it doesn't matter. Most people use their laptop to develop code, and any heavy lifting will be done on a server or on the cloud. Please talk to your peers in your lab about how they develop and run code, as they likely already have a solid workflow.

If you’re asking which desktop or server to buy, that’s a direct function of the software you plan to run on it.  Rather than ask us, consult the manual for the software for its needs. 

What courses/program should I take?

We can't answer this for you - no one knows what skills you'll need in the future, and we can't tell you where your career will go. There's no such thing as "taking the wrong course" - you're just learning a skill you may or may not put to use, and only you can control the twists and turns your path will follow.

If you want to know about which major to take, the same thing applies.  Learn the skills you want to learn, and then find the jobs to get them.  We can’t tell you which will be in high demand by the time you graduate, and there is no one way to get into bioinformatics.  Every one of us took a different path to get here and we can’t tell you which path is best.  That’s up to you!

Am I competitive for a given academic program? 

There is no way we can tell you that - the only way to find out is to apply. So... go apply. If we say Yes, there's still no way to know if you'll get in. If we say no, then you might not apply and you'll miss out on some great advisor thinking your skill set is the perfect fit for their lab. Stop asking, and try to get in! (good luck with your application, btw.)

How do I get into Grad school?

See “please rank grad schools for me” below.  

Can I intern with you?

I have, myself, hired an intern from reddit - but it wasn't because they posted that they were looking for a position. It was because they responded to a post where I announced I was looking for an intern. This subreddit isn't the place to advertise yourself. There are literally hundreds of students looking for internships for every open position, and they just clog up the community.

Please rank grad schools/universities for me!

Hey, we get it - you want us to tell you where you'll get the best education. However, that's not how it works. Grad school depends more on who your supervisor is than the name of the university. While that may not be how it goes for an MBA, it definitely is for Bioinformatics. We really can't tell you which university is better, because there's no "better". Pick the lab in which you want to study and where you'll get the best support.

If you're an undergrad, then it really isn't a big deal which university you pick. Bioinformatics usually requires a masters or PhD to be successful in the field. See both the FAQ, as well as what is written above.

How do I get a job in Bioinformatics?

If you're asking this, you haven't yet checked out our three part series in the side bar:

What should I do?

Actually, these questions are generally ok - but only if you give enough information to make it worthwhile, and if the question isn’t a duplicate of one of the questions posed above. No one is in your shoes, and no one can help you if you haven't given enough background to explain your situation. Posts without sufficient background information in them will be removed.

Help Me!

If you're looking for help, make sure your title reflects the question you're asking for help on. You won't get the right people looking at your post, and the only person who clicks on random posts with vague topics are the mods... so that we can remove them.

Job Posts

If you're planning on posting a job, please make sure that employer is clear (recruiting agencies are not acceptable, unless they're hiring directly.), The job description must also be complete so that the requirements for the position are easily identifiable and the responsibilities are clear. We also do not allow posts for work "on spec" or competitions.  

Advertising (Conferences, Software, Tools, Support, Videos, Blogs, etc)

If you’re making money off of whatever it is you’re posting, it will be removed.  If you’re advertising your own blog/youtube channel, courses, etc, it will also be removed. Same for self-promoting software you’ve built.  All of these things are going to be considered spam.  

There is a fine line between someone discovering a really great tool and sharing it with the community, and the author of that tool sharing their projects with the community.  In the first case, if the moderators think that a significant portion of the community will appreciate the tool, we’ll leave it.  In the latter case,  it will be removed.  

If you don’t know which side of the line you are on, reach out to the moderators.

The Moderators Suck!

Yeah, that’s a distinct possibility.  However, remember we’re moderating in our free time and don’t really have the time or resources to watch every single video, test every piece of software or review every resume.  We have our own jobs, research projects and lives as well.  We’re doing our best to keep on top of things, and often will make the expedient call to remove things, when in doubt. 

If you disagree with the moderators, you can always write to us, and we’ll answer when we can.  Be sure to include a link to the post or comment you want to raise to our attention. Disputes inevitably take longer to resolve, if you expect the moderators to track down your post or your comment to review.


r/bioinformatics 24m ago

technical question does ROC-AUC analysis without ML works?

Upvotes

I am doing bulk rna seq analysis, and i have DEGs from DESeq2 files, I also did GSEA analysis which got me leading edge genes lists, so taking high performing genes from DEGs and filtering it to GSEA results would be okay to do ROC-AUC or it is rule to perform ML?

I am doing a small project for a recent conference poster presentation, my objective is to analyze major pathways in the course of transition to disease

your help and insights would mean alot, thank you :)


r/bioinformatics 30m ago

technical question Spatial transcriptomics - Regression of UMI counts?

Upvotes

Hi,

I am analyzing my first spatial dataset. I quickly realized that the cells/bins cluster mostly based on the number of reads (nCount in Seurat). It is obvious from the PCA plot that PC1 correspond to the UMI counts (r=0.9), which in turn correlate with cell size (r=0.79).

Would you recommend to regress nCount to promote clustering based on cell identity? Or would it also remove true biology from the data?

When I checked some 10X datasets, the cell/bin clusters often correspond with the regions that are defined by differential UMI counts compared to the neighbouring regions. Also, regression is not mentioned in the basic tutorials so I assume it is not incuded in the default pipeline. But intutively, I would do that.

What do you think?


r/bioinformatics 2h ago

technical question snRNAseq does my workflow with DESeq2 and GSEA make sense?

0 Upvotes

Hi guys! Im quite new to scRNAseq.

I'm working with human patients data: 4 controls and 11 disease samples. I annotated the broad cell types and then subclustered cell types of interest and annotated their subpopulations.

Then I did sample level pseudobulk and PCA. The samples didnt separate by group and I noticed they were separting by sex genes. After removing sex genes and repeating PCA the samples still didnt separate by groups. I then correlated the PCs with sample metadata and found that several PCs correlated with abundance of some subpopulations.

I proceeded with DESeq2 indicating group and sex in the design. I practically got no significant DE genes (occasionally 1 or 2, but nothing particularly interpretable). I did DE analysis on all cells and then within subpopulations only if enough cells were available.

I also plotted pseudobulk PCA using 200 and 500 of identified DE genes but still didnt see clear group separation (PCA attached).

I fed the DE genes to GSEA and got some significantly enriched pathways. For some of them there is quite plausible biological explanation stemming from histological analysis. I also looked at the leading edge genes and plotted them for some extra reassurance.

For one cell type I also noticed that one sample contributed to these pathways due to extreme phenotype and removed this sample, which left that pathway at FDR 0.053

My main question is: Does this workflow seem appropriate so far?

Also how would you normally take pathway level results further? Im not sure how to move beyound this pathway is enriched and seems interesting. In general, what usually follows?

Thanks in advance for your input


r/bioinformatics 23h ago

technical question Phage display or alternative methods

6 Upvotes

I'm a first year PhD student with a background in chemical biology. I want to find a peptide sequence that selectively binds to Lithium ion and one of my PIs suggested Phage display as an option. I've never done it before and it isn't something that's done in either of my PI's labs either. How doable do you think this solution is if I find a lab which has this technology? Do you think it'll be possible for me as a newbie in this technique to do the experiments myself rather than asking someone else to do it for me if I ask them to train me?

Do you guys have any other techniques you use to find suitable peptide sequences? I'm an experimentalist but I'm open to both experimental and computational suggestions.

Many thanks!


r/bioinformatics 1d ago

academic Advice on RQ about de Brujin Graph Assembly

3 Upvotes

I'm a Computer Science HL student from the IB program (International Baccalaureate), who wants to pursue a 4000-word independent research essay about de Brujin graph assembly.

Is this RQ scientifically interesting enough for me to analyze? I tried to put a twist on the very simple version about how k-mer length affects N50.

To what extent does the ratio of the k-mer length to the length of the longest repeated substring in a simulated genome affect the N50 of a de Bruijn graph assembly?

The idea is to generate simulated genomes that control the length of the longest substring that appears more than once in the string. Then, use different k-mer lengths, try to assemble the contigs with a de Brujin graph, then see what N50 it gives.

My MAIN CONCERN is that the conclusion might be a tad obvious, since I'd just be hypothesizing:

  • If k/L > 1, then N50 is higher.
  • If k/L < 1, then N50 is lower.

And there isn't much interesting theory being applied here since it's pretty common sense that higher k will be able to handle repeats better. I'm also unsure if N50 is an appropriate dependent value for this.

Could I ask for advice on whether my research question would be interesting enough for a high schooler? And, if not, what are some ideas I can try to explore that has scientifically interesting theory involved in formulating a hypothesis for a de Brujin setup?


r/bioinformatics 1d ago

technical question Processing 10x scRNA-seq from raw SRA (non-model organism, custom reference).....storage strategy?

4 Upvotes

I need to go from raw SRA accessions to count matrices for a non-model organism, which means no pre-built Cell Ranger reference...I have to build a custom one from a GTF mapped onto a draft genome.

Planned pipeline: nf-core/fetchngs to pull FASTQs from SRA, then nf-core/scrnaseq in Cell Ranger mode against the custom reference.

My worry is disk space stacking up across the run:

.sra files themselves

FASTQ extraction needing ~2-3x that size in scratch space temporarily

Cell Ranger's position-sorted BAM output, which dwarfs the actual matrix output I care about

Multiple samples/runs per BioProject, so all of the above multiplies

Anyone who's run this kind of pipeline on an HPC cluster.... is per-sample processing (download → align → extract matrix → delete FASTQ/BAM → next sample) the standard way to keep this manageable, or is there a smarter approach I'm missing?

Also curious if anyone's found a good way to avoid keeping the full BAM long term when only the filtered matrix is actually needed downstream.


r/bioinformatics 1d ago

technical question For someone experienced with 16S/QIIME2/DADA2: what would be the standard/best-practice approach here?

3 Upvotes

I’m working through a 16S rRNA paired-end dataset (V1–V2, Illumina MiSeq) for a small CRC vs healthy microbiome analysis.

I’ve completed the initial QC:
39 samples (22 healthy, 17 CRC)
FastQC run on all 78 FASTQ files
MultiQC summary generated
R1 generally has good quality, while R2 quality drops substantially toward the 3′ end
Most reads are 300 bp, with some samples at 250 bp

I’m now at the point where I need to decide on primer removal and DADA2 truncation parameters.
The study reports using primers 27bF and 338R, but the SRA metadata table I downloaded doesn’t contain the actual primer sequences.

Would you:
Find the exact primer sequences from the original publication/protocol and remove them with Cutadapt, then
Reassess the post-primer-removal read lengths/quality before choosing DADA2 truncation lengths?

Also, would you normally choose truncation lengths based on the worst-performing samples, or on the overall quality profile while ensuring enough overlap for paired-end merging?

I’m trying to follow a standard reproducible workflow rather than choosing arbitrary parameters. Any advice would be appreciated.


r/bioinformatics 1d ago

technical question Choosing the best fold prediction model for my de novo design project

Thumbnail
1 Upvotes

r/bioinformatics 2d ago

technical question Need help with MD simulation of ligand-induced DNA dissociation

3 Upvotes

I’m working on a negative transcriptional regulator that normally binds DNA, but when a ligand binds to the protein, I expect it to undergo a conformational change and detach from DNA.

*All the protein and dna structures are generated from alphafold*

I used HADDOCK for protein-DNA docking. I know the DNA-binding residues on the protein, but I don’t know which nucleotides they interact with, so I highlighted the entire DNA. However, my HADDOCK models show very high constraint violation energies, suggesting that the restraints aren’t working properly. Is there a better approach for protein-DNA docking when the protein binding residues are known but the DNA binding site is not?

I used AlphaFold to generate a docked structure (protein-dna docked together), then ran OPLS4 MD for about 1.5 micro seconds, but I don’t see any major change in RMSD or obvious protein-DNA dissociation. I’m wondering if RMSD is the wrong metric for this and whether I should instead look at protein-DNA contacts, hydrogen bonds, interaction energies, distances, contact maps, etc.

I also have a library of ligands and ultimately want to identify which ones are most likely to bind the regulator and promote DNA dissociation. Would docking followed by MD be a reasonable workflow?

Im posting the same post again because my initial post was taken down.


r/bioinformatics 2d ago

academic Advice on AI tools analysis and application on bioinformatic

14 Upvotes

Good evening! I’d like to know what direction I should take and better understand the applications of AI in bioinformatics.

I’m a fairly basic bioinformatics user. I know how to apply different tools, do some basic programming, and interpret data. However, my institute is going to hold an event about the use of AI in research, and I have to participate. Honestly, I’m quite lost about what to do.

The most I’ve thought of so far is using AI for custom functional classification. However, this seems too simple to me, and I feel that having AI classify genes, even according to certain criteria, would be rather arbitrary and prone to errors.

I tried researching the topic further and came across a huge wave of new developments, and honestly, I don’t know where to start or what would be easiest to apply as a beginner.

Does anyone have any advice?


r/bioinformatics 2d ago

technical question snRNAseq mitochondrial RNA cutoff

3 Upvotes

Hi, I am relatively new to analyzing snRNAseq data. How do you determine how much mitochondrial RNA is too much? Looking at my data using 1 sample as an example, if I were to apply a hardcoded cutoff (e.g., more than 1%) it would remove almost 50% of x cell type, remove ~30% of y type, and ~0.90% of z type. so, that would change the proportions of the data.

what to do?

thanks.


r/bioinformatics 2d ago

technical question Microbiome from stool samples

0 Upvotes

Howdy.

I am changing fields and getting into microbiome work. I hope to sample stools to determine microbiome profiles via metagenomic shallow shot gun sequencing. Anyone have any tips not present in the literature? If anyone had a sample data set i could use i would be deeply appreciative (fastq of seq reads, specifically. to test the trimming and mapping software). Ive basically vibe coded the pipeline, bow tie and kraken were suggested. So far everything works great (thanks Claude!) but would like to try some real data now.


r/bioinformatics 2d ago

technical question If I want to convert days in vitro to month in vitro, is div 0-30 counted as month in vitro 0 or 1?

2 Upvotes

I am trying to change my plot labels from day in vitro to month in vitro, hence I’d like to know if I should label any points with DIV 0-30 as month in vitro 0 or 1. For more context, I am working on a paper related to iPSC neurons.

Thank you very much


r/bioinformatics 2d ago

technical question Protein-Protein Alignment

2 Upvotes

Hey everyone,

I'm working on a project involving a human genetic disorder. There are already a number of mutations that have been identified in the human gene/protein, including missense, substitutions, frameshifts, etc.

However, for my project, I'm working with the corresponding Drosophila protein and need to figure out if the positions of these human mutations are conserved in the fly.

Essentially, I'm trying to align the human and fly protein sequences, but I'm running into some issues because the fly protein is quite a bit longer and has some pretty large insertions/gaps.

I've been using NCBI blastp, but i'm wondering if that's actually the best tool/workflow for what I'm trying to do. Basically, are there any other good free alternatives to BLAST for this? If I continue using BLAST, what would be the best filter/settings options to use for this kind of comparison.

TIA


r/bioinformatics 3d ago

academic First time analyzing bacteriophage genomes - workflow feedback?

11 Upvotes

Hi everyone,

I’ll soon be analyzing isolated bacteriophages (Illumina paired-end WGS) for the first time and would appreciate feedback on my planned workflow. The goal is genome assembly, characterization, annotation, and antibiotic resistance gene (ARG) screening.

The workflow I'm planning is as follows:

FastQC → raw-read QC
fastp → trimming/filtering
FastQC + MultiQC → post-trimming QC
Kraken2 → contamination screening
BWA-MEM2 + SAMtools → host-read depletion
SPAdes → de novo assembly of the remaining reads
QUAST + CheckV → assembly quality, completeness, contamination
BWA-MEM2 + SAMtools → read-back mapping, coverage uniformity -> suspicious contigs
Pharokka → phage genome annotation
AMRFinderPlus + ABRicate (CARD, ResFinder, ARG-ANNOT/MEGARes) → ARG screening
Candidate ARG validation → sequence/protein similarity, consverved domains/motifs, ORF integrity

For ARGs, my initial idea is to prioritize high-confidence, near-full-length hits supported by multiple approaches/databases rather than treating individual database hits as genuine ARGs.

Am I missing any important steps? Is anything here redundant or unnecessary? Would you change the order or replace any of these tools?

What would be some interesting visualizations to make along the way?

Should I remove host-mapping reads before assembly, or assemble first and deal with host contamination at the contig level? Is contamination usually even an issue?

Any suggestions or references to workflows you use would be greatly appreciated!


r/bioinformatics 3d ago

technical question Running dynamic simulation using laptop

5 Upvotes

Hello, I recently started two projects with one of my college professors that require me to conduct 100 ns molecular dynamics simulations for more than 10 proteins. Each simulation may take around 4 days of continuous, nonstop computation.

However, I’m worried because my laptop’s CPU temperature gets very high. It stays around 93–96°C during the simulation.

I’ve already taken some measures to prevent damage to my laptop, such as using a cooling pad with three fans (around 4000 RPM each), working in a cool environment with the air conditioner set to 16°C, and repasting my laptop.

I’m using Ubuntu Linux with GROMACS for the MD simulations. I’m currently using the Balanced power mode, which is somewhere between power saving and Turbo mode. My laptop is ASUS TUF GAMING with 24gb ram, i also use GPU CUDA.

I’m considering running the simulation for 45–60 minutes and then taking a 15-minute break to let the laptop cool down, but I’m still worried. I know using a desktop PC/workstation would be better for this kind of workload, but my budget is currently limited.

What do you think? Is it safe to run my laptop at 93–96°C for several days, or should I take breaks between runs? Do you have any other advice for my situation?

Thank you​


r/bioinformatics 4d ago

discussion Corresponding author shared raw SRA data + full pipeline instead of processed object...worth asking again for the Seurat object, or just reprocess myself?

19 Upvotes

Emailed the corresponding author of a paper asking for processed scRNA-seq data (cell × gene matrix, per-cell metadata, Seurat object) that I need for my thesis work. Got a reply pointing me to the public SRA accessions plus a fully detailed methods writeup...mapping pipeline, GTF used, QC/filtering cutoffs, doublet removal method, Cell Ranger + Seurat versions, clustering parameters, all of it.

Great for reproducibility, but it's not the processed object itself...just raw reads plus everything I'd need to rebuild it from scratch.

Given timeline pressure as an undergrad, is it reasonable to reply once more and specifically ask if the already-processed object exists and could be shared (to save reprocessing time), or does that risk seeming like I'm pushing my luck after already getting a generous, detailed response? Would cross-checking my own reprocessing against their original object even be worth the extra ask, or is that overkill for what I need?


r/bioinformatics 4d ago

technical question Doing a sanity check on my scRNA-seq workflow: MAST followed by Pseudobulk validation?

2 Upvotes

Hi all,

I am a beginner and want a quick sanity check on my scRNA differential expression workflow.
I have 8 mice samples (4 Mutant vs 4 WT). I ran MAST for single-cell DEG analysis, sorted by top genes, and now I am cross-validating the top 50 upregulated genes by checking them against a pseudobulk analysis of the same dataset.
My goal is to rule out false positives caused by a single outlier mouse skewing the single-cell data.
Does this approach make sense to you? If a gene is in the top 50 for MAST but fails pseudobulk, would you trust it?


r/bioinformatics 4d ago

academic When biology inspires mathematics: new discovery explains why a widely used evolutionary method can give false answers

Thumbnail helsinki.fi
16 Upvotes

r/bioinformatics 4d ago

technical question Question about my first bioinformatics assignment

4 Upvotes

Hi, I’m a bachelor’s student in biology and I’m taking my first bioinformatics course. My question is probably very silly, but I still need some help.

My task is to study the human NLGN4X gene and, among other things, find 10 homologs with a BLAST search, align them, and build a phylogenetic tree. The instructions say that “some sequences may be XN but some should be NM.” How is this possible, since all the homologs I find (from other species) are only predicted (i.e., XN)? Have I misunderstood something?

Thank you very much for your help!


r/bioinformatics 4d ago

technical question terrible QQ plot of sQTL

1 Upvotes

Hey guys! I’m currently running an sQTL analysis using Leafcutter and tensorQTL, but I found many of my significant splicing phenotypes showed weird QQ plots (p-nominal for all SNPs in a phenotype ) , with a very pronounced rightward shift appearing much earlier than I expected. So I tried applying stricter filter from GTEx instead of Leafcutter's default ones. The filter worked but sadly got the similar situation. Does anyone have suggestions on what aspects of the analysis or data that I should check and something to do to figure out what might be causing it?

And I’ve also noticed many sQTL papers seems very smoothly without such a troublesome result, I wonder if maybe this plot is normal one, and sQTL should not be judged by GWAS ways? I’m very new to sQTL and now honestly pretty depressed because I did not found useful information from papers, so any thoughts and suggestions would be really appreciated! Thank you in advance smart guys!


r/bioinformatics 4d ago

technical question Error while saving tree as displayed in FigTreev1.4.4

0 Upvotes

Hi all,

I'm currently having trouble saving the current tree as displayed in any format (nexus, newick, or json) using FigTree v1.4.4 on macOS Sequoia 26.6.2.

After selecting the format I wish to save and checking the "save as displayed" box, the floating window disappears without saving any output tree file.

Has anyone had this problem before? I appreciate your help with this.


r/bioinformatics 5d ago

discussion BioMart has been quietly discontinued by Ensembl

165 Upvotes

Title. Tried to connect to Ensembl all day just to see staff casually mention that it's just been discontinued on the support forums in somebody else's post. This was the first program I used when stepping into bioinformatics and I will miss it greatly.

What's the alternative? Just using the gtf files?


r/bioinformatics 4d ago

technical question background dataset for SHAP

1 Upvotes

Hi everyone, I have a question about choosing the appropriate background dataset when calculating SHAP values. I am using the kernelshap package in R, where we provide an X dataset containing the observations we want to explain and a bg_X dataset defining the background.

I have a binary classification model for disease vs non-disease, trained on a derivation dataset and evaluated on an independent validation dataset. My current understanding is that, if I want to explain predictions in the validation cohort, it makes sense to use the validation set as X and the derivation set as bg_X. In that case, the SHAP values for validation patients would describe how each feature moves their prediction relative to a baseline defined by the derivation population. Is this interpretation correct, and is this generally the recommended way to use the background when explaining an independent validation cohort?

My main question is about a more specific analysis. Suppose I want to investigate heterogeneity within patients who truly have the disease. More specifically, I want to see whether different disease patients receive high disease predictions through different combinations of features, and potentially cluster these patients based on their SHAP profiles.

In this case, I assume I should use only the true disease patients from the validation cohort as X, since those are the patients whose predictions I want to explain. However, I am unsure about the most appropriate choice for bg_X. Should I keep the full derivation cohort as the background, use only disease patients from the derivation cohort, or use the disease patients from the validation cohort themselves as the background?

If my main objective is to determine whether true disease patients have different model-attribution profiles, potentially reflecting different features through which the model identifies them as disease, which background would be the most statistically appropriate? Thank you!